@ninjaxtools/slopdex 0.10.0 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/slopdex/SKILL.md +31 -14
- package/README.md +45 -18
- package/dist/cli.js +1121 -425
- package/dist/cli.js.map +1 -1
- package/dist/index.d.ts +86 -41
- package/dist/index.js +812 -370
- package/dist/index.js.map +1 -1
- package/docs/implementation.md +21 -19
- package/package.json +6 -1
package/docs/implementation.md
CHANGED
|
@@ -7,12 +7,12 @@ For installation, command examples, configuration, and result interpretation, se
|
|
|
7
7
|
| Module | Responsibility |
|
|
8
8
|
| --- | --- |
|
|
9
9
|
| `src/cli.ts` | Argument parsing, configuration, automatic refresh/recovery, diagnostics, and output selection. |
|
|
10
|
-
| `src/code-index.ts` | Index lifecycle, file preparation, Git/working-tree reconciliation, embedding/
|
|
10
|
+
| `src/code-index.ts` | Index lifecycle, file preparation, Git/working-tree reconciliation, embedding/description caching, and search facade. |
|
|
11
11
|
| `src/parser/` | Language dispatch, native Tree-sitter extraction, callable identity, and recoverable diagnostics. |
|
|
12
12
|
| `src/source-policy.ts`, `src/gitignore.ts` | Supported paths, built-in/config exclusions, nested ignore rules. |
|
|
13
13
|
| `src/git/repository.ts` | Git commits, trees, blobs, diffs, ancestry, and working-tree changes. |
|
|
14
|
-
| `src/embeddings/`, `src/
|
|
15
|
-
| `src/storage/database.ts` | SQLite schema,
|
|
14
|
+
| `src/embeddings/`, `src/descriptions/` | Provider requests and provider profiles. |
|
|
15
|
+
| `src/storage/database.ts` | SQLite schema, durable artifact caches, transactions, metadata, and vector queries. |
|
|
16
16
|
| `src/search/` | Analysis scoring selection and cross-index neighbor discovery. |
|
|
17
17
|
| `src/analysis/cohesion.ts` | Physical distance, gap scores, aggregate affinity, and groups. |
|
|
18
18
|
| `src/format.ts` | Human-readable search, cluster, and cohesion output. |
|
|
@@ -44,18 +44,20 @@ The CLI reconciles an index before most commands. Library callers choose when to
|
|
|
44
44
|
|
|
45
45
|
Git updates resolve the target commit, validate ancestry against the checkpoint, reconcile committed blobs, and optionally overlay current working-tree files when the target equals HEAD. Overlays account for staged, unstaged, untracked, renamed, and deleted paths. The saved checkpoint remains the committed base. Historical/other-branch targets are committed-only. Explicit `updateFiles` operations do not advance the checkpoint.
|
|
46
46
|
|
|
47
|
-
|
|
47
|
+
Tree-sitter results and each valid provider result are committed to content-addressed cache tables immediately. Generation checks detect concurrent logical index changes; working-tree updates also verify source and ignore-rule stability. A separate SQLite transaction atomically applies related file, function, diagnostic, description-reference, and checkpoint changes. Failed or aborted initialization retains the valid database and its cache rows so the next invocation resumes without repeating completed work.
|
|
48
48
|
|
|
49
|
-
Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable WAL, foreign keys, and
|
|
49
|
+
Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable WAL, foreign keys, and full synchronization. The schema contains:
|
|
50
50
|
|
|
51
|
-
- `metadata`: repository root, embedding and
|
|
51
|
+
- `metadata`: repository root, embedding and description profiles, generation, checkpoint, schema version, and feature/scan state.
|
|
52
52
|
- `files`: paths, content hashes, blob IDs, source mode, previous paths, language, and size.
|
|
53
|
-
- `functions`: identity, names, signatures, locations, source, provenance, and embedding/
|
|
54
|
-
- `embeddings
|
|
53
|
+
- `functions`: identity, names, signatures, locations, source, provenance, and embedding/description references.
|
|
54
|
+
- `embeddings`: code, description, and query vectors keyed by embedding profile, operation, and exact input.
|
|
55
|
+
- `description_cache`: generated description text keyed independently by description profile and complete source context.
|
|
56
|
+
- `parse_cache`: successful Tree-sitter extraction results keyed by parser strategy, path, and file-content hash.
|
|
55
57
|
- `callable_provenance`: first-seen committed source identity.
|
|
56
58
|
- `indexing_errors`: diagnostics associated with files.
|
|
57
59
|
|
|
58
|
-
Current schema version is `
|
|
60
|
+
Current schema version is `7`. Schema 6 is migrated in place by adding file-description state; earlier schemas require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. Forced rebuilds clear logical index state while retaining content-addressed caches when possible; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
|
|
59
61
|
|
|
60
62
|
### Diagnostics
|
|
61
63
|
|
|
@@ -63,7 +65,7 @@ Diagnostics cover parse errors, parser exceptions, extraction failures, read fai
|
|
|
63
65
|
|
|
64
66
|
Diagnostics commit with their corresponding file update. Updates retry failed files even when their Git blobs are unchanged; successful replacement, deletion, and exclusion clear failures. Standalone readers inspect saved diagnostics without constructing an embedding provider. The CLI's exit handler reports remaining failures for source and target indexes; version exits before registering that handler.
|
|
65
67
|
|
|
66
|
-
## Embeddings and
|
|
68
|
+
## Embeddings and descriptions
|
|
67
69
|
|
|
68
70
|
Embedding profiles consist of provider, model, dimensions, and strategy version. Cross-index analysis requires matching profiles. Vectors are validated and normalized before storage/search.
|
|
69
71
|
|
|
@@ -72,21 +74,21 @@ Embedding profiles consist of provider, model, dimensions, and strategy version.
|
|
|
72
74
|
|
|
73
75
|
Embedding inputs identify language, callable kind, qualified symbol, signature, documentation, and source. For Python, a function's first-statement docstring is included in a separate `documentation` section in addition to remaining part of the callable source.
|
|
74
76
|
|
|
75
|
-
Purpose generation uses OpenAI's Responses API with `gpt-5.6-sol
|
|
77
|
+
Purpose generation uses the AI SDK and strategy `callable-purpose-v2`. The default is OpenAI's Responses API with `gpt-5.6-sol`; OpenCode Zen and Go are also selectable, defaulting to `gpt-5.6-sol` and `gpt-5.6-luna`. OpenCode catalog models are routed through their published protocol using the OpenAI Responses, Anthropic Messages, Google Generative AI, or OpenAI-compatible adapter. Generation opens one conversation per file: stable instructions and complete file context form the prefix, the first request describes the file overall, then callable prompts and generated answers are appended sequentially in source order. This avoids repeating the file within a request and gives provider prompt caches an increasingly large reusable prefix. Responses requests use `store: false`. File and callable text are embedded with the configured embedding provider.
|
|
76
78
|
|
|
77
|
-
|
|
79
|
+
Description inputs include contextual and profile information so file-context, path, model, or generation-strategy changes invalidate relevant cached results. Description text identity is independent of the embedding profile, allowing an embedding-model change to reuse generation output while producing the required new vector. Every validated description and vector is cached before indexing continues. If generation resumes partway through a file, completed file/callable prompts and cached answers are replayed locally before the next request so the conversation prefix remains equivalent. Ordinary source updates retain the previous file description and its source hash, making staleness explicit without incurring automatic regeneration; `reindex-files` replaces stale file descriptions, optionally continuing through callable regeneration. `useDescriptions` persists the profile and enabled state; `disableDescriptions` turns automatic updates and description search/scoring off without deleting cached artifacts. Function references are attached only by the final logical transaction, so provider failure cannot expose partially updated callable records. Deleting a function removes it from description search but retains reusable cache rows.
|
|
78
80
|
|
|
79
81
|
## Similarity and analysis
|
|
80
82
|
|
|
81
|
-
`search` embeds a query and
|
|
83
|
+
`search` embeds a query and, when descriptions are complete, scores code, callable-description, and containing-file-description vectors. `search-description` scores callable and file descriptions. Analysis uses code-only cosine similarity unless all indexed callables and files have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
|
|
82
84
|
|
|
83
|
-
When
|
|
85
|
+
When descriptions are complete:
|
|
84
86
|
|
|
85
87
|
```text
|
|
86
|
-
similarity =
|
|
88
|
+
similarity = (codeSimilarity + descriptionSimilarity + fileDescriptionSimilarity) / 3
|
|
87
89
|
```
|
|
88
90
|
|
|
89
|
-
|
|
91
|
+
All component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and description profiles. Cohesion JSON uses schema version 3 for the three-component scoring contract.
|
|
90
92
|
|
|
91
93
|
Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
|
|
92
94
|
|
|
@@ -102,7 +104,7 @@ separationWeight = 1 - exp(-physicalDistance / 2)
|
|
|
102
104
|
cohesionGap = semanticWeight * separationWeight
|
|
103
105
|
```
|
|
104
106
|
|
|
105
|
-
Affinity ratios and mean distance are weighted by `semanticWeight`.
|
|
107
|
+
Affinity ratios and mean distance are weighted by `semanticWeight`. Cohesion metrics use all qualifying edges before output limiting. File reports aggregate internal, same-folder, and external affinity for selected-source files. Pairs rank by gap and tie-breakers; groups are connected components of reported pairs. Source/test classification is a path heuristic, not a dependency or call-graph analysis.
|
|
106
108
|
|
|
107
109
|
## Library API
|
|
108
110
|
|
|
@@ -126,12 +128,12 @@ try {
|
|
|
126
128
|
}
|
|
127
129
|
```
|
|
128
130
|
|
|
129
|
-
Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `JinaEmbeddingProvider`, `
|
|
131
|
+
Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `JinaEmbeddingProvider`, `OpenAIDescriptionProvider`, error types, and the contracts in `src/types.ts`. Standalone update/search helpers wrap the corresponding index methods.
|
|
130
132
|
|
|
131
133
|
- Use `updateFromWorkingTree()` when Git is unavailable. Unlike the CLI, the library does not automatically refresh before queries or fall back from Git.
|
|
132
134
|
- Set `sourceFilter.nameRegex` for source-only cross-search/cohesion filtering; combine it with `path` and a filter type (`all`, `changed-since`, or `uncommitted`). `changed-since` also accepts `uncommitted: true`.
|
|
133
135
|
- Top-level analysis `nameRegex` filters both sources and candidates. Query `SimilaritySearchOptions.nameRegex` filters result names before limiting.
|
|
134
|
-
- Call `await index.
|
|
136
|
+
- Call `await index.useDescriptions()`, then `await index.searchDescription({ query: "maintain the repository index" })`. Select a model via `descriptionProvider: new OpenAIDescriptionProvider({ model: "gpt-5.6-sol" })` in index options. Custom description providers implement both stateless file/callable methods and may add `startFile()` for contextual sessions.
|
|
135
137
|
- Inspect failures through `index.indexErrors()` or exported `readIndexErrors(indexPath)` without a provider. Records use `IndexingError`.
|
|
136
138
|
- Public cohesion reports retain indexed-function data; the CLI presents compact function references and includes source bodies only with `--include-source`.
|
|
137
139
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@ninjaxtools/slopdex",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.13.0",
|
|
4
4
|
"description": "Tree-sitter callable-level semantic indexing, search, and duplicate discovery for Python, JavaScript, TypeScript, Rust, Go, Java, and C",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"repository": {
|
|
@@ -43,6 +43,11 @@
|
|
|
43
43
|
"check:parser-parity": "node --import tsx scripts/check-parser-parity.mjs"
|
|
44
44
|
},
|
|
45
45
|
"dependencies": {
|
|
46
|
+
"@ai-sdk/anthropic": "^4.0.53",
|
|
47
|
+
"@ai-sdk/google": "^4.0.69",
|
|
48
|
+
"@ai-sdk/openai": "^4.0.66",
|
|
49
|
+
"@ai-sdk/openai-compatible": "^3.0.48",
|
|
50
|
+
"ai": "^7.0.99",
|
|
46
51
|
"ignore": "^7.0.9",
|
|
47
52
|
"install": "^0.13.0",
|
|
48
53
|
"js-tiktoken": "^1.0.21",
|