@ninjaxtools/slopdex 0.9.0 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/slopdex/SKILL.md +29 -72
- package/README.md +38 -58
- package/dist/cli.js +460 -300
- package/dist/cli.js.map +1 -1
- package/dist/index.d.ts +42 -40
- package/dist/index.js +351 -230
- package/dist/index.js.map +1 -1
- package/docs/implementation.md +24 -20
- package/package.json +1 -1
package/docs/implementation.md
CHANGED
|
@@ -7,12 +7,12 @@ For installation, command examples, configuration, and result interpretation, se
|
|
|
7
7
|
| Module | Responsibility |
|
|
8
8
|
| --- | --- |
|
|
9
9
|
| `src/cli.ts` | Argument parsing, configuration, automatic refresh/recovery, diagnostics, and output selection. |
|
|
10
|
-
| `src/code-index.ts` | Index lifecycle, file preparation, Git/working-tree reconciliation, embedding/
|
|
10
|
+
| `src/code-index.ts` | Index lifecycle, file preparation, Git/working-tree reconciliation, embedding/description caching, and search facade. |
|
|
11
11
|
| `src/parser/` | Language dispatch, native Tree-sitter extraction, callable identity, and recoverable diagnostics. |
|
|
12
12
|
| `src/source-policy.ts`, `src/gitignore.ts` | Supported paths, built-in/config exclusions, nested ignore rules. |
|
|
13
13
|
| `src/git/repository.ts` | Git commits, trees, blobs, diffs, ancestry, and working-tree changes. |
|
|
14
|
-
| `src/embeddings/`, `src/
|
|
15
|
-
| `src/storage/database.ts` | SQLite schema,
|
|
14
|
+
| `src/embeddings/`, `src/descriptions/` | Provider requests and provider profiles. |
|
|
15
|
+
| `src/storage/database.ts` | SQLite schema, durable artifact caches, transactions, metadata, and vector queries. |
|
|
16
16
|
| `src/search/` | Analysis scoring selection and cross-index neighbor discovery. |
|
|
17
17
|
| `src/analysis/cohesion.ts` | Physical distance, gap scores, aggregate affinity, and groups. |
|
|
18
18
|
| `src/format.ts` | Human-readable search, cluster, and cohesion output. |
|
|
@@ -44,18 +44,20 @@ The CLI reconciles an index before most commands. Library callers choose when to
|
|
|
44
44
|
|
|
45
45
|
Git updates resolve the target commit, validate ancestry against the checkpoint, reconcile committed blobs, and optionally overlay current working-tree files when the target equals HEAD. Overlays account for staged, unstaged, untracked, renamed, and deleted paths. The saved checkpoint remains the committed base. Historical/other-branch targets are committed-only. Explicit `updateFiles` operations do not advance the checkpoint.
|
|
46
46
|
|
|
47
|
-
|
|
47
|
+
Tree-sitter results and each valid provider result are committed to content-addressed cache tables immediately. Generation checks detect concurrent logical index changes; working-tree updates also verify source and ignore-rule stability. A separate SQLite transaction atomically applies related file, function, diagnostic, description-reference, and checkpoint changes. Failed or aborted initialization retains the valid database and its cache rows so the next invocation resumes without repeating completed work.
|
|
48
48
|
|
|
49
|
-
Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable WAL, foreign keys, and
|
|
49
|
+
Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable WAL, foreign keys, and full synchronization. The schema contains:
|
|
50
50
|
|
|
51
|
-
- `metadata`: repository root, embedding and
|
|
51
|
+
- `metadata`: repository root, embedding and description profiles, generation, checkpoint, schema version, and feature/scan state.
|
|
52
52
|
- `files`: paths, content hashes, blob IDs, source mode, previous paths, language, and size.
|
|
53
|
-
- `functions`: identity, names, signatures, locations, source, provenance, and embedding/
|
|
54
|
-
- `embeddings
|
|
53
|
+
- `functions`: identity, names, signatures, locations, source, provenance, and embedding/description references.
|
|
54
|
+
- `embeddings`: code, description, and query vectors keyed by embedding profile, operation, and exact input.
|
|
55
|
+
- `description_cache`: generated description text keyed independently by description profile and complete source context.
|
|
56
|
+
- `parse_cache`: successful Tree-sitter extraction results keyed by parser strategy, path, and file-content hash.
|
|
55
57
|
- `callable_provenance`: first-seen committed source identity.
|
|
56
58
|
- `indexing_errors`: diagnostics associated with files.
|
|
57
59
|
|
|
58
|
-
Current schema version is `
|
|
60
|
+
Current schema version is `6`. Earlier schemas are intentionally incompatible and require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. For a schema-6 index, `--force-reindex` clears logical index state while retaining content-addressed caches; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
|
|
59
61
|
|
|
60
62
|
### Diagnostics
|
|
61
63
|
|
|
@@ -63,28 +65,30 @@ Diagnostics cover parse errors, parser exceptions, extraction failures, read fai
|
|
|
63
65
|
|
|
64
66
|
Diagnostics commit with their corresponding file update. Updates retry failed files even when their Git blobs are unchanged; successful replacement, deletion, and exclusion clear failures. Standalone readers inspect saved diagnostics without constructing an embedding provider. The CLI's exit handler reports remaining failures for source and target indexes; version exits before registering that handler.
|
|
65
67
|
|
|
66
|
-
## Embeddings and
|
|
68
|
+
## Embeddings and descriptions
|
|
67
69
|
|
|
68
70
|
Embedding profiles consist of provider, model, dimensions, and strategy version. Cross-index analysis requires matching profiles. Vectors are validated and normalized before storage/search.
|
|
69
71
|
|
|
70
|
-
- OpenAI defaults to `text-embedding-3-large`, 3072 dimensions, strategy `callable-
|
|
71
|
-
- Jina defaults to `jina-embeddings-v4`, 1024 dimensions, strategy `callable-
|
|
72
|
+
- OpenAI defaults to `text-embedding-3-large`, 3072 dimensions, strategy `callable-v2`. Inputs are truncated to 8192 `cl100k_base` tokens.
|
|
73
|
+
- Jina defaults to `jina-embeddings-v4`, 1024 dimensions, strategy `callable-v2:code-query-passage`. Requests distinguish `code.passage` documents from `code.query` queries and enable truncation.
|
|
74
|
+
|
|
75
|
+
Embedding inputs identify language, callable kind, qualified symbol, signature, documentation, and source. For Python, a function's first-statement docstring is included in a separate `documentation` section in addition to remaining part of the callable source.
|
|
72
76
|
|
|
73
77
|
Purpose generation uses OpenAI's Responses API with `gpt-5.6-sol` and strategy `callable-purpose-v1`. The prompt asks for one to three sentences describing responsibility and visible relationships, using repository name, path, callable source, and full file context. Requests use `store: false`. Generated text is embedded with the configured embedding provider.
|
|
74
78
|
|
|
75
|
-
|
|
79
|
+
Description inputs include contextual and profile information so file-context, path, or model changes invalidate relevant cached results. Description text identity is independent of the embedding profile, allowing an embedding-model change to reuse generation output while producing the required new vector. Every validated description and vector is cached before indexing continues. `useDescriptions` persists the profile and enabled state; `disableDescriptions` turns automatic updates and description search/scoring off without deleting cached artifacts. Function references are attached only by the final logical transaction, so provider failure cannot expose partially updated callable records. Deleting a function removes it from description search but retains reusable cache rows.
|
|
76
80
|
|
|
77
81
|
## Similarity and analysis
|
|
78
82
|
|
|
79
|
-
`search` embeds a query and searches code vectors; `search-
|
|
83
|
+
`search` embeds a query and searches code vectors; `search-description` searches description vectors. Analysis uses code-only cosine similarity unless all indexed callables have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
|
|
80
84
|
|
|
81
|
-
When
|
|
85
|
+
When descriptions are complete:
|
|
82
86
|
|
|
83
87
|
```text
|
|
84
|
-
similarity = 0.5 * codeSimilarity + 0.5 *
|
|
88
|
+
similarity = 0.5 * codeSimilarity + 0.5 * descriptionSimilarity
|
|
85
89
|
```
|
|
86
90
|
|
|
87
|
-
Both component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and
|
|
91
|
+
Both component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and description profiles.
|
|
88
92
|
|
|
89
93
|
Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
|
|
90
94
|
|
|
@@ -100,7 +104,7 @@ separationWeight = 1 - exp(-physicalDistance / 2)
|
|
|
100
104
|
cohesionGap = semanticWeight * separationWeight
|
|
101
105
|
```
|
|
102
106
|
|
|
103
|
-
Affinity ratios and mean distance are weighted by `semanticWeight`.
|
|
107
|
+
Affinity ratios and mean distance are weighted by `semanticWeight`. Cohesion metrics use all qualifying edges before output limiting. File reports aggregate internal, same-folder, and external affinity for selected-source files. Pairs rank by gap and tie-breakers; groups are connected components of reported pairs. Source/test classification is a path heuristic, not a dependency or call-graph analysis.
|
|
104
108
|
|
|
105
109
|
## Library API
|
|
106
110
|
|
|
@@ -124,12 +128,12 @@ try {
|
|
|
124
128
|
}
|
|
125
129
|
```
|
|
126
130
|
|
|
127
|
-
Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `JinaEmbeddingProvider`, `
|
|
131
|
+
Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `JinaEmbeddingProvider`, `OpenAIDescriptionProvider`, error types, and the contracts in `src/types.ts`. Standalone update/search helpers wrap the corresponding index methods.
|
|
128
132
|
|
|
129
133
|
- Use `updateFromWorkingTree()` when Git is unavailable. Unlike the CLI, the library does not automatically refresh before queries or fall back from Git.
|
|
130
134
|
- Set `sourceFilter.nameRegex` for source-only cross-search/cohesion filtering; combine it with `path` and a filter type (`all`, `changed-since`, or `uncommitted`). `changed-since` also accepts `uncommitted: true`.
|
|
131
135
|
- Top-level analysis `nameRegex` filters both sources and candidates. Query `SimilaritySearchOptions.nameRegex` filters result names before limiting.
|
|
132
|
-
- Call `await index.
|
|
136
|
+
- Call `await index.useDescriptions()`, then `await index.searchDescription({ query: "maintain the repository index" })`. Select a model via `descriptionProvider: new OpenAIDescriptionProvider({ model: "gpt-5.6-sol" })` in index options. Custom providers implement `EmbeddingProvider` or `DescriptionProvider`.
|
|
133
137
|
- Inspect failures through `index.indexErrors()` or exported `readIndexErrors(indexPath)` without a provider. Records use `IndexingError`.
|
|
134
138
|
- Public cohesion reports retain indexed-function data; the CLI presents compact function references and includes source bodies only with `--include-source`.
|
|
135
139
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@ninjaxtools/slopdex",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.11.0",
|
|
4
4
|
"description": "Tree-sitter callable-level semantic indexing, search, and duplicate discovery for Python, JavaScript, TypeScript, Rust, Go, Java, and C",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"repository": {
|