@ninjaxtools/slopdex 0.13.0 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/slopdex/SKILL.md +33 -32
- package/README.md +45 -38
- package/dist/cli.js +583 -462
- package/dist/cli.js.map +1 -1
- package/dist/index.d.ts +91 -2
- package/dist/index.js +393 -27
- package/dist/index.js.map +1 -1
- package/docs/implementation.md +14 -8
- package/package.json +1 -1
package/docs/implementation.md
CHANGED
|
@@ -11,11 +11,11 @@ For installation, command examples, configuration, and result interpretation, se
|
|
|
11
11
|
| `src/parser/` | Language dispatch, native Tree-sitter extraction, callable identity, and recoverable diagnostics. |
|
|
12
12
|
| `src/source-policy.ts`, `src/gitignore.ts` | Supported paths, built-in/config exclusions, nested ignore rules. |
|
|
13
13
|
| `src/git/repository.ts` | Git commits, trees, blobs, diffs, ancestry, and working-tree changes. |
|
|
14
|
-
| `src/embeddings/`, `src/descriptions/` | Provider requests and provider profiles. |
|
|
14
|
+
| `src/embeddings/`, `src/descriptions/`, `src/rerankers/` | Provider requests and provider profiles. |
|
|
15
15
|
| `src/storage/database.ts` | SQLite schema, durable artifact caches, transactions, metadata, and vector queries. |
|
|
16
16
|
| `src/search/` | Analysis scoring selection and cross-index neighbor discovery. |
|
|
17
17
|
| `src/analysis/cohesion.ts` | Physical distance, gap scores, aggregate affinity, and groups. |
|
|
18
|
-
| `src/format.ts` | Human-readable search
|
|
18
|
+
| `src/format.ts` | Human-readable search and cluster output, plus library cohesion-report formatting. |
|
|
19
19
|
| `src/index.ts`, `src/types.ts` | Public exports and data contracts. |
|
|
20
20
|
|
|
21
21
|
## Callable extraction
|
|
@@ -52,12 +52,13 @@ Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable
|
|
|
52
52
|
- `files`: paths, content hashes, blob IDs, source mode, previous paths, language, and size.
|
|
53
53
|
- `functions`: identity, names, signatures, locations, source, provenance, and embedding/description references.
|
|
54
54
|
- `embeddings`: code, description, and query vectors keyed by embedding profile, operation, and exact input.
|
|
55
|
+
- `function_vectors`: synchronized `vec0` storage for filtered code-vector nearest-neighbor queries.
|
|
55
56
|
- `description_cache`: generated description text keyed independently by description profile and complete source context.
|
|
56
57
|
- `parse_cache`: successful Tree-sitter extraction results keyed by parser strategy, path, and file-content hash.
|
|
57
58
|
- `callable_provenance`: first-seen committed source identity.
|
|
58
59
|
- `indexing_errors`: diagnostics associated with files.
|
|
59
60
|
|
|
60
|
-
Current schema version is `
|
|
61
|
+
Current schema version is `8`. Schemas 6 and 7 are migrated in place by adding file-description state when necessary and building the nearest-neighbor vector table; earlier schemas require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. Forced rebuilds clear logical index state while retaining content-addressed caches when possible; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
|
|
61
62
|
|
|
62
63
|
### Diagnostics
|
|
63
64
|
|
|
@@ -80,7 +81,11 @@ Description inputs include contextual and profile information so file-context, p
|
|
|
80
81
|
|
|
81
82
|
## Similarity and analysis
|
|
82
83
|
|
|
83
|
-
`search` embeds a query and, when descriptions are complete, scores code, callable-description, and containing-file-description vectors. `search-description` scores callable and file descriptions. Analysis uses code-only cosine similarity unless all indexed callables and files have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
|
|
84
|
+
`search` embeds a query and, when descriptions are complete, scores code, callable-description, and containing-file-description vectors. `search-description` scores callable and file descriptions. Filter-compatible code-only searches up to sqlite-vec's 8,192-dimension limit use the synchronized `vec0` nearest-neighbor table; larger custom profiles plus fused, regex, upper-bound, and multi-path searches retain the exact scalar scoring path. Analysis uses code-only cosine similarity unless all indexed callables and files have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
|
|
85
|
+
|
|
86
|
+
An optional `Reranker` performs a second-stage pass for the two natural-language query methods. After applying name and similarity filters, Cohere and Jina retrieve five times the requested result limit. `OpenAILLMReranker` retrieves its configured candidate count (10 by default), or the result limit when larger. Every reranker receives the query plus candidate path, code metadata/source, and purpose description when available, then returns the requested number in relevance order. Results preserve the embedding/fused `similarity` and add `rerankScore`.
|
|
87
|
+
|
|
88
|
+
Cohere defaults to `rerank-v4.0-pro` with `COHERE_API_KEY`; Jina defaults to `jina-reranker-v3.5` with `JINA_API_KEY`. The OpenAI LLM path defaults to `gpt-5.6-luna`, sends a strict JSON schema through the Responses API with high reasoning, no reasoning summary, and `store: false`, and validates result cardinality, indexes, uniqueness, and 0-1 scores. Its prompt preserves descriptions before source and caps source-bearing candidate text at 12,000 tokens each and 80,000 tokens in aggregate. Reranking does not participate in index metadata or cross-search because it creates no persisted artifacts and cross-search is callable-to-callable analysis rather than natural-language retrieval.
|
|
84
89
|
|
|
85
90
|
When descriptions are complete:
|
|
86
91
|
|
|
@@ -90,7 +95,7 @@ similarity = (codeSimilarity + descriptionSimilarity + fileDescriptionSimilarity
|
|
|
90
95
|
|
|
91
96
|
All component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and description profiles. Cohesion JSON uses schema version 3 for the three-component scoring contract.
|
|
92
97
|
|
|
93
|
-
Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
|
|
98
|
+
Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. With the `cohesion` option, each selected match receives its physical path distance and matches are re-ranked by descending distance, then similarity. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
|
|
94
99
|
|
|
95
100
|
### Cohesion metrics
|
|
96
101
|
|
|
@@ -109,11 +114,12 @@ Affinity ratios and mean distance are weighted by `semanticWeight`. Cohesion met
|
|
|
109
114
|
## Library API
|
|
110
115
|
|
|
111
116
|
```ts
|
|
112
|
-
import { OpenAIEmbeddingProvider, openCodeIndex } from "@ninjaxtools/slopdex";
|
|
117
|
+
import { OpenAIEmbeddingProvider, OpenAILLMReranker, openCodeIndex } from "@ninjaxtools/slopdex";
|
|
113
118
|
|
|
114
119
|
const index = openCodeIndex({
|
|
115
120
|
rootDir: "/path/to/repository",
|
|
116
121
|
provider: new OpenAIEmbeddingProvider(),
|
|
122
|
+
reranker: new OpenAILLMReranker({ candidateCount: 10 }),
|
|
117
123
|
});
|
|
118
124
|
|
|
119
125
|
try {
|
|
@@ -128,14 +134,14 @@ try {
|
|
|
128
134
|
}
|
|
129
135
|
```
|
|
130
136
|
|
|
131
|
-
Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `
|
|
137
|
+
Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, embedding/description providers, `CohereReranker`, `JinaReranker`, `OpenAILLMReranker`, error types, and the contracts in `src/types.ts`. Standalone update/search helpers wrap the corresponding index methods.
|
|
132
138
|
|
|
133
139
|
- Use `updateFromWorkingTree()` when Git is unavailable. Unlike the CLI, the library does not automatically refresh before queries or fall back from Git.
|
|
134
140
|
- Set `sourceFilter.nameRegex` for source-only cross-search/cohesion filtering; combine it with `path` and a filter type (`all`, `changed-since`, or `uncommitted`). `changed-since` also accepts `uncommitted: true`.
|
|
135
141
|
- Top-level analysis `nameRegex` filters both sources and candidates. Query `SimilaritySearchOptions.nameRegex` filters result names before limiting.
|
|
136
142
|
- Call `await index.useDescriptions()`, then `await index.searchDescription({ query: "maintain the repository index" })`. Select a model via `descriptionProvider: new OpenAIDescriptionProvider({ model: "gpt-5.6-sol" })` in index options. Custom description providers implement both stateless file/callable methods and may add `startFile()` for contextual sessions.
|
|
137
143
|
- Inspect failures through `index.indexErrors()` or exported `readIndexErrors(indexPath)` without a provider. Records use `IndexingError`.
|
|
138
|
-
-
|
|
144
|
+
- `analyzeCohesion` remains a programmatic report API. The CLI exposes physical-distance review through `cross-search --cohesion` instead of a separate command.
|
|
139
145
|
|
|
140
146
|
## Build and development
|
|
141
147
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@ninjaxtools/slopdex",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.14.0",
|
|
4
4
|
"description": "Tree-sitter callable-level semantic indexing, search, and duplicate discovery for Python, JavaScript, TypeScript, Rust, Go, Java, and C",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"repository": {
|