@ninjaxtools/slopdex 0.13.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -11,11 +11,11 @@ For installation, command examples, configuration, and result interpretation, se
11
11
  | `src/parser/` | Language dispatch, native Tree-sitter extraction, callable identity, and recoverable diagnostics. |
12
12
  | `src/source-policy.ts`, `src/gitignore.ts` | Supported paths, built-in/config exclusions, nested ignore rules. |
13
13
  | `src/git/repository.ts` | Git commits, trees, blobs, diffs, ancestry, and working-tree changes. |
14
- | `src/embeddings/`, `src/descriptions/` | Provider requests and provider profiles. |
14
+ | `src/embeddings/`, `src/descriptions/`, `src/rerankers/` | Provider requests and provider profiles. |
15
15
  | `src/storage/database.ts` | SQLite schema, durable artifact caches, transactions, metadata, and vector queries. |
16
16
  | `src/search/` | Analysis scoring selection and cross-index neighbor discovery. |
17
17
  | `src/analysis/cohesion.ts` | Physical distance, gap scores, aggregate affinity, and groups. |
18
- | `src/format.ts` | Human-readable search, cluster, and cohesion output. |
18
+ | `src/format.ts` | Human-readable search and cluster output, plus library cohesion-report formatting. |
19
19
  | `src/index.ts`, `src/types.ts` | Public exports and data contracts. |
20
20
 
21
21
  ## Callable extraction
@@ -52,12 +52,13 @@ Storage uses Node's `node:sqlite` and `sqlite-vec`. Writable connections enable
52
52
  - `files`: paths, content hashes, blob IDs, source mode, previous paths, language, and size.
53
53
  - `functions`: identity, names, signatures, locations, source, provenance, and embedding/description references.
54
54
  - `embeddings`: code, description, and query vectors keyed by embedding profile, operation, and exact input.
55
+ - `function_vectors`: synchronized `vec0` storage for filtered code-vector nearest-neighbor queries.
55
56
  - `description_cache`: generated description text keyed independently by description profile and complete source context.
56
57
  - `parse_cache`: successful Tree-sitter extraction results keyed by parser strategy, path, and file-content hash.
57
58
  - `callable_provenance`: first-seen committed source identity.
58
59
  - `indexing_errors`: diagnostics associated with files.
59
60
 
60
- Current schema version is `7`. Schema 6 is migrated in place by adding file-description state; earlier schemas require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. Forced rebuilds clear logical index state while retaining content-addressed caches when possible; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
61
+ Current schema version is `8`. Schemas 6 and 7 are migrated in place by adding file-description state when necessary and building the nearest-neighbor vector table; earlier schemas require `--force-reindex`. Metadata validation rejects incompatible roots, embedding profiles, and unsupported schemas. Forced rebuilds clear logical index state while retaining content-addressed caches when possible; incompatible older databases are recreated. Enabled OpenAI description settings are preserved for the same repository where possible. Git divergence reconciliation is a separate operation controlled by `--rebuild-on-divergence`.
61
62
 
62
63
  ### Diagnostics
63
64
 
@@ -80,7 +81,11 @@ Description inputs include contextual and profile information so file-context, p
80
81
 
81
82
  ## Similarity and analysis
82
83
 
83
- `search` embeds a query and, when descriptions are complete, scores code, callable-description, and containing-file-description vectors. `search-description` scores callable and file descriptions. Analysis uses code-only cosine similarity unless all indexed callables and files have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
84
+ `search` embeds a query and, when descriptions are complete, scores code, callable-description, and containing-file-description vectors. `search-description` scores callable and file descriptions. Filter-compatible code-only searches up to sqlite-vec's 8,192-dimension limit use the synchronized `vec0` nearest-neighbor table; larger custom profiles plus fused, regex, upper-bound, and multi-path searches retain the exact scalar scoring path. Analysis uses code-only cosine similarity unless all indexed callables and files have enabled description embeddings. Cross-index analysis requires completeness on both sides. Description-generator models may differ across indexes even though embedding profiles must match.
85
+
86
+ An optional `Reranker` performs a second-stage pass for the two natural-language query methods. After applying name and similarity filters, Cohere and Jina retrieve five times the requested result limit. `OpenAILLMReranker` retrieves its configured candidate count (10 by default), or the result limit when larger. Every reranker receives the query plus candidate path, code metadata/source, and purpose description when available, then returns the requested number in relevance order. Results preserve the embedding/fused `similarity` and add `rerankScore`.
87
+
88
+ Cohere defaults to `rerank-v4.0-pro` with `COHERE_API_KEY`; Jina defaults to `jina-reranker-v3.5` with `JINA_API_KEY`. The OpenAI LLM path defaults to `gpt-5.6-luna`, sends a strict JSON schema through the Responses API with high reasoning, no reasoning summary, and `store: false`, and validates result cardinality, indexes, uniqueness, and 0-1 scores. Its prompt preserves descriptions before source and caps source-bearing candidate text at 12,000 tokens each and 80,000 tokens in aggregate. Reranking does not participate in index metadata or cross-search because it creates no persisted artifacts and cross-search is callable-to-callable analysis rather than natural-language retrieval.
84
89
 
85
90
  When descriptions are complete:
86
91
 
@@ -90,7 +95,7 @@ similarity = (codeSimilarity + descriptionSimilarity + fileDescriptionSimilarity
90
95
 
91
96
  All component scores and the average are computed in one SQLite query. Name, line-count, and path exclusions plus similarity bounds apply before ranking/limiting. Range upper bounds are exclusive. Combined JSON includes component scores; analysis metadata records mode, weights, and description profiles. Cohesion JSON uses schema version 3 for the three-component scoring contract.
92
97
 
93
- Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
98
+ Cross-search selects sources, queries neighbors per source, and deduplicates unordered same-index pairs unless symmetric results are requested. Self-matches are excluded in same-index queries. Same-file exclusion uses canonical roots and file identity to handle aliases. With the `cohesion` option, each selected match receives its physical path distance and matches are re-ranked by descending distance, then similarity. Cluster formatting builds connected components from emitted matches and sorts by member count, then name; transitive connectivity does not imply all-to-all similarity.
94
99
 
95
100
  ### Cohesion metrics
96
101
 
@@ -109,11 +114,12 @@ Affinity ratios and mean distance are weighted by `semanticWeight`. Cohesion met
109
114
  ## Library API
110
115
 
111
116
  ```ts
112
- import { OpenAIEmbeddingProvider, openCodeIndex } from "@ninjaxtools/slopdex";
117
+ import { OpenAIEmbeddingProvider, OpenAILLMReranker, openCodeIndex } from "@ninjaxtools/slopdex";
113
118
 
114
119
  const index = openCodeIndex({
115
120
  rootDir: "/path/to/repository",
116
121
  provider: new OpenAIEmbeddingProvider(),
122
+ reranker: new OpenAILLMReranker({ candidateCount: 10 }),
117
123
  });
118
124
 
119
125
  try {
@@ -128,14 +134,14 @@ try {
128
134
  }
129
135
  ```
130
136
 
131
- Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, `JinaEmbeddingProvider`, `OpenAIDescriptionProvider`, error types, and the contracts in `src/types.ts`. Standalone update/search helpers wrap the corresponding index methods.
137
+ Exports include `CodeIndex`, `crossSearch`, `analyzeCohesion`, `cohesionLocation`, embedding/description providers, `CohereReranker`, `JinaReranker`, `OpenAILLMReranker`, error types, and the contracts in `src/types.ts`. Standalone update/search helpers wrap the corresponding index methods.
132
138
 
133
139
  - Use `updateFromWorkingTree()` when Git is unavailable. Unlike the CLI, the library does not automatically refresh before queries or fall back from Git.
134
140
  - Set `sourceFilter.nameRegex` for source-only cross-search/cohesion filtering; combine it with `path` and a filter type (`all`, `changed-since`, or `uncommitted`). `changed-since` also accepts `uncommitted: true`.
135
141
  - Top-level analysis `nameRegex` filters both sources and candidates. Query `SimilaritySearchOptions.nameRegex` filters result names before limiting.
136
142
  - Call `await index.useDescriptions()`, then `await index.searchDescription({ query: "maintain the repository index" })`. Select a model via `descriptionProvider: new OpenAIDescriptionProvider({ model: "gpt-5.6-sol" })` in index options. Custom description providers implement both stateless file/callable methods and may add `startFile()` for contextual sessions.
137
143
  - Inspect failures through `index.indexErrors()` or exported `readIndexErrors(indexPath)` without a provider. Records use `IndexingError`.
138
- - Public cohesion reports retain indexed-function data; the CLI presents compact function references and includes source bodies only with `--include-source`.
144
+ - `analyzeCohesion` remains a programmatic report API. The CLI exposes physical-distance review through `cross-search --cohesion` instead of a separate command.
139
145
 
140
146
  ## Build and development
141
147
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ninjaxtools/slopdex",
3
- "version": "0.13.0",
3
+ "version": "0.15.0",
4
4
  "description": "Tree-sitter callable-level semantic indexing, search, and duplicate discovery for Python, JavaScript, TypeScript, Rust, Go, Java, and C",
5
5
  "license": "MIT",
6
6
  "repository": {