@titan-design/code-graph 0.1.1 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +350 -15
- package/dist/change-coupling-CyqHgRsm.d.ts +73 -0
- package/dist/chunk-PFI5XMUG.js +332 -0
- package/dist/chunk-PFI5XMUG.js.map +1 -0
- package/dist/history/index.d.ts +56 -0
- package/dist/history/index.js +29 -0
- package/dist/history/index.js.map +1 -0
- package/dist/index.d.ts +941 -38
- package/dist/index.js +2982 -231
- package/dist/index.js.map +1 -1
- package/package.json +9 -4
package/README.md
CHANGED
|
@@ -5,8 +5,8 @@ symbols, external packages, and the import / re-export / reference edges between
|
|
|
5
5
|
index-time source metrics, in one SQLite file built from `@titan-design/store-sqlite` kit
|
|
6
6
|
tables and refreshed incrementally.
|
|
7
7
|
|
|
8
|
-
Tier 2 of the titan-platform DAG. Depends on `store-sqlite`, `
|
|
9
|
-
|
|
8
|
+
Tier 2 of the titan-platform DAG. Depends on `store-sqlite`, `code-parser` (tree-sitter WASM
|
|
9
|
+
parsing and the file filter), `ts-morph`, `web-tree-sitter` for node types, and `embed` plus `retrieval` for similar-symbol search. Extracted from codewatch's `@codewatch/graph` (TP-9), split along the seam the
|
|
10
10
|
audit identified: that package did the job of both a store and a code graph.
|
|
11
11
|
|
|
12
12
|
```ts
|
|
@@ -23,7 +23,8 @@ listEdges(store, snapshotId); // imports / re-exports, references on request
|
|
|
23
23
|
|
|
24
24
|
## What was extracted, and what was not
|
|
25
25
|
|
|
26
|
-
In: the parser (tree-sitter WASM for TypeScript, TSX and Python
|
|
26
|
+
In: the parser (tree-sitter WASM for TypeScript, TSX and Python, since moved to
|
|
27
|
+
`@titan-design/code-parser` and still re-exported here), the walk, the ts-morph
|
|
27
28
|
extractor and its symbol layer, role classification, generated-file detection, id aliasing
|
|
28
29
|
across git renames, the three-tier incremental reuse, and the metrics computed at index time
|
|
29
30
|
(degree, utilization, loc, cyclomatic, cognitive, nesting, class count, lcom4, per-symbol
|
|
@@ -31,16 +32,22 @@ complexity). `lcom.ts` came along despite being an analysis: `source-metrics.ts`
|
|
|
31
32
|
directly and lcom4 is a pure function of a file's bytes, so it belongs with the metrics that
|
|
32
33
|
carry forward under reuse.
|
|
33
34
|
|
|
35
|
+
Ported later (TP-123, TP-124), strictly as codewatch had them: the rules engine
|
|
36
|
+
(`check*.ts`, now `src/check/`) and the snapshot diff plus check diff (`diff.ts`,
|
|
37
|
+
`check-diff.ts`, now `src/diff/`). See [Checks and diffs](#checks-and-diffs).
|
|
38
|
+
Ported with TP-127 and TP-128: dead code, growth risk, PageRank, relevance, and symbol
|
|
39
|
+
coupling, now `src/analysis/`. See [Graph analyses](#graph-analyses).
|
|
40
|
+
Ported with TP-133: the test linker and the Istanbul coverage overlay, also in `src/analysis/`,
|
|
41
|
+
and test-coverage ownership. See [Test linking and coverage](#test-linking-and-coverage).
|
|
42
|
+
|
|
43
|
+
Ported with TP-250: package partition quality (`src/analysis/partition-quality.ts`) and
|
|
44
|
+
snapshot pruning (`src/prune.ts`). See [Partition quality](#partition-quality) and
|
|
45
|
+
[Pruning snapshots](#pruning-snapshots).
|
|
46
|
+
|
|
34
47
|
Deferred, all of it still in codewatch, all of it a follow-up on this package rather than a
|
|
35
48
|
change to it:
|
|
36
49
|
|
|
37
|
-
-
|
|
38
|
-
- Git-history mining: churn, ownership, change coupling, symbol coupling, test coverage
|
|
39
|
-
linking. `buildIndexerMetrics` used to fold these into the same pass; here it computes only
|
|
40
|
-
what a file's own bytes and the assembled graph determine.
|
|
41
|
-
- Graph analyses over a finished snapshot: communities, pagerank, partition quality,
|
|
42
|
-
relevance, conventions, coverage overlay, dead code, growth risk, patterns, prune,
|
|
43
|
-
test-linker, diff, reuse-delta reporting, embeddings.
|
|
50
|
+
- Graph analyses over a finished snapshot: communities, conventions, reuse-delta reporting.
|
|
44
51
|
|
|
45
52
|
Python support is new here rather than ported. codewatch walked TypeScript only; the parser
|
|
46
53
|
already had the grammar. The Python extractor is deliberately narrower than the ts-morph one:
|
|
@@ -49,16 +56,24 @@ same tree-sitter declaration walk that feeds complexity.
|
|
|
49
56
|
|
|
50
57
|
## The id scheme
|
|
51
58
|
|
|
52
|
-
|
|
53
|
-
codewatch's CLI
|
|
59
|
+
File, module and external ids are preserved exactly from codewatch, because this repo's own
|
|
60
|
+
`dag:check` consumes them through codewatch's CLI. Symbol ids diverge from codewatch since
|
|
61
|
+
index version 0.14.0:
|
|
54
62
|
|
|
55
63
|
- A **file** id is its path relative to the git toplevel, in posix form:
|
|
56
64
|
`packages/registry/src/index.ts`. Ids root at the git toplevel even when you walk a
|
|
57
65
|
subtree, so importers across subtrees share one id space.
|
|
58
66
|
- A **module** id is the file id minus its extension: `packages/registry/src/index`. Its
|
|
59
67
|
parent is the directory above it.
|
|
60
|
-
- A **symbol** id hangs under its declaring file as `<fileId>#<
|
|
61
|
-
|
|
68
|
+
- A **symbol** id hangs under its declaring file as `<fileId>#<qualifiedName>`, split on the
|
|
69
|
+
first `#`. A top-level declaration's qualified name is its own name
|
|
70
|
+
(`src/a.ts#createThing`). A member or nested declaration is prefixed by its enclosing named
|
|
71
|
+
scopes, joined with `.`: `src/a.ts#Job.run`, `src/a.ts#outer.helper`, and
|
|
72
|
+
`src/a.ts#handlers.onClick` for a method of `const handlers = {…}`. Anonymous scopes, such
|
|
73
|
+
as a callback argument or an unbound class expression, add no segment, so ids do not depend
|
|
74
|
+
on declaration order. One scope binds a name once: a getter/setter pair, a Python property's
|
|
75
|
+
accessors, and overloads each share one node. Index versions before 0.14.0 keyed members by
|
|
76
|
+
bare name, so same-named methods in one file collapsed into one node (TP-182).
|
|
62
77
|
- An **external** id is `npm:<package>` (scope-aware) or the `node:` builtin verbatim.
|
|
63
78
|
|
|
64
79
|
`NodeKind`, `EdgeKind`, and the `role` vocabulary are unchanged. So is the property the DAG
|
|
@@ -67,11 +82,54 @@ package's source file, not to an `npm:` external, by remapping the `dist/*.d.ts`
|
|
|
67
82
|
ts-morph resolves back onto `src/`. That remap needs the target package built, which is why
|
|
68
83
|
`pnpm build` precedes both `pnpm test` and `dag:check`.
|
|
69
84
|
|
|
85
|
+
## Identity across renames
|
|
86
|
+
|
|
87
|
+
Added in index version 0.15.0 (TP-187). Git rename detection writes `id_alias` rows mapping a
|
|
88
|
+
file's and its module's old id to the new one. A file rename also writes an alias for every
|
|
89
|
+
symbol present on both sides, so `a.ts#Job.run` follows `a.ts` to `b.ts#Job.run`. A symbol
|
|
90
|
+
renamed inside a file is not followed.
|
|
91
|
+
|
|
92
|
+
Each snapshot records its alias base in `attrs.aliasBase`: the snapshot whose commit its
|
|
93
|
+
aliases were computed against. A new index of a ref diffs against that ref's newest committed
|
|
94
|
+
snapshot, falling back to the newest committed snapshot of any ref. The bases form a tree, so
|
|
95
|
+
ids can be carried between any two snapshots that share an ancestor, forward or backward.
|
|
96
|
+
|
|
97
|
+
```ts
|
|
98
|
+
import { aliasChain, priorSnapshotForRef, resolveAlias } from "@titan-design/code-graph";
|
|
99
|
+
|
|
100
|
+
resolveAlias(store, "src/job.ts", 4, { fromSnapshotId: 1 });
|
|
101
|
+
// { id: 'core/worker.ts', reason: 'move',
|
|
102
|
+
// hops: [ src/job.ts -> src/task.ts, src/task.ts -> src/worker.ts, src/worker.ts -> core/worker.ts ] }
|
|
103
|
+
resolveAlias(store, "src/job.ts#Job.run", 4, { fromSnapshotId: 1 }).id; // 'core/worker.ts#Job.run'
|
|
104
|
+
aliasChain(store, 1, 4).resolve; // one resolver for many ids
|
|
105
|
+
priorSnapshotForRef(store, "main", { before: 7, repoRoot: "." }); // the snapshot main denotes
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
Without `fromSnapshotId`, `resolveAlias` walks from the root of the target's lineage, so an id
|
|
109
|
+
from any ancestor resolves. Each snapshot's aliases are applied once, in order: they come from
|
|
110
|
+
one git diff, so `a.ts` to `b.ts` plus `b.ts` to `a.ts` in one snapshot is a swap, not a
|
|
111
|
+
cycle. Walks stop after 10,000 snapshots. Two snapshots with no common base fall back to the
|
|
112
|
+
to-snapshot's own aliases, which is what `diffSnapshots` did before 0.15.0.
|
|
113
|
+
|
|
114
|
+
`priorSnapshotForRef` prefers the snapshot of the commit git resolves the ref to, and
|
|
115
|
+
otherwise the newest snapshot labelled with that ref. A snapshot written before 0.15.0 records
|
|
116
|
+
no base; its base is inferred as the newest earlier snapshot with a commit, which is the one
|
|
117
|
+
the 0.14.0 indexer diffed against. Nothing is written on read, so a 0.14.0 store opened
|
|
118
|
+
read-only diffs and checks across renames too.
|
|
119
|
+
|
|
70
120
|
## The three reuse tiers
|
|
71
121
|
|
|
72
122
|
Every run writes a fingerprint per file: a content hash and a comment/whitespace-insensitive
|
|
73
123
|
hash of its parse structure. The next run diffs against the most recent snapshot carrying the
|
|
74
|
-
same `INDEX_VERSION` and sorts each file into one tier.
|
|
124
|
+
same `INDEX_VERSION` and sorts each file into one tier. `INDEX_VERSION` is bumped whenever a
|
|
125
|
+
metric can change for the same bytes, not only when the node or edge shape changes: 0.12.0
|
|
126
|
+
marks `.tsx` files moving to the tsx grammar, which changed their complexity metrics and
|
|
127
|
+
symbol spans; 0.13.0 marks the dead-code and growth-risk metrics joining the carry-forward set;
|
|
128
|
+
0.14.0 marks symbol ids qualified by their enclosing scopes (TP-182); 0.15.0 marks symbol
|
|
129
|
+
aliases on a file rename and the recorded alias base (TP-187). A snapshot from an older version
|
|
130
|
+
is never reused, so the first run after an upgrade is a full index. The first 0.14.0 run after
|
|
131
|
+
an older snapshot writes `requalify` id aliases from each bare-name id to its qualified
|
|
132
|
+
successor, only where exactly one declaration in the file carries that name.
|
|
75
133
|
|
|
76
134
|
| Tier | Trigger | Work skipped |
|
|
77
135
|
|---|---|---|
|
|
@@ -88,6 +146,10 @@ always recomputed over the whole assembled graph, so a heavily-reused run and a
|
|
|
88
146
|
|
|
89
147
|
`openCodeGraph` opens the database with the kit's pragmas and runs three migrations.
|
|
90
148
|
|
|
149
|
+
It refuses a database stamped past `SCHEMA_VERSION` with `SchemaTooNewError`: a code graph
|
|
150
|
+
database belongs to one build, so a higher version means a newer build already moved the
|
|
151
|
+
schema and this one would query columns that are gone.
|
|
152
|
+
|
|
91
153
|
Kit tables: `snapshot` (the snapshot registry), `node` (`entity_snap`, keyed
|
|
92
154
|
`(snapshot_id, id)`), `blob_cache` (content-addressed, for the embedding and summary caches
|
|
93
155
|
an analysis layer will want). Migration 2 adds four columns the kit shapes do not carry but
|
|
@@ -103,3 +165,276 @@ The symbol layer is hidden by default. `listNodes` drops `symbol` nodes and `lis
|
|
|
103
165
|
`references` edges unless you ask for them, so a caller reasoning about module structure sees
|
|
104
166
|
the graph it expects and does not have one import of thirty names read as thirty
|
|
105
167
|
dependencies.
|
|
168
|
+
|
|
169
|
+
## Targeted reads and the metric catalogue
|
|
170
|
+
|
|
171
|
+
Added for the read API (TP-183). Three store reads answer one node or one metric without
|
|
172
|
+
loading a snapshot, and a catalogue describes every metric name the package writes.
|
|
173
|
+
|
|
174
|
+
```ts
|
|
175
|
+
import { aggregateMetrics, describeMetric, listEdgesTouching, listMetricsForNode } from "@titan-design/code-graph";
|
|
176
|
+
|
|
177
|
+
listMetricsForNode(store, snapshotId, "packages/registry/src/invoke.ts"); // every metric on one node
|
|
178
|
+
listEdgesTouching(store, snapshotId, "packages/registry/src/invoke.ts"); // edges in and out; references on request
|
|
179
|
+
aggregateMetrics(store, snapshotId, { name: "loc" }); // [{ name: "loc", nodeKind: "file", count, sum, min, max }]
|
|
180
|
+
describeMetric("churn_90d"); // { unit: "lines", rollup: "sum", direction: "neutral", absent: "zero", window: "90d", … }
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
Three more store reads answer a report's questions without loading a snapshot:
|
|
184
|
+
`listMetricNames(snapshotId)` for the distinct names stored, `topByMetric({ snapshotId,
|
|
185
|
+
metric, limit, kind })` for the highest-valued nodes joined to their kind and role, and
|
|
186
|
+
`replaceMetricsByName(snapshotId, name, metrics)`, which swaps one metric's whole row set in
|
|
187
|
+
a single transaction so a re-ingested overlay such as coverage never accumulates stale rows.
|
|
188
|
+
|
|
189
|
+
`computeDeepAst({ filePath, absPath, symbolName })` reads structure too heavy to persist
|
|
190
|
+
(params, return type, class members) from the working tree on demand, and returns null when
|
|
191
|
+
the file is unreadable.
|
|
192
|
+
|
|
193
|
+
- **Reads.** `listMetricsForNode` searches the metric primary key, `listEdgesTouching` the edge
|
|
194
|
+
primary key plus `idx_edge_dst`, and `aggregateMetrics` `idx_metric_name`. None scans a
|
|
195
|
+
table, and no index or migration was added, so any existing store serves them. A self-loop
|
|
196
|
+
comes back once. `aggregateMetrics` groups by metric name and node kind, because
|
|
197
|
+
`utilization` and the degree metrics sit on several kinds; `count` skips null values.
|
|
198
|
+
- **Catalogue.** `METRIC_CATALOGUE` holds one `MetricDescriptor` per name: `unit`,
|
|
199
|
+
`appliesTo` (node kinds), `rollup` (`sum`, `max`, `mean`, or `none`), `direction`
|
|
200
|
+
(`higher-worse`, `lower-worse`, or `neutral`), `absent`, `source`, and `description`.
|
|
201
|
+
Windowed names are templates such as `churn_{w}`; `describeMetric` resolves a stored name
|
|
202
|
+
such as `churn_90d` or `test_bus_factor_lifetime` to a concrete descriptor, and returns
|
|
203
|
+
null for a name nothing writes.
|
|
204
|
+
- **`rollup: "none"`** means no rollup reproduces the group's own value. Summing fan-in counts
|
|
205
|
+
a directory's internal edges, and summing per-file commit counts counts a shared commit
|
|
206
|
+
more than once. A reader must not synthesize those.
|
|
207
|
+
- **`absent`** says what a missing row means. `zero`: the writer is sparse, so count the node
|
|
208
|
+
as zero (dead code, growth risk, churn, `linked_test_count`). `exclude`: the metric does not
|
|
209
|
+
apply or was not measured (a max over no functions, `coverage_pct` before ingest), so leave
|
|
210
|
+
the node out of means, percentiles, and ranks.
|
|
211
|
+
- **Completeness is tested.** `catalogue-completeness.test.ts` indexes a fixture repo with
|
|
212
|
+
history and a coverage overlay, and fails, naming the metric, when a stored name has no
|
|
213
|
+
descriptor, when a descriptor's unit or node kinds disagree with the rows, or when a
|
|
214
|
+
descriptor matches nothing.
|
|
215
|
+
- **Browser-safe.** The catalogue imports only `types.ts`. The `code-graph-catalogue-pure`
|
|
216
|
+
and `code-graph-catalogue-no-node` rules in `.codewatch/check.json` hold it there, so a
|
|
217
|
+
later browser subpath can re-export it. It ships from the root export today, which does
|
|
218
|
+
pull in Node.
|
|
219
|
+
|
|
220
|
+
## Checks and diffs
|
|
221
|
+
|
|
222
|
+
The rules engine turns a snapshot into pass/fail against a `check.json`. Six rule types:
|
|
223
|
+
`metric-max`, `metric-min`, `metric-product-max`, `forbid-import`, `layered-deps`, and
|
|
224
|
+
`no-internal-only-barrels`. Severity defaults to `error`; only new errors fail a check.
|
|
225
|
+
|
|
226
|
+
```ts
|
|
227
|
+
import { checkSnapshot, loadCheckRules, openCodeGraph } from "@titan-design/code-graph";
|
|
228
|
+
|
|
229
|
+
const store = openCodeGraph(".codewatch/graph.db");
|
|
230
|
+
const rules = await loadCheckRules(".codewatch/check.json", { onWarn: console.warn });
|
|
231
|
+
const { result } = checkSnapshot(store, { snapshot: "head", baseline: "main", rules });
|
|
232
|
+
result.passed; // false only when a violation is an error and absent from the baseline
|
|
233
|
+
```
|
|
234
|
+
|
|
235
|
+
`snapshot` and `baseline` take a numeric snapshot id or a ref name; a ref resolves to its
|
|
236
|
+
newest snapshot. `runChecks(store, { snapshotId, rules, baselineSnapshotId })` is the same
|
|
237
|
+
engine on ids, and `validateRules(json)` validates an already-parsed rules object.
|
|
238
|
+
|
|
239
|
+
The baseline is a ratchet. A violation whose key (rule id, node id, and destination id for
|
|
240
|
+
edge rules) also fires on the baseline snapshot is marked `isCarryover` and counts as
|
|
241
|
+
carryover, so existing debt does not block a change but new debt does. The baseline's node ids
|
|
242
|
+
are first carried through the alias chain into the checked snapshot (`rebasedViolationKey`),
|
|
243
|
+
so a moved file's violations carry over instead of reading as one resolved plus one new.
|
|
244
|
+
Unmoved ids key exactly as `violationKey` always did. Deprecated metric and
|
|
245
|
+
role spellings in a rules file (`lines`, `tests`) heal to their canonical names with a
|
|
246
|
+
warning through `onWarn` instead of failing validation.
|
|
247
|
+
|
|
248
|
+
`diffSnapshots(store, { fromSnapshotId, toSnapshotId })` reports added, removed and renamed
|
|
249
|
+
nodes, added and removed edges, and metric deltas on nodes present in both. Ids follow the
|
|
250
|
+
alias chain between the two snapshots, across every rename in between, so a move reads as a
|
|
251
|
+
rename rather than a delete plus an add, and its edges do not churn. Metric and edge-kind spellings are
|
|
252
|
+
canonicalised before comparing.
|
|
253
|
+
|
|
254
|
+
`diffCheckResults(store, { fromSnapshotId, toSnapshotId, rules })` runs the rules on both
|
|
255
|
+
snapshots and buckets each violation as new, resolved, or unchanged; unchanged metric
|
|
256
|
+
violations are further split into worsened and improved by value. From-side ids follow the
|
|
257
|
+
alias chain, as in the ratchet.
|
|
258
|
+
|
|
259
|
+
`scripts/dag-check-self.mjs` in the repo root runs this repo's DAG check on this engine
|
|
260
|
+
instead of codewatch's CLI.
|
|
261
|
+
|
|
262
|
+
The glob matching the rules use is exported too, because a CLI filters its own `--include`
|
|
263
|
+
and `--exclude` flags with the same semantics: `patternToRegex` (a `*` glob when the pattern
|
|
264
|
+
holds one, a case-sensitive substring otherwise), `compilePatterns`, and `matchesAny`.
|
|
265
|
+
|
|
266
|
+
## Similar symbols
|
|
267
|
+
|
|
268
|
+
Ported from codewatch's `embeddings.ts` (TP-129). It answers "does something like this
|
|
269
|
+
already exist?" before you write it. The embedder is injected; this package never builds
|
|
270
|
+
one and never talks to a network on its own.
|
|
271
|
+
|
|
272
|
+
```ts
|
|
273
|
+
import { OllamaEmbedder } from "@titan-design/embed";
|
|
274
|
+
import { findSimilarCapability, tryEmbedSnapshot } from "@titan-design/code-graph";
|
|
275
|
+
|
|
276
|
+
const embedder = new OllamaEmbedder();
|
|
277
|
+
const attempt = await tryEmbedSnapshot(store, snapshotId, embedder);
|
|
278
|
+
// { ok: true, result: { symbols, embedded, withPurpose, newlyEmbedded, reused, model } }
|
|
279
|
+
// or { ok: false, model, error } when the backend is down; the snapshot is unaffected
|
|
280
|
+
const { candidates, coverage } = await findSimilarCapability(store, snapshotId, "parse a duration", embedder);
|
|
281
|
+
```
|
|
282
|
+
|
|
283
|
+
- **Corpus.** Exported symbols with a signature, minus test, fixture and generated files.
|
|
284
|
+
The embedded text is `signature -- purpose` (the docstring), never the body.
|
|
285
|
+
- **Storage.** Vectors go in `blob_cache` under namespace `code-graph/symbol-embedding`,
|
|
286
|
+
keyed by `embedder.model` and the SHA-256 of the text. They are not snapshot-scoped, so
|
|
287
|
+
re-embedding unchanged text costs zero embed calls. `embedSnapshot` throws on a backend
|
|
288
|
+
failure; `tryEmbedSnapshot` reports it.
|
|
289
|
+
- **Query.** `vectorRetriever` over a `BruteForceVectorIndex` from `retrieval`. The query is
|
|
290
|
+
embedded with role `query` and the symbol texts with role `document`; the embedder applies
|
|
291
|
+
the matching prefix (`search_query: ` / `search_document: ` for nomic models).
|
|
292
|
+
- **Results are candidates, not verdicts.** Each has a cosine score, and `coverage` says how
|
|
293
|
+
many symbols were searchable and how many carry purpose text. There is deliberately no
|
|
294
|
+
co-location filter.
|
|
295
|
+
- **Python symbols are not in the corpus yet.** The Python extractor records no signature or
|
|
296
|
+
docstring, so no Python symbol passes the corpus filter.
|
|
297
|
+
- **Prefixes are part of the cache key.** `embedder.model` includes a hash of the embedder's
|
|
298
|
+
prefix table, so two prefix configurations never share a vector. codewatch used no prefix;
|
|
299
|
+
`new OllamaEmbedder({ prefixes: { document: "", query: "" } })` reproduces its rankings.
|
|
300
|
+
|
|
301
|
+
## Graph analyses
|
|
302
|
+
|
|
303
|
+
Dead-code and growth-risk metrics are computed at index time, like the source metrics, and
|
|
304
|
+
carry forward for unchanged files. Both are sparse: a file gets a row only when a count is
|
|
305
|
+
above zero.
|
|
306
|
+
|
|
307
|
+
- Dead code, TypeScript only: `unreachable_statements` (after a `return`, `throw`, `break`
|
|
308
|
+
or `continue` in the same block), `unused_locals`, and `unused_params` (trailing run only).
|
|
309
|
+
- Growth risk, TypeScript and Python: `loop_depth` (at 2 or more), `recursive_functions`,
|
|
310
|
+
and `search_in_loop` (`.includes`, `.find` and similar inside a loop). These are smells,
|
|
311
|
+
not complexity bounds. Recursion and search match TypeScript call nodes only, so Python
|
|
312
|
+
files get `loop_depth` alone, as in codewatch.
|
|
313
|
+
|
|
314
|
+
PageRank, relevance, and symbol coupling run at query time over one snapshot:
|
|
315
|
+
|
|
316
|
+
```ts
|
|
317
|
+
import {
|
|
318
|
+
snapshotPageRank,
|
|
319
|
+
snapshotRelevance,
|
|
320
|
+
snapshotSymbolCoupling,
|
|
321
|
+
} from "@titan-design/code-graph";
|
|
322
|
+
|
|
323
|
+
snapshotPageRank(store, snapshotId); // global centrality over the file-level graph
|
|
324
|
+
snapshotPageRank(store, snapshotId, { personalization: new Map([[fileId, 1]]) }); // seeded
|
|
325
|
+
snapshotRelevance(store, snapshotId, [fileId]); // seeded over symmetrized edges
|
|
326
|
+
snapshotSymbolCoupling(store, snapshotId); // symbol pairs co-imported by 2+ files
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
The pure `computePageRank`, `computeRelevance`, `computeSymbolConsumers`, and
|
|
330
|
+
`computeSymbolCoupling` take node and edge arrays instead of a store.
|
|
331
|
+
|
|
332
|
+
### Partition quality
|
|
333
|
+
|
|
334
|
+
`computePartitionQuality` scores a package partition of the file graph: per-package cohesion,
|
|
335
|
+
Martin instability, an abstractness proxy (the share of `role: "types"` files), a layer label
|
|
336
|
+
(`top`, `middle`, `foundation`), pair coupling by intensity (`edges / files(from)`), and a
|
|
337
|
+
Newman-Girvan modularity Q over the whole partition.
|
|
338
|
+
|
|
339
|
+
```ts
|
|
340
|
+
import { computePartitionQuality, invertBuckets } from "@titan-design/code-graph";
|
|
341
|
+
|
|
342
|
+
const result = computePartitionQuality({ packages, fileByPackage, nodes, edges });
|
|
343
|
+
result.modularityQ; // 0.77 on titan-platform's packages/*
|
|
344
|
+
invertBuckets(fileByPackage); // file id -> package id, skipping the "" unassigned bucket
|
|
345
|
+
```
|
|
346
|
+
|
|
347
|
+
`resolveBarrels: true` rewrites an edge landing on a `role: "barrel"` file to the files it
|
|
348
|
+
re-exports, transitively. It over-attributes: one import of one name through a barrel becomes
|
|
349
|
+
one edge per re-export target, which is why it is off by default.
|
|
350
|
+
|
|
351
|
+
### Pruning snapshots
|
|
352
|
+
|
|
353
|
+
`planPrune` keeps the most recent `keep` snapshots (default 10) plus every snapshot whose ref
|
|
354
|
+
is in `keepRefs`; `runPrune` deletes the rest and reports row counts before and after.
|
|
355
|
+
|
|
356
|
+
```ts
|
|
357
|
+
import { planPrune, runPrune } from "@titan-design/code-graph";
|
|
358
|
+
|
|
359
|
+
planPrune(store, { keep: 10, keepRefs: ["main"] }); // { keep, remove }
|
|
360
|
+
runPrune(store, { keep: 10, vacuum: true }); // { plan, rowsBefore, rowsAfter, vacuumed }
|
|
361
|
+
```
|
|
362
|
+
|
|
363
|
+
The domain tables declare no foreign key, so `CodeGraphStore.deleteSnapshots` clears each of
|
|
364
|
+
`SNAPSHOT_SCOPED_TABLES` itself rather than relying on a cascade. `blob_cache` is
|
|
365
|
+
content-addressed and not snapshot-scoped, so a prune never drops a cached embedding.
|
|
366
|
+
|
|
367
|
+
## Git history
|
|
368
|
+
|
|
369
|
+
Ported in TP-126, strictly as codewatch had it: churn over rolling windows, first-seen dates,
|
|
370
|
+
ownership and bus factor, and change coupling. The engine lives in `src/history/` and is
|
|
371
|
+
published on its own subpath. Its API is path-based: repo-relative posix paths in, plain
|
|
372
|
+
records out, no node ids or snapshots.
|
|
373
|
+
|
|
374
|
+
```ts
|
|
375
|
+
import { computeChangeCoupling, couplingFor, loadChurnEntries } from "@titan-design/code-graph/history";
|
|
376
|
+
|
|
377
|
+
const entries = loadChurnEntries({ repoRoot: ".", windowDays: 90 }) ?? []; // null outside git
|
|
378
|
+
const { pairs, skippedLargeCommits } = computeChangeCoupling(entries);
|
|
379
|
+
couplingFor(pairs, "packages/code-graph/src/indexer.ts"); // partners by co-edit count
|
|
380
|
+
```
|
|
381
|
+
|
|
382
|
+
`indexPaths` writes history metrics on file nodes by default, through the
|
|
383
|
+
`history-metrics.ts` adapter, with codewatch's names: `churn_{w}`, `churn_{w}_commits`,
|
|
384
|
+
`churn_{w}_authors` and `recency_{w}` for each window, `file_age_days`, and `bus_factor_{w}`
|
|
385
|
+
plus `top_author_share_{w}` for the primary window. Windows default to 30, 90 and 180 days
|
|
386
|
+
plus `churnWindowDays` (the primary, default 30); `churnWindows` replaces the defaults and
|
|
387
|
+
`lifetime: true` adds an all-history window with its own ownership. `computeChurn: false`
|
|
388
|
+
turns all of it off. Outside git, or without a git binary, the index simply has no history
|
|
389
|
+
metrics.
|
|
390
|
+
|
|
391
|
+
The adapter is exported from the package root, not from `./history`, because it speaks
|
|
392
|
+
`GraphMetric` and the seam below does not. A product that indexes on its own terms calls
|
|
393
|
+
`loadHistoryMetrics(nodes, idRoot, options)` for both the metric rows and the primary-window
|
|
394
|
+
churn entries they were built from (`LoadedHistory`), with `HistoryMetricsOptions`,
|
|
395
|
+
`DEFAULT_CHURN_WINDOWS` (`[30, 90, 180]`), `resolveChurnWindows`, `windowSuffix` (`30d`,
|
|
396
|
+
`lifetime`) and `computeRecencyWindows` alongside it. Node ids are the history engine's
|
|
397
|
+
repo-relative paths, so both must be rooted at the same `idRoot`.
|
|
398
|
+
|
|
399
|
+
Change coupling is not stored. It is computed on demand from `loadChurnEntries`, as
|
|
400
|
+
codewatch's `graph coupled` command did.
|
|
401
|
+
|
|
402
|
+
**The seam.** Nothing under `src/history/` imports the rest of code-graph; the rest may import
|
|
403
|
+
it. The `code-graph-history-seam` rule in `.codewatch/check.json` fails `pnpm dag:check` on
|
|
404
|
+
any such import. That keeps a later extraction into `@titan-design/git-history` a directory
|
|
405
|
+
move plus an import-path change.
|
|
406
|
+
|
|
407
|
+
**Behaviour kept from codewatch, gaps included.** A rename is followed only inside the commit
|
|
408
|
+
that made it, so a file's churn before the rename stays on its old path and is dropped.
|
|
409
|
+
First-seen dates come from a separate `--no-renames` pass, which makes a renamed file look
|
|
410
|
+
younger. `--since` resolves against git's clock while window slicing uses `nowEpoch`. History
|
|
411
|
+
metrics are recomputed on every index and never carried forward under reuse.
|
|
412
|
+
|
|
413
|
+
## Test linking and coverage
|
|
414
|
+
|
|
415
|
+
Ported in TP-133, strictly as codewatch had it. `linkTestsToSources` pairs each test file
|
|
416
|
+
with non-test files in two passes. Pass 1 uses path conventions: it strips a `.test` or `.spec`
|
|
417
|
+
infix and collapses a `__tests__/`, `test/` or `tests/` segment. Pass 2 gives a test that
|
|
418
|
+
pass 1 left unpaired its strongest co-edited non-test partner, with at least 2 shared commits.
|
|
419
|
+
|
|
420
|
+
`indexPaths` writes `linked_test_count` on each linked source, with or without git. With git
|
|
421
|
+
history on it also writes `test_bus_factor_{w}` and `test_top_author_share_{w}` for the
|
|
422
|
+
primary window. These summarize churn authorship across all tests linked to a source, so a
|
|
423
|
+
file can be well spread in production code and a single-author silo in its tests. All three
|
|
424
|
+
are recomputed on every index and never carried forward.
|
|
425
|
+
|
|
426
|
+
`computeTestCoverageOwnership` lives in the `history-metrics.ts` adapter, not in
|
|
427
|
+
`src/history/`. It needs test links, and the seam forbids history from importing them.
|
|
428
|
+
|
|
429
|
+
```ts
|
|
430
|
+
import { attributeCoverage } from "@titan-design/code-graph";
|
|
431
|
+
|
|
432
|
+
// fileIdOf maps an absolute path to a file id, or null to skip; spans come from symbol nodes' attrs.
|
|
433
|
+
const metrics = attributeCoverage(istanbulReport, fileIdOf, symbolSpansByFile);
|
|
434
|
+
```
|
|
435
|
+
|
|
436
|
+
`attributeCoverage` turns an Istanbul `coverage-final.json` into `coverage_pct` metrics: one
|
|
437
|
+
per file (covered functions over total functions) and one per symbol, matched by line-range
|
|
438
|
+
containment to the innermost symbol. Coverage depends on which tests ran, not on file bytes,
|
|
439
|
+
so the index never writes or carries it. The caller stores it on the snapshot it measured,
|
|
440
|
+
as codewatch's `graph coverage` command does.
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* A churn window is either a rolling N-day window or `"lifetime"` = all of git
|
|
3
|
+
* history (no `--since` bound). Lifetime mode lets an unfamiliar repo be audited
|
|
4
|
+
* cold, where any recent rolling slice is thin relative to the repo's whole
|
|
5
|
+
* life, so a windowed view reads an established repo as nearly churn-free.
|
|
6
|
+
*/
|
|
7
|
+
type ChurnWindow = number | "lifetime";
|
|
8
|
+
|
|
9
|
+
interface ChurnEntry {
|
|
10
|
+
commit: string;
|
|
11
|
+
/** Author identity: git author email (%ae), stable across name spelling drift. */
|
|
12
|
+
author: string;
|
|
13
|
+
/** Committer time (%ct), epoch seconds; lets one wide log be sliced per window. */
|
|
14
|
+
epoch: number;
|
|
15
|
+
/** Repo-relative posix path, rebased onto the caller's root. */
|
|
16
|
+
filePath: string;
|
|
17
|
+
added: number;
|
|
18
|
+
deleted: number;
|
|
19
|
+
}
|
|
20
|
+
interface LoadChurnOptions {
|
|
21
|
+
/** Directory paths are made relative to; entries outside it are dropped. */
|
|
22
|
+
repoRoot: string;
|
|
23
|
+
windowDays?: ChurnWindow;
|
|
24
|
+
}
|
|
25
|
+
/**
|
|
26
|
+
* Parse the last `windowDays` of git history into ChurnEntry[] rebased onto
|
|
27
|
+
* `repoRoot`. Returns null if git isn't available; [] if no commits matched.
|
|
28
|
+
* Used both for churn metrics and for change-coupling.
|
|
29
|
+
*/
|
|
30
|
+
declare function loadChurnEntries(options: LoadChurnOptions): ChurnEntry[] | null;
|
|
31
|
+
declare function parseChurnLog(text: string): ChurnEntry[];
|
|
32
|
+
declare function resolveRenamedPath(rawPath: string): string;
|
|
33
|
+
/** Entries whose commit landed within the last `windowDays` of `nowEpoch`. */
|
|
34
|
+
declare function entriesWithin(entries: readonly ChurnEntry[], windowDays: number, nowEpoch: number): ChurnEntry[];
|
|
35
|
+
|
|
36
|
+
interface CoEditPair {
|
|
37
|
+
fileA: string;
|
|
38
|
+
fileB: string;
|
|
39
|
+
/** Commits in which both files changed. */
|
|
40
|
+
count: number;
|
|
41
|
+
/** Most recent commits (up to maxCommitsPerPair). */
|
|
42
|
+
commits: string[];
|
|
43
|
+
}
|
|
44
|
+
interface ComputeChangeCouplingOptions {
|
|
45
|
+
/** Skip pairs that co-occur in fewer than this many commits. Default 2. */
|
|
46
|
+
minCount?: number;
|
|
47
|
+
/** Truncate the commits sample per pair. Default 10. */
|
|
48
|
+
maxCommitsPerPair?: number;
|
|
49
|
+
/**
|
|
50
|
+
* Skip commits touching more than this many files (sweeping refactors, mass
|
|
51
|
+
* renames). Default 50: beyond that the O(n²) pair explosion is signal-poor.
|
|
52
|
+
*/
|
|
53
|
+
largeCommitThreshold?: number;
|
|
54
|
+
/** Filter file paths to this set before pairing. */
|
|
55
|
+
knownPaths?: ReadonlySet<string>;
|
|
56
|
+
}
|
|
57
|
+
interface ChangeCouplingResult {
|
|
58
|
+
pairs: CoEditPair[];
|
|
59
|
+
/** Commits dropped by largeCommitThreshold, so callers can say when sweeps suppressed the signal. */
|
|
60
|
+
skippedLargeCommits: number;
|
|
61
|
+
/** Threshold that was applied (echoed for diagnostic messages). */
|
|
62
|
+
largeCommitThreshold: number;
|
|
63
|
+
}
|
|
64
|
+
interface CouplingPartner {
|
|
65
|
+
partner: string;
|
|
66
|
+
count: number;
|
|
67
|
+
commits: string[];
|
|
68
|
+
}
|
|
69
|
+
declare function computeChangeCoupling(entries: readonly ChurnEntry[], options?: ComputeChangeCouplingOptions): ChangeCouplingResult;
|
|
70
|
+
/** Files most coupled to `seed`, by co-edit count desc then path; the seed itself is excluded. */
|
|
71
|
+
declare function couplingFor(pairs: readonly CoEditPair[], seed: string): CouplingPartner[];
|
|
72
|
+
|
|
73
|
+
export { type ChurnEntry as C, type LoadChurnOptions as L, type ChurnWindow as a, type ChangeCouplingResult as b, type CoEditPair as c, type ComputeChangeCouplingOptions as d, type CouplingPartner as e, computeChangeCoupling as f, couplingFor as g, entriesWithin as h, loadChurnEntries as l, parseChurnLog as p, resolveRenamedPath as r };
|