@titan-design/code-graph 0.2.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +194 -13
- package/dist/index.d.ts +483 -5
- package/dist/index.js +1429 -49
- package/dist/index.js.map +1 -1
- package/package.json +3 -3
package/README.md
CHANGED
|
@@ -40,11 +40,17 @@ coupling, now `src/analysis/`. See [Graph analyses](#graph-analyses).
|
|
|
40
40
|
Ported with TP-133: the test linker and the Istanbul coverage overlay, also in `src/analysis/`,
|
|
41
41
|
and test-coverage ownership. See [Test linking and coverage](#test-linking-and-coverage).
|
|
42
42
|
|
|
43
|
-
|
|
44
|
-
|
|
43
|
+
Ported with TP-250: package partition quality (`src/analysis/partition-quality.ts`) and
|
|
44
|
+
snapshot pruning (`src/prune.ts`). See [Partition quality](#partition-quality) and
|
|
45
|
+
[Pruning snapshots](#pruning-snapshots).
|
|
45
46
|
|
|
46
|
-
|
|
47
|
-
|
|
47
|
+
Ported with TP-130: community detection and the convention layer (`src/conventions/`). See
|
|
48
|
+
[Conventions](#conventions). codewatch keeps the CLI command, the `claude -p` summarizer, and
|
|
49
|
+
the MCP and read-API wiring.
|
|
50
|
+
|
|
51
|
+
Deferred, still in codewatch, a follow-up on this package rather than a change to it:
|
|
52
|
+
|
|
53
|
+
- Reuse-delta reporting over a finished snapshot.
|
|
48
54
|
|
|
49
55
|
Python support is new here rather than ported. codewatch walked TypeScript only; the parser
|
|
50
56
|
already had the grammar. The Python extractor is deliberately narrower than the ts-morph one:
|
|
@@ -79,6 +85,41 @@ package's source file, not to an `npm:` external, by remapping the `dist/*.d.ts`
|
|
|
79
85
|
ts-morph resolves back onto `src/`. That remap needs the target package built, which is why
|
|
80
86
|
`pnpm build` precedes both `pnpm test` and `dag:check`.
|
|
81
87
|
|
|
88
|
+
## Identity across renames
|
|
89
|
+
|
|
90
|
+
Added in index version 0.15.0 (TP-187). Git rename detection writes `id_alias` rows mapping a
|
|
91
|
+
file's and its module's old id to the new one. A file rename also writes an alias for every
|
|
92
|
+
symbol present on both sides, so `a.ts#Job.run` follows `a.ts` to `b.ts#Job.run`. A symbol
|
|
93
|
+
renamed inside a file is not followed.
|
|
94
|
+
|
|
95
|
+
Each snapshot records its alias base in `attrs.aliasBase`: the snapshot whose commit its
|
|
96
|
+
aliases were computed against. A new index of a ref diffs against that ref's newest committed
|
|
97
|
+
snapshot, falling back to the newest committed snapshot of any ref. The bases form a tree, so
|
|
98
|
+
ids can be carried between any two snapshots that share an ancestor, forward or backward.
|
|
99
|
+
|
|
100
|
+
```ts
|
|
101
|
+
import { aliasChain, priorSnapshotForRef, resolveAlias } from "@titan-design/code-graph";
|
|
102
|
+
|
|
103
|
+
resolveAlias(store, "src/job.ts", 4, { fromSnapshotId: 1 });
|
|
104
|
+
// { id: 'core/worker.ts', reason: 'move',
|
|
105
|
+
// hops: [ src/job.ts -> src/task.ts, src/task.ts -> src/worker.ts, src/worker.ts -> core/worker.ts ] }
|
|
106
|
+
resolveAlias(store, "src/job.ts#Job.run", 4, { fromSnapshotId: 1 }).id; // 'core/worker.ts#Job.run'
|
|
107
|
+
aliasChain(store, 1, 4).resolve; // one resolver for many ids
|
|
108
|
+
priorSnapshotForRef(store, "main", { before: 7, repoRoot: "." }); // the snapshot main denotes
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Without `fromSnapshotId`, `resolveAlias` walks from the root of the target's lineage, so an id
|
|
112
|
+
from any ancestor resolves. Each snapshot's aliases are applied once, in order: they come from
|
|
113
|
+
one git diff, so `a.ts` to `b.ts` plus `b.ts` to `a.ts` in one snapshot is a swap, not a
|
|
114
|
+
cycle. Walks stop after 10,000 snapshots. Two snapshots with no common base fall back to the
|
|
115
|
+
to-snapshot's own aliases, which is what `diffSnapshots` did before 0.15.0.
|
|
116
|
+
|
|
117
|
+
`priorSnapshotForRef` prefers the snapshot of the commit git resolves the ref to, and
|
|
118
|
+
otherwise the newest snapshot labelled with that ref. A snapshot written before 0.15.0 records
|
|
119
|
+
no base; its base is inferred as the newest earlier snapshot with a commit, which is the one
|
|
120
|
+
the 0.14.0 indexer diffed against. Nothing is written on read, so a 0.14.0 store opened
|
|
121
|
+
read-only diffs and checks across renames too.
|
|
122
|
+
|
|
82
123
|
## The three reuse tiers
|
|
83
124
|
|
|
84
125
|
Every run writes a fingerprint per file: a content hash and a comment/whitespace-insensitive
|
|
@@ -87,10 +128,11 @@ same `INDEX_VERSION` and sorts each file into one tier. `INDEX_VERSION` is bumpe
|
|
|
87
128
|
metric can change for the same bytes, not only when the node or edge shape changes: 0.12.0
|
|
88
129
|
marks `.tsx` files moving to the tsx grammar, which changed their complexity metrics and
|
|
89
130
|
symbol spans; 0.13.0 marks the dead-code and growth-risk metrics joining the carry-forward set;
|
|
90
|
-
0.14.0 marks symbol ids qualified by their enclosing scopes (TP-182).
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
131
|
+
0.14.0 marks symbol ids qualified by their enclosing scopes (TP-182); 0.15.0 marks symbol
|
|
132
|
+
aliases on a file rename and the recorded alias base (TP-187). A snapshot from an older version
|
|
133
|
+
is never reused, so the first run after an upgrade is a full index. The first 0.14.0 run after
|
|
134
|
+
an older snapshot writes `requalify` id aliases from each bare-name id to its qualified
|
|
135
|
+
successor, only where exactly one declaration in the file carries that name.
|
|
94
136
|
|
|
95
137
|
| Tier | Trigger | Work skipped |
|
|
96
138
|
|---|---|---|
|
|
@@ -127,6 +169,57 @@ The symbol layer is hidden by default. `listNodes` drops `symbol` nodes and `lis
|
|
|
127
169
|
the graph it expects and does not have one import of thirty names read as thirty
|
|
128
170
|
dependencies.
|
|
129
171
|
|
|
172
|
+
## Targeted reads and the metric catalogue
|
|
173
|
+
|
|
174
|
+
Added for the read API (TP-183). Three store reads answer one node or one metric without
|
|
175
|
+
loading a snapshot, and a catalogue describes every metric name the package writes.
|
|
176
|
+
|
|
177
|
+
```ts
|
|
178
|
+
import { aggregateMetrics, describeMetric, listEdgesTouching, listMetricsForNode } from "@titan-design/code-graph";
|
|
179
|
+
|
|
180
|
+
listMetricsForNode(store, snapshotId, "packages/registry/src/invoke.ts"); // every metric on one node
|
|
181
|
+
listEdgesTouching(store, snapshotId, "packages/registry/src/invoke.ts"); // edges in and out; references on request
|
|
182
|
+
aggregateMetrics(store, snapshotId, { name: "loc" }); // [{ name: "loc", nodeKind: "file", count, sum, min, max }]
|
|
183
|
+
describeMetric("churn_90d"); // { unit: "lines", rollup: "sum", direction: "neutral", absent: "zero", window: "90d", … }
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
Three more store reads answer a report's questions without loading a snapshot:
|
|
187
|
+
`listMetricNames(snapshotId)` for the distinct names stored, `topByMetric({ snapshotId,
|
|
188
|
+
metric, limit, kind })` for the highest-valued nodes joined to their kind and role, and
|
|
189
|
+
`replaceMetricsByName(snapshotId, name, metrics)`, which swaps one metric's whole row set in
|
|
190
|
+
a single transaction so a re-ingested overlay such as coverage never accumulates stale rows.
|
|
191
|
+
|
|
192
|
+
`computeDeepAst({ filePath, absPath, symbolName })` reads structure too heavy to persist
|
|
193
|
+
(params, return type, class members) from the working tree on demand, and returns null when
|
|
194
|
+
the file is unreadable.
|
|
195
|
+
|
|
196
|
+
- **Reads.** `listMetricsForNode` searches the metric primary key, `listEdgesTouching` the edge
|
|
197
|
+
primary key plus `idx_edge_dst`, and `aggregateMetrics` `idx_metric_name`. None scans a
|
|
198
|
+
table, and no index or migration was added, so any existing store serves them. A self-loop
|
|
199
|
+
comes back once. `aggregateMetrics` groups by metric name and node kind, because
|
|
200
|
+
`utilization` and the degree metrics sit on several kinds; `count` skips null values.
|
|
201
|
+
- **Catalogue.** `METRIC_CATALOGUE` holds one `MetricDescriptor` per name: `unit`,
|
|
202
|
+
`appliesTo` (node kinds), `rollup` (`sum`, `max`, `mean`, or `none`), `direction`
|
|
203
|
+
(`higher-worse`, `lower-worse`, or `neutral`), `absent`, `source`, and `description`.
|
|
204
|
+
Windowed names are templates such as `churn_{w}`; `describeMetric` resolves a stored name
|
|
205
|
+
such as `churn_90d` or `test_bus_factor_lifetime` to a concrete descriptor, and returns
|
|
206
|
+
null for a name nothing writes.
|
|
207
|
+
- **`rollup: "none"`** means no rollup reproduces the group's own value. Summing fan-in counts
|
|
208
|
+
a directory's internal edges, and summing per-file commit counts counts a shared commit
|
|
209
|
+
more than once. A reader must not synthesize those.
|
|
210
|
+
- **`absent`** says what a missing row means. `zero`: the writer is sparse, so count the node
|
|
211
|
+
as zero (dead code, growth risk, churn, `linked_test_count`). `exclude`: the metric does not
|
|
212
|
+
apply or was not measured (a max over no functions, `coverage_pct` before ingest), so leave
|
|
213
|
+
the node out of means, percentiles, and ranks.
|
|
214
|
+
- **Completeness is tested.** `catalogue-completeness.test.ts` indexes a fixture repo with
|
|
215
|
+
history and a coverage overlay, and fails, naming the metric, when a stored name has no
|
|
216
|
+
descriptor, when a descriptor's unit or node kinds disagree with the rows, or when a
|
|
217
|
+
descriptor matches nothing.
|
|
218
|
+
- **Browser-safe.** The catalogue imports only `types.ts`. The `code-graph-catalogue-pure`
|
|
219
|
+
and `code-graph-catalogue-no-node` rules in `.codewatch/check.json` hold it there, so a
|
|
220
|
+
later browser subpath can re-export it. It ships from the root export today, which does
|
|
221
|
+
pull in Node.
|
|
222
|
+
|
|
130
223
|
## Checks and diffs
|
|
131
224
|
|
|
132
225
|
The rules engine turns a snapshot into pass/fail against a `check.json`. Six rule types:
|
|
@@ -148,23 +241,31 @@ engine on ids, and `validateRules(json)` validates an already-parsed rules objec
|
|
|
148
241
|
|
|
149
242
|
The baseline is a ratchet. A violation whose key (rule id, node id, and destination id for
|
|
150
243
|
edge rules) also fires on the baseline snapshot is marked `isCarryover` and counts as
|
|
151
|
-
carryover, so existing debt does not block a change but new debt does.
|
|
244
|
+
carryover, so existing debt does not block a change but new debt does. The baseline's node ids
|
|
245
|
+
are first carried through the alias chain into the checked snapshot (`rebasedViolationKey`),
|
|
246
|
+
so a moved file's violations carry over instead of reading as one resolved plus one new.
|
|
247
|
+
Unmoved ids key exactly as `violationKey` always did. Deprecated metric and
|
|
152
248
|
role spellings in a rules file (`lines`, `tests`) heal to their canonical names with a
|
|
153
249
|
warning through `onWarn` instead of failing validation.
|
|
154
250
|
|
|
155
251
|
`diffSnapshots(store, { fromSnapshotId, toSnapshotId })` reports added, removed and renamed
|
|
156
|
-
nodes, added and removed edges, and metric deltas on nodes present in both.
|
|
157
|
-
|
|
158
|
-
than a delete plus an add, and its edges do not churn. Metric and edge-kind spellings are
|
|
252
|
+
nodes, added and removed edges, and metric deltas on nodes present in both. Ids follow the
|
|
253
|
+
alias chain between the two snapshots, across every rename in between, so a move reads as a
|
|
254
|
+
rename rather than a delete plus an add, and its edges do not churn. Metric and edge-kind spellings are
|
|
159
255
|
canonicalised before comparing.
|
|
160
256
|
|
|
161
257
|
`diffCheckResults(store, { fromSnapshotId, toSnapshotId, rules })` runs the rules on both
|
|
162
258
|
snapshots and buckets each violation as new, resolved, or unchanged; unchanged metric
|
|
163
|
-
violations are further split into worsened and improved by value.
|
|
259
|
+
violations are further split into worsened and improved by value. From-side ids follow the
|
|
260
|
+
alias chain, as in the ratchet.
|
|
164
261
|
|
|
165
262
|
`scripts/dag-check-self.mjs` in the repo root runs this repo's DAG check on this engine
|
|
166
263
|
instead of codewatch's CLI.
|
|
167
264
|
|
|
265
|
+
The glob matching the rules use is exported too, because a CLI filters its own `--include`
|
|
266
|
+
and `--exclude` flags with the same semantics: `patternToRegex` (a `*` glob when the pattern
|
|
267
|
+
holds one, a case-sensitive substring otherwise), `compilePatterns`, and `matchesAny`.
|
|
268
|
+
|
|
168
269
|
## Similar symbols
|
|
169
270
|
|
|
170
271
|
Ported from codewatch's `embeddings.ts` (TP-129). It answers "does something like this
|
|
@@ -231,6 +332,78 @@ snapshotSymbolCoupling(store, snapshotId); // symbol pairs co-imported by 2+ fil
|
|
|
231
332
|
The pure `computePageRank`, `computeRelevance`, `computeSymbolConsumers`, and
|
|
232
333
|
`computeSymbolCoupling` take node and edge arrays instead of a store.
|
|
233
334
|
|
|
335
|
+
### Partition quality
|
|
336
|
+
|
|
337
|
+
`computePartitionQuality` scores a package partition of the file graph: per-package cohesion,
|
|
338
|
+
Martin instability, an abstractness proxy (the share of `role: "types"` files), a layer label
|
|
339
|
+
(`top`, `middle`, `foundation`), pair coupling by intensity (`edges / files(from)`), and a
|
|
340
|
+
Newman-Girvan modularity Q over the whole partition.
|
|
341
|
+
|
|
342
|
+
```ts
|
|
343
|
+
import { computePartitionQuality, invertBuckets } from "@titan-design/code-graph";
|
|
344
|
+
|
|
345
|
+
const result = computePartitionQuality({ packages, fileByPackage, nodes, edges });
|
|
346
|
+
result.modularityQ; // 0.77 on titan-platform's packages/*
|
|
347
|
+
invertBuckets(fileByPackage); // file id -> package id, skipping the "" unassigned bucket
|
|
348
|
+
```
|
|
349
|
+
|
|
350
|
+
`resolveBarrels: true` rewrites an edge landing on a `role: "barrel"` file to the files it
|
|
351
|
+
re-exports, transitively. It over-attributes: one import of one name through a barrel becomes
|
|
352
|
+
one edge per re-export target, which is why it is off by default.
|
|
353
|
+
|
|
354
|
+
### Pruning snapshots
|
|
355
|
+
|
|
356
|
+
`planPrune` keeps the most recent `keep` snapshots (default 10) plus every snapshot whose ref
|
|
357
|
+
is in `keepRefs`; `runPrune` deletes the rest and reports row counts before and after.
|
|
358
|
+
|
|
359
|
+
```ts
|
|
360
|
+
import { planPrune, runPrune } from "@titan-design/code-graph";
|
|
361
|
+
|
|
362
|
+
planPrune(store, { keep: 10, keepRefs: ["main"] }); // { keep, remove }
|
|
363
|
+
runPrune(store, { keep: 10, vacuum: true }); // { plan, rowsBefore, rowsAfter, vacuumed }
|
|
364
|
+
```
|
|
365
|
+
|
|
366
|
+
The domain tables declare no foreign key, so `CodeGraphStore.deleteSnapshots` clears each of
|
|
367
|
+
`SNAPSHOT_SCOPED_TABLES` itself rather than relying on a cascade. `blob_cache` is
|
|
368
|
+
content-addressed and not snapshot-scoped, so a prune never drops a cached embedding.
|
|
369
|
+
|
|
370
|
+
## Conventions
|
|
371
|
+
|
|
372
|
+
"How does this repo do X, and where does code like this belong?" Ported from codewatch's
|
|
373
|
+
unmerged C-88 branch (TP-130). The layer cuts the barrel-resolved file graph into a few coarse
|
|
374
|
+
areas, has an injected summarizer describe each one, and ranks areas against a question by
|
|
375
|
+
embedding similarity. Verified against this release on this repo's `packages/`, with a fake
|
|
376
|
+
summarizer:
|
|
377
|
+
|
|
378
|
+
```ts
|
|
379
|
+
import { findConventions, getConventionMap, summarizeConventions } from "@titan-design/code-graph";
|
|
380
|
+
|
|
381
|
+
const summarizer = { model: "claude:sonnet", summarize: (prompt: string) => callYourLlm(prompt) };
|
|
382
|
+
await summarizeConventions(store, snapshotId, summarizer);
|
|
383
|
+
// { coverage: { files: 713, grouped: 513, areas: 24, summarized: 24 }, newlySummarized: 24, reused: 0, … }
|
|
384
|
+
// a second run: { newlySummarized: 0, reused: 24 }
|
|
385
|
+
|
|
386
|
+
getConventionMap(store, snapshotId, "claude:sonnet"); // the same areas, stored summaries only
|
|
387
|
+
await findConventions(store, snapshotId, "how are CLI commands registered?", embedder, "claude:sonnet");
|
|
388
|
+
// { matches: [ { label, summary, files: [ …up to 5 ], size, score }, … up to 3 ] }
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
- **The cut.** `detectCommunities` is greedy modularity (Clauset-Newman-Moore), not Leiden. It
|
|
392
|
+
is deterministic without a seed: ties resolve by sorted id. `targetCount` keeps merging past
|
|
393
|
+
the natural modularity stop until that many communities remain, and a size cap of twice the
|
|
394
|
+
ideal share keeps a dense repo from collapsing into one area. The default target is one area
|
|
395
|
+
per 25 files, clamped to 6..40. Areas under `minSize` (default 3) files are left unsummarized.
|
|
396
|
+
Disconnected components never merge, so the component count is the floor.
|
|
397
|
+
- **Only the coarse level is summarized.** LLM cost is one call per area, and each summary is
|
|
398
|
+
stored in `blob_cache` under `code-graph/community-summary`, keyed by `summarizer.model` and a
|
|
399
|
+
hash of the prompt. The prompt carries the member files and their key exported signatures, so
|
|
400
|
+
an area whose membership and signatures did not change is a cache hit in any snapshot. On
|
|
401
|
+
this repo a finer cut (`targetCount: 60`) still reused 8 of its 35 areas.
|
|
402
|
+
- **The package ships no LLM client.** `Summarizer` is `{ model, summarize(prompt) }`; the
|
|
403
|
+
product supplies it. `getConventionMap` and `findConventions` never call it. `findConventions`
|
|
404
|
+
throws when no summary is stored for the model, and returns candidates with scores, not
|
|
405
|
+
verdicts. Summary vectors go through the same embedding cache as similar symbols.
|
|
406
|
+
|
|
234
407
|
## Git history
|
|
235
408
|
|
|
236
409
|
Ported in TP-126, strictly as codewatch had it: churn over rolling windows, first-seen dates,
|
|
@@ -255,6 +428,14 @@ plus `churnWindowDays` (the primary, default 30); `churnWindows` replaces the de
|
|
|
255
428
|
turns all of it off. Outside git, or without a git binary, the index simply has no history
|
|
256
429
|
metrics.
|
|
257
430
|
|
|
431
|
+
The adapter is exported from the package root, not from `./history`, because it speaks
|
|
432
|
+
`GraphMetric` and the seam below does not. A product that indexes on its own terms calls
|
|
433
|
+
`loadHistoryMetrics(nodes, idRoot, options)` for both the metric rows and the primary-window
|
|
434
|
+
churn entries they were built from (`LoadedHistory`), with `HistoryMetricsOptions`,
|
|
435
|
+
`DEFAULT_CHURN_WINDOWS` (`[30, 90, 180]`), `resolveChurnWindows`, `windowSuffix` (`30d`,
|
|
436
|
+
`lifetime`) and `computeRecencyWindows` alongside it. Node ids are the history engine's
|
|
437
|
+
repo-relative paths, so both must be rooted at the same `idRoot`.
|
|
438
|
+
|
|
258
439
|
Change coupling is not stored. It is computed on demand from `loadChurnEntries`, as
|
|
259
440
|
codewatch's `graph coupled` command did.
|
|
260
441
|
|