@titan-design/code-graph 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -40,11 +40,14 @@ coupling, now `src/analysis/`. See [Graph analyses](#graph-analyses).
40
40
  Ported with TP-133: the test linker and the Istanbul coverage overlay, also in `src/analysis/`,
41
41
  and test-coverage ownership. See [Test linking and coverage](#test-linking-and-coverage).
42
42
 
43
+ Ported with TP-250: package partition quality (`src/analysis/partition-quality.ts`) and
44
+ snapshot pruning (`src/prune.ts`). See [Partition quality](#partition-quality) and
45
+ [Pruning snapshots](#pruning-snapshots).
46
+
43
47
  Deferred, all of it still in codewatch, all of it a follow-up on this package rather than a
44
48
  change to it:
45
49
 
46
- - Graph analyses over a finished snapshot: communities, partition quality, conventions,
47
- patterns, prune, reuse-delta reporting.
50
+ - Graph analyses over a finished snapshot: communities, conventions, reuse-delta reporting.
48
51
 
49
52
  Python support is new here rather than ported. codewatch walked TypeScript only; the parser
50
53
  already had the grammar. The Python extractor is deliberately narrower than the ts-morph one:
@@ -79,6 +82,41 @@ package's source file, not to an `npm:` external, by remapping the `dist/*.d.ts`
79
82
  ts-morph resolves back onto `src/`. That remap needs the target package built, which is why
80
83
  `pnpm build` precedes both `pnpm test` and `dag:check`.
81
84
 
85
+ ## Identity across renames
86
+
87
+ Added in index version 0.15.0 (TP-187). Git rename detection writes `id_alias` rows mapping a
88
+ file's and its module's old id to the new one. A file rename also writes an alias for every
89
+ symbol present on both sides, so `a.ts#Job.run` follows `a.ts` to `b.ts#Job.run`. A symbol
90
+ renamed inside a file is not followed.
91
+
92
+ Each snapshot records its alias base in `attrs.aliasBase`: the snapshot whose commit its
93
+ aliases were computed against. A new index of a ref diffs against that ref's newest committed
94
+ snapshot, falling back to the newest committed snapshot of any ref. The bases form a tree, so
95
+ ids can be carried between any two snapshots that share an ancestor, forward or backward.
96
+
97
+ ```ts
98
+ import { aliasChain, priorSnapshotForRef, resolveAlias } from "@titan-design/code-graph";
99
+
100
+ resolveAlias(store, "src/job.ts", 4, { fromSnapshotId: 1 });
101
+ // { id: 'core/worker.ts', reason: 'move',
102
+ // hops: [ src/job.ts -> src/task.ts, src/task.ts -> src/worker.ts, src/worker.ts -> core/worker.ts ] }
103
+ resolveAlias(store, "src/job.ts#Job.run", 4, { fromSnapshotId: 1 }).id; // 'core/worker.ts#Job.run'
104
+ aliasChain(store, 1, 4).resolve; // one resolver for many ids
105
+ priorSnapshotForRef(store, "main", { before: 7, repoRoot: "." }); // the snapshot main denotes
106
+ ```
107
+
108
+ Without `fromSnapshotId`, `resolveAlias` walks from the root of the target's lineage, so an id
109
+ from any ancestor resolves. Each snapshot's aliases are applied once, in order: they come from
110
+ one git diff, so `a.ts` to `b.ts` plus `b.ts` to `a.ts` in one snapshot is a swap, not a
111
+ cycle. Walks stop after 10,000 snapshots. Two snapshots with no common base fall back to the
112
+ to-snapshot's own aliases, which is what `diffSnapshots` did before 0.15.0.
113
+
114
+ `priorSnapshotForRef` prefers the snapshot of the commit git resolves the ref to, and
115
+ otherwise the newest snapshot labelled with that ref. A snapshot written before 0.15.0 records
116
+ no base; its base is inferred as the newest earlier snapshot with a commit, which is the one
117
+ the 0.14.0 indexer diffed against. Nothing is written on read, so a 0.14.0 store opened
118
+ read-only diffs and checks across renames too.
119
+
82
120
  ## The three reuse tiers
83
121
 
84
122
  Every run writes a fingerprint per file: a content hash and a comment/whitespace-insensitive
@@ -87,10 +125,11 @@ same `INDEX_VERSION` and sorts each file into one tier. `INDEX_VERSION` is bumpe
87
125
  metric can change for the same bytes, not only when the node or edge shape changes: 0.12.0
88
126
  marks `.tsx` files moving to the tsx grammar, which changed their complexity metrics and
89
127
  symbol spans; 0.13.0 marks the dead-code and growth-risk metrics joining the carry-forward set;
90
- 0.14.0 marks symbol ids qualified by their enclosing scopes (TP-182). A snapshot from before
91
- 0.14.0 is never reused, so re-index. The first 0.14.0 run after an older snapshot writes
92
- `requalify` id aliases from each bare-name id to its qualified successor, only where exactly one
93
- declaration in the file carries that name.
128
+ 0.14.0 marks symbol ids qualified by their enclosing scopes (TP-182); 0.15.0 marks symbol
129
+ aliases on a file rename and the recorded alias base (TP-187). A snapshot from an older version
130
+ is never reused, so the first run after an upgrade is a full index. The first 0.14.0 run after
131
+ an older snapshot writes `requalify` id aliases from each bare-name id to its qualified
132
+ successor, only where exactly one declaration in the file carries that name.
94
133
 
95
134
  | Tier | Trigger | Work skipped |
96
135
  |---|---|---|
@@ -127,6 +166,57 @@ The symbol layer is hidden by default. `listNodes` drops `symbol` nodes and `lis
127
166
  the graph it expects and does not have one import of thirty names read as thirty
128
167
  dependencies.
129
168
 
169
+ ## Targeted reads and the metric catalogue
170
+
171
+ Added for the read API (TP-183). Three store reads answer one node or one metric without
172
+ loading a snapshot, and a catalogue describes every metric name the package writes.
173
+
174
+ ```ts
175
+ import { aggregateMetrics, describeMetric, listEdgesTouching, listMetricsForNode } from "@titan-design/code-graph";
176
+
177
+ listMetricsForNode(store, snapshotId, "packages/registry/src/invoke.ts"); // every metric on one node
178
+ listEdgesTouching(store, snapshotId, "packages/registry/src/invoke.ts"); // edges in and out; references on request
179
+ aggregateMetrics(store, snapshotId, { name: "loc" }); // [{ name: "loc", nodeKind: "file", count, sum, min, max }]
180
+ describeMetric("churn_90d"); // { unit: "lines", rollup: "sum", direction: "neutral", absent: "zero", window: "90d", … }
181
+ ```
182
+
183
+ Three more store reads answer a report's questions without loading a snapshot:
184
+ `listMetricNames(snapshotId)` for the distinct names stored, `topByMetric({ snapshotId,
185
+ metric, limit, kind })` for the highest-valued nodes joined to their kind and role, and
186
+ `replaceMetricsByName(snapshotId, name, metrics)`, which swaps one metric's whole row set in
187
+ a single transaction so a re-ingested overlay such as coverage never accumulates stale rows.
188
+
189
+ `computeDeepAst({ filePath, absPath, symbolName })` reads structure too heavy to persist
190
+ (params, return type, class members) from the working tree on demand, and returns null when
191
+ the file is unreadable.
192
+
193
+ - **Reads.** `listMetricsForNode` searches the metric primary key, `listEdgesTouching` the edge
194
+ primary key plus `idx_edge_dst`, and `aggregateMetrics` `idx_metric_name`. None scans a
195
+ table, and no index or migration was added, so any existing store serves them. A self-loop
196
+ comes back once. `aggregateMetrics` groups by metric name and node kind, because
197
+ `utilization` and the degree metrics sit on several kinds; `count` skips null values.
198
+ - **Catalogue.** `METRIC_CATALOGUE` holds one `MetricDescriptor` per name: `unit`,
199
+ `appliesTo` (node kinds), `rollup` (`sum`, `max`, `mean`, or `none`), `direction`
200
+ (`higher-worse`, `lower-worse`, or `neutral`), `absent`, `source`, and `description`.
201
+ Windowed names are templates such as `churn_{w}`; `describeMetric` resolves a stored name
202
+ such as `churn_90d` or `test_bus_factor_lifetime` to a concrete descriptor, and returns
203
+ null for a name nothing writes.
204
+ - **`rollup: "none"`** means no rollup reproduces the group's own value. Summing fan-in counts
205
+ a directory's internal edges, and summing per-file commit counts counts a shared commit
206
+ more than once. A reader must not synthesize those.
207
+ - **`absent`** says what a missing row means. `zero`: the writer is sparse, so count the node
208
+ as zero (dead code, growth risk, churn, `linked_test_count`). `exclude`: the metric does not
209
+ apply or was not measured (a max over no functions, `coverage_pct` before ingest), so leave
210
+ the node out of means, percentiles, and ranks.
211
+ - **Completeness is tested.** `catalogue-completeness.test.ts` indexes a fixture repo with
212
+ history and a coverage overlay, and fails, naming the metric, when a stored name has no
213
+ descriptor, when a descriptor's unit or node kinds disagree with the rows, or when a
214
+ descriptor matches nothing.
215
+ - **Browser-safe.** The catalogue imports only `types.ts`. The `code-graph-catalogue-pure`
216
+ and `code-graph-catalogue-no-node` rules in `.codewatch/check.json` hold it there, so a
217
+ later browser subpath can re-export it. It ships from the root export today, which does
218
+ pull in Node.
219
+
130
220
  ## Checks and diffs
131
221
 
132
222
  The rules engine turns a snapshot into pass/fail against a `check.json`. Six rule types:
@@ -148,23 +238,31 @@ engine on ids, and `validateRules(json)` validates an already-parsed rules objec
148
238
 
149
239
  The baseline is a ratchet. A violation whose key (rule id, node id, and destination id for
150
240
  edge rules) also fires on the baseline snapshot is marked `isCarryover` and counts as
151
- carryover, so existing debt does not block a change but new debt does. Deprecated metric and
241
+ carryover, so existing debt does not block a change but new debt does. The baseline's node ids
242
+ are first carried through the alias chain into the checked snapshot (`rebasedViolationKey`),
243
+ so a moved file's violations carry over instead of reading as one resolved plus one new.
244
+ Unmoved ids key exactly as `violationKey` always did. Deprecated metric and
152
245
  role spellings in a rules file (`lines`, `tests`) heal to their canonical names with a
153
246
  warning through `onWarn` instead of failing validation.
154
247
 
155
248
  `diffSnapshots(store, { fromSnapshotId, toSnapshotId })` reports added, removed and renamed
156
- nodes, added and removed edges, and metric deltas on nodes present in both. The
157
- to-snapshot's `id_alias` rows carry a renamed file across, so a move reads as a rename rather
158
- than a delete plus an add, and its edges do not churn. Metric and edge-kind spellings are
249
+ nodes, added and removed edges, and metric deltas on nodes present in both. Ids follow the
250
+ alias chain between the two snapshots, across every rename in between, so a move reads as a
251
+ rename rather than a delete plus an add, and its edges do not churn. Metric and edge-kind spellings are
159
252
  canonicalised before comparing.
160
253
 
161
254
  `diffCheckResults(store, { fromSnapshotId, toSnapshotId, rules })` runs the rules on both
162
255
  snapshots and buckets each violation as new, resolved, or unchanged; unchanged metric
163
- violations are further split into worsened and improved by value.
256
+ violations are further split into worsened and improved by value. From-side ids follow the
257
+ alias chain, as in the ratchet.
164
258
 
165
259
  `scripts/dag-check-self.mjs` in the repo root runs this repo's DAG check on this engine
166
260
  instead of codewatch's CLI.
167
261
 
262
+ The glob matching the rules use is exported too, because a CLI filters its own `--include`
263
+ and `--exclude` flags with the same semantics: `patternToRegex` (a `*` glob when the pattern
264
+ holds one, a case-sensitive substring otherwise), `compilePatterns`, and `matchesAny`.
265
+
168
266
  ## Similar symbols
169
267
 
170
268
  Ported from codewatch's `embeddings.ts` (TP-129). It answers "does something like this
@@ -231,6 +329,41 @@ snapshotSymbolCoupling(store, snapshotId); // symbol pairs co-imported by 2+ fil
231
329
  The pure `computePageRank`, `computeRelevance`, `computeSymbolConsumers`, and
232
330
  `computeSymbolCoupling` take node and edge arrays instead of a store.
233
331
 
332
+ ### Partition quality
333
+
334
+ `computePartitionQuality` scores a package partition of the file graph: per-package cohesion,
335
+ Martin instability, an abstractness proxy (the share of `role: "types"` files), a layer label
336
+ (`top`, `middle`, `foundation`), pair coupling by intensity (`edges / files(from)`), and a
337
+ Newman-Girvan modularity Q over the whole partition.
338
+
339
+ ```ts
340
+ import { computePartitionQuality, invertBuckets } from "@titan-design/code-graph";
341
+
342
+ const result = computePartitionQuality({ packages, fileByPackage, nodes, edges });
343
+ result.modularityQ; // 0.77 on titan-platform's packages/*
344
+ invertBuckets(fileByPackage); // file id -> package id, skipping the "" unassigned bucket
345
+ ```
346
+
347
+ `resolveBarrels: true` rewrites an edge landing on a `role: "barrel"` file to the files it
348
+ re-exports, transitively. It over-attributes: one import of one name through a barrel becomes
349
+ one edge per re-export target, which is why it is off by default.
350
+
351
+ ### Pruning snapshots
352
+
353
+ `planPrune` keeps the most recent `keep` snapshots (default 10) plus every snapshot whose ref
354
+ is in `keepRefs`; `runPrune` deletes the rest and reports row counts before and after.
355
+
356
+ ```ts
357
+ import { planPrune, runPrune } from "@titan-design/code-graph";
358
+
359
+ planPrune(store, { keep: 10, keepRefs: ["main"] }); // { keep, remove }
360
+ runPrune(store, { keep: 10, vacuum: true }); // { plan, rowsBefore, rowsAfter, vacuumed }
361
+ ```
362
+
363
+ The domain tables declare no foreign key, so `CodeGraphStore.deleteSnapshots` clears each of
364
+ `SNAPSHOT_SCOPED_TABLES` itself rather than relying on a cascade. `blob_cache` is
365
+ content-addressed and not snapshot-scoped, so a prune never drops a cached embedding.
366
+
234
367
  ## Git history
235
368
 
236
369
  Ported in TP-126, strictly as codewatch had it: churn over rolling windows, first-seen dates,
@@ -255,6 +388,14 @@ plus `churnWindowDays` (the primary, default 30); `churnWindows` replaces the de
255
388
  turns all of it off. Outside git, or without a git binary, the index simply has no history
256
389
  metrics.
257
390
 
391
+ The adapter is exported from the package root, not from `./history`, because it speaks
392
+ `GraphMetric` and the seam below does not. A product that indexes on its own terms calls
393
+ `loadHistoryMetrics(nodes, idRoot, options)` for both the metric rows and the primary-window
394
+ churn entries they were built from (`LoadedHistory`), with `HistoryMetricsOptions`,
395
+ `DEFAULT_CHURN_WINDOWS` (`[30, 90, 180]`), `resolveChurnWindows`, `windowSuffix` (`30d`,
396
+ `lifetime`) and `computeRecencyWindows` alongside it. Node ids are the history engine's
397
+ repo-relative paths, so both must be rooted at the same `idRoot`.
398
+
258
399
  Change coupling is not stored. It is computed on demand from `loadChurnEntries`, as
259
400
  codewatch's `graph coupled` command did.
260
401