@titan-design/code-graph 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -5,8 +5,8 @@ symbols, external packages, and the import / re-export / reference edges between
5
5
  index-time source metrics, in one SQLite file built from `@titan-design/store-sqlite` kit
6
6
  tables and refreshed incrementally.
7
7
 
8
- Tier 2 of the titan-platform DAG. Depends on `store-sqlite`, `ts-morph`, and the tree-sitter
9
- WASM grammars. Extracted from codewatch's `@codewatch/graph` (TP-9), split along the seam the
8
+ Tier 2 of the titan-platform DAG. Depends on `store-sqlite`, `code-parser` (tree-sitter WASM
9
+ parsing and the file filter), `ts-morph`, `web-tree-sitter` for node types, and `embed` plus `retrieval` for similar-symbol search. Extracted from codewatch's `@codewatch/graph` (TP-9), split along the seam the
10
10
  audit identified: that package did the job of both a store and a code graph.
11
11
 
12
12
  ```ts
@@ -23,7 +23,8 @@ listEdges(store, snapshotId); // imports / re-exports, references on request
23
23
 
24
24
  ## What was extracted, and what was not
25
25
 
26
- In: the parser (tree-sitter WASM for TypeScript, TSX and Python), the walk, the ts-morph
26
+ In: the parser (tree-sitter WASM for TypeScript, TSX and Python, since moved to
27
+ `@titan-design/code-parser` and still re-exported here), the walk, the ts-morph
27
28
  extractor and its symbol layer, role classification, generated-file detection, id aliasing
28
29
  across git renames, the three-tier incremental reuse, and the metrics computed at index time
29
30
  (degree, utilization, loc, cyclomatic, cognitive, nesting, class count, lcom4, per-symbol
@@ -31,16 +32,19 @@ complexity). `lcom.ts` came along despite being an analysis: `source-metrics.ts`
31
32
  directly and lcom4 is a pure function of a file's bytes, so it belongs with the metrics that
32
33
  carry forward under reuse.
33
34
 
35
+ Ported later (TP-123, TP-124), strictly as codewatch had them: the rules engine
36
+ (`check*.ts`, now `src/check/`) and the snapshot diff plus check diff (`diff.ts`,
37
+ `check-diff.ts`, now `src/diff/`). See [Checks and diffs](#checks-and-diffs).
38
+ Ported with TP-127 and TP-128: dead code, growth risk, PageRank, relevance, and symbol
39
+ coupling, now `src/analysis/`. See [Graph analyses](#graph-analyses).
40
+ Ported with TP-133: the test linker and the Istanbul coverage overlay, also in `src/analysis/`,
41
+ and test-coverage ownership. See [Test linking and coverage](#test-linking-and-coverage).
42
+
34
43
  Deferred, all of it still in codewatch, all of it a follow-up on this package rather than a
35
44
  change to it:
36
45
 
37
- - The rules engine (`check*.ts`) that turns a snapshot into pass/fail against a config.
38
- - Git-history mining: churn, ownership, change coupling, symbol coupling, test coverage
39
- linking. `buildIndexerMetrics` used to fold these into the same pass; here it computes only
40
- what a file's own bytes and the assembled graph determine.
41
- - Graph analyses over a finished snapshot: communities, pagerank, partition quality,
42
- relevance, conventions, coverage overlay, dead code, growth risk, patterns, prune,
43
- test-linker, diff, reuse-delta reporting, embeddings.
46
+ - Graph analyses over a finished snapshot: communities, partition quality, conventions,
47
+ patterns, prune, reuse-delta reporting.
44
48
 
45
49
  Python support is new here rather than ported. codewatch walked TypeScript only; the parser
46
50
  already had the grammar. The Python extractor is deliberately narrower than the ts-morph one:
@@ -49,16 +53,24 @@ same tree-sitter declaration walk that feeds complexity.
49
53
 
50
54
  ## The id scheme
51
55
 
52
- Preserved exactly from codewatch, because this repo's own `dag:check` consumes it through
53
- codewatch's CLI:
56
+ File, module and external ids are preserved exactly from codewatch, because this repo's own
57
+ `dag:check` consumes them through codewatch's CLI. Symbol ids diverge from codewatch since
58
+ index version 0.14.0:
54
59
 
55
60
  - A **file** id is its path relative to the git toplevel, in posix form:
56
61
  `packages/registry/src/index.ts`. Ids root at the git toplevel even when you walk a
57
62
  subtree, so importers across subtrees share one id space.
58
63
  - A **module** id is the file id minus its extension: `packages/registry/src/index`. Its
59
64
  parent is the directory above it.
60
- - A **symbol** id hangs under its declaring file as `<fileId>#<name>`. `#` is legal in
61
- neither a posix path nor a JS identifier, so the first one is the split.
65
+ - A **symbol** id hangs under its declaring file as `<fileId>#<qualifiedName>`, split on the
66
+ first `#`. A top-level declaration's qualified name is its own name
67
+ (`src/a.ts#createThing`). A member or nested declaration is prefixed by its enclosing named
68
+ scopes, joined with `.`: `src/a.ts#Job.run`, `src/a.ts#outer.helper`, and
69
+ `src/a.ts#handlers.onClick` for a method of `const handlers = {…}`. Anonymous scopes, such
70
+ as a callback argument or an unbound class expression, add no segment, so ids do not depend
71
+ on declaration order. One scope binds a name once: a getter/setter pair, a Python property's
72
+ accessors, and overloads each share one node. Index versions before 0.14.0 keyed members by
73
+ bare name, so same-named methods in one file collapsed into one node (TP-182).
62
74
  - An **external** id is `npm:<package>` (scope-aware) or the `node:` builtin verbatim.
63
75
 
64
76
  `NodeKind`, `EdgeKind`, and the `role` vocabulary are unchanged. So is the property the DAG
@@ -71,7 +83,14 @@ ts-morph resolves back onto `src/`. That remap needs the target package built, w
71
83
 
72
84
  Every run writes a fingerprint per file: a content hash and a comment/whitespace-insensitive
73
85
  hash of its parse structure. The next run diffs against the most recent snapshot carrying the
74
- same `INDEX_VERSION` and sorts each file into one tier.
86
+ same `INDEX_VERSION` and sorts each file into one tier. `INDEX_VERSION` is bumped whenever a
87
+ metric can change for the same bytes, not only when the node or edge shape changes: 0.12.0
88
+ marks `.tsx` files moving to the tsx grammar, which changed their complexity metrics and
89
+ symbol spans; 0.13.0 marks the dead-code and growth-risk metrics joining the carry-forward set;
90
+ 0.14.0 marks symbol ids qualified by their enclosing scopes (TP-182). A snapshot from before
91
+ 0.14.0 is never reused, so re-index. The first 0.14.0 run after an older snapshot writes
92
+ `requalify` id aliases from each bare-name id to its qualified successor, only where exactly one
93
+ declaration in the file carries that name.
75
94
 
76
95
  | Tier | Trigger | Work skipped |
77
96
  |---|---|---|
@@ -88,6 +107,10 @@ always recomputed over the whole assembled graph, so a heavily-reused run and a
88
107
 
89
108
  `openCodeGraph` opens the database with the kit's pragmas and runs three migrations.
90
109
 
110
+ It refuses a database stamped past `SCHEMA_VERSION` with `SchemaTooNewError`: a code graph
111
+ database belongs to one build, so a higher version means a newer build already moved the
112
+ schema and this one would query columns that are gone.
113
+
91
114
  Kit tables: `snapshot` (the snapshot registry), `node` (`entity_snap`, keyed
92
115
  `(snapshot_id, id)`), `blob_cache` (content-addressed, for the embedding and summary caches
93
116
  an analysis layer will want). Migration 2 adds four columns the kit shapes do not carry but
@@ -103,3 +126,174 @@ The symbol layer is hidden by default. `listNodes` drops `symbol` nodes and `lis
103
126
  `references` edges unless you ask for them, so a caller reasoning about module structure sees
104
127
  the graph it expects and does not have one import of thirty names read as thirty
105
128
  dependencies.
129
+
130
+ ## Checks and diffs
131
+
132
+ The rules engine turns a snapshot into pass/fail against a `check.json`. Six rule types:
133
+ `metric-max`, `metric-min`, `metric-product-max`, `forbid-import`, `layered-deps`, and
134
+ `no-internal-only-barrels`. Severity defaults to `error`; only new errors fail a check.
135
+
136
+ ```ts
137
+ import { checkSnapshot, loadCheckRules, openCodeGraph } from "@titan-design/code-graph";
138
+
139
+ const store = openCodeGraph(".codewatch/graph.db");
140
+ const rules = await loadCheckRules(".codewatch/check.json", { onWarn: console.warn });
141
+ const { result } = checkSnapshot(store, { snapshot: "head", baseline: "main", rules });
142
+ result.passed; // false only when a violation is an error and absent from the baseline
143
+ ```
144
+
145
+ `snapshot` and `baseline` take a numeric snapshot id or a ref name; a ref resolves to its
146
+ newest snapshot. `runChecks(store, { snapshotId, rules, baselineSnapshotId })` is the same
147
+ engine on ids, and `validateRules(json)` validates an already-parsed rules object.
148
+
149
+ The baseline is a ratchet. A violation whose key (rule id, node id, and destination id for
150
+ edge rules) also fires on the baseline snapshot is marked `isCarryover` and counts as
151
+ carryover, so existing debt does not block a change but new debt does. Deprecated metric and
152
+ role spellings in a rules file (`lines`, `tests`) heal to their canonical names with a
153
+ warning through `onWarn` instead of failing validation.
154
+
155
+ `diffSnapshots(store, { fromSnapshotId, toSnapshotId })` reports added, removed and renamed
156
+ nodes, added and removed edges, and metric deltas on nodes present in both. The
157
+ to-snapshot's `id_alias` rows carry a renamed file across, so a move reads as a rename rather
158
+ than a delete plus an add, and its edges do not churn. Metric and edge-kind spellings are
159
+ canonicalised before comparing.
160
+
161
+ `diffCheckResults(store, { fromSnapshotId, toSnapshotId, rules })` runs the rules on both
162
+ snapshots and buckets each violation as new, resolved, or unchanged; unchanged metric
163
+ violations are further split into worsened and improved by value.
164
+
165
+ `scripts/dag-check-self.mjs` in the repo root runs this repo's DAG check on this engine
166
+ instead of codewatch's CLI.
167
+
168
+ ## Similar symbols
169
+
170
+ Ported from codewatch's `embeddings.ts` (TP-129). It answers "does something like this
171
+ already exist?" before you write it. The embedder is injected; this package never builds
172
+ one and never talks to a network on its own.
173
+
174
+ ```ts
175
+ import { OllamaEmbedder } from "@titan-design/embed";
176
+ import { findSimilarCapability, tryEmbedSnapshot } from "@titan-design/code-graph";
177
+
178
+ const embedder = new OllamaEmbedder();
179
+ const attempt = await tryEmbedSnapshot(store, snapshotId, embedder);
180
+ // { ok: true, result: { symbols, embedded, withPurpose, newlyEmbedded, reused, model } }
181
+ // or { ok: false, model, error } when the backend is down; the snapshot is unaffected
182
+ const { candidates, coverage } = await findSimilarCapability(store, snapshotId, "parse a duration", embedder);
183
+ ```
184
+
185
+ - **Corpus.** Exported symbols with a signature, minus test, fixture and generated files.
186
+ The embedded text is `signature -- purpose` (the docstring), never the body.
187
+ - **Storage.** Vectors go in `blob_cache` under namespace `code-graph/symbol-embedding`,
188
+ keyed by `embedder.model` and the SHA-256 of the text. They are not snapshot-scoped, so
189
+ re-embedding unchanged text costs zero embed calls. `embedSnapshot` throws on a backend
190
+ failure; `tryEmbedSnapshot` reports it.
191
+ - **Query.** `vectorRetriever` over a `BruteForceVectorIndex` from `retrieval`. The query is
192
+ embedded with role `query` and the symbol texts with role `document`; the embedder applies
193
+ the matching prefix (`search_query: ` / `search_document: ` for nomic models).
194
+ - **Results are candidates, not verdicts.** Each has a cosine score, and `coverage` says how
195
+ many symbols were searchable and how many carry purpose text. There is deliberately no
196
+ co-location filter.
197
+ - **Python symbols are not in the corpus yet.** The Python extractor records no signature or
198
+ docstring, so no Python symbol passes the corpus filter.
199
+ - **Prefixes are part of the cache key.** `embedder.model` includes a hash of the embedder's
200
+ prefix table, so two prefix configurations never share a vector. codewatch used no prefix;
201
+ `new OllamaEmbedder({ prefixes: { document: "", query: "" } })` reproduces its rankings.
202
+
203
+ ## Graph analyses
204
+
205
+ Dead-code and growth-risk metrics are computed at index time, like the source metrics, and
206
+ carry forward for unchanged files. Both are sparse: a file gets a row only when a count is
207
+ above zero.
208
+
209
+ - Dead code, TypeScript only: `unreachable_statements` (after a `return`, `throw`, `break`
210
+ or `continue` in the same block), `unused_locals`, and `unused_params` (trailing run only).
211
+ - Growth risk, TypeScript and Python: `loop_depth` (at 2 or more), `recursive_functions`,
212
+ and `search_in_loop` (`.includes`, `.find` and similar inside a loop). These are smells,
213
+ not complexity bounds. Recursion and search match TypeScript call nodes only, so Python
214
+ files get `loop_depth` alone, as in codewatch.
215
+
216
+ PageRank, relevance, and symbol coupling run at query time over one snapshot:
217
+
218
+ ```ts
219
+ import {
220
+ snapshotPageRank,
221
+ snapshotRelevance,
222
+ snapshotSymbolCoupling,
223
+ } from "@titan-design/code-graph";
224
+
225
+ snapshotPageRank(store, snapshotId); // global centrality over the file-level graph
226
+ snapshotPageRank(store, snapshotId, { personalization: new Map([[fileId, 1]]) }); // seeded
227
+ snapshotRelevance(store, snapshotId, [fileId]); // seeded over symmetrized edges
228
+ snapshotSymbolCoupling(store, snapshotId); // symbol pairs co-imported by 2+ files
229
+ ```
230
+
231
+ The pure `computePageRank`, `computeRelevance`, `computeSymbolConsumers`, and
232
+ `computeSymbolCoupling` take node and edge arrays instead of a store.
233
+
234
+ ## Git history
235
+
236
+ Ported in TP-126, strictly as codewatch had it: churn over rolling windows, first-seen dates,
237
+ ownership and bus factor, and change coupling. The engine lives in `src/history/` and is
238
+ published on its own subpath. Its API is path-based: repo-relative posix paths in, plain
239
+ records out, no node ids or snapshots.
240
+
241
+ ```ts
242
+ import { computeChangeCoupling, couplingFor, loadChurnEntries } from "@titan-design/code-graph/history";
243
+
244
+ const entries = loadChurnEntries({ repoRoot: ".", windowDays: 90 }) ?? []; // null outside git
245
+ const { pairs, skippedLargeCommits } = computeChangeCoupling(entries);
246
+ couplingFor(pairs, "packages/code-graph/src/indexer.ts"); // partners by co-edit count
247
+ ```
248
+
249
+ `indexPaths` writes history metrics on file nodes by default, through the
250
+ `history-metrics.ts` adapter, with codewatch's names: `churn_{w}`, `churn_{w}_commits`,
251
+ `churn_{w}_authors` and `recency_{w}` for each window, `file_age_days`, and `bus_factor_{w}`
252
+ plus `top_author_share_{w}` for the primary window. Windows default to 30, 90 and 180 days
253
+ plus `churnWindowDays` (the primary, default 30); `churnWindows` replaces the defaults and
254
+ `lifetime: true` adds an all-history window with its own ownership. `computeChurn: false`
255
+ turns all of it off. Outside git, or without a git binary, the index simply has no history
256
+ metrics.
257
+
258
+ Change coupling is not stored. It is computed on demand from `loadChurnEntries`, as
259
+ codewatch's `graph coupled` command did.
260
+
261
+ **The seam.** Nothing under `src/history/` imports the rest of code-graph; the rest may import
262
+ it. The `code-graph-history-seam` rule in `.codewatch/check.json` fails `pnpm dag:check` on
263
+ any such import. That keeps a later extraction into `@titan-design/git-history` a directory
264
+ move plus an import-path change.
265
+
266
+ **Behaviour kept from codewatch, gaps included.** A rename is followed only inside the commit
267
+ that made it, so a file's churn before the rename stays on its old path and is dropped.
268
+ First-seen dates come from a separate `--no-renames` pass, which makes a renamed file look
269
+ younger. `--since` resolves against git's clock while window slicing uses `nowEpoch`. History
270
+ metrics are recomputed on every index and never carried forward under reuse.
271
+
272
+ ## Test linking and coverage
273
+
274
+ Ported in TP-133, strictly as codewatch had it. `linkTestsToSources` pairs each test file
275
+ with non-test files in two passes. Pass 1 uses path conventions: it strips a `.test` or `.spec`
276
+ infix and collapses a `__tests__/`, `test/` or `tests/` segment. Pass 2 gives a test that
277
+ pass 1 left unpaired its strongest co-edited non-test partner, with at least 2 shared commits.
278
+
279
+ `indexPaths` writes `linked_test_count` on each linked source, with or without git. With git
280
+ history on it also writes `test_bus_factor_{w}` and `test_top_author_share_{w}` for the
281
+ primary window. These summarize churn authorship across all tests linked to a source, so a
282
+ file can be well spread in production code and a single-author silo in its tests. All three
283
+ are recomputed on every index and never carried forward.
284
+
285
+ `computeTestCoverageOwnership` lives in the `history-metrics.ts` adapter, not in
286
+ `src/history/`. It needs test links, and the seam forbids history from importing them.
287
+
288
+ ```ts
289
+ import { attributeCoverage } from "@titan-design/code-graph";
290
+
291
+ // fileIdOf maps an absolute path to a file id, or null to skip; spans come from symbol nodes' attrs.
292
+ const metrics = attributeCoverage(istanbulReport, fileIdOf, symbolSpansByFile);
293
+ ```
294
+
295
+ `attributeCoverage` turns an Istanbul `coverage-final.json` into `coverage_pct` metrics: one
296
+ per file (covered functions over total functions) and one per symbol, matched by line-range
297
+ containment to the innermost symbol. Coverage depends on which tests ran, not on file bytes,
298
+ so the index never writes or carries it. The caller stores it on the snapshot it measured,
299
+ as codewatch's `graph coverage` command does.
@@ -0,0 +1,73 @@
1
+ /**
2
+ * A churn window is either a rolling N-day window or `"lifetime"` = all of git
3
+ * history (no `--since` bound). Lifetime mode lets an unfamiliar repo be audited
4
+ * cold, where any recent rolling slice is thin relative to the repo's whole
5
+ * life, so a windowed view reads an established repo as nearly churn-free.
6
+ */
7
+ type ChurnWindow = number | "lifetime";
8
+
9
+ interface ChurnEntry {
10
+ commit: string;
11
+ /** Author identity: git author email (%ae), stable across name spelling drift. */
12
+ author: string;
13
+ /** Committer time (%ct), epoch seconds; lets one wide log be sliced per window. */
14
+ epoch: number;
15
+ /** Repo-relative posix path, rebased onto the caller's root. */
16
+ filePath: string;
17
+ added: number;
18
+ deleted: number;
19
+ }
20
+ interface LoadChurnOptions {
21
+ /** Directory paths are made relative to; entries outside it are dropped. */
22
+ repoRoot: string;
23
+ windowDays?: ChurnWindow;
24
+ }
25
+ /**
26
+ * Parse the last `windowDays` of git history into ChurnEntry[] rebased onto
27
+ * `repoRoot`. Returns null if git isn't available; [] if no commits matched.
28
+ * Used both for churn metrics and for change-coupling.
29
+ */
30
+ declare function loadChurnEntries(options: LoadChurnOptions): ChurnEntry[] | null;
31
+ declare function parseChurnLog(text: string): ChurnEntry[];
32
+ declare function resolveRenamedPath(rawPath: string): string;
33
+ /** Entries whose commit landed within the last `windowDays` of `nowEpoch`. */
34
+ declare function entriesWithin(entries: readonly ChurnEntry[], windowDays: number, nowEpoch: number): ChurnEntry[];
35
+
36
+ interface CoEditPair {
37
+ fileA: string;
38
+ fileB: string;
39
+ /** Commits in which both files changed. */
40
+ count: number;
41
+ /** Most recent commits (up to maxCommitsPerPair). */
42
+ commits: string[];
43
+ }
44
+ interface ComputeChangeCouplingOptions {
45
+ /** Skip pairs that co-occur in fewer than this many commits. Default 2. */
46
+ minCount?: number;
47
+ /** Truncate the commits sample per pair. Default 10. */
48
+ maxCommitsPerPair?: number;
49
+ /**
50
+ * Skip commits touching more than this many files (sweeping refactors, mass
51
+ * renames). Default 50: beyond that the O(n²) pair explosion is signal-poor.
52
+ */
53
+ largeCommitThreshold?: number;
54
+ /** Filter file paths to this set before pairing. */
55
+ knownPaths?: ReadonlySet<string>;
56
+ }
57
+ interface ChangeCouplingResult {
58
+ pairs: CoEditPair[];
59
+ /** Commits dropped by largeCommitThreshold, so callers can say when sweeps suppressed the signal. */
60
+ skippedLargeCommits: number;
61
+ /** Threshold that was applied (echoed for diagnostic messages). */
62
+ largeCommitThreshold: number;
63
+ }
64
+ interface CouplingPartner {
65
+ partner: string;
66
+ count: number;
67
+ commits: string[];
68
+ }
69
+ declare function computeChangeCoupling(entries: readonly ChurnEntry[], options?: ComputeChangeCouplingOptions): ChangeCouplingResult;
70
+ /** Files most coupled to `seed`, by co-edit count desc then path; the seed itself is excluded. */
71
+ declare function couplingFor(pairs: readonly CoEditPair[], seed: string): CouplingPartner[];
72
+
73
+ export { type ChurnEntry as C, type LoadChurnOptions as L, type ChurnWindow as a, type ChangeCouplingResult as b, type CoEditPair as c, type ComputeChangeCouplingOptions as d, type CouplingPartner as e, computeChangeCoupling as f, couplingFor as g, entriesWithin as h, loadChurnEntries as l, parseChurnLog as p, resolveRenamedPath as r };
@@ -0,0 +1,332 @@
1
+ // src/history/log.ts
2
+ import { realpathSync } from "fs";
3
+ import * as path from "path";
4
+
5
+ // src/history/git.ts
6
+ import { execFileSync } from "child_process";
7
+ function discoveryEnv() {
8
+ const env = { ...process.env };
9
+ delete env.GIT_DIR;
10
+ delete env.GIT_INDEX_FILE;
11
+ delete env.GIT_WORK_TREE;
12
+ return env;
13
+ }
14
+ function runGit(cwd, args) {
15
+ try {
16
+ return execFileSync("git", [...args], {
17
+ cwd,
18
+ encoding: "utf-8",
19
+ stdio: ["ignore", "pipe", "ignore"],
20
+ env: discoveryEnv()
21
+ }).trim();
22
+ } catch {
23
+ return null;
24
+ }
25
+ }
26
+ function detectGitToplevel(cwd) {
27
+ return runGit(cwd, ["rev-parse", "--show-toplevel"]);
28
+ }
29
+ function runGitLarge(cwd, args, maxBuffer) {
30
+ try {
31
+ return execFileSync("git", [...args], {
32
+ cwd,
33
+ encoding: "utf-8",
34
+ maxBuffer,
35
+ stdio: ["ignore", "pipe", "ignore"],
36
+ env: discoveryEnv()
37
+ });
38
+ } catch {
39
+ return null;
40
+ }
41
+ }
42
+
43
+ // src/history/window.ts
44
+ var DAY_SECONDS = 86400;
45
+ function windowCutoff(windowDays, nowEpoch) {
46
+ return nowEpoch - windowDays * DAY_SECONDS;
47
+ }
48
+
49
+ // src/history/log.ts
50
+ var DEFAULT_WINDOW_DAYS = 30;
51
+ var COMMIT_HASH_RE = /^[0-9a-f]{7,40}$/;
52
+ var NUMSTAT_FIRST_RE = /^(\d+|-)$/;
53
+ function loadChurnEntries(options) {
54
+ const windowDays = options.windowDays ?? DEFAULT_WINDOW_DAYS;
55
+ const gitRoot = detectGitToplevel(options.repoRoot);
56
+ if (gitRoot === null) return null;
57
+ const log = runChurnLog(options.repoRoot, windowDays);
58
+ if (log === null) return null;
59
+ const canonicalRoot = canonicalize(options.repoRoot);
60
+ return parseChurnLog(log).flatMap((entry) => {
61
+ const rel = rebasePath(entry.filePath, gitRoot, canonicalRoot);
62
+ return rel === null ? [] : [{ ...entry, filePath: rel }];
63
+ });
64
+ }
65
+ function rebasePath(gitPath, gitRoot, rootDir) {
66
+ const rel = path.relative(rootDir, path.resolve(gitRoot, gitPath));
67
+ if (rel.startsWith("..") || path.isAbsolute(rel)) return null;
68
+ return rel.split(path.sep).join("/");
69
+ }
70
+ function canonicalize(p) {
71
+ try {
72
+ return realpathSync(p);
73
+ } catch {
74
+ return path.resolve(p);
75
+ }
76
+ }
77
+ function parseChurnLog(text) {
78
+ const out = [];
79
+ let header = null;
80
+ for (const rawLine of text.split("\n")) {
81
+ const line = rawLine.trimEnd();
82
+ if (!line) continue;
83
+ const parts = line.split(" ");
84
+ if (parts.length === 3 && header && NUMSTAT_FIRST_RE.test(parts[0]) && NUMSTAT_FIRST_RE.test(parts[1])) {
85
+ out.push({ ...header, ...parseNumstat(parts), filePath: resolveRenamedPath(parts[2]) });
86
+ continue;
87
+ }
88
+ if (parts.length >= 2 && COMMIT_HASH_RE.test(parts[0])) {
89
+ header = { commit: parts[0], author: parts[1], epoch: parts.length >= 3 ? Number(parts[2]) : 0 };
90
+ }
91
+ }
92
+ return out;
93
+ }
94
+ function parseNumstat(parts) {
95
+ return {
96
+ added: parts[0] === "-" ? 0 : Number(parts[0]),
97
+ deleted: parts[1] === "-" ? 0 : Number(parts[1])
98
+ };
99
+ }
100
+ function resolveRenamedPath(rawPath) {
101
+ const braceMatch = /^(.*)\{(.*) => (.*)\}(.*)$/.exec(rawPath);
102
+ if (braceMatch) {
103
+ const [, prefix, , newSeg, suffix] = braceMatch;
104
+ return `${prefix}${newSeg}${suffix}`.replace(/\/+/g, "/");
105
+ }
106
+ const arrow = rawPath.indexOf(" => ");
107
+ if (arrow >= 0) return rawPath.slice(arrow + 4);
108
+ return rawPath;
109
+ }
110
+ function entriesWithin(entries, windowDays, nowEpoch) {
111
+ const cutoff = windowCutoff(windowDays, nowEpoch);
112
+ return entries.filter((e) => e.epoch >= cutoff);
113
+ }
114
+ function runChurnLog(repoRoot, windowDays) {
115
+ const sinceArgs = windowDays === "lifetime" ? [] : [`--since=${windowDays}.days.ago`];
116
+ const format = "--pretty=format:%H%x09%ae%x09%ct";
117
+ return runGitLarge(repoRoot, ["log", ...sinceArgs, "--no-merges", "--numstat", "-M", format], 64 * 1024 * 1024);
118
+ }
119
+
120
+ // src/history/first-seen.ts
121
+ function loadFileFirstSeen(options) {
122
+ const gitRoot = detectGitToplevel(options.repoRoot);
123
+ if (gitRoot === null) return null;
124
+ const log = runFirstSeenLog(options.repoRoot);
125
+ if (log === null) return null;
126
+ const canonicalRoot = canonicalize(options.repoRoot);
127
+ const firstSeen = /* @__PURE__ */ new Map();
128
+ for (const { epoch, gitPath } of parseFirstSeenLog(log)) {
129
+ const rel = rebasePath(resolveRenamedPath(gitPath), gitRoot, canonicalRoot);
130
+ if (rel === null) continue;
131
+ if (options.knownPaths && !options.knownPaths.has(rel)) continue;
132
+ if (!firstSeen.has(rel)) firstSeen.set(rel, epoch);
133
+ }
134
+ return firstSeen;
135
+ }
136
+ function parseFirstSeenLog(text) {
137
+ const out = [];
138
+ let epoch = 0;
139
+ for (const rawLine of text.split("\n")) {
140
+ const line = rawLine.trimEnd();
141
+ if (!line) continue;
142
+ const parts = line.split(" ");
143
+ if (parts.length === 2 && COMMIT_HASH_RE.test(parts[0])) {
144
+ epoch = Number(parts[1]);
145
+ continue;
146
+ }
147
+ if (epoch) out.push({ epoch, gitPath: line });
148
+ }
149
+ return out;
150
+ }
151
+ function runFirstSeenLog(repoRoot) {
152
+ const args = ["log", "--reverse", "--no-merges", "--no-renames", "--name-only", "--pretty=format:%H%x09%ct"];
153
+ return runGitLarge(repoRoot, args, 128 * 1024 * 1024);
154
+ }
155
+
156
+ // src/history/churn.ts
157
+ function aggregateChurn(entries, knownPaths) {
158
+ const lines = /* @__PURE__ */ new Map();
159
+ const commits = /* @__PURE__ */ new Map();
160
+ const authors = /* @__PURE__ */ new Map();
161
+ for (const e of entries) {
162
+ if (knownPaths && !knownPaths.has(e.filePath)) continue;
163
+ lines.set(e.filePath, (lines.get(e.filePath) ?? 0) + e.added + e.deleted);
164
+ setAdd(commits, e.filePath, e.commit);
165
+ setAdd(authors, e.filePath, e.author);
166
+ }
167
+ const out = /* @__PURE__ */ new Map();
168
+ for (const [filePath, total] of lines) {
169
+ out.set(filePath, {
170
+ lines: total,
171
+ commits: commits.get(filePath).size,
172
+ authors: authors.get(filePath).size
173
+ });
174
+ }
175
+ return out;
176
+ }
177
+ function aggregateChurnWindows(entries, windowsDays, nowEpoch, knownPaths) {
178
+ const out = /* @__PURE__ */ new Map();
179
+ for (const windowDays of windowsDays) {
180
+ const within = windowDays === "lifetime" ? entries : entriesWithin(entries, windowDays, nowEpoch);
181
+ out.set(windowDays, aggregateChurn(within, knownPaths));
182
+ }
183
+ return out;
184
+ }
185
+ function setAdd(map, key, value) {
186
+ let set = map.get(key);
187
+ if (!set) {
188
+ set = /* @__PURE__ */ new Set();
189
+ map.set(key, set);
190
+ }
191
+ set.add(value);
192
+ }
193
+
194
+ // src/history/ownership.ts
195
+ var DEFAULT_BUS_FACTOR_THRESHOLD = 0.5;
196
+ function computeOwnership(entries, options = {}) {
197
+ const threshold = options.busFactorThreshold ?? DEFAULT_BUS_FACTOR_THRESHOLD;
198
+ const out = /* @__PURE__ */ new Map();
199
+ for (const [filePath, byAuthor] of authorLinesByPath(entries, options.knownPaths)) {
200
+ const summary = summarizeOwnership(byAuthor, threshold);
201
+ if (summary !== null) out.set(filePath, summary);
202
+ }
203
+ return out;
204
+ }
205
+ function authorLinesByPath(entries, known) {
206
+ const out = /* @__PURE__ */ new Map();
207
+ for (const e of entries) {
208
+ if (known && !known.has(e.filePath)) continue;
209
+ const lines = e.added + e.deleted;
210
+ if (lines === 0) continue;
211
+ let byAuthor = out.get(e.filePath);
212
+ if (!byAuthor) {
213
+ byAuthor = /* @__PURE__ */ new Map();
214
+ out.set(e.filePath, byAuthor);
215
+ }
216
+ byAuthor.set(e.author, (byAuthor.get(e.author) ?? 0) + lines);
217
+ }
218
+ return out;
219
+ }
220
+ function summarizeOwnership(byAuthor, threshold = DEFAULT_BUS_FACTOR_THRESHOLD) {
221
+ let total = 0;
222
+ for (const v of byAuthor.values()) total += v;
223
+ if (total === 0) return null;
224
+ const sorted = [...byAuthor.values()].sort((a, b) => b - a);
225
+ const topAuthorShare = sorted[0] / total;
226
+ const busFactor = minAuthorsToReach(sorted, total, threshold);
227
+ return { authors: sorted.length, topAuthorShare, busFactor };
228
+ }
229
+ function minAuthorsToReach(sortedDesc, total, threshold) {
230
+ let acc = 0;
231
+ for (let i = 0; i < sortedDesc.length; i++) {
232
+ acc += sortedDesc[i];
233
+ if (acc / total >= threshold) return i + 1;
234
+ }
235
+ return sortedDesc.length;
236
+ }
237
+
238
+ // src/history/change-coupling.ts
239
+ var DEFAULT_MIN_COUNT = 2;
240
+ var DEFAULT_MAX_COMMITS_PER_PAIR = 10;
241
+ var DEFAULT_LARGE_COMMIT_THRESHOLD = 50;
242
+ function computeChangeCoupling(entries, options = {}) {
243
+ const maxCommits = options.maxCommitsPerPair ?? DEFAULT_MAX_COMMITS_PER_PAIR;
244
+ const largeThreshold = options.largeCommitThreshold ?? DEFAULT_LARGE_COMMIT_THRESHOLD;
245
+ const pairs = /* @__PURE__ */ new Map();
246
+ let skippedLargeCommits = 0;
247
+ for (const [commit, files] of groupByCommit(entries, options.knownPaths)) {
248
+ if (files.length < 2) continue;
249
+ if (files.length > largeThreshold) {
250
+ skippedLargeCommits++;
251
+ continue;
252
+ }
253
+ accumulatePairs(commit, files, pairs, maxCommits);
254
+ }
255
+ return {
256
+ pairs: finalize(pairs, options.minCount ?? DEFAULT_MIN_COUNT),
257
+ skippedLargeCommits,
258
+ largeCommitThreshold: largeThreshold
259
+ };
260
+ }
261
+ function groupByCommit(entries, known) {
262
+ const out = /* @__PURE__ */ new Map();
263
+ for (const e of entries) {
264
+ if (known && !known.has(e.filePath)) continue;
265
+ let files = out.get(e.commit);
266
+ if (!files) {
267
+ files = [];
268
+ out.set(e.commit, files);
269
+ }
270
+ if (!files.includes(e.filePath)) files.push(e.filePath);
271
+ }
272
+ return out;
273
+ }
274
+ function accumulatePairs(commit, files, pairs, maxCommits) {
275
+ const sorted = [...files].sort();
276
+ for (let i = 0; i < sorted.length; i++) {
277
+ for (let j = i + 1; j < sorted.length; j++) {
278
+ const key = `${sorted[i]} ${sorted[j]}`;
279
+ let acc = pairs.get(key);
280
+ if (!acc) {
281
+ acc = { count: 0, commits: [] };
282
+ pairs.set(key, acc);
283
+ }
284
+ acc.count++;
285
+ if (acc.commits.length < maxCommits) acc.commits.push(commit);
286
+ }
287
+ }
288
+ }
289
+ function finalize(pairs, minCount) {
290
+ const out = [];
291
+ for (const [key, acc] of pairs) {
292
+ if (acc.count < minCount) continue;
293
+ const [a, b] = key.split(" ");
294
+ out.push({ fileA: a, fileB: b, count: acc.count, commits: acc.commits });
295
+ }
296
+ return out.sort(comparePairs);
297
+ }
298
+ function comparePairs(a, b) {
299
+ if (b.count !== a.count) return b.count - a.count;
300
+ if (a.fileA !== b.fileA) return a.fileA < b.fileA ? -1 : 1;
301
+ return a.fileB < b.fileB ? -1 : a.fileB > b.fileB ? 1 : 0;
302
+ }
303
+ function couplingFor(pairs, seed) {
304
+ const out = [];
305
+ for (const p of pairs) {
306
+ if (p.fileA === seed) out.push({ partner: p.fileB, count: p.count, commits: p.commits });
307
+ else if (p.fileB === seed) out.push({ partner: p.fileA, count: p.count, commits: p.commits });
308
+ }
309
+ out.sort((a, b) => {
310
+ if (b.count !== a.count) return b.count - a.count;
311
+ return a.partner < b.partner ? -1 : a.partner > b.partner ? 1 : 0;
312
+ });
313
+ return out;
314
+ }
315
+
316
+ export {
317
+ runGit,
318
+ detectGitToplevel,
319
+ loadChurnEntries,
320
+ parseChurnLog,
321
+ resolveRenamedPath,
322
+ entriesWithin,
323
+ loadFileFirstSeen,
324
+ aggregateChurn,
325
+ aggregateChurnWindows,
326
+ computeOwnership,
327
+ authorLinesByPath,
328
+ summarizeOwnership,
329
+ computeChangeCoupling,
330
+ couplingFor
331
+ };
332
+ //# sourceMappingURL=chunk-PFI5XMUG.js.map