delivery-friction-analyzer 0.16.2 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +17 -17
- package/docs/contracts/friction-report.md +10 -7
- package/docs/contracts/source-bundle.md +85 -15
- package/docs/contracts/target-repository.md +3 -3
- package/docs/reference/github-data-inventory.md +2 -2
- package/docs/reference/release-automation.md +1 -1
- package/package.json +1 -1
- package/release-log.md +16 -0
- package/schemas/{github-source-bundle.schema.json → source-bundle.schema.json} +31 -17
- package/src/cli/analyze-github.js +11 -4
- package/src/collect/github-source-bundle.js +9 -5
- package/src/contracts/target-repository.js +1 -1
- package/src/report/evidence-artifacts.js +12 -5
- package/src/report/friction-report.js +5 -2
package/README.md
CHANGED
|
@@ -17,40 +17,40 @@ The analyzer runs locally with your GitHub credentials. Generated artifacts pres
|
|
|
17
17
|
- GitHub CLI (`gh`) installed and authenticated with access to the target repository.
|
|
18
18
|
- A repository profile JSON for the repository you want to analyze. Interactive setup can create a starter profile for you.
|
|
19
19
|
|
|
20
|
-
For public repositories, ordinary read access is usually enough. Private repositories need a `gh` token with enough read access for the requested
|
|
20
|
+
For public repositories, ordinary read access is usually enough. Private repositories need a `gh` token with enough read access for the requested source families. With a classic PAT, that usually means the `repo` scope. With a fine-grained token or GitHub App, grant read permissions for repository metadata and contents, pull requests, Actions, and checks where available. Missing or partial source coverage is recorded in the generated methodology and coverage artifacts instead of being treated as complete data.
|
|
21
21
|
|
|
22
22
|
## Quickstart
|
|
23
23
|
|
|
24
|
-
###
|
|
24
|
+
### Guided setup
|
|
25
25
|
|
|
26
|
-
From this repository, install dependencies and
|
|
26
|
+
From this repository, install dependencies and let interactive setup create or confirm the repository profile for a GitHub repository you want to measure:
|
|
27
27
|
|
|
28
28
|
```sh
|
|
29
29
|
npm install
|
|
30
30
|
npm run analyze:github -- \
|
|
31
|
-
--repo
|
|
31
|
+
--repo owner/name \
|
|
32
32
|
--limit 30 \
|
|
33
|
-
--profile
|
|
34
|
-
--out reports/
|
|
33
|
+
--profile profiles/owner-name.json \
|
|
34
|
+
--out reports/owner-name \
|
|
35
|
+
--interactive \
|
|
36
|
+
--dry-run
|
|
35
37
|
```
|
|
36
38
|
|
|
37
|
-
|
|
39
|
+
If the profile path does not exist, interactive setup can create a minimal `repository-profile.v1` profile. `--dry-run` validates repository access, profile JSON, output directory writability, and a small sample of GitHub API coverage without writing the full report bundle. When the profile looks right, rerun the command without `--dry-run`.
|
|
38
40
|
|
|
39
|
-
###
|
|
41
|
+
### Run the analysis
|
|
40
42
|
|
|
41
|
-
|
|
43
|
+
After you have a repository profile, run the analyzer against the target repository:
|
|
42
44
|
|
|
43
45
|
```sh
|
|
44
46
|
npm run analyze:github -- \
|
|
45
|
-
--interactive \
|
|
46
47
|
--repo owner/name \
|
|
47
48
|
--limit 30 \
|
|
48
49
|
--profile profiles/owner-name.json \
|
|
49
|
-
--out reports/owner-name
|
|
50
|
-
--dry-run
|
|
50
|
+
--out reports/owner-name
|
|
51
51
|
```
|
|
52
52
|
|
|
53
|
-
|
|
53
|
+
Open `reports/owner-name/friction-report.md` first. It is the main human-readable report. Use the JSON and CSV files when you want to audit a finding, compare PRs, or build follow-up analysis.
|
|
54
54
|
|
|
55
55
|
To run the CLI from another project with the npm package, pass the same choices as explicit flags:
|
|
56
56
|
|
|
@@ -78,7 +78,7 @@ Profiles can define:
|
|
|
78
78
|
|
|
79
79
|
For a new repository, the easiest path is the guided `--interactive --dry-run` command in Quickstart. If the profile path does not exist, interactive setup asks whether to create it, then writes a minimal `repository-profile.v1` profile with user-provided workflow context and optional release PR title rules. It may create the output directory and briefly write then remove a temporary probe file to confirm writability. When interactive setup saves or generates a profile during a dry run, the completion output prints the saved profile path so you can inspect and edit it before a full run.
|
|
80
80
|
|
|
81
|
-
Use `
|
|
81
|
+
Use `docs/reference/repository-profile.md` and `schemas/repository-profile.schema.json` when you prefer to create a profile by hand. Existing profiles and fixtures in this repository are internal validation examples; copy them only if their repository-specific assumptions match your target.
|
|
82
82
|
|
|
83
83
|
## Outputs
|
|
84
84
|
|
|
@@ -94,7 +94,7 @@ Use these when you want to audit, automate, or build follow-up analysis:
|
|
|
94
94
|
- `normalized.json`: normalized repository, PR, file, review, and validation entities.
|
|
95
95
|
- `source-bundle.json`: collected source data for auditability. Its canonical
|
|
96
96
|
analyzer contract is documented in `docs/contracts/source-bundle.md` and
|
|
97
|
-
checked by `schemas/
|
|
97
|
+
checked by `schemas/source-bundle.schema.json`; it is not a full GitHub
|
|
98
98
|
API payload schema.
|
|
99
99
|
|
|
100
100
|
When CSV exports are enabled, the bundle also includes spreadsheet-friendly evidence files:
|
|
@@ -102,7 +102,7 @@ When CSV exports are enabled, the bundle also includes spreadsheet-friendly evid
|
|
|
102
102
|
- `pr-metrics.csv`: per-PR metrics for spreadsheet review.
|
|
103
103
|
- `bottleneck-examples.csv`: representative bottleneck examples.
|
|
104
104
|
- `comment-sources.csv`: review-comment source breakdowns.
|
|
105
|
-
- `collection-coverage.csv`:
|
|
105
|
+
- `collection-coverage.csv`: source-family coverage diagnostics.
|
|
106
106
|
|
|
107
107
|
Each ranked bottleneck example includes source references, workflow-run conclusions, review-thread source information, comment-source breakdowns, and a dominance note when one PR contributes most of the displayed signal.
|
|
108
108
|
|
|
@@ -163,7 +163,7 @@ The current product focus is a maintainer workflow:
|
|
|
163
163
|
|
|
164
164
|
The product should eventually combine GitHub delivery friction with token and model usage, but GitHub-only analytics remain the active validation surface.
|
|
165
165
|
|
|
166
|
-
`hannasdev/mcp-writing` remains
|
|
166
|
+
`hannasdev/mcp-writing` remains an internal validation target and fixture source, not a public tutorial path or product-specific scope.
|
|
167
167
|
|
|
168
168
|
The existing metrics-summary-only report command remains available for fixture and advanced workflows:
|
|
169
169
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Friction Report Contract
|
|
2
2
|
|
|
3
|
-
Milestone 3 introduced `friction-report.v1`, a deterministic report generated from a `friction-metrics.v1` repository metrics summary. Milestone 4 added configured workflow context. Milestone 5 adds sanitized contributor-source metadata for configured `.all-contributorsrc` coverage without raw contributor file contents or individual rankings. The report layer does not fetch
|
|
3
|
+
Milestone 3 introduced `friction-report.v1`, a deterministic report generated from a `friction-metrics.v1` repository metrics summary. Milestone 4 added configured workflow context. Milestone 5 adds sanitized contributor-source metadata for configured `.all-contributorsrc` coverage without raw contributor file contents or individual rankings. The report layer does not fetch source data, mutate repositories, rank individuals, or depend on services beyond the source collection path that produced the metrics summary.
|
|
4
4
|
|
|
5
5
|
## Outputs
|
|
6
6
|
|
|
@@ -19,16 +19,19 @@ Local artifact generation is available from an existing metrics summary:
|
|
|
19
19
|
node src/report/generate-report.js --metrics-summary fixtures/github/mcp-writing/metrics-summary.golden.json --json-out fixtures/github/mcp-writing/reports/friction-report.golden.json --markdown-out fixtures/github/mcp-writing/reports/friction-report.golden.md
|
|
20
20
|
```
|
|
21
21
|
|
|
22
|
-
The command reads local `friction-metrics.v1` JSON and writes deterministic `friction-report.v1` JSON and Markdown files. It does not fetch
|
|
22
|
+
The command reads local `friction-metrics.v1` JSON and writes deterministic `friction-report.v1` JSON and Markdown files. It does not fetch source data, mutate the analyzed repository, or write CSV/methodology companion artifacts.
|
|
23
23
|
|
|
24
24
|
## Report Shape
|
|
25
25
|
|
|
26
26
|
- `reportVersion`: report contract version.
|
|
27
27
|
- `metricVersion`: source metrics contract version.
|
|
28
|
+
- `source`: optional source-bundle provenance copied into live-generated report
|
|
29
|
+
artifacts, including `kind` and a human-readable label.
|
|
28
30
|
- `targetRepository`: analyzed repository identity; live analysis sample size is encoded as `targetRepository.analysisPullRequestLimit` from collection metadata.
|
|
29
31
|
- `analysisFilter`: optional metadata for explicit filters applied before metrics computation, including excluded PR classes and before/after PR counts.
|
|
30
32
|
- `configuredWorkflow`: optional user-configured workflow context from the repository profile. It is not observed GitHub evidence and does not change scoring, ranking, CSV exports, or PR class matching.
|
|
31
33
|
- `contributorSource`: optional sanitized contributor-source metadata from the repository profile and collection path. It records source type, path, coverage status, parsed hint count, and guardrail note. It does not include raw contributor file contents or contributor rankings.
|
|
34
|
+
- `collectionCoverage`: optional source collection coverage metadata copied into live-generated report JSON artifacts. `collectionCoverage.status` records aggregate coverage, and `collectionCoverage.sourceFamilies` lists source-family entries with `family`, `status`, `attempts`, `source`, `diagnostics`, and `downstreamImpact`. This uses the generic `source-bundle.v1` `coverage.sourceFamilies` name; legacy GitHub `apiFamilies` is not a live `friction-report.v1` field.
|
|
32
35
|
- `summary`: repository totals and top bottleneck identifiers.
|
|
33
36
|
- `coverage`: PR-open diff, workflow-run, and review-thread coverage counts plus caveats.
|
|
34
37
|
- `commentSources`: total and source-grouped review comments for Copilot, human, bot, scanner, author replies, and unknown sources.
|
|
@@ -99,7 +102,7 @@ The M3 report contract supports these recommendation categories:
|
|
|
99
102
|
|
|
100
103
|
## Coverage And Confidence
|
|
101
104
|
|
|
102
|
-
Reports must label unavailable or partial
|
|
105
|
+
Reports must label unavailable or partial source evidence instead of inferring unavailable values from merge-time data. Final/current PR metadata can come from source-bundle PR evidence, but PR-open diff growth remains unavailable unless an open-time snapshot or equivalent captured state exists. Workflow coverage and review-thread sources are summarized separately.
|
|
103
106
|
|
|
104
107
|
Representative examples should carry enough source evidence to trace a report claim back to generated artifacts. Validation examples should name the workflow-run source and conclusions. Review churn examples should name the review-thread source, review decision evidence, and comment sources. PR class evidence should be visible in representative bottleneck examples so readers can distinguish workflow populations such as release, dependency, development, or repository-specific classes. When `reviewThreads` is zero, review decision evidence should make clean human approval distinguishable from unavailable review evidence and from observed absence of human review. When displayed examples are dominated by one PR or one PR class, the report should say so instead of implying a repository-wide pattern from an outlier or workflow population.
|
|
105
108
|
|
|
@@ -119,7 +122,7 @@ Full live analysis writes `methodology.md` as a hybrid artifact: stable explanat
|
|
|
119
122
|
- contributor-source context when configured, including source type, path, coverage status, and parsed hint count, without raw contributor contents or rankings;
|
|
120
123
|
- profile suggestions when PR class, file/path, or workflow-context profile evidence crosses deterministic fallback thresholds, or an explicit no-threshold note when none were triggered;
|
|
121
124
|
- requested and collected PR counts;
|
|
122
|
-
- collection coverage status and
|
|
125
|
+
- collection coverage status and source-family diagnostics from `collectionCoverage.sourceFamilies`;
|
|
123
126
|
- scoring, ranking, dominance, sensitivity, and limitation explanations;
|
|
124
127
|
- generated artifact names and artifact-sensitivity guidance.
|
|
125
128
|
|
|
@@ -138,11 +141,11 @@ Minimum CSV column groups:
|
|
|
138
141
|
- `pr-metrics.csv`: PR number, title, URL, PR class, PR class source/rule evidence, changed lines, non-generated changed lines, review comments, review threads, review decision, human reviewer count, human approval / changes-requested booleans, failed checks, failed workflow runs, cancelled workflow runs, post-review commits, review-thread source, workflow-run source/coverage, and main ranking scores.
|
|
139
142
|
- `bottleneck-examples.csv`: bottleneck identity, recommendation category, PR identity, score/value, changed lines, validation counts, review counts, comment-source counts, workflow/review source and coverage labels, dominance, and source labels.
|
|
140
143
|
- `comment-sources.csv`: source name, total comments, bot/scanner classification, human/author classification, and share of all comments.
|
|
141
|
-
- `collection-coverage.csv`:
|
|
144
|
+
- `collection-coverage.csv`: source family, status, attempts, source label, diagnostics, and downstream impact.
|
|
142
145
|
|
|
143
146
|
Empty CSV cells mean unavailable or not applicable. Numeric zero should be used only for observed or computed zero counts. Count columns that depend on optional GitHub coverage should keep source or coverage labels nearby so spreadsheet readers can tell unavailable evidence apart from observed zeroes. CSVs must not include raw comment bodies, raw workflow logs, tokens, secret-bearing environment details, or individual contributor/reviewer rankings.
|
|
144
147
|
|
|
145
|
-
Contributor-source coverage appears in `collection-coverage.csv` as the `contributor_source`
|
|
148
|
+
Contributor-source coverage appears in `collection-coverage.csv` as the `contributor_source` source family when configured. CSVs may include aggregate comment-source counts influenced by contributor hints, but they must not include raw `.all-contributorsrc` contents, contributor names, contributor login lists, or person rankings.
|
|
146
149
|
|
|
147
150
|
## Optional Downstream Narrative Drafting
|
|
148
151
|
|
|
@@ -153,7 +156,7 @@ Use `friction-report.json` as the structured source of truth for report identity
|
|
|
153
156
|
- `pr-metrics.csv` for analyzed PR rows, class labels, review/validation counts, and ranking scores.
|
|
154
157
|
- `bottleneck-examples.csv` for representative examples tied to bottleneck identity, recommendation category, evidence sources, dominance, and score/value.
|
|
155
158
|
- `comment-sources.csv` for source-grouped comment totals and classification flags.
|
|
156
|
-
- `collection-coverage.csv` for
|
|
159
|
+
- `collection-coverage.csv` for source-family coverage, attempts, source labels, diagnostics, and downstream impact.
|
|
157
160
|
|
|
158
161
|
A guarded narrative-drafting workflow should:
|
|
159
162
|
|
|
@@ -1,15 +1,40 @@
|
|
|
1
1
|
# Source Bundle Contract
|
|
2
2
|
|
|
3
|
-
Schema: `schemas/
|
|
3
|
+
Schema: `schemas/source-bundle.schema.json`.
|
|
4
4
|
|
|
5
|
-
`source-bundle.json` is the analyzer's canonical
|
|
6
|
-
|
|
7
|
-
|
|
5
|
+
`source-bundle.json` is the analyzer's canonical source-evidence artifact. It is
|
|
6
|
+
the shared analyzer contract for live GitHub collection and bundled tutorial
|
|
7
|
+
sample data, not a raw provider dump. Both source kinds feed the same
|
|
8
|
+
normalization, metrics, report, methodology, and CSV pipeline, so they share one
|
|
9
|
+
contract with explicit provenance instead of pretending sample evidence is live
|
|
10
|
+
GitHub evidence.
|
|
8
11
|
|
|
9
|
-
The `
|
|
12
|
+
The current product contract is `source-bundle.v1`.
|
|
10
13
|
|
|
11
|
-
|
|
12
|
-
|
|
14
|
+
## Provenance
|
|
15
|
+
|
|
16
|
+
Every bundle records the source boundary:
|
|
17
|
+
|
|
18
|
+
- `schemaVersion`: must be `source-bundle.v1`.
|
|
19
|
+
- `source.kind`: `github` for live GitHub evidence or `sample` for bundled
|
|
20
|
+
synthetic sample evidence.
|
|
21
|
+
- `source.label`: human-readable label used by generated artifacts where source
|
|
22
|
+
provenance is shown.
|
|
23
|
+
- `collector`: collector identity and provider details. Collector names are not
|
|
24
|
+
constrained to a GitHub-only value.
|
|
25
|
+
|
|
26
|
+
Sample data must use labels that identify it as sample or synthetic evidence.
|
|
27
|
+
Live GitHub data should use labels such as `GitHub live collection` and
|
|
28
|
+
provider-specific source labels where the evidence came from GitHub APIs or
|
|
29
|
+
`gh`.
|
|
30
|
+
|
|
31
|
+
## Analyzer Evidence
|
|
32
|
+
|
|
33
|
+
The schema covers collector-owned fields:
|
|
34
|
+
|
|
35
|
+
- collection metadata, source provenance, target repository, repository
|
|
36
|
+
metadata, PR selection, source-family coverage diagnostics, and language
|
|
37
|
+
distribution context;
|
|
13
38
|
- optional sanitized contributor-source metadata: source type, repository path,
|
|
14
39
|
coverage/status diagnostics, and parsed hint count;
|
|
15
40
|
- pull request fields consumed downstream: identity, title, author object, URL,
|
|
@@ -25,11 +50,56 @@ than pretending zero runs were observed.
|
|
|
25
50
|
|
|
26
51
|
Contributor-source metadata must not persist raw contributor file contents or
|
|
27
52
|
parsed login lists. Parsed contributor hints may be used transiently during a
|
|
28
|
-
run, but generated artifacts keep only
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
53
|
+
run, but generated artifacts keep only sanitized metadata and `hintCount`.
|
|
54
|
+
|
|
55
|
+
## Coverage
|
|
56
|
+
|
|
57
|
+
Top-level `coverage.status` is retained as the aggregate source coverage status.
|
|
58
|
+
`coverage.sourceFamilies` lists the source-family entries that make up the
|
|
59
|
+
aggregate status. Each entry preserves the existing coverage-entry shape:
|
|
60
|
+
|
|
61
|
+
- `family`
|
|
62
|
+
- `source`
|
|
63
|
+
- `status`
|
|
64
|
+
- `attempts`
|
|
65
|
+
- `diagnostics`
|
|
66
|
+
- `downstreamImpact`
|
|
67
|
+
|
|
68
|
+
The `source` label is generic. Live entries may name GitHub REST, GraphQL, or
|
|
69
|
+
`gh` sources; sample entries may name bundled tutorial evidence. The per-PR
|
|
70
|
+
`pullRequests[*].coverage` object is retained in place with keys such as
|
|
71
|
+
`prOpenDiff`, `reviewThreads`, and `workflowRuns`.
|
|
72
|
+
|
|
73
|
+
## Field Migration From Legacy GitHub Bundles
|
|
74
|
+
|
|
75
|
+
`github-source-bundle.v1` is a legacy artifact shape. Current generated bundles
|
|
76
|
+
must use `source-bundle.v1`. There is no silent compatibility adapter in the
|
|
77
|
+
runtime; legacy artifacts need an explicit migration before they can validate
|
|
78
|
+
against the current schema.
|
|
79
|
+
|
|
80
|
+
| Legacy `github-source-bundle.v1` field | `source-bundle.v1` treatment |
|
|
81
|
+
| --- | --- |
|
|
82
|
+
| `schemaVersion` | Rename value to `source-bundle.v1`. |
|
|
83
|
+
| `collectedAt` | Retained as the materialization timestamp for live and sample bundles. |
|
|
84
|
+
| `collector` | Retained and generalized; `collector.name` is no longer constrained to a GitHub-only value. |
|
|
85
|
+
| `source` | New required object with `kind`, `label`, and optional metadata. |
|
|
86
|
+
| `targetRepository` | Retained. Sample bundles use fictional placeholder repository identity. |
|
|
87
|
+
| `repositoryMetadata` | Retained. Sample metadata must be fictional and clearly synthetic. |
|
|
88
|
+
| `selection` | Retained and generalized; selection strategies and source labels do not have to be GitHub-only. |
|
|
89
|
+
| `coverage.status` | Retained as the aggregate status over `coverage.sourceFamilies`. |
|
|
90
|
+
| `coverage.apiFamilies` | Renamed to `coverage.sourceFamilies`; entry shape is preserved. |
|
|
91
|
+
| `languageDistribution` | Retained with generalized source labels. |
|
|
92
|
+
| `contributorSource` | Retained as optional sanitized metadata with generalized source labels. |
|
|
93
|
+
| `pullRequests` | Retained as GitHub-shaped PR evidence consumed by normalization. |
|
|
94
|
+
| `pullRequests[*].coverage` | Retained as the per-PR coverage object with current per-family keys. |
|
|
95
|
+
| `raw` | Retained as the only subtree for provider- or sample-specific raw details. |
|
|
96
|
+
|
|
97
|
+
No analyzer-owned evidence field from `github-source-bundle.v1` is intentionally
|
|
98
|
+
removed in this migration.
|
|
99
|
+
|
|
100
|
+
The schema remains strict for analyzer-owned wrapper objects and mapped
|
|
101
|
+
entities. It does not attempt to schema every raw provider field, and normal
|
|
102
|
+
upstream GitHub API additions should not require schema updates unless the
|
|
103
|
+
collector maps them into canonical bundle fields. Future raw or
|
|
104
|
+
provider-specific payloads must live under an explicit `raw` subtree so
|
|
105
|
+
downstream normalization stays tied to canonical fields.
|
|
@@ -11,15 +11,15 @@ The local analyzer accepts a target repository and a pull request sample size. T
|
|
|
11
11
|
- `defaultBranch`: expected default branch for merge-base and branch lookup context.
|
|
12
12
|
- `visibility`: `public`, `private`, or `unknown`.
|
|
13
13
|
- `analysisPullRequestLimit`: latest merged pull request count from 1 to 100, supplied by the live CLI as `--limit`.
|
|
14
|
-
- `isValidationTarget`: optional flag for fixture-source repositories such as `hannasdev/mcp-writing`.
|
|
14
|
+
- `isValidationTarget`: optional metadata flag for internal validation or fixture-source repositories such as `hannasdev/mcp-writing`. It does not bypass target repository validation.
|
|
15
15
|
|
|
16
16
|
Schema: `schemas/target-repository.schema.json`. Live analysis selection is latest-N merged pull requests, not a rolling day window.
|
|
17
17
|
|
|
18
18
|
## Product Repository Separation
|
|
19
19
|
|
|
20
|
-
The validator rejects a target repository that exactly matches the configured product repository. For this repository, the product repository is `hannasdev/delivery-friction-analyzer`;
|
|
20
|
+
The validator rejects a target repository that exactly matches the configured product repository. For this repository, the product repository is `hannasdev/delivery-friction-analyzer`; `hannasdev/mcp-writing` is an internal validation target and fixture source.
|
|
21
21
|
|
|
22
|
-
Live GitHub analysis enforces this separation before GitHub collection starts. If `--repo` names this tool's product repository, the command fails before provider calls, tells you to choose the repository you want to measure with `--repo owner/name`, and confirms that no GitHub data was collected. The product repository identity is repo-local implementation configuration, not a public CLI option.
|
|
22
|
+
Live GitHub analysis enforces this separation before GitHub collection starts. If `--repo` names this tool's product repository, the command fails before provider calls, explains that the guard prevents accidental self-analysis during normal live runs rather than protecting already readable GitHub data, tells you to choose the repository you want to measure with `--repo owner/name`, and confirms that no GitHub data was collected. `--validation-target` only marks output metadata and does not bypass this guard. The product repository identity is repo-local implementation configuration, not a public CLI option.
|
|
23
23
|
|
|
24
24
|
## Degraded Behavior
|
|
25
25
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# GitHub Data Inventory
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Internal validation inventory was checked against `hannasdev/mcp-writing` on 2026-06-08 with `gh` authenticated as `hannasdev`. This repository is fixture and calibration context, not the public tutorial target.
|
|
4
4
|
|
|
5
5
|
## API Fields
|
|
6
6
|
|
|
@@ -54,6 +54,6 @@ Availability decision:
|
|
|
54
54
|
|
|
55
55
|
Fixtures store compact, redacted source-shaped data rather than full raw payloads.
|
|
56
56
|
Live `source-bundle.json` artifacts use the analyzer-owned
|
|
57
|
-
`
|
|
57
|
+
`source-bundle.v1` contract in `schemas/source-bundle.schema.json`;
|
|
58
58
|
that schema covers the canonical fields consumed downstream, not the full GitHub
|
|
59
59
|
REST or GraphQL API response shape.
|
|
@@ -9,7 +9,7 @@ The npm package allowlist includes:
|
|
|
9
9
|
- runtime source in `src`;
|
|
10
10
|
- JSON schemas in `schemas`;
|
|
11
11
|
- public contract and reference docs in `docs/contracts` and `docs/reference`;
|
|
12
|
-
- the
|
|
12
|
+
- the internal validation repository profile at `fixtures/github/mcp-writing/profile.json`;
|
|
13
13
|
- `LICENSE`;
|
|
14
14
|
- `README.md`;
|
|
15
15
|
- `release-log.md`.
|
package/package.json
CHANGED
package/release-log.md
CHANGED
|
@@ -2,6 +2,22 @@
|
|
|
2
2
|
|
|
3
3
|
## Unreleased
|
|
4
4
|
|
|
5
|
+
### 2026-06-26 — Generic Source Bundle Report Wording
|
|
6
|
+
|
|
7
|
+
- What changed: Report Markdown, methodology-adjacent contract docs, and CSV coverage expectations now use generic source-bundle and source-family wording for `source-bundle.v1`, including `collectionCoverage.status` and `collectionCoverage.sourceFamilies`.
|
|
8
|
+
- Why it matters: Maintainers can review live GitHub output and future bundled sample output without stale GitHub-only labels implying that every report source is live API evidence.
|
|
9
|
+
- Who is affected: Maintainers reviewing generated reports, `friction-report.v1` JSON, contract docs, or source coverage CSV exports during the tutorial sample migration.
|
|
10
|
+
- Action needed: Update any scripts or downstream tooling that parse source coverage fields from `apiFamilies`/`api_family` to `sourceFamilies`/`source_family`.
|
|
11
|
+
- PR: #60
|
|
12
|
+
|
|
13
|
+
### 2026-06-23 — Tutorial Target Guidance
|
|
14
|
+
|
|
15
|
+
- What changed: Public quickstart guidance no longer depends on `hannasdev/mcp-writing`, and product-repository rejection plus `--validation-target` help now make the analysis boundary clearer.
|
|
16
|
+
- Why it matters: First-time users and maintainers get a clearer path for learning the CLI without confusing internal validation history, product-repository guardrails, or validation-target metadata for an override.
|
|
17
|
+
- Who is affected: Users running the CLI for the first time and maintainers reviewing first-run documentation or help text.
|
|
18
|
+
- Action needed: None.
|
|
19
|
+
- PR: #59
|
|
20
|
+
|
|
5
21
|
### 2026-06-21 — Done Initiative Hygiene Guard
|
|
6
22
|
|
|
7
23
|
- What changed: `npm test` now enforces completed initiative documentation hygiene, and `AGENTS.md` documents how shipped, deferred, future-decision, backlog-linked, and intentionally omitted items should be recorded in done initiative docs.
|
|
@@ -1,12 +1,13 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
3
|
-
"$id": "https://delivery-friction-analyzer.local/schemas/
|
|
4
|
-
"title": "
|
|
3
|
+
"$id": "https://delivery-friction-analyzer.local/schemas/source-bundle.schema.json",
|
|
4
|
+
"title": "SourceBundle",
|
|
5
5
|
"type": "object",
|
|
6
6
|
"additionalProperties": false,
|
|
7
7
|
"required": [
|
|
8
8
|
"schemaVersion",
|
|
9
9
|
"collectedAt",
|
|
10
|
+
"source",
|
|
10
11
|
"collector",
|
|
11
12
|
"targetRepository",
|
|
12
13
|
"repositoryMetadata",
|
|
@@ -16,14 +17,27 @@
|
|
|
16
17
|
"pullRequests"
|
|
17
18
|
],
|
|
18
19
|
"properties": {
|
|
19
|
-
"schemaVersion": { "const": "
|
|
20
|
+
"schemaVersion": { "const": "source-bundle.v1" },
|
|
20
21
|
"collectedAt": { "type": "string" },
|
|
22
|
+
"source": {
|
|
23
|
+
"type": "object",
|
|
24
|
+
"additionalProperties": false,
|
|
25
|
+
"required": ["kind", "label"],
|
|
26
|
+
"properties": {
|
|
27
|
+
"kind": { "enum": ["github", "sample"] },
|
|
28
|
+
"label": { "type": "string", "minLength": 1 },
|
|
29
|
+
"metadata": {
|
|
30
|
+
"type": "object",
|
|
31
|
+
"additionalProperties": true
|
|
32
|
+
}
|
|
33
|
+
}
|
|
34
|
+
},
|
|
21
35
|
"collector": {
|
|
22
36
|
"type": "object",
|
|
23
37
|
"additionalProperties": false,
|
|
24
38
|
"required": ["name", "provider"],
|
|
25
39
|
"properties": {
|
|
26
|
-
"name": { "
|
|
40
|
+
"name": { "type": "string", "minLength": 1 },
|
|
27
41
|
"provider": { "type": "string", "minLength": 1 }
|
|
28
42
|
}
|
|
29
43
|
},
|
|
@@ -48,7 +62,7 @@
|
|
|
48
62
|
"additionalProperties": false,
|
|
49
63
|
"required": ["strategy", "requestedLimit", "collectedCount", "source"],
|
|
50
64
|
"properties": {
|
|
51
|
-
"strategy": { "
|
|
65
|
+
"strategy": { "type": "string", "minLength": 1 },
|
|
52
66
|
"requestedLimit": { "type": "integer", "minimum": 1, "maximum": 100 },
|
|
53
67
|
"collectedCount": { "type": "integer", "minimum": 0, "maximum": 100 },
|
|
54
68
|
"source": { "type": "string", "minLength": 1 }
|
|
@@ -57,12 +71,12 @@
|
|
|
57
71
|
"coverage": {
|
|
58
72
|
"type": "object",
|
|
59
73
|
"additionalProperties": false,
|
|
60
|
-
"required": ["status", "
|
|
74
|
+
"required": ["status", "sourceFamilies"],
|
|
61
75
|
"properties": {
|
|
62
76
|
"status": { "$ref": "#/$defs/coverageStatus" },
|
|
63
|
-
"
|
|
77
|
+
"sourceFamilies": {
|
|
64
78
|
"type": "array",
|
|
65
|
-
"items": { "$ref": "#/$defs/
|
|
79
|
+
"items": { "$ref": "#/$defs/sourceCoverageEntry" }
|
|
66
80
|
}
|
|
67
81
|
}
|
|
68
82
|
},
|
|
@@ -71,7 +85,7 @@
|
|
|
71
85
|
"additionalProperties": false,
|
|
72
86
|
"required": ["source", "bytesByLanguage", "coverage"],
|
|
73
87
|
"properties": {
|
|
74
|
-
"source": { "
|
|
88
|
+
"source": { "type": "string", "minLength": 1 },
|
|
75
89
|
"bytesByLanguage": {
|
|
76
90
|
"type": "object",
|
|
77
91
|
"additionalProperties": { "type": "integer", "minimum": 0 }
|
|
@@ -83,7 +97,7 @@
|
|
|
83
97
|
"type": "object",
|
|
84
98
|
"properties": {
|
|
85
99
|
"family": { "const": "languages" },
|
|
86
|
-
"source": { "
|
|
100
|
+
"source": { "type": "string", "minLength": 1 }
|
|
87
101
|
}
|
|
88
102
|
}
|
|
89
103
|
]
|
|
@@ -104,7 +118,7 @@
|
|
|
104
118
|
"type": "object",
|
|
105
119
|
"properties": {
|
|
106
120
|
"family": { "const": "contributor_source" },
|
|
107
|
-
"source": { "
|
|
121
|
+
"source": { "type": "string", "minLength": 1 }
|
|
108
122
|
}
|
|
109
123
|
}
|
|
110
124
|
]
|
|
@@ -150,7 +164,7 @@
|
|
|
150
164
|
"downstreamImpact": { "type": ["string", "null"] }
|
|
151
165
|
}
|
|
152
166
|
},
|
|
153
|
-
"
|
|
167
|
+
"sourceCoverageEntry": {
|
|
154
168
|
"allOf": [
|
|
155
169
|
{ "$ref": "#/$defs/coverageEntry" },
|
|
156
170
|
{
|
|
@@ -319,7 +333,7 @@
|
|
|
319
333
|
"additionalProperties": false,
|
|
320
334
|
"required": ["source", "totalCount", "nodes"],
|
|
321
335
|
"properties": {
|
|
322
|
-
"source": { "
|
|
336
|
+
"source": { "type": "string", "minLength": 1 },
|
|
323
337
|
"totalCount": { "type": "integer", "minimum": 0 },
|
|
324
338
|
"nodes": {
|
|
325
339
|
"type": "array",
|
|
@@ -378,7 +392,7 @@
|
|
|
378
392
|
"additionalProperties": false,
|
|
379
393
|
"required": ["source", "totalCount", "conclusions", "runs"],
|
|
380
394
|
"properties": {
|
|
381
|
-
"source": { "
|
|
395
|
+
"source": { "type": "string", "minLength": 1 },
|
|
382
396
|
"totalCount": { "type": ["integer", "null"], "minimum": 0 },
|
|
383
397
|
"conclusions": {
|
|
384
398
|
"type": "object",
|
|
@@ -434,7 +448,7 @@
|
|
|
434
448
|
"type": "object",
|
|
435
449
|
"properties": {
|
|
436
450
|
"family": { "const": "pr_open_diff" },
|
|
437
|
-
"source": { "
|
|
451
|
+
"source": { "type": "string", "minLength": 1 },
|
|
438
452
|
"status": { "enum": ["available", "partial", "unavailable", "rate_limited"] }
|
|
439
453
|
}
|
|
440
454
|
}
|
|
@@ -447,7 +461,7 @@
|
|
|
447
461
|
"type": "object",
|
|
448
462
|
"properties": {
|
|
449
463
|
"family": { "const": "review_threads" },
|
|
450
|
-
"source": { "
|
|
464
|
+
"source": { "type": "string", "minLength": 1 }
|
|
451
465
|
}
|
|
452
466
|
}
|
|
453
467
|
]
|
|
@@ -459,7 +473,7 @@
|
|
|
459
473
|
"type": "object",
|
|
460
474
|
"properties": {
|
|
461
475
|
"family": { "const": "workflow_runs" },
|
|
462
|
-
"source": { "
|
|
476
|
+
"source": { "type": "string", "minLength": 1 }
|
|
463
477
|
}
|
|
464
478
|
}
|
|
465
479
|
]
|
|
@@ -125,7 +125,7 @@ Options:
|
|
|
125
125
|
--dry-run Validate inputs and sample GitHub coverage without writing artifacts.
|
|
126
126
|
--no-dry-run Disable dry-run mode when a preset enabled it.
|
|
127
127
|
--metadata-only Alias for --dry-run.
|
|
128
|
-
--validation-target Mark
|
|
128
|
+
--validation-target Mark output metadata as an internal validation run; does not bypass target validation.
|
|
129
129
|
--no-validation-target Disable validation-target mode when a preset enabled it.
|
|
130
130
|
--exclude-pr-class <cls> Exclude a PR class from normalized, metrics, report, methodology, and CSV artifacts. Repeat or comma-separate values.
|
|
131
131
|
--csv Enable curated CSV evidence exports when a preset disabled them.
|
|
@@ -1400,16 +1400,22 @@ export async function writeAnalysisArtifacts(outDir, paths, artifacts, { disable
|
|
|
1400
1400
|
}
|
|
1401
1401
|
|
|
1402
1402
|
function collectionCoverageMarkdown(sourceBundle) {
|
|
1403
|
+
const source = sourceBundle.source;
|
|
1404
|
+
const sourceLabel = source?.label
|
|
1405
|
+
? `${source.label}${source.kind ? ` (${source.kind})` : ""}`
|
|
1406
|
+
: "not recorded";
|
|
1403
1407
|
const lines = [
|
|
1404
1408
|
"",
|
|
1405
1409
|
"## Collection Coverage",
|
|
1406
1410
|
"",
|
|
1411
|
+
`Source: ${sourceLabel}`,
|
|
1412
|
+
"",
|
|
1407
1413
|
`Overall collection coverage: ${sourceBundle.coverage.status}`,
|
|
1408
1414
|
"",
|
|
1409
|
-
"
|
|
1415
|
+
"Source families:",
|
|
1410
1416
|
];
|
|
1411
1417
|
|
|
1412
|
-
for (const family of sourceBundle.coverage.
|
|
1418
|
+
for (const family of sourceBundle.coverage.sourceFamilies ?? []) {
|
|
1413
1419
|
const diagnostics = (family.diagnostics ?? []).length ? `; diagnostics: ${family.diagnostics.join(" | ")}` : "";
|
|
1414
1420
|
const impact = family.downstreamImpact ? `; impact: ${family.downstreamImpact}` : "";
|
|
1415
1421
|
lines.push(`- ${family.family}: ${family.status} (${family.attempts ?? 1} attempt(s))${diagnostics}${impact}`);
|
|
@@ -1422,6 +1428,7 @@ function collectionCoverageMarkdown(sourceBundle) {
|
|
|
1422
1428
|
function attachCollectionCoverage(report, sourceBundle) {
|
|
1423
1429
|
return {
|
|
1424
1430
|
...report,
|
|
1431
|
+
source: sourceBundle.source,
|
|
1425
1432
|
collectionCoverage: sourceBundle.coverage,
|
|
1426
1433
|
artifactSensitivity: "Generated artifacts may include repository names, PR URLs, titles, file paths, comment metadata, contributor-source metadata, curated CSV evidence, and coverage diagnostics. Raw contributor file contents and individual contributor rankings are not emitted. Treat artifacts as local/private unless intentionally shared.",
|
|
1427
1434
|
};
|
|
@@ -1607,7 +1614,7 @@ function coverageLine(family) {
|
|
|
1607
1614
|
}
|
|
1608
1615
|
|
|
1609
1616
|
function coverageCaveats(coverage) {
|
|
1610
|
-
return (coverage?.
|
|
1617
|
+
return (coverage?.sourceFamilies ?? []).filter(family => {
|
|
1611
1618
|
const diagnostics = (family.diagnostics ?? []).filter(Boolean);
|
|
1612
1619
|
return family.status !== "available" || diagnostics.length > 0;
|
|
1613
1620
|
});
|
|
@@ -18,7 +18,7 @@ import {
|
|
|
18
18
|
redactDiagnostic,
|
|
19
19
|
} from "./coverage.js";
|
|
20
20
|
|
|
21
|
-
export const
|
|
21
|
+
export const SOURCE_BUNDLE_VERSION = "source-bundle.v1";
|
|
22
22
|
|
|
23
23
|
const REPOSITORY_SLUG = /^([A-Za-z0-9_.-]+)\/([A-Za-z0-9_.-]+)$/;
|
|
24
24
|
function parseRepositoryInput(input) {
|
|
@@ -507,7 +507,7 @@ export async function collectGitHubSourceBundle({
|
|
|
507
507
|
pullRequests.push(pr);
|
|
508
508
|
}
|
|
509
509
|
|
|
510
|
-
const
|
|
510
|
+
const sourceFamilies = [
|
|
511
511
|
repositoryCoverage,
|
|
512
512
|
languagesAttempt.coverage,
|
|
513
513
|
inventoryCoverage,
|
|
@@ -542,8 +542,12 @@ export async function collectGitHubSourceBundle({
|
|
|
542
542
|
];
|
|
543
543
|
|
|
544
544
|
return {
|
|
545
|
-
schemaVersion:
|
|
545
|
+
schemaVersion: SOURCE_BUNDLE_VERSION,
|
|
546
546
|
collectedAt,
|
|
547
|
+
source: {
|
|
548
|
+
kind: "github",
|
|
549
|
+
label: "GitHub live collection",
|
|
550
|
+
},
|
|
547
551
|
collector: {
|
|
548
552
|
name: "github-live-collector",
|
|
549
553
|
provider: provider.kind ?? "custom",
|
|
@@ -557,8 +561,8 @@ export async function collectGitHubSourceBundle({
|
|
|
557
561
|
source: "gh pr list --state merged --search \"is:merged sort:merged-desc\"",
|
|
558
562
|
},
|
|
559
563
|
coverage: {
|
|
560
|
-
status: buildCoverageSummary(
|
|
561
|
-
|
|
564
|
+
status: buildCoverageSummary(sourceFamilies),
|
|
565
|
+
sourceFamilies,
|
|
562
566
|
},
|
|
563
567
|
languageDistribution: {
|
|
564
568
|
source: "rest:/repos/{owner}/{repo}/languages",
|
|
@@ -99,5 +99,5 @@ export function productRepositoryTargetError(input) {
|
|
|
99
99
|
const repository = typeof input?.owner === "string" && typeof input?.name === "string"
|
|
100
100
|
? `${input.owner}/${input.name}`
|
|
101
101
|
: "the requested repository";
|
|
102
|
-
return `Cannot analyze ${repository} because it is this tool's product repository
|
|
102
|
+
return `Cannot analyze ${repository} because it is this tool's product repository. The guard prevents accidental self-analysis during normal live runs; it is not a data-security boundary. Choose a different repository with --repo owner/name. No GitHub data was collected.`;
|
|
103
103
|
}
|
|
@@ -247,17 +247,17 @@ function commentSourcesCsv(report, analysisFilter) {
|
|
|
247
247
|
|
|
248
248
|
function collectionCoverageCsv(collectionCoverage, analysisFilter) {
|
|
249
249
|
const headers = [
|
|
250
|
-
"
|
|
250
|
+
"source_family",
|
|
251
251
|
"status",
|
|
252
252
|
"attempts",
|
|
253
253
|
"source",
|
|
254
254
|
"diagnostics",
|
|
255
255
|
"downstream_impact",
|
|
256
256
|
];
|
|
257
|
-
const rows = [...(collectionCoverage?.
|
|
257
|
+
const rows = [...(collectionCoverage?.sourceFamilies ?? [])]
|
|
258
258
|
.sort((left, right) => String(left.family).localeCompare(String(right.family)))
|
|
259
259
|
.map(family => ({
|
|
260
|
-
|
|
260
|
+
source_family: family.family,
|
|
261
261
|
status: family.status,
|
|
262
262
|
attempts: family.attempts ?? 1,
|
|
263
263
|
source: family.source,
|
|
@@ -283,8 +283,13 @@ function repositoryLabel(report) {
|
|
|
283
283
|
: "unknown repository";
|
|
284
284
|
}
|
|
285
285
|
|
|
286
|
+
function sourceLabel(source) {
|
|
287
|
+
if (!source?.label) return "not recorded";
|
|
288
|
+
return `${source.label}${source.kind ? ` (${source.kind})` : ""}`;
|
|
289
|
+
}
|
|
290
|
+
|
|
286
291
|
function formatCoverageFamilies(collectionCoverage) {
|
|
287
|
-
const families = collectionCoverage?.
|
|
292
|
+
const families = collectionCoverage?.sourceFamilies ?? [];
|
|
288
293
|
if (!families.length) return "- No collection coverage families were recorded.";
|
|
289
294
|
return families
|
|
290
295
|
.map(family => {
|
|
@@ -419,12 +424,14 @@ export function renderRepositoryFrictionMethodology({
|
|
|
419
424
|
}) {
|
|
420
425
|
const selection = sourceBundle?.selection ?? {};
|
|
421
426
|
const collectionCoverage = sourceBundle?.coverage ?? report.collectionCoverage;
|
|
427
|
+
const source = sourceBundle?.source ?? report.source;
|
|
422
428
|
|
|
423
429
|
return `${[
|
|
424
430
|
`# Methodology: ${repositoryLabel(report)}`,
|
|
425
431
|
"",
|
|
426
432
|
`Report version: ${report.reportVersion}`,
|
|
427
433
|
`Metric version: ${report.metricVersion}`,
|
|
434
|
+
`Source: ${sourceLabel(source)}`,
|
|
428
435
|
`Repository: ${repositoryLabel(report)}`,
|
|
429
436
|
`Profile path: ${profilePath ?? "not recorded"}`,
|
|
430
437
|
`Requested pull requests: ${selection.requestedLimit ?? "unknown"}`,
|
|
@@ -434,7 +441,7 @@ export function renderRepositoryFrictionMethodology({
|
|
|
434
441
|
"",
|
|
435
442
|
"## What This Analysis Uses",
|
|
436
443
|
"",
|
|
437
|
-
"The analyzer
|
|
444
|
+
"The analyzer normalizes source-bundle pull request evidence through the supplied profile, computes transparent component metrics, and renders a repository-level report. It does not inspect local working trees, mutate repositories, rank people, or apply recommendations automatically.",
|
|
438
445
|
"",
|
|
439
446
|
"## Pull Request Selection",
|
|
440
447
|
"",
|
|
@@ -1624,6 +1624,9 @@ export function renderRepositoryFrictionMarkdown(report) {
|
|
|
1624
1624
|
"",
|
|
1625
1625
|
`Report version: ${report.reportVersion}`,
|
|
1626
1626
|
`Metric version: ${report.metricVersion}`,
|
|
1627
|
+
...(report.source?.label
|
|
1628
|
+
? [`Source: ${report.source.label}${report.source.kind ? ` (${report.source.kind})` : ""}`]
|
|
1629
|
+
: []),
|
|
1627
1630
|
`Pull requests analyzed: ${report.summary?.pullRequests ?? "unknown"}`,
|
|
1628
1631
|
"",
|
|
1629
1632
|
...analysisFilterLines,
|
|
@@ -1641,7 +1644,7 @@ export function renderRepositoryFrictionMarkdown(report) {
|
|
|
1641
1644
|
"",
|
|
1642
1645
|
"## How To Read This Report",
|
|
1643
1646
|
"",
|
|
1644
|
-
"- Observed evidence is measured from
|
|
1647
|
+
"- Observed evidence is measured from source-bundle evidence and repository-profile classifications.",
|
|
1645
1648
|
"- Interpretation is the analyzer's explanation of what the observed evidence suggests.",
|
|
1646
1649
|
"- Recommendation is a workflow intervention to consider; the report does not modify repositories.",
|
|
1647
1650
|
"- Confidence and caveats call out outliers, missing coverage, and evidence-quality limits before you act.",
|
|
@@ -1731,7 +1734,7 @@ export function renderRepositoryFrictionMarkdown(report) {
|
|
|
1731
1734
|
: "- No profile suggestion thresholds were triggered by this report's PR class, role, functional-surface, or workflow-coverage evidence.",
|
|
1732
1735
|
"- Bottlenecks are ranked by their strongest representative observed signal, with stable category order only used to break ties.",
|
|
1733
1736
|
"- Recommendations are inferred from transparent component evidence and representative PR examples; they are not automated changes.",
|
|
1734
|
-
"- Missing or partial
|
|
1737
|
+
"- Missing or partial source evidence remains visible in coverage tables rather than being inferred from unrelated fields.",
|
|
1735
1738
|
...(hasConfiguredWorkflowContext(report.configuredWorkflow)
|
|
1736
1739
|
? ["- Configured workflow context is user-configured repository-profile context; it does not change scoring, ranking, CSV exports, or PR class matching."]
|
|
1737
1740
|
: []),
|