knodin 0.7.6 → 0.8.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (130) hide show
  1. package/README.md +19 -7
  2. package/benchmarks/competitors/SYNTHESIS.md +66 -0
  3. package/dist/bin/cli.js +2164 -108
  4. package/dist/bin/launcher.js +25 -3
  5. package/dist/src/agent-integration.js +304 -0
  6. package/dist/src/artifact-refresh.js +82 -0
  7. package/dist/src/cli-args.js +292 -0
  8. package/dist/src/cli-model.js +384 -0
  9. package/dist/src/codeflow-replay.js +81 -0
  10. package/dist/src/compact-structural.js +96 -0
  11. package/dist/src/compare.js +39 -0
  12. package/dist/src/competitive-cold-mcp.js +40 -0
  13. package/dist/src/competitive-constraints.js +21 -0
  14. package/dist/src/competitive-manifest.js +411 -0
  15. package/dist/src/competitive-measurement.js +183 -0
  16. package/dist/src/competitive-runner.js +487 -0
  17. package/dist/src/competitive-sandbox.js +108 -0
  18. package/dist/src/context-export.js +423 -0
  19. package/dist/src/context.js +102 -0
  20. package/dist/src/deterministic-random.js +34 -0
  21. package/dist/src/diagnostics-write-helper.js +473 -0
  22. package/dist/src/diagnostics.js +1476 -0
  23. package/dist/src/docs-sections.js +141 -0
  24. package/dist/src/doctor.js +382 -0
  25. package/dist/src/engine/ann-hnsw.js +261 -0
  26. package/dist/src/engine/embeddings.js +193 -0
  27. package/dist/src/engine/file-walker.js +49 -0
  28. package/dist/src/engine/git-history.js +289 -0
  29. package/dist/src/engine/index.js +14238 -0
  30. package/dist/src/engine/perf.js +115 -0
  31. package/dist/src/engine/prune.js +112 -0
  32. package/dist/src/engine/sarif-import.js +341 -0
  33. package/dist/src/engine/scip-import.js +423 -0
  34. package/dist/src/engine/source-policy.js +85 -0
  35. package/dist/src/engine/sqlite.js +71 -0
  36. package/dist/src/engine/state-paths.js +175 -0
  37. package/dist/src/engine/symbol-delete.js +58 -0
  38. package/dist/src/execution-profile.js +208 -0
  39. package/dist/src/failure-diagnosis.js +655 -0
  40. package/dist/src/fleet.js +7 -0
  41. package/dist/src/git-executable.js +31 -0
  42. package/dist/src/graph-layout.js +173 -0
  43. package/dist/src/graph-query-health.js +115 -0
  44. package/dist/src/hook-manager-integration.js +156 -0
  45. package/dist/src/index-activity.js +126 -0
  46. package/dist/src/init-progress-worker.js +106 -2
  47. package/dist/src/init-progress.js +155 -0
  48. package/dist/src/init.js +1295 -0
  49. package/dist/src/lifecycle-health.js +282 -0
  50. package/dist/src/lsp-readonly.js +217 -0
  51. package/dist/src/mcp-graph-worker.js +69 -0
  52. package/dist/src/mcp-reliability.js +154 -0
  53. package/dist/src/mcp-worker-supervisor.js +350 -0
  54. package/dist/src/mirror.js +290 -0
  55. package/dist/src/node-runtime.js +157 -0
  56. package/dist/src/output-compression.js +630 -0
  57. package/dist/src/output-telemetry.js +368 -0
  58. package/dist/src/pr-triage.js +638 -0
  59. package/dist/src/progressive-evidence.js +477 -0
  60. package/dist/src/pure-compression-cli.js +102 -0
  61. package/dist/src/relationship-adapters.js +377 -0
  62. package/dist/src/release-attestation.js +533 -0
  63. package/dist/src/release-preflight.js +513 -0
  64. package/dist/src/repair-lease.js +85 -0
  65. package/dist/src/repair-progress-worker.js +120 -2
  66. package/dist/src/repair-progress.js +262 -0
  67. package/dist/src/repository-init-process.js +177 -0
  68. package/dist/src/repository-management.js +1261 -0
  69. package/dist/src/response-budget.js +196 -0
  70. package/dist/src/server.js +217 -0
  71. package/dist/src/structural-fast-path.js +344 -0
  72. package/dist/src/structural-snapshot.js +37 -0
  73. package/dist/src/system-config.js +638 -0
  74. package/dist/src/terminal-help.js +83 -0
  75. package/dist/src/tools/knodin-tools.js +1640 -0
  76. package/dist/src/update-ceremony.js +162 -0
  77. package/dist/src/update-policy.js +944 -0
  78. package/dist/src/update-trust.js +504 -0
  79. package/dist/src/version.js +13 -0
  80. package/dist/src/visualization.js +515 -0
  81. package/dist/src/wait-for-fresh.js +98 -0
  82. package/dist/src/worktree-lifecycle.js +234 -0
  83. package/docs/BEHAVIORAL-CONTRACT.md +72 -0
  84. package/docs/CLI.md +20 -1
  85. package/docs/COMPARISON.md +403 -0
  86. package/docs/COMPETITIVE-LANDSCAPE-2026-08.md +267 -0
  87. package/docs/CONTAINED-EXECUTION.md +77 -0
  88. package/docs/DIAGNOSTICS.md +80 -0
  89. package/docs/GIT-HISTORY-REVIEW.md +39 -0
  90. package/docs/HANDOFF.md +180 -0
  91. package/docs/INSTALLATION.md +21 -18
  92. package/docs/MCP.md +59 -8
  93. package/docs/PROGRESSIVE-EVIDENCE.md +37 -0
  94. package/docs/PT-ACCESS-RECOMMENDATION.md +89 -0
  95. package/docs/RELEASE-0.3-EVIDENCE.md +73 -0
  96. package/docs/REPOSITORIES-AND-WORKTREES.md +18 -6
  97. package/docs/SCIP-IMPORT.md +62 -0
  98. package/docs/SIGNED-UPDATES.md +151 -0
  99. package/docs/TELEMETRY.md +46 -0
  100. package/docs/TOKEN-OPTIMIZER-SCORECARD.md +79 -0
  101. package/docs/assets/knodin-favicon.svg +4 -0
  102. package/docs/releases/0.3.0.md +46 -0
  103. package/docs/releases/0.4.0.md +68 -0
  104. package/docs/releases/0.4.1.md +28 -0
  105. package/docs/releases/0.4.2.md +27 -0
  106. package/docs/releases/0.4.3.md +23 -0
  107. package/docs/releases/0.5.0.md +29 -0
  108. package/docs/releases/0.5.1.md +17 -0
  109. package/docs/releases/0.6.0.md +18 -0
  110. package/docs/releases/0.7.0.md +24 -0
  111. package/docs/releases/0.7.1.md +21 -0
  112. package/docs/releases/0.7.2.md +21 -0
  113. package/docs/releases/0.7.3.md +23 -0
  114. package/docs/releases/0.7.4.md +17 -0
  115. package/docs/releases/0.7.5.md +20 -0
  116. package/docs/releases/0.8.0.md +74 -0
  117. package/docs/releases/0.8.2.md +34 -0
  118. package/package.json +127 -4
  119. package/roadmap/competitive-roadmap.md +3801 -0
  120. package/schemas/release-attestation-v1.schema.json +210 -0
  121. package/schemas/support-bundle-v2.schema.json +212 -0
  122. package/dist/chunks/chunk-DMQAGX77.js +0 -654
  123. package/dist/chunks/chunk-F4Z3Z766.js +0 -4
  124. package/dist/chunks/chunk-SIJAQVSX.js +0 -3
  125. package/dist/chunks/chunk-X6M4HUUE.js +0 -2
  126. package/dist/chunks/chunk-YPRMY2LP.js +0 -8
  127. package/dist/chunks/pure-compression-cli-4TA2TQD5.js +0 -5
  128. package/dist/chunks/server-7EDF4CBY.js +0 -14
  129. package/dist/chunks/structural-fast-path-KD5KQSPX.js +0 -4
  130. package/docs/releases/0.7.6.md +0 -25
@@ -0,0 +1,3801 @@
1
+ # knodin — Competitive Roadmap
2
+
3
+ This is the authoritative active work queue after the original parity program
4
+ closed on 2026-07-21. The completed R1–R48 history, including measured no-go and
5
+ deferred performance decisions, is preserved in
6
+ [`archive/parity-roadmap-2026-07-21.md`](archive/parity-roadmap-2026-07-21.md).
7
+
8
+ ## Product constraint
9
+
10
+ knodin wins by being local, low-latency, and compact: one MCP tool, no hosted
11
+ service, no required credentials, and no source-code egress. Competitor features
12
+ belong here only when live evidence shows that they improve correctness or a
13
+ core developer loop without undermining that constraint.
14
+
15
+ ## Evidence policy
16
+
17
+ Every competitive item must link to a reproducible result under
18
+ `benchmarks/competitors/`. A competitor having a feature is not sufficient by
19
+ itself. Items use four dispositions:
20
+
21
+ - **Must close** — correctness, safety, or a prerequisite for the stated product
22
+ position.
23
+ - **Should close** — meaningful capability or efficiency improvement with a
24
+ bounded implementation path.
25
+ - **Evaluate** — evidence is incomplete or implementation cost may outweigh the
26
+ benefit; do not implement before its gate passes.
27
+ - **Won't pursue** — conflicts with the product constraint or duplicates a
28
+ better knodin mechanism.
29
+
30
+ The first evidence set is
31
+ [`GitNexus vs knodin`](../benchmarks/competitors/gitnexus-vs-knodin.md), backed
32
+ by 49 local cases and their raw MCP results. Later competitor bake-offs append
33
+ new items or strengthen existing ones; they do not reopen the archived roadmap.
34
+
35
+ ## Execution order
36
+
37
+ | ID | Priority | Disposition | Item | Depends on |
38
+ |---|---:|---|---|---|
39
+ | C1 | P0 | Must close | Declare a permissive license | — |
40
+ | C2 | P0 | Must close | Stable symbol identity and ambiguity-safe resolution | — |
41
+ | C3 | P0 | Must close | Enforce truthful response budgets and minimal modes | — |
42
+ | C4 | P1 | Must close | Symbol-level directional impact analysis | C2, C3 |
43
+ | C5 | P1 | Should close | Explicit git diff scopes | C3 |
44
+ | C6 | P1 | Should close | Deterministic import-cycle analysis | C2 |
45
+ | C7 | P2 | Should close | MCP tool and handler mapping | C2 |
46
+ | C8 | P2 | Evaluate | Local statement-level PDG | C2, C3 |
47
+ | C9 | P2 | Evaluate | API contract and response-shape analysis | C2, C3 |
48
+ | C10 | P3 | Evaluate | Bounded expert graph-query DSL | C2, C3 |
49
+ | C11 | P0 | Must close | Call/reference precision and structural identity | C2 |
50
+ | C12 | P1 | Should close | Uniform filters, pagination, and compact drill-downs | C2, C3 |
51
+ | C13 | P1 | Should close | Typed traversal and architecture facets | C2, C3, C11 |
52
+ | C14 | P2 | Should close | Portable budgeted context export | C3, C5 |
53
+ | C15 | P2 | Evaluate | LSP diagnostics and guarded symbol editing | C2, C11 |
54
+ | C16 | P2 | Should close | Index health, repair, and output telemetry | C3 |
55
+ | C17 | P3 | Evaluate | DFS and feature-path navigation | C2, C3, C11 |
56
+ | C18 | P1 | Should close | Unified competitive regression harness | — |
57
+ | C19 | P2 | Evaluate | Salesforce LWC bundles and Apex/data bridges | C2, C3, C11 |
58
+ | C20 | P2 | Evaluate | Salesforce metadata and Flow/automation model | C2, C3, C11 |
59
+ | C21 | P3 | Evaluate | Aura bundles and Visualforce/Aura/LWC interop | C19 |
60
+ | C22 | P3 | Evaluate | LWR and Experience Cloud topology | C19 |
61
+ | C23 | P2 | Evaluate | Apex platform-entry and asynchronous semantics | C2, C3, C11 |
62
+ | C24 | P0 | Must close | Competitive correctness contract and replay baseline | C18 |
63
+ | C25 | P0 | Must close | Prove symbol identity and reference precision leadership | C24, C2, C11 |
64
+ | C26 | P1 | Must close | Prove diff-review and directional traversal leadership | C24, C4, C5, C13 |
65
+ | C27 | P1 | Must close | Prove architecture and community-analysis leadership | C24, C3, C12, C13 |
66
+ | C28 | P1 | Must close | Prove dead-code and semantic-search quality | C24, C3, C11, C12 |
67
+ | C29 | P1 | Must close | Prove context-packing and export leadership | C24, C3, C5, C14 |
68
+ | C30 | P2 | Evaluate | Prove guarded editing and diagnostics leadership | C24, C2, C11, C15 |
69
+ | C31 | P1 | Must close | Prove index lifecycle and output-telemetry leadership | C24, C3, C16, C18 |
70
+ | C32 | P2 | Evaluate | Prove local API and statement-flow analysis leadership | C24, C2, C3, C8, C9 |
71
+ | C33 | P0 | Must close | Preserve local privacy and one-tool schema economy | C24 |
72
+ | C34 | P1 | Must close | Keep local graph indexes current across repository lifecycle events | C16, C24, C33 |
73
+ | C35 | P2 | implemented | Produce local interactive architecture and call-flow visualizations | C3, C13, C24, C33 |
74
+ | C36 | P0 | implemented | Instrument the current warm-operation performance paths | C18, C24 |
75
+ | C37 | P0 | implemented | Reuse generation-scoped graph analytics and traversal snapshots | C36 |
76
+ | C38 | P1 | implemented | Compose orientation context from one shared snapshot | C37 |
77
+ | C39 | P1 | implemented | Separate fast truthful status from explicit deep audit | C36 |
78
+ | C40 | P1 | implemented | Make context packing and artifact access linear and cache-safe | C36 |
79
+ | C41 | P0 | evaluated — deferred | Replay like-for-like performance after the implementation batch | C37, C38, C39, C40 |
80
+ | C42 | P0 | implemented | Enforce comparable competitive cases and correctness oracles | C41 |
81
+ | C43 | P0 | implemented | Measure cold/warm p50, p95, and independent process RSS | C42 |
82
+ | C44 | P0 | implemented | Restore low-latency diff review without weakening evidence | C43 |
83
+ | C45 | P1 | implemented | Add facet-selective architecture fast paths | C43 |
84
+ | C46 | P1 | implemented | Cache search state and hydrate only the requested page | C43 |
85
+ | C47 | P1 | implemented | Reduce status, repair, and measured lifecycle resource cost | C43 |
86
+ | C48 | P2 | evaluated — deferred | Decide an optional local project-memory boundary | C42 |
87
+ | C49 | P1 | implemented | Keep immutable runs lint-safe and make the full suite terminate | — |
88
+ | C50 | P0 | implemented | Make the single-worker full-suite gate start reliably | C49 |
89
+ | C51 | P2 | evaluated — rejected for production | Native training-free embedding quantization | C46 |
90
+ | C52 | P0 | implemented | Package repository lifecycle initialization and refresh | C16, C34, C50 |
91
+ | C53 | P0 | implemented | Make repair observable, cancellable, and resumable | C47, C52 |
92
+ | C54 | P0 | implemented | Certify the local 0.1.0 release artifact | C50, C52, C53 |
93
+ | C55 | P0 | implemented | Pin and audit the authorized Token Optimizer source | C24 |
94
+ | C56 | P0 | implemented | Replay all seven Token Optimizer workflows on shared fixtures | C55 |
95
+ | C57 | P0 | implemented | Add bounded recoverable diagnostic-output compression | C3, C55 |
96
+ | C58 | P0 | evaluated — rejected | Gate command execution on reviewed portable containment | C57 |
97
+ | C59 | P1 | implemented | Connect diagnostic failures to source-evidenced graph context | C57 |
98
+ | C60 | P1 | parked-external-evidence — 2026-08-02 (previous: active — legacy GHES 0.3.0 live; external mirror governance, workload identity, and native certification pending) | Complete two-minute cross-client and GHES onboarding | C52, C55 |
99
+ | C61 | P1 | implemented | Deliver private real-token ROI telemetry and dashboard | C16, C55 |
100
+ | C62 | P0 | parked-external-evidence — 2026-08-02 (previous: active — verifier implemented; production ceremony pending) | Define and implement signed update metadata trust | C55 |
101
+ | C63 | P0 | parked-external-evidence — 2026-08-02 (previous: active — signed-only client implemented; production activation gated) | Add safe update status, check, explain, apply, and rollback policy | C62 |
102
+ | C64 | P0 | parked-external-evidence — 2026-08-02 (previous: active — implementation and adversarial fixtures complete; live release evidence pending) | Produce provenance, SBOM, and cross-channel release attestation | C62 |
103
+ | C65 | P0 | parked-external-evidence — 2026-08-02 (previous: active — bounded multi-step root recovery implemented; production drills pending) | Prove compromised-channel, rollback, freeze, and recovery defenses | C63, C64 |
104
+ | C66 | P0 | parked-external-evidence — 2026-08-02 (previous: active — scorecard published; terminal refresh awaits active dependencies) | Publish an evidence-linked Token Optimizer scorecard | C56, C57, C59, C60, C61, C65, C68, C75 |
105
+ | C67 | P0 | parked-external-evidence — 2026-08-02 (previous: Must close) | Certify and distribute the completed competitive release | C60, C62, C63, C64, C65, C66, C75 |
106
+ | C68 | P0 | implemented | Close Token Optimizer structural output and latency gaps | C56 |
107
+ | C69 | P0 | implemented | Stream large repository inventories with per-repository degradation | C52 |
108
+ | C70 | P0 | implemented | Select portfolio repositories before inventory and reject unknown selectors | C69 |
109
+ | C71 | P0 | implemented | Make portfolio init bounded, resumable, and truthful in dry-run mode | C69, C70 |
110
+ | C72 | P0 | implemented | Make configure/init lifecycle claims atomic and drain compatible queued events | C52 |
111
+ | C73 | P1 | implemented | Distinguish portfolio doctor and installed-versus-active MCP states | C60, C69 |
112
+ | C74 | P1 | implemented | Bound and classify Salesforce metadata architecture candidates | C69 |
113
+ | C75 | P0 | certified with recorded degradations | Certify the 61-repository portfolio under the one-GB memory ceiling | C69, C70, C71, C72, C73, C74 |
114
+ | C76 | P1 | implemented | Replace duplicated CLI parsing/help with one responsive declarative command model | C73 |
115
+ | C77 | P2 | evaluated — retained knodin unchanged | Replay CodeFlow edge-provenance and architecture-export claims | C24, C27, C42 |
116
+ | C78 | P1 | evaluated — retained C8 unchanged | Add bounded source-to-sink resource reachability | C2, C3, C8, C24 |
117
+ | C79 | P1 | implemented | Evaluate or add bounded repository applicability signals | C52, C69 |
118
+ | C80 | P0 | implemented | Route retrieval through structural evidence before embeddings | C3, C24, C42, C43 |
119
+ | C81 | P0 | evaluated — retained knodin unchanged | Run a powered end-to-end benchmark with a diagnose arm | C42, C43, C59, C80 |
120
+ | C82 | P0 | implemented | Productize the retained compress diagnose workflow | C59, C81 |
121
+ | C83 | P0 | implemented | Deliver progressive evidence with a verified hash handshake | C80, C82 |
122
+ | C84 | P1 | implemented | Import optional SCIP evidence without a live-LSP dependency | C24, C80, C83 |
123
+ | C85 | P1 | implemented | Add bounded, itemized git-history review signals | C24, C44, C81 |
124
+ | C86 | P2 | evaluated — retained MiniLM unchanged | Evaluate static embeddings against retained MiniLM | C46, C51, C80, C81 |
125
+ | C87 | P0 | implemented | Add a truthful structural cold-start fast path | C56, C68 |
126
+ | C88 | P1 | implemented — macOS certified; Linux/Windows unavailable | Add profile-based contained execution | C57, C58, C59 |
127
+ | C89 | P0 | implemented | Harden MCP request reliability and recovery | C3, C16, C53 |
128
+ | C90 | P0 | implemented | Produce previewable privacy-safe support bundles | C16, C59, C89 |
129
+ | C91 | P0 | parked-external-evidence — 2026-08-05 (previous: owner-approved adoption priority) | Certify boring macOS installation and Node runtime handoff | C52, C73, C89, C90 |
130
+ | C92 | P0 | parked-external-evidence — 2026-08-06 (previous: proposed) | Unify fail-closed five-channel release orchestration | C64, C91 |
131
+ | C93 | P1 | implemented | Measure engineering task outcomes with and without knodin | C24, C81 |
132
+ | C94 | P1 | parked-external-evidence — 2026-08-06 (previous: blocked) | Make first use self-explanatory and demonstrable | C89, C90, C91 |
133
+ | C95 | P1 | implemented | Codify and replay the defensible behavioral contract | C3, C24, C83 |
134
+ | C96 | P2 | evaluated — deferred | Gate every new capability on a narrow outcome replay | C24, C93, C95 |
135
+
136
+ ## Roadmap completion contract
137
+
138
+ This section is the control plane for completing this roadmap with a persistent
139
+ goal prompt. The execution-order table remains the full historical ledger. The
140
+ owner split remaining work on 2026-08-02 into Track A, active feature work whose
141
+ acceptance evidence is local and checked in, and Track B, production-release and
142
+ trusted-distribution work parked until its external evidence arrives. Parking
143
+ does not close, weaken, or simulate any gate.
144
+
145
+ Run the roadmap loop with:
146
+
147
+ ```text
148
+ Complete Track A in roadmap/competitive-roadmap.md in dependency order. Preserve its
149
+ product constraint, evidence policy, limitations, user-owned changes, and
150
+ checked-in evidence. Work autonomously on every safe local step. For each item,
151
+ run its acceptance gates and the repository verification bar, update its Status
152
+ and evidence truthfully, and remove it from Track A only when its
153
+ terminal condition is met. Never turn an evaluation into an implementation or
154
+ a superiority claim unless its gate authorizes that outcome. When an external
155
+ credential, production channel, independent custodian, or platform is required,
156
+ route the item to Track B only by an explicit owner decision; do not simulate
157
+ production evidence. Continue until `npm run check:roadmap -- --complete`
158
+ passes. That command proves Track A completion only and reports Track B parked;
159
+ it is not a production-release-readiness claim.
160
+ ```
161
+
162
+ ### Status and closure rules
163
+
164
+ - `implemented`, `done`, `closed`, `certified`, `evaluated — deferred`,
165
+ `evaluated — rejected`, and `won't pursue` are terminal only when the item
166
+ links checked-in evidence satisfying its acceptance or evaluation gate.
167
+ - `proposed`, `active`, `permission-gated`, and `blocked` are non-terminal.
168
+ - An `Evaluate` item closes by recording one allowed disposition: retain the
169
+ current product unchanged, defer/reject the idea with evidence, or create a
170
+ separately numbered implementation item. Evaluation never silently expands
171
+ production scope.
172
+ - P0/P1/P2/P3 affect ordering, not whether an active item must be resolved.
173
+ Track A completion does not close Track B or the entire roadmap.
174
+ - External work is fail-closed. A missing credential, custodian, production
175
+ channel, Windows host, or approval may remain parked with exact evidence; it
176
+ may not produce a synthetic pass. Parked rows remain in Track B.
177
+ - Proposals in ADRs, landscape documents, task systems, or competitor reports
178
+ are outside this roadmap's completion boundary until assigned a C-number in
179
+ the execution-order table. Repository applicability signals are C79 and ADR
180
+ 006 items 1–7 are promoted intact as C80–C86;
181
+ later ADR 006 proposals remain outside the boundary until separately approved.
182
+ - Closing an item requires updating its item-level `Status`, evidence links,
183
+ limitations, the execution-order disposition if needed, and its ledger in
184
+ the same change. Do not rewrite immutable raw benchmark artifacts.
185
+ - `npm run check:roadmap -- --publish-ready` is the ordinary 0.x publication
186
+ guard. It requires Track A to be empty while retaining and reporting Track B.
187
+ The owner clarified on 2026-08-03 that early Docusign usage is a rollout
188
+ cohort, not a separate artifact channel; ordinary releases may publish while
189
+ Track B remains parked. `npm run check:roadmap -- --release-ready` remains the
190
+ stricter GA/trusted-distribution guard and fails unless both tracks are empty.
191
+ - While Track B remains open, documentation and positioning must not claim
192
+ cross-platform certification, trusted distribution, or production-release
193
+ readiness. All existing qualifications and recorded limitations remain in
194
+ force.
195
+
196
+ ### Track A — Active local-work ledger
197
+
198
+ Track A contains only feature or local-evidence work whose acceptance can be
199
+ proved by checked-in tests and replays. Rows are ordered by dependency and then
200
+ priority. The owner approved repository applicability signals as the top item
201
+ on 2026-08-03, followed by ADR 006 items 1–7 intact as C80–C86;
202
+ later ADR proposals remain outside the completion boundary until separately
203
+ approved and numbered.
204
+
205
+ | ID | State | Previous state | Depends on | Exit evidence |
206
+ |---|---|---|---|---|
207
+
208
+ ### Track B — Parked external-evidence ledger
209
+
210
+ Track B parked (10 items, external). The owner parked C60–C67 on 2026-08-02,
211
+ parked C91 after its code-readiness increment on 2026-08-05, and parked C92
212
+ and C94 on 2026-08-06 because their acceptance depends on C91 evidence.
213
+ Every former state, dependency, exit condition, acceptance criterion, and
214
+ limitation remains binding. Re-enter this track when either the documented C60
215
+ packet or C62 packet arrives; completing one packet does not waive the other.
216
+ The durable handoff index is
217
+ `docs/evidence/handoff/README.md`, the C60 packet is
218
+ `docs/evidence/c60-autonomous-preparation-2026-08-02.md`, and the C62 packet is
219
+ `docs/ROOT-CEREMONY-RUNBOOK.md`.
220
+
221
+ | ID | State | Previous state | Depends on | Exit evidence |
222
+ |---|---|---|---|---|
223
+ | C60 | parked-external-evidence — 2026-08-02 | active | C52, C55 | Mirror automation, timed onboarding, and Linux/Windows certification evidence |
224
+ | C62 | parked-external-evidence — 2026-08-02 | active | C55 | Offline root ceremony, independent pin review, and recovery receipt |
225
+ | C63 | parked-external-evidence — 2026-08-02 | active | C62 | Production activation/rollback evidence or explicit safe deferred policy |
226
+ | C64 | parked-external-evidence — 2026-08-02 | active | C62 | Verified release-candidate provenance/SBOM bundles and attestation draft |
227
+ | C65 | parked-external-evidence — 2026-08-02 | active | C63, C64 | Production ceremony and channel/runner/manager defense drills |
228
+ | C66 | parked-external-evidence — 2026-08-02 | active | C56, C57, C59, C60, C61, C65, C68, C75 | Terminal scorecard refresh linked to closed dependencies |
229
+ | C67 | parked-external-evidence — 2026-08-02 | proposed | C60, C62, C63, C64, C65, C66, C75 | Certified release record and distribution receipts |
230
+ | C91 | parked-external-evidence — 2026-08-05 | owner-approved adoption priority | C52, C73, C89, C90 | Native macOS manager/client/Homebrew/Artifactory receipts plus transactional doctor repairs |
231
+ | C92 | parked-external-evidence — 2026-08-06 | owner-approved release priority | C64, C91 | Dry-run preflight, retry replay, immutable artifact proof, and receipt schema |
232
+ | C94 | parked-external-evidence — 2026-08-06 | owner-approved clarity priority | C89, C90, C91 | First-use copy tests and a checked-in five-minute real-repository demonstration |
233
+
234
+ ### Goal-loop ordering
235
+
236
+ 1. Keep Track A active work independent of release authority and prove every
237
+ item with local, checked-in acceptance evidence.
238
+ 2. Re-enter Track B when either C60 or C62's documented evidence packet arrives.
239
+ Execute C60 and C62 in parallel where people and platforms permit.
240
+ 3. After C62, produce C64's release-candidate artifacts. Close C63 with its
241
+ fail-closed deferred-activation policy, then use that terminal client as the
242
+ subject of C65's adversarial drills. C65, not C63, authorizes any later
243
+ production activation.
244
+ 4. Produce C66's terminal prerelease snapshot only from terminal dependency
245
+ evidence. C67 may append release receipts without reopening C66.
246
+ 5. Execute C67 last. GA and trusted-distribution claims remain prohibited until
247
+ Track B is empty and `npm run check:roadmap -- --release-ready` passes.
248
+ Ordinary 0.x publication requires `npm run check:roadmap -- --publish-ready`.
249
+ 6. Keep C92 and C94 parked behind C91. Re-enter C91 only when the named native macOS
250
+ environments and credentialed channels are available and the transactional
251
+ doctor repair contract is implemented; deterministic fixtures alone do not
252
+ satisfy certification.
253
+
254
+ ---
255
+
256
+ ## C1 — Declare a permissive license
257
+
258
+ - Status: implemented
259
+ - Priority: P0
260
+ - Disposition: Must close
261
+ - Evidence: the GitNexus bake-off found GitNexus explicitly licensed under
262
+ PolyForm Noncommercial 1.0.0 while knodin has neither a `LICENSE` file nor a
263
+ `package.json` license field. “Local and free” is not a legally supportable
264
+ marketing claim until this is resolved.
265
+ - Touches: `LICENSE`, `package.json`, `README.md`, `docs/COMPARISON.md`
266
+
267
+ ### Acceptance criteria
268
+
269
+ 1. The repository owner selects and adds an OSI-approved permissive license;
270
+ implementation must not guess the owner's legal choice.
271
+ 2. `package.json` declares the matching SPDX identifier.
272
+ 3. README and comparison copy use only claims permitted by that license.
273
+ 4. Package contents include the license after `bun pm pack --dry-run` or the
274
+ repository's equivalent packaging check.
275
+
276
+ ---
277
+
278
+ ## C2 — Stable symbol identity and ambiguity-safe resolution
279
+
280
+ - Status: implemented
281
+ - Priority: P0
282
+ - Disposition: Must close
283
+ - Evidence: GitNexus's `trace(main, createEngine)` returned three ranked `main`
284
+ candidates and required disambiguation. knodin silently returned
285
+ `main → createEngine`, with no indication of which definition it selected.
286
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
287
+ symbol/query/explain/rename/path tests, persisted schema if IDs must survive
288
+ rebuilds
289
+
290
+ ### Required behavior
291
+
292
+ 1. Every symbol-facing result exposes a stable repository-scoped identity.
293
+ 2. `explain`, structured queries, rename preview/apply, and shortest path accept
294
+ optional identity, file, and kind selectors.
295
+ 3. A bare-name request with multiple viable definitions returns an explicit
296
+ ambiguity result with ranked candidates; it never silently merges or chooses
297
+ definitions when that can change the answer.
298
+ 4. Existing unambiguous bare-name calls remain source-compatible.
299
+
300
+ ### Acceptance criteria
301
+
302
+ 1. A fixture with at least three same-named functions reproduces the GitNexus
303
+ `main` case and proves no conflated path is returned.
304
+ 2. Selecting each candidate by identity returns only that candidate's real
305
+ callers, callees, impact, and paths.
306
+ 3. Rename by identity edits only the selected definition and its resolved
307
+ references.
308
+ 4. Migration/rebuild behavior for persisted IDs is deterministic and tested.
309
+
310
+ ---
311
+
312
+ ## C3 — Enforce truthful response budgets and minimal modes
313
+
314
+ - Status: implemented
315
+ - Priority: P0
316
+ - Disposition: Must close
317
+ - Evidence: knodin's minimal symbol context still returned 22,843 bytes because
318
+ source was always present, and minimal `map` returned 254,102 bytes. GitNexus
319
+ was usually more compact when source was not requested. Conversely,
320
+ GitNexus's source-inclusive query returned 606,235 bytes, demonstrating why a
321
+ hard cap—not convention—is required.
322
+ - Reinforced by Graphify's 417-byte top-10 hubs versus knodin's 258,628-byte
323
+ mapped response, Aider's 425-byte mean at a 128-token budget versus knodin's
324
+ roughly 190 KB mapped mean, and Repomix/code2prompt's explicit token metrics.
325
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, response-size and
326
+ operation tests, docs
327
+
328
+ ### Required behavior
329
+
330
+ 1. Every operation has a tested serialized-byte ceiling and reports truncation
331
+ plus total counts when the ceiling binds.
332
+ 2. Minimal `explain` omits source unless explicitly requested; standard explain
333
+ retains edit-ready source within its cap.
334
+ 3. Minimal `map` returns counts, aggregates, and bounded top-N summaries only;
335
+ it never leaks full community member arrays or full edge lists.
336
+ 4. Caps are applied after serialization-aware accounting so nested content
337
+ cannot bypass them.
338
+ 5. Callers may choose token, byte, and item budgets; all three produce honest
339
+ totals, truncation state, and a continuation/drill-down route.
340
+ 6. Hub, bridge, community, traversal, search, source, and export operations honor
341
+ their requested top-N or budget without constructing a full map response.
342
+
343
+ ### Acceptance criteria
344
+
345
+ 1. The GitNexus benchmark's `context-name-min` and `cypher-communities`
346
+ counterparts remain below documented budgets.
347
+ 2. Large-symbol and large-map fixtures assert both byte ceilings and honest
348
+ `truncated`/total-count metadata.
349
+ 3. No standard response silently loses edit-critical source without a clear
350
+ continuation or drill-down route.
351
+
352
+ ---
353
+
354
+ ## C4 — Symbol-level directional impact analysis
355
+
356
+ - Status: implemented
357
+ - Priority: P1
358
+ - Disposition: Must close
359
+ - DependsOn: C2, C3
360
+ - Evidence: GitNexus distinguished upstream/downstream/both, depths 1/3/6,
361
+ tests, relation confidence, processes, and modules for `createEngine`.
362
+ knodin's closest public impact query was file-level and returned the same
363
+ broad 31-row set for every directional comparison.
364
+ - Reinforced by codebase-memory-mcp's directional typed trace and argument-level
365
+ data-flow evidence, and Graphify's relation-filtered direct neighbors.
366
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
367
+ impact/query tests and docs
368
+
369
+ ### Required behavior
370
+
371
+ 1. Impact accepts a stable symbol selector or an explicit file-level mode.
372
+ 2. Symbol mode supports direction, maximum depth, relation kinds, confidence
373
+ threshold, test inclusion, result limit, argument/data-flow evidence, and
374
+ compact summary.
375
+ 3. Results preserve depth and edge evidence and identify affected flows and
376
+ communities without presenting heuristic reach as exact.
377
+ 4. File-level impact remains available and is labeled distinctly.
378
+
379
+ ### Acceptance criteria
380
+
381
+ 1. The `createEngine` fixture produces meaningfully different upstream and
382
+ downstream results.
383
+ 2. Depth limits are monotonic and enforced; confidence/relation filters have
384
+ targeted fixtures.
385
+ 3. Ambiguous symbols use C2's candidate response rather than conflated reach.
386
+ 4. Compact output satisfies C3's byte budget.
387
+
388
+ ---
389
+
390
+ ## C5 — Explicit git diff scopes
391
+
392
+ - Status: implemented
393
+ - Priority: P1
394
+ - Disposition: Should close
395
+ - DependsOn: C3
396
+ - Evidence: GitNexus supports `unstaged`, `staged`, `all`, and `compare`.
397
+ knodin cannot request staged-only review, and its result can omit changed
398
+ files when no indexed symbol maps to the diff.
399
+ - Reinforced by code-review-graph's caller-supplied changed-file review and
400
+ code2prompt's working-tree diff, branch diff, and branch log exports.
401
+ - Touches: git diff parsing, engine review contract, MCP/CLI schemas, review
402
+ tests and docs
403
+
404
+ ### Acceptance criteria
405
+
406
+ 1. Review accepts `unstaged|staged|all|compare`; compare requires or defaults a
407
+ documented base ref.
408
+ 2. Each scope has a fixture containing different staged and unstaged edits.
409
+ 3. Results always report changed files, including non-code/unindexed files,
410
+ separately from mapped changed symbols.
411
+ 4. Existing `base` behavior has a documented compatibility mapping.
412
+ 5. Review accepts an explicit file list or revision pair without requiring the
413
+ checkout to contain an equivalent live diff.
414
+
415
+ ---
416
+
417
+ ## C6 — Deterministic import-cycle analysis
418
+
419
+ - Status: implemented
420
+ - Priority: P1
421
+ - Disposition: Should close
422
+ - DependsOn: C2
423
+ - Evidence: GitNexus exposes a bounded read-only cycle check; knodin has no
424
+ direct equivalent despite already storing import dependencies.
425
+ - Touches: `src/engine/index.ts`, query schema/docs, cycle fixtures
426
+
427
+ ### Acceptance criteria
428
+
429
+ 1. A structured query returns canonical deterministic file-import cycles.
430
+ 2. Equivalent rotations of one cycle are deduplicated.
431
+ 3. Limits and truncation metadata prevent exponential output.
432
+ 4. Fixtures cover no cycle, one cycle, overlapping cycles, and type-only/import
433
+ edge policy.
434
+
435
+ ---
436
+
437
+ ## C7 — MCP tool and handler mapping
438
+
439
+ - Status: implemented
440
+ - Priority: P2
441
+ - Disposition: Should close
442
+ - DependsOn: C2
443
+ - Evidence: GitNexus correctly detected knodin's `knodin` MCP tool and its
444
+ source file. knodin has endpoint patterns but no MCP registration-to-handler
445
+ map.
446
+ - Touches: language extraction, schema, structured queries, MCP fixtures/docs
447
+
448
+ ### Acceptance criteria
449
+
450
+ 1. Index common TypeScript MCP SDK registration forms and associate tool name,
451
+ description, schema declaration, handler symbol, and file.
452
+ 2. A query supports all tools and lookup by tool name.
453
+ 3. Dynamic/unresolved registrations are labeled heuristic rather than exact.
454
+ 4. The repository's own `knodin` registration is found without matching prose
455
+ examples or benchmark fixture strings.
456
+
457
+ ---
458
+
459
+ ## C8 — On-demand local statement-level flow analysis
460
+
461
+ - Status: implemented (bounded on-demand analysis; no persisted PDG)
462
+ - Priority: P2
463
+ - Disposition: Evaluate
464
+ - DependsOn: C2, C3
465
+ - Evidence: GitNexus's opt-in local PDG returned concrete control-dependence and
466
+ reaching-definition rows, including variable-filtered `db` flows. knodin has
467
+ no statement-level answer. GitNexus also produced large 14–42 KB PDG
468
+ responses, so copying its output contract would conflict with knodin's
469
+ compactness goal.
470
+ - Touches if approved: parser/extractor architecture, persisted schema,
471
+ performance harness, query surface, multi-language fixtures
472
+ - Evaluation evidence: [`C8 summary and methodology`](../benchmarks/evaluations/c8-pdg/summary-methodology.md)
473
+ - Decision: the disposable prototype passed the mechanical gate at 100% labeled
474
+ fixture precision, with three real repositories measured and responses below
475
+ the C3 ceiling. The supported implementation is the narrower on-demand
476
+ `flow_analysis` query for one selected TS/JS symbol, with optional variable
477
+ filtering, compact source evidence, C3 budgets, and explicit unresolved cases;
478
+ it does not persist a broad PDG.
479
+
480
+ ### Evaluation gate
481
+
482
+ 1. Prototype TS/JS control-flow and reaching-definitions on disposable code,
483
+ not the production schema.
484
+ 2. Measure clean-index time, incremental-index time, database growth, peak RSS,
485
+ query latency, precision, and response size against at least three real
486
+ repositories.
487
+ 3. Require ≥90% precision on a labeled control/data-flow fixture and a bounded
488
+ response below C3's ceiling.
489
+ 4. Approve implementation only if the value cannot be delivered more cheaply
490
+ through compiler/LSP facts or a narrower on-demand analysis.
491
+
492
+ ---
493
+
494
+ ## C9 — Evaluate API contract and response-shape analysis
495
+
496
+ - Status: implemented
497
+ - Priority: P2
498
+ - Disposition: Evaluate
499
+ - DependsOn: C2, C3
500
+ - Evidence: GitNexus exposes route maps, response-shape checking, and API impact,
501
+ but this repository had zero detected HTTP routes. The bake-off proves the
502
+ surface exists, not that it is correct or useful. knodin now persists
503
+ source-evidenced literal Express/Fastify route and fetch/Axios client facts,
504
+ with a bounded `api_contract_mismatches` query.
505
+ - Touches: route extraction, schema/shape model, cross-file/client usage
506
+ resolution, fixtures and query surface
507
+ - Evaluation evidence: [`C9 summary`](../benchmarks/evaluations/c9-api/summary.md)
508
+ - Decision: the evaluation gate passed. GitNexus found the actionable `profile`
509
+ response mismatch that current handler, endpoint, and impact queries missed.
510
+ The production implementation shipped after the checked-in fixture produced
511
+ 3 true positives, 0 false positives, and 0 false negatives (precision 1.00,
512
+ recall 1.00); dynamic routes remain unresolved and absolute-origin clients do
513
+ not link to local endpoints. See [`C9 production acceptance`](../benchmarks/evaluations/c9-api/production-acceptance.md).
514
+
515
+ ### Evaluation gate
516
+
517
+ 1. Build a local fixture with multiple verbs on one route, middleware, typed and
518
+ inferred response bodies, internal clients, mismatches, and dynamic routes.
519
+ 2. Run GitNexus and knodin endpoint primitives against the same fixture and
520
+ manually label precision/recall.
521
+ 3. Approve implementation only if shape analysis finds actionable breakage that
522
+ current handler/endpoint/impact queries miss.
523
+
524
+ ---
525
+
526
+ ## C10 — Evaluate a bounded expert graph-query DSL
527
+
528
+ - Status: closed; general DSL will not be pursued
529
+ - Priority: P3
530
+ - Disposition: Evaluate
531
+ - DependsOn: C2, C3
532
+ - Evidence: GitNexus's parameterized Cypher expressed caller, community
533
+ aggregate, and custom dead-code queries outside a fixed enum. knodin's fixed
534
+ patterns are easier to route and safer but cannot answer arbitrary structural
535
+ questions. GitNexus's naive custom dead-code query also returned many false
536
+ positives, illustrating the cost of unconstrained power.
537
+ - Touches if reopened: parser/validator, query planner, response budget,
538
+ read-only safety tests and docs
539
+ - Evaluation evidence: [`C10 summary`](../benchmarks/evaluations/c10-dsl/summary.md)
540
+ - Decision: no DSL is justified in this roadmap slice. Three bounded fixed-pattern
541
+ candidates covered 12 of 13 provenance-backed questions (92.3%), so the DSL
542
+ prototype gate is a no-go.
543
+
544
+ ### Evaluation gate
545
+
546
+ 1. Collect at least ten real questions that existing patterns cannot answer.
547
+ 2. Determine whether two or three new fixed patterns cover most demand.
548
+ 3. If not, design a read-only bounded DSL with explicit node/edge allowlists,
549
+ mandatory limits, traversal-depth caps, and no raw SQL/Cypher execution.
550
+ 4. Approve only with deterministic cost bounds and C3-compliant output.
551
+
552
+ ---
553
+
554
+ ## C11 — Call/reference precision and structural identity
555
+
556
+ - Status: implemented
557
+ - Priority: P0
558
+ - Disposition: Must close
559
+ - DependsOn: C2
560
+ - Evidence: Graphify found four credible call neighbors for `createEngine` while
561
+ knodin included a known name-collision false positive; Serena returned precise
562
+ LSP references and structural `KnodinEngine` implementations knodin's nominal
563
+ inheritance query missed. grepai and codebase-memory-mcp also demonstrated
564
+ that imports/files must not be mislabeled as callers.
565
+
566
+ ### Acceptance criteria
567
+
568
+ 1. Labeled fixtures distinguish call, import, usage, containment, inheritance,
569
+ and TypeScript structural implementation edges.
570
+ 2. Callers/callees reach at least 95% precision and recall on those fixtures;
571
+ built-ins and same-name symbols never create cross-definition edges.
572
+ 3. Structural implementations are queryable separately from nominal inheritors.
573
+ 4. Tests cover top-level calls, aliases, barrels, object-literal interface
574
+ implementations, scripts, tests, and ambiguous names.
575
+
576
+ ---
577
+
578
+ ## C12 — Uniform filters, pagination, and compact drill-downs
579
+
580
+ - Status: implemented
581
+ - Priority: P1
582
+ - Disposition: Should close
583
+ - DependsOn: C2, C3
584
+ - Evidence: code-review-graph filters semantic kind, large-code threshold/kind,
585
+ flow/community sort and detail; codebase-memory-mcp adds label, qualified name,
586
+ file, relationship, degree, total/offset pagination; Claude Context filters
587
+ extensions. Graphify honors top-N hubs and relation filters.
588
+
589
+ ### Acceptance criteria
590
+
591
+ 1. Search accepts language/extension, kind, path, test/production, and source
592
+ inclusion filters plus offset/cursor pagination with total and `hasMore`.
593
+ 2. Large-code queries accept explicit line/complexity thresholds, symbol kind,
594
+ and path; no requested parameter is silently ignored.
595
+ 3. Hub, bridge, community, flow, and traversal drill-downs honor top-N, sort,
596
+ relation, and detail controls under C3.
597
+ 4. Parameter-combination tests prove filters compose rather than overwrite one
598
+ another.
599
+
600
+ ---
601
+
602
+ ## C13 — Typed traversal and architecture facets
603
+
604
+ - Status: implemented
605
+ - Priority: P1
606
+ - Disposition: Should close
607
+ - DependsOn: C2, C3, C11
608
+ - Evidence: codebase-memory-mcp beat knodin on inbound/outbound/both traces,
609
+ edge types, argument expressions, tests, risk labels, packages, layers,
610
+ boundaries, hotspots, and path scoping. Graphify adds direct relation-filtered
611
+ neighbors.
612
+
613
+ ### Acceptance criteria
614
+
615
+ 1. Traversal supports direction and explicit edge types with per-hop evidence.
616
+ 2. Call/data-flow results preserve argument expressions where extraction is
617
+ compiler-grounded and label heuristic facts honestly.
618
+ 3. Architecture exposes independently selectable packages, layers, boundaries,
619
+ hotspots, entry points, languages, and path scope.
620
+ 4. Results satisfy C11 precision fixtures and C3 budgets.
621
+
622
+ ---
623
+
624
+ ## C14 — Portable budgeted context export
625
+
626
+ - Status: implemented
627
+ - Priority: P2
628
+ - Disposition: Should close
629
+ - DependsOn: C3, C5
630
+ - Evidence: Repomix, code2prompt, and Aider provide portable source-shaped
631
+ context, deterministic formatting, token accounting, file policies, ranged
632
+ retrieval, mention/chat awareness, and Git context. These are complementary
633
+ to knodin's graph rather than replacements for it.
634
+
635
+ ### Acceptance criteria
636
+
637
+ 1. Export supports Markdown, JSON, and XML with relative paths, optional line
638
+ numbers/tree, deterministic ordering, and explicit tokenizer cost estimate.
639
+ 2. Include/exclude rules and per-file `full|summary|structure-only` policies are
640
+ composable with a hard token/byte budget.
641
+ 3. Already-present/chat files can be excluded from duplicated source context.
642
+ 4. Optional diff/log sections use C5 scopes; packed artifacts support bounded
643
+ ranged reads and exact regex grep.
644
+
645
+ ---
646
+
647
+ ## C15 — Evaluate LSP diagnostics and guarded symbol editing
648
+
649
+ - Status: implemented (read-only TypeScript language-service queries)
650
+ - Priority: P2
651
+ - Disposition: Evaluate
652
+ - DependsOn: C2, C11
653
+ - Evidence: Serena successfully exposed diagnostics, declarations,
654
+ implementations, symbol-body replacement, before/after insertion, safe delete,
655
+ and guarded multi-file replacement. knodin's verified rename is safer but much
656
+ narrower. The 2026-07-22 disposable evaluation exercised TypeScript Language
657
+ Server, Pyright, JDTLS, Java/Javac, .NET compilation, and the official Roslyn
658
+ C# Language Server with guarded TypeScript previews and rollback. The final
659
+ raw result records a C# `textDocument/diagnostic` response and
660
+ `textDocument/implementation` navigation to the fixture's line-10
661
+ implementation; the passing compiler check remains corroboration rather than a
662
+ substitute for LSP evidence. knodin now ships optional local TypeScript
663
+ diagnostics, definition, declaration, and implementation queries with no
664
+ daemon, subprocess, workspace mutation, or index initialization. Missing or
665
+ unsupported adapters return explicit unavailable results; guarded editing and
666
+ cross-language mutation remain unapproved.
667
+
668
+ ### Evaluation gate
669
+
670
+ 1. Prototype a local TypeScript/LSP adapter for diagnostics, implementation, and
671
+ edit previews without adding a required daemon.
672
+ 2. Require all edits to produce a preview and reuse rename's compiler check and
673
+ rollback guarantees; never expose unguarded filesystem mutation.
674
+ 3. Measure startup/RSS/latency and verify behavior across TS, Python, Java, and
675
+ C# fixtures before approving a general editing surface.
676
+
677
+ ---
678
+
679
+ ## C16 — Index health, repair, and output telemetry
680
+
681
+ - Status: implemented
682
+ - Priority: P2
683
+ - Disposition: Should close
684
+ - DependsOn: C3
685
+ - Evidence: codebase-memory-mcp exposes portable lifecycle modes and schema;
686
+ Claude Context exposes clear/status lifecycle; mcp-codebase-index statically
687
+ defines health/repair; grepai records output/token savings. knodin has strong
688
+ freshness internals but limited public health and budget telemetry.
689
+
690
+ ### Acceptance criteria
691
+
692
+ 1. Status reports schema/model/version, file and symbol coverage, orphaned or
693
+ missing records, last successful reconciliation, and actionable repair steps.
694
+ 2. A local repair command can rebuild only damaged/missing state and verify the
695
+ result without deleting healthy data.
696
+ 3. Benchmark telemetry records operation, latency, serialized bytes, estimated
697
+ tokens, truncation, and detail mode without recording source content.
698
+ 4. All telemetry is local, opt-in for persistence, and documented as such.
699
+
700
+ ---
701
+
702
+ ## C17 — Evaluate DFS and feature-path navigation
703
+
704
+ - Status: implemented
705
+ - Priority: P3
706
+ - Disposition: Evaluate
707
+ - DependsOn: C2, C3, C11
708
+ - Evidence: Graphify and code-review-graph expose DFS; grepai adds deterministic
709
+ functional feature paths. `query feature_path <symbol>` now supplies a compact,
710
+ deterministic downstream DFS over C2/C11-resolved references. It includes
711
+ file/line edge evidence, cycle guards, depth/node caps, and explicit
712
+ truncation without adding persistent identity state. The focused fixture-led
713
+ regression in `feature-path.spec.ts` proves a branching result observably
714
+ differs from BFS while remaining stable after harmless line shifts.
715
+
716
+ ### Delivered behavior
717
+
718
+ 1. `feature_path` follows only resolved downstream references in stable DFS
719
+ order; `traverse` remains the existing selectable-direction BFS neighborhood.
720
+ 2. Standard-detail non-root nodes include C11-resolved identity; every hop
721
+ includes source file/line, kind, provenance, and confidence evidence.
722
+ 3. Cycles are visited once, depth is clamped to 1–6, the item cap is enforced,
723
+ and a larger reachable graph reports `truncated: true`.
724
+ 4. The feature is source-only: dynamic dispatch, reflection, unresolved calls,
725
+ and runtime execution order remain explicitly outside the contract.
726
+
727
+ ---
728
+
729
+ ## C18 — Unified competitive regression harness
730
+
731
+ - Status: implemented
732
+ - Priority: P1
733
+ - Disposition: Should close
734
+ - Evidence: the first program produced 12 independent protocols and 713 paired
735
+ cases. Without one version-pinned entry point, improvements to C1-C17 could
736
+ silently lose coverage or be compared against a different tool installation.
737
+ - Touches: `scripts/competitive-bakeoff.ts`, `src/competitive-manifest.ts`,
738
+ `src/competitive-runner.ts`, `benchmarks/competitors/README.md`
739
+
740
+ ### Acceptance criteria
741
+
742
+ 1. One manifest selects competitors directly or by C1-C17 roadmap coverage and
743
+ records exact versions, prerequisites, preparation, commands, and blockers.
744
+ 2. Runs refuse dirty checkouts by default, validate installed versions, isolate
745
+ fixture mutations, and preserve immutable timestamp-and-commit artifacts.
746
+ 3. Heterogeneous raw results normalize into one summary and compare against the
747
+ checked-in baseline with a configurable latency tolerance and nonzero
748
+ regression exit status.
749
+ 4. List, prerequisite-check, and dry-run modes are non-mutating; commands use
750
+ argument arrays rather than shell interpolation.
751
+ 5. Unit tests cover selection, immutable history, normalization, blockers, and
752
+ material-regression detection.
753
+
754
+ ---
755
+
756
+ ## Salesforce scope boundary and remaining limitations
757
+
758
+ The baseline parses Apex and Visualforce and now includes the bounded C19–C23
759
+ implementations for LWC, selected metadata/Flow automation, Aura interop, LWR
760
+ topology, and Apex platform entry/asynchronous semantics. Their individual
761
+ sections and fixtures define the supported static evidence. This remains
762
+ partial Salesforce coverage, not proof that every deployed-org dependency or
763
+ runtime dispatch is resolved. Static dead-code candidates still require
764
+ corroboration against deployed-org and platform dependency data before deletion.
765
+
766
+ OmniStudio and industry-cloud packages are intentionally out of scope: promote
767
+ them only after a concrete, legally usable Salesforce repository demonstrates a
768
+ workflow that C19–C23 do not cover.
769
+
770
+ Their proposed shared `src/engine/index.ts` footprint is a collision, so they
771
+ must be scheduled sequentially even where their logical dependencies are
772
+ independent.
773
+
774
+ ---
775
+
776
+ ## C19 — Evaluate Salesforce LWC bundles and Apex/data bridges
777
+
778
+ - Status: implemented
779
+ - Priority: P2
780
+ - Disposition: Evaluate
781
+ - DependsOn: C2, C3, C11
782
+ - Evidence: the version-pinned, source-only Salesforce DX fixture now records
783
+ 13/13 exact bundle, component, Apex, schema, wire-adapter, and Lightning Data
784
+ Service labels (100% precision, 100% recall) in
785
+ `benchmarks/evaluations/c19-salesforce-lwc/raw-results-20260722T132842614Z.json`.
786
+ The replay copies the fixture to a temporary directory, needs no org login,
787
+ credentials, network access, or source egress, and records one-file LWC/Apex
788
+ reindex plus C3-bounded response measurements. Dynamic and namespaced forms
789
+ remain omitted rather than exact claims. The raw result hashes the evaluated
790
+ engine, runner, and label revision in addition to the fixture. This authorizes no broader Salesforce
791
+ product commitment beyond the evaluated static source edges.
792
+ - Touches if approved: `src/engine/index.ts`, `src/__tests__/unit/salesforce.spec.ts`,
793
+ `src/__tests__/unit/lwc.spec.ts` (new),
794
+ `benchmarks/evaluations/c19-salesforce-lwc/` (new),
795
+ `src/competitive-manifest.ts`
796
+
797
+ ### Evaluation gate
798
+
799
+ 1. Build a self-contained Salesforce DX fixture with LWC JavaScript, HTML, CSS,
800
+ and `*.js-meta.xml` bundle members; include component-to-component imports,
801
+ `@salesforce/apex`, `@salesforce/schema`, wire adapters, and Lightning Data
802
+ Service usage alongside resolvable Apex methods.
803
+ 2. Label bundle containment plus LWC-to-Apex, LWC-to-object/field, and
804
+ component-to-component edges. Require at least 95% precision and 90% recall;
805
+ unresolved, dynamic, or namespaced references must be omitted or labeled
806
+ heuristic rather than reported as exact.
807
+ 3. Re-index a one-file LWC and one-file Apex change, measure index/response
808
+ cost against C3, and prove the analysis needs neither an org login nor source
809
+ egress.
810
+
811
+ ---
812
+
813
+ ## C20 — Evaluate Salesforce metadata and Flow/automation model
814
+
815
+ - Status: implemented
816
+ - Priority: P2
817
+ - Disposition: Evaluate
818
+ - DependsOn: C2, C3, C11
819
+ - Evidence: the version-pinned source-only Salesforce DX fixture records 15/15
820
+ exact custom-object/field/record-type, permission-set, Flow, workflow, and
821
+ approval-process labels (100% precision and recall) in
822
+ `benchmarks/evaluations/c20-salesforce-metadata/raw-results-20260722T134732728Z.json`.
823
+ The replay copies the fixture to a temporary directory, runs a metadata-only
824
+ one-file update, stays within C3's response budget, and needs no org login,
825
+ deployment, credentials, network access, or source egress. Formula and
826
+ runtime-selected targets remain explicit unknowns rather than exact edges.
827
+ This authorizes no broader Salesforce product commitment beyond the evaluated
828
+ static source relationships.
829
+ - Touches if approved: `src/engine/index.ts`,
830
+ `src/__tests__/unit/salesforce-metadata.spec.ts` (new),
831
+ `benchmarks/evaluations/c20-salesforce-metadata/` (new),
832
+ `src/competitive-manifest.ts`
833
+
834
+ ### Evaluation gate
835
+
836
+ 1. Build a version-pinned Salesforce DX fixture containing custom objects and
837
+ fields, record types, permission sets, a record-triggered Flow, a screen or
838
+ autolaunched Flow, workflow/approval metadata, and a Flow-invoked Apex
839
+ action.
840
+ 2. Label only statically declared metadata edges: Flow/automation to Apex,
841
+ object, field, permission, and referenced UI component. Require at least
842
+ 95% precision and 90% recall; dynamic formulas and runtime-selected targets
843
+ must remain explicit unknowns.
844
+ 3. Demonstrate that a metadata-only change updates the affected topology without
845
+ requiring deployment to an org, and keep output within C3's budget.
846
+
847
+ ---
848
+
849
+ ## C21 — Evaluate Aura bundles and Visualforce/Aura/LWC interop
850
+
851
+ - Status: implemented
852
+ - Priority: P3
853
+ - Disposition: Evaluate
854
+ - DependsOn: C19
855
+ - Evidence: the version-pinned, source-only Salesforce DX fixture records 10/10
856
+ exact Aura containment, Aura-to-Apex, Aura-to-LWC, and Visualforce controller
857
+ labels (100% precision and recall) in
858
+ `benchmarks/evaluations/c21-salesforce-interop/raw-results-20260722T161801709Z.json`.
859
+ The replay copies its fixture to a temporary directory, needs no org login,
860
+ credentials, network access, or source egress, and verifies source evidence
861
+ for every claimed cross-stack edge. knodin supports these static source
862
+ relationships; `$A` runtime lookup strings remain omitted rather than exact
863
+ claims.
864
+ - Touches: `src/engine/index.ts`,
865
+ `src/__tests__/unit/visualforce.spec.ts`,
866
+ `src/__tests__/unit/salesforce-interop.spec.ts` (new),
867
+ `benchmarks/evaluations/c21-salesforce-interop/` (new),
868
+ `src/competitive-manifest.ts`
869
+
870
+ ### Evaluation gate
871
+
872
+ 1. Build a fixture with Aura component, application, interface, event, and
873
+ design members; Visualforce page/component/controller bindings; and a
874
+ supported Aura-to-LWC or Visualforce-to-Lightning bridge.
875
+ 2. Label containment and cross-stack edges, including Aura-to-Apex and the
876
+ declared bridge target. Require at least 95% precision and 90% recall, while
877
+ `$A` runtime lookup strings and unresolvable expressions remain heuristic.
878
+ 3. Confirm that unchanged Apex/Visualforce behavior retains its current test
879
+ results and that cross-stack results expose the source evidence for every
880
+ claimed edge.
881
+
882
+ ---
883
+
884
+ ## C22 — Evaluate LWR and Experience Cloud topology
885
+
886
+ - Status: implemented
887
+ - Priority: P3
888
+ - Disposition: Evaluate
889
+ - DependsOn: C19
890
+ - Evidence: the version-pinned, source-only LWR and Experience Cloud fixture
891
+ records 9/9 exact site, route, view, LWC/Aura, theme, and SVG asset labels
892
+ (100% precision and recall) in
893
+ `benchmarks/evaluations/c22-salesforce-experience/raw-results-20260722T162757428Z.json`.
894
+ The replay copies its fixture to a temporary directory, measures a route-only
895
+ reindex, stays within C3's 65,536-byte map budget, and needs no org login,
896
+ credentials, network access, source egress, server, or deployment. knodin
897
+ supports this static source topology; runtime route expressions remain
898
+ unresolved rather than fabricated.
899
+ - Touches: `src/engine/index.ts`,
900
+ `src/__tests__/unit/salesforce-experience.spec.ts` (new),
901
+ `benchmarks/evaluations/c22-salesforce-experience/` (new),
902
+ `src/competitive-manifest.ts`
903
+
904
+ ### Evaluation gate
905
+
906
+ 1. Build a source-only fixture containing LWR configuration plus the supported
907
+ Experience Cloud metadata for one site, its routes, views, theme/assets, and
908
+ declared LWC or Aura targets.
909
+ 2. Label site-to-route-to-view-to-component and static-asset relationships.
910
+ Require at least 95% precision and 90% recall; route values assembled at
911
+ runtime must be labeled unresolved rather than fabricated.
912
+ 3. Measure a route/view-only incremental update and prove that topology output
913
+ is compact, locally computed, and does not require serving or deploying the
914
+ site.
915
+
916
+ ---
917
+
918
+ ## C23 — Evaluate Apex platform-entry and asynchronous semantics
919
+
920
+ - Status: implemented
921
+ - Priority: P2
922
+ - Disposition: Evaluate
923
+ - DependsOn: C2, C3, C11
924
+ - Evidence: the version-pinned, source-only Salesforce DX fixture records 20/20
925
+ exact REST/SOAP/invocable/trigger, async callback, enqueue, subscription, and
926
+ Salesforce Function labels (100% precision and recall) in
927
+ `benchmarks/evaluations/c23-apex-platform-semantics/raw-results-20260722T140720946Z.json`.
928
+ The replay copies the fixture to a temporary directory, reindexes one Apex
929
+ file, keeps map responses under C3's 65,536-byte budget, and requires no org
930
+ login, credentials, network access, or source egress. Transaction ordering,
931
+ bulk behavior, retries, and runtime payload shape remain unknown.
932
+ - Touches if approved: `src/engine/index.ts`,
933
+ `src/__tests__/unit/salesforce.spec.ts`,
934
+ `src/__tests__/unit/apex-platform-semantics.spec.ts` (new),
935
+ `benchmarks/evaluations/c23-apex-platform-semantics/` (new),
936
+ `src/competitive-manifest.ts`
937
+
938
+ ### Evaluation gate
939
+
940
+ 1. Build a version-pinned fixture covering `@RestResource` and verb handlers,
941
+ `webService`/SOAP, `@InvocableMethod`, trigger object/event declarations,
942
+ `@future`, Queueable, Batchable, Schedulable, Platform Event or change-event
943
+ subscribers, and statically resolvable Salesforce Function invocations.
944
+ 2. Label entry-point, enqueue, callback, subscription, and invocation edges;
945
+ require at least 95% precision and 90% recall. Never infer transaction order,
946
+ bulk behavior, retries, or runtime payload shape from source alone.
947
+ 3. Validate that each result carries an exact-versus-heuristic confidence label,
948
+ stays within C3's response budget, and is produced with no org credentials or
949
+ network access.
950
+
951
+ ---
952
+
953
+ ## Competitive leadership program
954
+
955
+ The competitive audit is evidence, not a second roadmap. C24-C33 turn every
956
+ row of [`COMPETITIVE-AUDIT.md`](../benchmarks/competitors/COMPETITIVE-AUDIT.md)
957
+ into a measurable delivery contract. A row is complete only when knodin meets
958
+ or exceeds the stated criterion against the pinned competitor and shared
959
+ fixture. A missing local prerequisite, incompatible license, hosted-service
960
+ requirement, or resource failure is a reproducible blocker, never a pass.
961
+
962
+ All replays must record the knodin commit, competitor version, command,
963
+ fixture revision, warm/cold mode, requested scope, response budget, and oracle
964
+ result. Correctness wins require a labeled oracle; latency, response size, or
965
+ feature presence alone cannot establish a win. A replay that proves a loss
966
+ must add a follow-on implementation item before the parent may close.
967
+
968
+ ### Product-claim gate
969
+
970
+ Every roadmap item that changes a user-facing claim must state and verify all
971
+ four parts of the adoption case:
972
+
973
+ 1. **Value:** the named developer or agent workflow and the concrete failure or
974
+ delay it removes.
975
+ 2. **ROI:** a labeled correctness, latency, size, safety, or effort measure
976
+ with its baseline and measurement window; do not substitute a vendor claim
977
+ or feature count.
978
+ 3. **Tomorrow:** a local, credential-free adoption path that is compatible with
979
+ existing Git, editor, and test workflows, including explicit unavailable or
980
+ degraded states.
981
+ 4. **Defensibility:** the evidence, integration, or outcome that remains hard
982
+ to substitute after the underlying model capability becomes commonplace.
983
+
984
+ “Uses AI” and an unmeasured capability are never roadmap completion signals.
985
+ An item without a measurable adoption case remains evaluation-only; a replay
986
+ that finds a quality gap updates the roadmap and prioritizes the smallest
987
+ workflow correction before any new surface area is added.
988
+
989
+ | Audit row | Leadership item | Win target |
990
+ |---|---|---|
991
+ | Exact symbol identity and ambiguity | C25 | Correct selected identity or explicit ambiguity for every labeled duplicate-name case |
992
+ | Caller/callee/reference accuracy | C25 | ≥95% precision and recall; zero cross-definition false positives |
993
+ | Diff review and change scope | C26 | Exact changed-file/symbol oracle across each supported scope |
994
+ | Blast radius and traversal | C26 | Oracle-checked typed edges, depth, filters, and truncation |
995
+ | Architecture and communities | C27 | Oracle-checked topology plus deterministic bounded drill-downs |
996
+ | Dead code | C28 | Reported precision/recall and no unlabelled heuristic claim |
997
+ | Semantic orientation/search | C28 | Frozen relevance judgments with reported nDCG and recall@k |
998
+ | Code/context packing | C29 | Exact inclusion/range/token oracle under equivalent budgets |
999
+ | Source editing and diagnostics | C30 | Preview, compile/test, and rollback oracle on disposable fixtures |
1000
+ | Resource/index lifecycle | C31 | Comparable cold/warm/repair/RSS/disk evidence with successful recovery oracle |
1001
+ | API/PDG/specialized analysis | C32 | Precision/recall and bounded source-evidence oracle |
1002
+ | Local privacy and schema economy | C33 | Zero required credentials/egress/hosted service and one top-level MCP tool |
1003
+
1004
+ ### C24 — Competitive correctness contract and replay baseline
1005
+
1006
+ - Status: implemented
1007
+ - Priority: P0
1008
+ - Disposition: Must close
1009
+ - DependsOn: C18
1010
+ - Evidence: the audit records 713 paired cases but explicitly says that most
1011
+ current comparisons lack a labeled correctness oracle. Existing label sets
1012
+ under `benchmarks/evaluations/` are fragmented and are not represented in
1013
+ the normalized competitive-run summary.
1014
+ - Competitors: all locally runnable entries in `COMPETITIVE_MANIFEST`; keep
1015
+ `mcp-codebase-index` blocked while it requires hosted Gemini embeddings.
1016
+ - Touches: `src/competitive-runner.ts`, `src/competitive-manifest.ts`,
1017
+ `scripts/competitive-bakeoff.ts`,
1018
+ `src/__tests__/unit/competitive-runner.spec.ts`,
1019
+ `benchmarks/evaluations/`, `benchmarks/competitors/COMPETITIVE-AUDIT.md`,
1020
+ `benchmarks/competitors/SYNTHESIS.md`, this roadmap
1021
+ - What: normalize competitor cases into a version-pinned, oracle-aware replay
1022
+ contract and reconcile every historical audit claim with the current code.
1023
+
1024
+ ### Acceptance criteria
1025
+
1026
+ 1. The normalized summary separately records correctness-oracle status,
1027
+ exact-match/precision/recall/false-positive results where applicable, and
1028
+ performance/resource measurements; unavailable fields remain explicit nulls.
1029
+ 2. Every audit row has a fixture inventory that either reuses a checked-in
1030
+ labeled fixture or identifies the missing fixture as a blocked prerequisite.
1031
+ 3. Every competitor run records pinned version, command, fixture revision,
1032
+ mode, budgets, and reproducible blockers; a partial run cannot be reported
1033
+ as a competitive win.
1034
+ 4. The audit labels all existing rows `implemented—needs replay`, `unverified`,
1035
+ `demonstrated gap`, or `non-goal` instead of treating historical results as
1036
+ current product state.
1037
+
1038
+ ### C25 — Prove symbol identity and reference precision leadership
1039
+
1040
+ - Status: implemented
1041
+ - Priority: P0
1042
+ - Disposition: Must close
1043
+ - DependsOn: C24, C2, C11
1044
+ - Competitors: GitNexus, Serena, codebase-memory-mcp, CodeGraph
1045
+ - Touches: `benchmarks/evaluations/competitive-symbols/`,
1046
+ `src/__tests__/unit/symbol-identity.spec.ts`,
1047
+ `src/__tests__/unit/reference-precision.spec.ts`,
1048
+ `src/competitive-manifest.ts`, relevant competitor bake-off scripts,
1049
+ `src/engine/index.ts` and `src/tools/knodin-tools.ts` only if replay fails
1050
+ - What: make exact target selection, ambiguity safety, and callers/callees/
1051
+ references objectively comparable on shared duplicate-name and structural
1052
+ implementation fixtures.
1053
+ - Replay evidence:
1054
+ `benchmarks/evaluations/competitive-symbols/raw-results-20260723.json`.
1055
+ knodin achieved 100% precision and recall with zero cross-definition false
1056
+ positives under the shared response budget; every pinned competitor completed
1057
+ against the same fixture and its oracle score is preserved.
1058
+
1059
+ ### Acceptance criteria
1060
+
1061
+ 1. A labeled fixture covers duplicate names, aliases, barrels, imports versus
1062
+ calls, structural implementations, tests, and at least one unresolved case.
1063
+ 2. knodin returns the selected stable identity or an explicit ambiguity result;
1064
+ it never silently conflates viable candidates.
1065
+ 3. knodin meets at least 95% precision and recall for labeled resolved
1066
+ call/reference edges, with zero cross-definition false positives.
1067
+ 4. The pinned competitors and knodin run against the same fixture and budget;
1068
+ a win requires oracle parity or superiority, not a smaller response.
1069
+
1070
+ ### C26 — Prove diff-review and directional traversal leadership
1071
+
1072
+ - Status: implemented
1073
+ - Priority: P1
1074
+ - Disposition: Must close
1075
+ - DependsOn: C24, C4, C5, C13
1076
+ - Competitors: GitNexus, code-review-graph, Graphify, codebase-memory-mcp
1077
+ - Touches: `benchmarks/evaluations/competitive-review/`,
1078
+ `src/__tests__/unit/review-scopes.spec.ts`,
1079
+ `src/__tests__/unit/impact.spec.ts`,
1080
+ `src/__tests__/unit/typed-traversal.spec.ts`, relevant bake-off scripts,
1081
+ engine/tool/CLI code only if replay fails
1082
+ - What: compare staged, unstaged, untracked, and revision-pair review results
1083
+ plus typed directional traversal using source-evidenced expected edges.
1084
+ - Replay evidence:
1085
+ `benchmarks/evaluations/competitive-review/raw-results-20260723.json`.
1086
+ knodin matched every diff-scope, unindexed-file, directional traversal, and
1087
+ edge-evidence oracle under the frozen budget. GitNexus reached traversal
1088
+ parity; Graphify and code-review-graph did not, and the unavailable
1089
+ codebase-memory query adapter remains an explicit blocker rather than a win.
1090
+
1091
+ ### Acceptance criteria
1092
+
1093
+ 1. The fixture labels changed files, mapped symbols, unindexed files, direct
1094
+ and transitive typed edges, depth, and expected truncation.
1095
+ 2. knodin supports every documented diff scope and reports all changed files
1096
+ independently of index coverage.
1097
+ 3. Upstream/downstream/both traversal obeys selected relation, confidence,
1098
+ depth, test, and budget controls with oracle-checked edge evidence.
1099
+ 4. A replay declares success only when knodin meets the oracle and normalized
1100
+ budget; any competitor prerequisite failure is preserved as a blocker.
1101
+
1102
+ ### C27 — Prove architecture and community-analysis leadership
1103
+
1104
+ - Status: implemented
1105
+ - Priority: P1
1106
+ - Disposition: Must close
1107
+ - DependsOn: C24, C3, C12, C13
1108
+ - Competitors: GitNexus, Graphify, codebase-memory-mcp, code-review-graph
1109
+ - Touches: `benchmarks/evaluations/competitive-architecture/`,
1110
+ `src/__tests__/unit/map.spec.ts`, `src/__tests__/unit/uniform-filters.spec.ts`,
1111
+ `src/__tests__/unit/typed-traversal.spec.ts`, relevant bake-off scripts,
1112
+ engine/tool code only if replay fails
1113
+ - What: test boundaries, layers, packages, hubs, bridges, communities,
1114
+ relation filters, pagination, and compact drill-downs against a labeled
1115
+ module-topology fixture.
1116
+ - Replay evidence:
1117
+ `benchmarks/evaluations/competitive-architecture/raw-results-20260723.json`.
1118
+ knodin met the labeled topology oracle, deterministic pagination, and C3
1119
+ budget. Pinned GitNexus, Graphify, and code-review-graph replays completed
1120
+ against the same fixture; codebase-memory remains an explicit policy blocker.
1121
+
1122
+ ### Acceptance criteria
1123
+
1124
+ 1. The fixture labels module boundaries, allowed cross-boundary edges,
1125
+ packages/layers, hub/bridge expectations, and known absent relationships.
1126
+ 2. knodin produces deterministic, filterable, paginated results with explicit
1127
+ totals and C3-compliant byte/token/item budgets.
1128
+ 3. Architecture claims retain exact-versus-heuristic confidence and never turn
1129
+ an inferred grouping into a precise boundary claim.
1130
+ 4. A pinned replay proves oracle quality and budget parity or better.
1131
+
1132
+ ### C28 — Prove dead-code and semantic-search quality
1133
+
1134
+ - Status: implemented
1135
+ - Priority: P1
1136
+ - Disposition: Must close
1137
+ - DependsOn: C24, C3, C11, C12
1138
+ - Competitors: GitNexus, CodeGraph, grepai, Claude Context
1139
+ - Touches: `benchmarks/evaluations/competitive-search/`,
1140
+ `src/__tests__/unit/query.spec.ts`, `src/__tests__/unit/search.spec.ts`,
1141
+ `src/__tests__/unit/response-budget.spec.ts`, relevant bake-off scripts,
1142
+ engine/tool code only if replay fails
1143
+ - What: measure dead-code precision/recall and semantic-orientation relevance
1144
+ using known-live/known-dead and judged-query fixtures.
1145
+ - Replay evidence:
1146
+ `benchmarks/evaluations/competitive-search/raw-results-20260723T214300000Z.json`
1147
+ preserves the frozen local oracle and exact version checks. The pinned
1148
+ GitNexus 1.6.9 adapter normalizes ranked function definitions and its
1149
+ read-only no-incoming-`CALLS` Cypher heuristic to the same `file::symbol`
1150
+ labels. knodin records dead-code precision/recall of 1.0/1.0 versus
1151
+ GitNexus's 0.667/1.0, and mean semantic nDCG/recall@3 of 0.855/1.0 versus
1152
+ 0.677/1.0. Both stay within the 32 KiB response budget with no source egress.
1153
+ The remaining CodeGraph, grepai, and Claude Context capability blockers are
1154
+ retained honestly but no longer prevent a comparable pinned competitor replay.
1155
+
1156
+ ### Acceptance criteria
1157
+
1158
+ 1. Dead-code labels cover exports, imports, reflection-like unresolved uses,
1159
+ tests, barrels, and known live/dead symbols; results report precision,
1160
+ recall, and false positives.
1161
+ 2. Search has frozen relevance judgments and reports nDCG and recall@k under a
1162
+ fixed source and response budget.
1163
+ 3. knodin labels heuristic dead-code or semantic evidence honestly and offers a
1164
+ bounded drill-down for excluded evidence.
1165
+ 4. The replay establishes a win only with oracle parity or better and no
1166
+ regression to C3 budgets.
1167
+
1168
+ ### C29 — Prove context-packing and export leadership
1169
+
1170
+ - Status: implemented
1171
+ - Priority: P1
1172
+ - Disposition: Must close
1173
+ - DependsOn: C24, C3, C5, C14
1174
+ - Competitors: Repomix, Aider repo-map, code2prompt
1175
+ - Touches: `benchmarks/evaluations/competitive-context/`,
1176
+ `src/__tests__/unit/context-export.spec.ts`,
1177
+ `src/__tests__/unit/response-budget.spec.ts`, relevant bake-off scripts,
1178
+ `src/context-export.ts` and tool/CLI code only if replay fails
1179
+ - What: compare deterministic packing, exact token accounting, file policy,
1180
+ incremental/ranged retrieval, and optional Git context on a shared corpus.
1181
+ - Replay evidence:
1182
+ `benchmarks/evaluations/competitive-context/raw-results-20260723.json`.
1183
+ knodin met the shared inclusion, exclusion, range, Git, token, and byte
1184
+ oracle for deterministic Markdown, JSON, and XML exports. Repomix and
1185
+ code2prompt completed the bounded corpus replay with their unavailable
1186
+ controls recorded; Aider is explicitly blocked because it has no comparable
1187
+ artifact surface.
1188
+
1189
+ ### Acceptance criteria
1190
+
1191
+ 1. The fixture labels required/excluded source, tree, line ranges, Git changes,
1192
+ and expected token/byte bounds for each request.
1193
+ 2. knodin produces deterministic Markdown, JSON, and XML exports with relative
1194
+ paths and a declared tokenizer estimate.
1195
+ 3. Include/exclude and per-file policies compose with hard budgets; ranged
1196
+ reads and exact regex grep return verifiable source positions.
1197
+ 4. The competitor replay compares equivalent scope and budget and preserves
1198
+ any unavailable command or format as an explicit blocker.
1199
+
1200
+ ### C30 — Prove guarded editing and diagnostics leadership
1201
+
1202
+ - Status: implemented (evaluation found no production-surface gap)
1203
+ - Priority: P2
1204
+ - Disposition: Evaluate
1205
+ - DependsOn: C24, C2, C11, C15
1206
+ - Competitors: Serena, code-review-graph
1207
+ - Touches: `benchmarks/evaluations/competitive-editing/`,
1208
+ `src/__tests__/unit/lsp-readonly.spec.ts`,
1209
+ `src/__tests__/unit/rename-apply.spec.ts`, relevant bake-off scripts,
1210
+ `src/lsp-readonly.ts` and editing surfaces only if evaluation approves them
1211
+ - What: establish whether additional local, guarded edit previews deliver
1212
+ enough value beyond compiler-verified rename and read-only TypeScript
1213
+ language-service queries.
1214
+ - Replay evidence:
1215
+ `benchmarks/evaluations/competitive-editing/raw-results-20260723.json`.
1216
+ The disposable fixture verified diagnostic and navigation goldens,
1217
+ preview-first compiler/test validation, rejected-edit rollback, and explicit
1218
+ unavailable competitor adapters. It found no actionable gap requiring a new
1219
+ production editing surface.
1220
+
1221
+ ### Acceptance criteria
1222
+
1223
+ 1. Fixtures include diagnostic goldens, definition/implementation navigation,
1224
+ valid edits, rejected edits, rollback, and compile/test verification.
1225
+ 2. Any proposed mutation is preview-first, bounded to a disposable copy, and
1226
+ produces a compiler/test result before it can be judged successful.
1227
+ 3. No daemon, credential, network access, or unguarded workspace mutation is
1228
+ required; unsupported language adapters return explicit unavailable results.
1229
+ 4. Promote an implementation only if the competitor comparison finds an
1230
+ actionable, locally deliverable gap that rename and read-only queries cannot
1231
+ cover.
1232
+
1233
+ ### C31 — Prove index lifecycle and output-telemetry leadership
1234
+
1235
+ - Status: done
1236
+ - Priority: P1
1237
+ - Disposition: Must close
1238
+ - DependsOn: C24, C3, C16, C18
1239
+ - Competitors: codebase-memory-mcp, grepai, Claude Context
1240
+ - Touches: `benchmarks/evaluations/competitive-lifecycle/`,
1241
+ `src/__tests__/unit/index-health.spec.ts`,
1242
+ `src/__tests__/unit/perf.spec.ts`, `src/competitive-runner.ts`, relevant
1243
+ bake-off scripts, engine/tool/CLI code only if replay fails
1244
+ - What: compare cold index, incremental index, corruption repair, stale-index
1245
+ recovery, output telemetry, peak RSS, disk, and failure behavior locally.
1246
+ - Replay evidence:
1247
+ `benchmarks/evaluations/competitive-lifecycle/raw-results-20260723T220325Z.json`
1248
+ records disposable four-phase replays for pinned codebase-memory-mcp and
1249
+ grepai with verified versions and behavioral lifecycle oracles. Both expose
1250
+ stale indexed state after a unique symbol replacement, recover through their
1251
+ documented refresh/reindex mechanism, make the replacement queryable, remove
1252
+ the old symbol, and preserve the healthy symbol. Native telemetry inspection
1253
+ deterministically establishes the remaining capability gaps:
1254
+ codebase-memory reports graph/index counts and repository state but lacks
1255
+ detailed per-query output telemetry; grepai reports aggregate query/token
1256
+ savings but lacks per-query latency, truncation, and detail mode. knodin
1257
+ supplies those opt-in local fields. Claude Context remains a separately
1258
+ classified Ollama/Milvus setup-resource blocker rather than an inferred win.
1259
+
1260
+ ### Acceptance criteria
1261
+
1262
+ 1. The fixture records cold and warm timing, peak RSS, database/disk growth,
1263
+ failure injection, repair outcome, and telemetry fields without source
1264
+ content.
1265
+ 2. knodin detects stale/damaged state, provides an actionable repair, and
1266
+ verifies the repaired result without deleting healthy state.
1267
+ 3. Persisted telemetry is opt-in, local, and includes operation, latency,
1268
+ serialized bytes, estimated tokens, truncation, and detail mode.
1269
+ 4. A replay compares equivalent lifecycle phases; setup or resource failures
1270
+ remain distinct from correctness and performance outcomes.
1271
+
1272
+ ### C32 — Prove local API and statement-flow analysis leadership
1273
+
1274
+ - Status: implemented (evaluation found no production-surface gap)
1275
+ - Priority: P2
1276
+ - Disposition: Evaluate
1277
+ - DependsOn: C24, C2, C3, C8, C9
1278
+ - Competitors: GitNexus
1279
+ - Touches: `benchmarks/evaluations/competitive-api-flow/`,
1280
+ `src/__tests__/unit/flow-analysis.spec.ts`,
1281
+ `src/__tests__/unit/api-contracts.spec.ts`, relevant bake-off scripts,
1282
+ engine/tool code only if evaluation proves a gap
1283
+ - What: replay bounded API contract mismatch and statement-level flow results
1284
+ against a labeled fixture, retaining knodin's compact local-only contract.
1285
+ - Replay evidence:
1286
+ `benchmarks/evaluations/competitive-api-flow/raw-results-20260723.json`.
1287
+ The labeled local replay produced 1.00 precision and recall with zero false
1288
+ positives for both bounded API mismatches and variable-filtered statement
1289
+ flow. It verifies source evidence, an explicit C3 continuation at the flow
1290
+ budget, and the one-gateway/local-only constraints. No broader persisted PDG
1291
+ or API surface is justified by this evaluation.
1292
+
1293
+ ### Acceptance criteria
1294
+
1295
+ 1. The fixture labels true API mismatches, true control/data-flow rows,
1296
+ variable filtering, unresolved cases, and expected source evidence.
1297
+ 2. knodin reports precision, recall, false positives, unresolved cases,
1298
+ truncation, and continuation routes under C3 budgets.
1299
+ 3. The comparison never treats broad persisted PDG output, raw Cypher, or
1300
+ hosted analysis as required parity; it evaluates only proven local workflows.
1301
+ 4. Promote broader implementation only when it finds actionable breakage that
1302
+ existing bounded API/flow queries miss.
1303
+
1304
+ ### C33 — Preserve local privacy and one-tool schema economy
1305
+
1306
+ - Status: implemented
1307
+ - Priority: P0
1308
+ - Disposition: Must close
1309
+ - DependsOn: C24
1310
+ - Competitors: all; `mcp-codebase-index` and Claude Context supply negative
1311
+ evidence where hosted services or local service dependencies are required.
1312
+ - Touches: `src/tools/knodin-tools.ts`, `src/server.ts`,
1313
+ `src/competitive-manifest.ts`, `src/__tests__/unit/mcp-tools.spec.ts`,
1314
+ `benchmarks/evaluations/competitive-constraints/`, this roadmap
1315
+ - What: make the local/no-auth/no-egress and one-gateway constraints measurable
1316
+ release gates for every competitive claim.
1317
+
1318
+ ### Acceptance criteria
1319
+
1320
+ 1. A clean local replay requires no credentials, hosted service, network
1321
+ embeddings, source-code egress, or mandatory external daemon.
1322
+ 2. The MCP surface remains one `knodin` gateway; a capability addition must
1323
+ document schema-token cost and cannot create a top-level tool.
1324
+ 3. Any optional local service is reported as optional and unavailable states
1325
+ are explicit rather than silently degraded success.
1326
+ 4. No item C24-C32 can close unless its replay records this constraint check.
1327
+
1328
+ ### C34 — Keep local graph indexes current across repository lifecycle events
1329
+
1330
+ - Status: implemented
1331
+ - Priority: P1
1332
+ - Disposition: Must close
1333
+ - DependsOn: C16, C24, C33
1334
+ - Competitors: GitNexus, Graphify, codebase-memory-mcp
1335
+ - Touches: `bin/cli.ts`, `src/engine/index.ts`, `lefthook.yml`,
1336
+ `templates/hooks/`, `src/__tests__/unit/freshness.spec.ts`,
1337
+ `src/__tests__/unit/watcher-lifecycle.spec.ts`,
1338
+ `benchmarks/evaluations/competitive-lifecycle/`, docs
1339
+ - What: make fresh, reconciled, stale, and unavailable graph states explicit
1340
+ across file edits, branch switches, pulls, rebases, merges, and a restarted
1341
+ local process; provide safe, local refresh commands for every supported graph
1342
+ artifact without pretending an external index is current.
1343
+ - Evidence: `knodin refresh-artifacts [checkout|merge|code-change]` is an
1344
+ explicit opt-in, 30-second-bounded refresh for locally installed GitNexus and
1345
+ Graphify. It prefers the repository's GitNexus runner when present, records
1346
+ success, failure, or skipped state without source content, and is never wired
1347
+ into a commit hook. The lifecycle fixtures prove restart, replacement,
1348
+ branch-switch, merge, watcher-loss, and external-tool states; the replay
1349
+ records freshness latency, RSS, disk, and verified state.
1350
+
1351
+ ### Acceptance criteria
1352
+
1353
+ 1. knodin continues to reconcile its persisted index at cold start and before
1354
+ answers after offline drift, reporting `fresh`, `reconciled`, or `unknown`.
1355
+ 2. The repository exposes a documented local command or opt-in hook path to
1356
+ refresh GitNexus after checkout/merge and Graphify after code changes; it
1357
+ records success, failure, or skipped state and never blocks a commit on an
1358
+ unbounded rebuild.
1359
+ 3. A fixture proves correct behavior for edit, deletion, branch switch, merge,
1360
+ pull/rebase-equivalent replacement, watcher loss, and restart; no answer may
1361
+ silently claim an unverified index is fresh.
1362
+ 4. The lifecycle replay reports freshness latency and resource cost separately
1363
+ from query quality and preserves any unavailable external tool as a blocker.
1364
+
1365
+ ### C35 — Produce local interactive architecture and call-flow visualizations
1366
+
1367
+ - Status: implemented
1368
+ - Priority: P2
1369
+ - Disposition: Should close
1370
+ - DependsOn: C3, C13, C24, C33
1371
+ - Competitors: Graphify, GitNexus
1372
+ - Touches: `bin/cli.ts`, `src/visualization.ts`,
1373
+ `src/__tests__/unit/visualization.spec.ts`,
1374
+ `benchmarks/evaluations/competitive-visualization/`, docs
1375
+ - What: evaluate and, if the evidence justifies it, generate deterministic
1376
+ local HTML architecture and call-flow artifacts from knodin's bounded graph
1377
+ facts without adding a new MCP tool, hosted renderer, source egress, or a
1378
+ mandatory long-lived service.
1379
+ - Evaluation evidence:
1380
+ `benchmarks/evaluations/competitive-visualization/raw-results-20260723.json`
1381
+ records Graphify's pinned community-graph, tree, and call-flow artifacts and
1382
+ GitNexus's explicit lack of a deterministic static-artifact command. knodin's
1383
+ checked-in evaluation artifact is self-contained, deterministic, bounded,
1384
+ source-commit/index-freshness tagged, and source-evidenced. `knodin visualize
1385
+ <entry> --output <path.html>` now exposes that bounded artifact as a local
1386
+ CLI-only export with 1–6 depth, a 4–64 KiB generation ceiling, safe in-repo
1387
+ output paths, and stable identity/file/kind entry disambiguation. It adds no
1388
+ MCP operation, hosted renderer, credential, telemetry, or source egress.
1389
+
1390
+ ### Acceptance criteria
1391
+
1392
+ 1. The evaluation compares Graphify's community graph, tree, and call-flow
1393
+ artifacts with a knodin artifact generated from the same pinned fixture.
1394
+ 2. Generated HTML is self-contained or uses only local static assets, records
1395
+ its source commit/index freshness, and links every displayed edge to bounded
1396
+ source evidence or an explicit heuristic label.
1397
+ 3. The artifact supports at least subsystem drill-down, hub/bridge inspection,
1398
+ and bounded call-flow navigation while respecting C3 output and generation
1399
+ budgets.
1400
+ 4. It remains an optional CLI/export capability behind the existing `knodin`
1401
+ gateway: no new top-level MCP tool, network service, credential, or source
1402
+ egress is introduced.
1403
+
1404
+ ### C36 — Instrument the current warm-operation performance paths
1405
+
1406
+ - Status: implemented
1407
+ - Priority: P0
1408
+ - Disposition: Must close
1409
+ - DependsOn: C18, C24
1410
+ - Evidence: opt-in, source-free phase accounting now attributes freshness,
1411
+ graph analytics, traversal snapshots, architecture facets, context
1412
+ composition, status audit, pack walking/serialization, and artifact
1413
+ read/grep. `perf.spec.ts` verifies the complete phase record without latency
1414
+ thresholds.
1415
+ - Touches: `src/engine/perf.ts`, `src/engine/index.ts`, `src/context.ts`,
1416
+ `src/context-export.ts`, targeted unit tests
1417
+ - What: add opt-in phase accounting around the current production paths without
1418
+ changing response contracts or enabling a competitor replay.
1419
+
1420
+ ### Acceptance criteria
1421
+
1422
+ 1. Phase names separately cover freshness, graph analytics, traversal snapshot,
1423
+ architecture facets, context composition, status audit, pack walk,
1424
+ serialization, and artifact read/grep.
1425
+ 2. Instrumentation remains inert unless a performance session is active and
1426
+ records no source content.
1427
+ 3. Targeted tests verify phase accounting without wall-clock thresholds.
1428
+ 4. No competitor or before/after comparison is run until C37-C40 are complete
1429
+ and the repository owner explicitly approves the replay.
1430
+
1431
+ ### C37 — Reuse generation-scoped graph analytics and traversal snapshots
1432
+
1433
+ - Status: implemented
1434
+ - Priority: P0
1435
+ - Disposition: Must close
1436
+ - DependsOn: C36
1437
+ - Evidence: minimal and standard map now share one generation-keyed analytics
1438
+ snapshot; traversal uses a compact signed adjacency plus cached definition
1439
+ metadata; architecture membership/cohesion/coupling aggregation is one-pass.
1440
+ Option-keyed analytics/minimal-map caches and per-repository traversal/status
1441
+ caches have explicit entry ceilings rather than unbounded process growth.
1442
+ `typed-traversal.spec.ts` proves both map and traversal snapshots invalidate
1443
+ after an incremental index generation.
1444
+ - Touches: `src/engine/index.ts`, `src/__tests__/unit/map.spec.ts`,
1445
+ `src/__tests__/unit/typed-traversal.spec.ts`,
1446
+ `src/__tests__/unit/uniform-filters.spec.ts`
1447
+ - What: cache immutable graph analytics, bidirectional typed adjacency, and
1448
+ symbol metadata by repository/federation configuration plus index generation.
1449
+
1450
+ ### Acceptance criteria
1451
+
1452
+ 1. Minimal and standard map modes share one generation-scoped analytics source
1453
+ while retaining their existing response shapes and truthful totals.
1454
+ 2. Traversal performs bounded BFS over cached adjacency with no per-result
1455
+ symbol lookup and preserves identity, direction, filters, evidence, and
1456
+ deterministic ordering.
1457
+ 3. Architecture facets and coupling use one-pass identity-to-community indexes
1458
+ rather than repeated community-by-edge searches.
1459
+ 4. Index, reconcile, watcher flush, repair, and close invalidate every affected
1460
+ cache; tests prove no stale result survives a generation change.
1461
+
1462
+ ### C38 — Compose orientation context from one shared snapshot
1463
+
1464
+ - Status: implemented
1465
+ - Priority: P1
1466
+ - Disposition: Should close
1467
+ - DependsOn: C37
1468
+ - Evidence: `buildKnodinContext` establishes the standard map snapshot once,
1469
+ then composes stats, flows, and optional review over that initialized
1470
+ generation. `context-composition.spec.ts` verifies at most one probe and the
1471
+ existing compact caps/contracts without timing assertions.
1472
+ - Touches: `src/context.ts`, `src/engine/index.ts`,
1473
+ `src/__tests__/unit/knodin-tools.spec.ts`,
1474
+ `src/__tests__/unit/freshness.spec.ts`
1475
+ - What: acquire one freshness-qualified engine snapshot, reuse C37 analytics,
1476
+ and compose independent context fields without repeating graph construction.
1477
+
1478
+ ### Acceptance criteria
1479
+
1480
+ 1. One context request performs at most one freshness probe and one graph
1481
+ analytics construction for its repository generation.
1482
+ 2. Stats, communities, hubs, flows, and optional risk retain their current
1483
+ contract, sorting, caps, and evidence.
1484
+ 3. Independent work may run concurrently only after initialization is complete;
1485
+ no SQLite writer concurrency or duplicate engine singleton is introduced.
1486
+ 4. Tests use counters and cache state, not timing thresholds.
1487
+
1488
+ ### C39 — Separate fast truthful status from explicit deep audit
1489
+
1490
+ - Status: implemented
1491
+ - Priority: P1
1492
+ - Disposition: Should close
1493
+ - DependsOn: C36
1494
+ - Evidence: warm status reuses a generation/database-fingerprint keyed deep
1495
+ audit only after the existing bounded drift probe verifies the source state.
1496
+ Responses identify `deep-audit` versus
1497
+ `cached-after-freshness-probe` and include `verifiedAt`; `--deep` and MCP
1498
+ `statusAudit: deep` force the complete audit. `index-health.spec.ts` proves
1499
+ warm reuse, explicit deep mode, and file-change invalidation.
1500
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
1501
+ `src/__tests__/unit/index-health.spec.ts`
1502
+ - What: return an invalidation-safe cached health snapshot for warm status and
1503
+ expose the existing filesystem/database verification as an explicit deep
1504
+ audit, with verification age and mode reported truthfully.
1505
+
1506
+ ### Acceptance criteria
1507
+
1508
+ 1. Default warm status never claims a fresh audit when it is returning cached
1509
+ evidence; it reports verification mode and timestamp.
1510
+ 2. Deep audit preserves current missing, damaged, orphaned, coverage, and repair
1511
+ behavior.
1512
+ 3. Index generation, watcher changes, repair, schema change, and relevant file
1513
+ events invalidate the cached snapshot.
1514
+ 4. Existing callers remain source-compatible and a caller can explicitly
1515
+ request the deep audit.
1516
+
1517
+ ### C40 — Make context packing and artifact access linear and cache-safe
1518
+
1519
+ - Status: implemented
1520
+ - Priority: P1
1521
+ - Disposition: Should close
1522
+ - DependsOn: C36
1523
+ - Evidence: glob/policy matchers are bounded and compiled once, walking uses one
1524
+ accumulator, exact per-file serialization deltas replace accumulated
1525
+ reserialization, and artifact lines use a two-entry/four-MiB-per-file
1526
+ path/mtime/size cache. Context-export and response-budget tests preserve
1527
+ formats, policies, ordering, hard budgets, safe paths, reads, grep, and cache
1528
+ invalidation.
1529
+ - Touches: `src/context-export.ts`,
1530
+ `src/__tests__/unit/context-export.spec.ts`,
1531
+ `src/__tests__/unit/response-budget.spec.ts`
1532
+ - What: precompile selection policy, prune excluded directories, account bytes
1533
+ incrementally, serialize once, and reuse a bounded artifact line index keyed
1534
+ by path/mtime/size.
1535
+
1536
+ ### Acceptance criteria
1537
+
1538
+ 1. Candidate walking and pattern matching are linear in visited paths plus
1539
+ patterns; patterns are compiled once per request.
1540
+ 2. Packing does not rebuild the entire accumulated artifact per candidate and
1541
+ still enforces exact response byte/token ceilings.
1542
+ 3. Artifact read/grep reuse a bounded cache only when path, size, and mtime
1543
+ match; mutation and replacement invalidate it.
1544
+ 4. Markdown, JSON, XML, policies, tree/Git sections, deterministic ordering,
1545
+ ranged reads, regex grep, and telemetry remain byte-for-byte compatible in
1546
+ golden fixtures.
1547
+
1548
+ ### C41 — Replay like-for-like performance after the implementation batch
1549
+
1550
+ - Status: evaluated — deferred
1551
+ - Priority: P0
1552
+ - Disposition: Evaluate
1553
+ - DependsOn: C37, C38, C39, C40
1554
+ - Evidence: the terminal bounded replay and strengthened-oracle verdict are
1555
+ preserved at
1556
+ `benchmarks/competitors/runs/20260802T203822521Z-3e27547d095e/`. Its
1557
+ `c41-verdict.json` verifies all five competitors, explicit warm/cold p50/p95,
1558
+ independently attributed two-sided RSS, and grepai cached/deep lifecycle
1559
+ rows. Under the strengthened labels, 19 mapped rows passed, 10 failed, four
1560
+ grepai trace rows were unavailable, and six intentionally unmatched Repomix
1561
+ rows remained incomparable. Graphify passed its direct-call rows but not hub
1562
+ rows; codebase-memory passed six of ten search/traversal/package rows;
1563
+ Repomix passed seven pack/read/grep rows; code2prompt passed both pack rows;
1564
+ and grepai produced no oracle-qualified latency row because its local index
1565
+ remained empty. The bounded local-Ollama indexing attempts at
1566
+ `20260802T205026646Z-18c1107af915/`,
1567
+ `20260802T205625288Z-e97db4c32eb6/` preserve that setup limitation. Earlier
1568
+ runs retain the sandbox-launch, timeout, and uninitialized-index failures
1569
+ rather than overwriting them.
1570
+ - Evaluation disposition: retain the current product unchanged and defer any
1571
+ aggregate performance claim or optimization response. Only oracle-passing
1572
+ exact rows may support operation-specific observations; failures,
1573
+ unavailable rows, and setup limits remain separate dimensions and cannot
1574
+ become wins.
1575
+ - Touches if approved: competitor harness mappings, immutable run artifacts,
1576
+ `benchmarks/competitors/COMPETITIVE-AUDIT.md`, this roadmap
1577
+ - What: after explicit repository-owner permission, run warm/cold p50/p95/RSS
1578
+ measurements using equivalent operations, budgets, fixtures, and correctness
1579
+ oracles. A user instruction to complete this roadmap constitutes permission
1580
+ for the checked-in, zero-spend, local replay only; network access, installation
1581
+ of unpinned software, publication, or external mutation still requires its own
1582
+ authority.
1583
+
1584
+ ### Evaluation gate
1585
+
1586
+ 1. Graphify traversal/map rows use the same direction, relation, depth, item,
1587
+ token, and evidence scope.
1588
+ 2. codebase-memory architecture/search rows use the same facets, pagination,
1589
+ source inclusion, and correctness labels.
1590
+ 3. code2prompt and Repomix compare with knodin `pack`, packed read, and packed
1591
+ grep—not semantic `context`, `search`, or `file_summary`.
1592
+ 4. grepai status compares cached summary with cached summary and deep audit with
1593
+ an equivalent verified lifecycle phase.
1594
+ 5. Results record correctness before latency and treat RSS/setup failures as
1595
+ separate dimensions.
1596
+
1597
+ ### C42 — Enforce comparable competitive cases and correctness oracles
1598
+
1599
+ - Status: implemented
1600
+ - Priority: P0
1601
+ - DependsOn: C41
1602
+ - Evidence: `src/competitive-contract.ts` and the competitor mappings enforce
1603
+ typed equivalence/oracle dispositions; the Repomix replay qualifies 18
1604
+ comparable rows and excludes wrong-fast, partial, unverified, and
1605
+ incomparable rows from wins.
1606
+ - Touches: `src/competitive-manifest.ts`, `src/competitive-runner.ts`,
1607
+ `scripts/competitive-bakeoff.ts`, competitor adapters, labeled competitive
1608
+ fixtures and runner tests
1609
+ - What: define one typed case contract for equivalent operation scope, budgets,
1610
+ fixtures, and correctness. Incomparable or wrong-fast rows must never count as
1611
+ latency wins.
1612
+
1613
+ ### Acceptance criteria
1614
+
1615
+ 1. Every timed pair records use case, fixture revision, mode, scope, budgets,
1616
+ and a comparable/incomparable disposition with a reason.
1617
+ 2. Graphify, codebase-memory, Repomix, code2prompt, and grepai mappings satisfy
1618
+ C41's exact equivalence rules; unmatched rows are excluded explicitly.
1619
+ 3. Comparable rows carry checked-in correctness oracles, and summaries
1620
+ distinguish oracle pass, oracle failure, unavailable, and incomparable.
1621
+ 4. Partial runs and rows without passing oracles cannot produce a competitive
1622
+ win or aggregate performance claim.
1623
+
1624
+ ### C43 — Measure cold/warm p50, p95, and independent process RSS
1625
+
1626
+ - Status: implemented
1627
+ - Priority: P0
1628
+ - DependsOn: C42
1629
+ - Evidence: `benchmarks/competitors/MEASUREMENT-CONTRACT.md` and shared sampling
1630
+ record three warmups, at least 20 warm samples, raw values, p50/p95/range,
1631
+ independently attributed RSS, descendant cleanup, and atomic partial
1632
+ summaries across all 11 adapters.
1633
+ - Touches: shared competitive measurement helper, competitor adapters,
1634
+ `scripts/competitive-bakeoff.ts`, lifecycle runner, normalized summaries and
1635
+ tests
1636
+ - What: replace copied three-sample loops and aggregate process-tree ceilings
1637
+ with a shared statistical and resource-measurement contract.
1638
+
1639
+ ### Acceptance criteria
1640
+
1641
+ 1. Warm and cold samples are separate and retain raw samples, warm-up count,
1642
+ p50, p95, minimum, and maximum; the documented warm sample minimum supports a
1643
+ meaningful p95.
1644
+ 2. knodin and competitor RSS are attributed independently by process and phase;
1645
+ combined harness RSS is never assigned to either product.
1646
+ 3. Per-call and per-competitor timeouts reap descendants, preserve partial
1647
+ evidence, identify the failed phase/case, and still write a terminal summary.
1648
+ 4. Regression checks compare matching modes and enforce both p50 and p95
1649
+ tolerances after correctness passes.
1650
+
1651
+ ### C44 — Restore low-latency diff review without weakening evidence
1652
+
1653
+ - Status: implemented
1654
+ - Priority: P0
1655
+ - DependsOn: C43
1656
+ - Evidence: `benchmarks/evaluations/competitive-review/` records an
1657
+ oracle-qualified bounded-gateway replay at 53.018 ms p50, 89.449 ms p95, and
1658
+ 733 bytes; review tests preserve path-scoped churn/risk, diff scopes, flows,
1659
+ cache invalidation, and immutable evidence provenance.
1660
+ - Touches: `src/engine/index.ts`, `src/engine/perf.ts`, review/flow tests and
1661
+ competitive-review evidence
1662
+ - What: remove per-file Git history processes and whole-repository repeated work
1663
+ from warm review while preserving changed-file, risk, flow, and source
1664
+ evidence.
1665
+
1666
+ ### Acceptance criteria
1667
+
1668
+ 1. Review phase telemetry separates diff discovery, churn, parsing, flow lookup,
1669
+ and serialization.
1670
+ 2. Churn/history is collected in one bounded pass or a repository/HEAD cache,
1671
+ and affected flows use a generation-scoped file-to-flow lookup.
1672
+ 3. All diff scopes and existing correctness fixtures remain unchanged.
1673
+ 4. The oracle-qualified warm minimal review meets the recorded C43 regression
1674
+ threshold on both p50 and p95.
1675
+
1676
+ ### C45 — Add facet-selective architecture fast paths
1677
+
1678
+ - Status: implemented
1679
+ - Priority: P1
1680
+ - DependsOn: C43
1681
+ - Evidence: `benchmarks/evaluations/competitive-architecture/` records an
1682
+ oracle-qualified bounded-gateway replay at 0.082 ms p50 and 0.097 ms p95;
1683
+ architecture tests prove facet-selective construction, zero minimal edge
1684
+ materialization, mutation isolation, budgets, and generation/close
1685
+ invalidation.
1686
+ - Touches: `src/engine/index.ts`, architecture/map tests and
1687
+ competitive-architecture evidence
1688
+ - What: compute only requested architecture facets from reusable
1689
+ generation-scoped metadata instead of building a complete standard map.
1690
+
1691
+ ### Acceptance criteria
1692
+
1693
+ 1. Language, package, layer, entry-point, boundary, coupling, and hotspot
1694
+ requests construct only their required inputs.
1695
+ 2. Minimal facet requests do not build full edge lists or hydrate unrelated
1696
+ source/symbol detail.
1697
+ 3. Identity, filters, pagination, totals, budgets, and cache invalidation remain
1698
+ correct.
1699
+ 4. Equivalent architecture rows pass their oracle and C43 thresholds.
1700
+
1701
+ ### C46 — Cache search state and hydrate only the requested page
1702
+
1703
+ - Status: implemented
1704
+ - Priority: P1
1705
+ - DependsOn: C43
1706
+ - Evidence: `benchmarks/evaluations/competitive-search/` records a
1707
+ same-oracle 500-symbol offset-page improvement from 5.309/5.865 ms to
1708
+ 0.310/0.481 ms p50/p95; tests preserve exact-default search, global ordering,
1709
+ totals, pagination, filters, ANN gating, federation, corruption, and budgets.
1710
+ - Touches: `src/engine/index.ts`, search/tool contracts, search and ANN tests,
1711
+ competitive-search evidence
1712
+ - What: reuse decoded generation-scoped search metadata, rank before hydration,
1713
+ and perform source/caller work only for the requested page.
1714
+
1715
+ ### Acceptance criteria
1716
+
1717
+ 1. Cheap filters run before vector scoring and source hydration; community
1718
+ lookup does not build a full map.
1719
+ 2. Exact search remains the default unless the existing ANN recall gate passes.
1720
+ 3. Global ordering, offset pagination, totals, corruption handling, federation,
1721
+ and response budgets remain deterministic.
1722
+ 4. A separately specified opaque cursor may be added only if it binds query,
1723
+ filters, order, and generation and truthfully rejects stale cursors.
1724
+
1725
+ ### C47 — Reduce status, repair, and measured lifecycle resource cost
1726
+
1727
+ - Status: implemented
1728
+ - Priority: P1
1729
+ - DependsOn: C43
1730
+ - Evidence: `benchmarks/evaluations/competitive-lifecycle/` records isolated,
1731
+ ownership-labeled lifecycle stages and immutable provenance; health tests
1732
+ prove lease expiry guards, single-flight deduplicated batch repair, one
1733
+ invalidation, and final forced deep verification.
1734
+ - Touches: `src/engine/index.ts`, `src/engine/perf.ts`, lifecycle runner,
1735
+ freshness/index-health/watcher tests
1736
+ - What: eliminate duplicate freshness/audit work, batch repair indexing, and
1737
+ reduce only resource retention demonstrated by isolated C43 measurements.
1738
+
1739
+ ### Acceptance criteria
1740
+
1741
+ 1. Cached status reuses a valid freshness lease and deep-audit snapshot without
1742
+ claiming unverified freshness; lease expiry still performs the missed-watcher
1743
+ guard.
1744
+ 2. Repair reuses valid audit evidence, deduplicates paths, batches indexing and
1745
+ orphan cleanup, invalidates once, and completes one final deep verification.
1746
+ 3. The lifecycle runner records isolated baseline, post-index, post-query,
1747
+ post-close, peak RSS, and disk measurements for each product.
1748
+ 4. Cache/parser changes are accepted only when attribution proves a retained
1749
+ owner and lifecycle correctness remains intact.
1750
+
1751
+ ### C48 — Decide an optional local project-memory boundary
1752
+
1753
+ - Status: evaluated — deferred
1754
+ - Priority: P2
1755
+ - DependsOn: C42
1756
+ - Decision: defer production work. The evaluated workflows do not yet
1757
+ demonstrate value beyond repository-owned documentation, while a memory
1758
+ surface would add a second source of truth, retention/privacy obligations,
1759
+ and an output-budget surface. See
1760
+ [`docs/adr/004-optional-project-memory.md`](../docs/adr/004-optional-project-memory.md)
1761
+ and the
1762
+ [C48 evaluation](../benchmarks/evaluations/c48-project-memory/evaluation.md).
1763
+ - Touches: roadmap/ADR and Serena memory evaluation evidence; no production
1764
+ code because the evaluation gate did not pass
1765
+ - What: evaluate concrete durable-note workflows without turning generalized
1766
+ agent memory into a core graph/index requirement.
1767
+
1768
+ ### Evaluation gate
1769
+
1770
+ 1. Record user workflows, path containment, retention/backup, privacy, output
1771
+ budgets, and RSS/disk cost against Serena's memory operations.
1772
+ 2. Prefer an optional local `.reckon/notes/` adapter that does not affect graph
1773
+ identity, indexing, freshness, or the one-tool gateway count.
1774
+ 3. Approve production work only when a labeled workflow demonstrates value
1775
+ beyond ordinary repository documentation.
1776
+ 4. If approved, require atomic writes, traversal protection, deterministic
1777
+ list/read/edit/rename/delete behavior, and explicit size budgets.
1778
+
1779
+ ### C49 — Keep immutable runs lint-safe and make the full suite terminate
1780
+
1781
+ - Status: implemented
1782
+ - Priority: P1
1783
+ - DependsOn: —
1784
+ - Evidence: Biome excludes only immutable `benchmarks/competitors/runs/**`;
1785
+ handle diagnostics and teardown coverage close watchers, databases, queues,
1786
+ and child processes. The full coverage suite terminates naturally within the
1787
+ documented six-minute ceiling (474/474 in the final C47 gate).
1788
+ - Touches: `biome.json`, `package.json`, Vitest configuration/setup,
1789
+ watcher/freshness teardown tests, competitive-run documentation
1790
+ - What: exclude immutable generated evidence from formatting gates while
1791
+ diagnosing and fixing the open handle that prevents the full suite from
1792
+ exiting.
1793
+
1794
+ ### Acceptance criteria
1795
+
1796
+ 1. Biome ignores `benchmarks/competitors/runs/**` but continues checking
1797
+ generators, baselines, labels, schemas, and source.
1798
+ 2. A hanging-process diagnostic command identifies active handles without
1799
+ masking them through force-exit.
1800
+ 3. Watchers, databases, queues, and spawned test children are asserted closed in
1801
+ teardown, including injected-failure paths.
1802
+ 4. `bun run lint` and the full test suite terminate naturally within the
1803
+ repository's documented ceiling.
1804
+
1805
+ ### C50 — Make the single-worker full-suite gate start reliably
1806
+
1807
+ - Status: implemented
1808
+ - Priority: P0
1809
+ - DependsOn: C49
1810
+ - Touches: Vitest configuration/setup, embedding test doubles and teardown,
1811
+ gate documentation, and a worker-start regression fixture
1812
+ - What: eliminate repeated native ONNX initialization or worker churn that can
1813
+ exhaust Vitest's 600-second pool ceiling before all test files start, without
1814
+ skipping files, forcing exit, weakening isolation blindly, or changing
1815
+ production embedding behavior.
1816
+
1817
+ ### Acceptance criteria
1818
+
1819
+ 1. A diagnostic reproduces and attributes the failed fork-worker startup for
1820
+ `knodin.test.ts`, `map.spec.ts`, and the lifecycle runner instead of
1821
+ increasing the timeout.
1822
+ 2. Tests use an explicit deterministic embedder where semantic model quality is
1823
+ not under test; dedicated embedding tests retain production-path coverage.
1824
+ 3. All test files start, the full suite passes, and each fresh-worker shard
1825
+ exits naturally under its bounded five-minute test ceiling with immutable
1826
+ run artifacts present in the checkout.
1827
+ 4. Lint, typecheck, build, coverage thresholds, native-resource teardown, and
1828
+ production local-embedding behavior remain unchanged.
1829
+
1830
+ ### Evidence and corrected diagnosis
1831
+
1832
+ - Root cause (corrected): the handoff attributed the startup failures to repeated
1833
+ ONNX init, but a diagnostic reproduced them with the three ONNX-free canary
1834
+ files alone. The real cause is `isolate: true` (Vitest default): the pool
1835
+ re-forks a **fresh process per test file** — the config's stated "keep one
1836
+ long-lived worker" intent was never achieved — and under host load a per-file
1837
+ fork spawn intermittently exceeds Vitest's 60 s worker-start handshake
1838
+ (`START_TIMEOUT`), failing a neighbouring file. Actual test work is ~200 s
1839
+ (well under the ceiling); the overage was pure per-file fork churn.
1840
+ - Fix: `isolate: false` + `singleFork` (one reused process — the documented
1841
+ intent, not a blind isolation weakening), so native modules init once and only
1842
+ one fork is ever spawned. Cross-file safety is preserved by explicit guards:
1843
+ the shared deterministic lexical embedder is installed by setup and restored by
1844
+ every teardown; a per-test env/cwd snapshot-restore plus `restoreMocks`/
1845
+ `unstubEnvs`; the real ONNX model is isolated to `embeddings.spec` (loaded once,
1846
+ disposed after) and guarded by a `getRealEmbedderLoadCount()` regression; the
1847
+ six competitive benchmark `runner.spec.ts` files spawn their sub-benchmarks
1848
+ asynchronously (no synchronous `execFileSync` blocking the worker); and removed
1849
+ per-test repo state (SQLite handles + FSWatcher/FSEvent descriptors) is evicted
1850
+ each test.
1851
+ - Verified: the gate uses 16 strictly sequential, fresh-worker shards so
1852
+ native SQLite/watcher state cannot accumulate across the whole suite. The
1853
+ two semantic-ranking specs pass on the deterministic lexical embedder; lint,
1854
+ typecheck, and build pass. Final integration certification passed all 87
1855
+ files/503 tests in 1,471.79 seconds with 597,917,696 bytes peak RSS, and its
1856
+ merged 503-test coverage run passed at 89.5% statements, 79.25% branches,
1857
+ 90.62% functions, and 91.8% lines (901,840,896 bytes peak RSS). No tests or
1858
+ files are quarantined or skipped to obtain these results.
1859
+ - Release certification still reruns the fully accumulated integration checkout
1860
+ under C54; this distinction prevents cumulative-branch evidence from being
1861
+ mislabeled as a final integration-branch run.
1862
+
1863
+ ### C51 — Native training-free embedding quantization (memory-efficient vector search)
1864
+
1865
+ - Status: evaluated — rejected for production
1866
+ - Priority: P2
1867
+ - DependsOn: C46
1868
+ - Touches: `src/engine/embeddings.ts`, `src/engine/index.ts` (exact scan + ANN
1869
+ path), the `symbol_embeddings` storage, and a recall-gate fixture
1870
+ - What: add an opt-in, in-process, data-oblivious quantization of the 384-dim
1871
+ symbol embeddings — normalize → fixed deterministic rotation → precomputed
1872
+ bucketing (the TurboQuant concept, implemented natively in TypeScript) —
1873
+ shrinking stored vectors ~8× (4-bit) and speeding the cosine scan, with float32
1874
+ remaining authoritative and the safe default. This is the memory-efficient
1875
+ vector-search lever that keeps the zero-auth / local / no-egress / one-process
1876
+ guarantees intact, versus competitors that reach for a hosted vector DB. It is
1877
+ explicitly **not** a sidecar, external index, Rust dependency, or hosted store.
1878
+ See [ADR 005](../docs/adr/005-native-embedding-quantization.md).
1879
+ - Evidence: `benchmarks/evaluations/c51-quantization/raw-results.json` binds a
1880
+ 2,048-vector real MiniLM fixture and preserves per-query top-5/top-10 overlap,
1881
+ rank correlation, estimated database/resident payload bytes, conversion time,
1882
+ 3+20 warm and five cold samples, peak RSS, constraints, gates, and disposition. The fixed-seed
1883
+ block-Hadamard 4-bit prototype used 20 unique queries per tier and achieved
1884
+ mean top-10 recall 96%, 93.5%, and 92% at 128, 512, and 2,048 vectors. It
1885
+ therefore failed the >=95% gate at two engagement sizes.
1886
+ Payloads were 8× smaller, but warm p50 was not materially better and
1887
+ conversion-inclusive cold p50 was worse at every size; 99,876,864-byte peak
1888
+ process RSS did not establish an attributable end-to-end RSS win.
1889
+ - Disposition: reject production quantization from this design. Float32 remains
1890
+ authoritative and unchanged. Representation shrink without the recall gate
1891
+ and a material end-to-end latency or memory win does not authorize production
1892
+ work, so no separately numbered implementation item is proposed.
1893
+
1894
+ ### Acceptance criteria
1895
+
1896
+ Evaluation outcome: criteria 1 and 4 hold for the prototype; criterion 3 fails
1897
+ at 512 and 2,048 vectors. Criterion 2 is deliberately not implemented because
1898
+ the failed evaluation does not authorize a production path or environment gate.
1899
+
1900
+ 1. Encode/decode is deterministic (hard-coded rotation seed) and reproducible
1901
+ across machines, with no trained or per-repo codebook and no calibration pass.
1902
+ 2. Quantized codes are opt-in behind an env gate (default off); float32 stays the
1903
+ source of truth and is rebuildable/invalidated by the existing generation
1904
+ counter.
1905
+ 3. A committed real-vector fixture proves at least 95% mean top-10 recall against
1906
+ the float32 exact scan at every size where the path could engage, mirroring
1907
+ the R18/R48 gate. Report top-5/top-10 overlap, rank correlation, database and
1908
+ resident bytes, conversion time, cold/warm p50/p95, and peak RSS. Eligibility
1909
+ additionally requires a material end-to-end latency or memory win; an inner-
1910
+ loop-only improvement is insufficient.
1911
+ 4. No external index, sidecar, Rust/native dependency, hosted DB, or source/vector
1912
+ egress; the embedding model, its dimensionality, and its normalization are
1913
+ unchanged.
1914
+
1915
+ ### C52 — Package repository lifecycle initialization and refresh
1916
+
1917
+ - Status: implemented
1918
+ - Priority: P0
1919
+ - DependsOn: C16, C34, C50
1920
+ - Touches: `src/init.ts`, `src/repository-management.ts`, the deprecated
1921
+ `src/fleet.ts` compatibility adapter, `src/indexable-paths.ts`,
1922
+ `bin/cli.ts`, `src/engine/index.ts`, `scripts/pack-install-smoke.ts`,
1923
+ repository-init/CLI acceptance tests, package build output
1924
+ - What: make the compiled package usable across existing repositories without a
1925
+ knodin source checkout. Resolve the installed executable in hooks, share one
1926
+ canonical indexable-path policy, refresh after commit/checkout/merge/rewrite,
1927
+ and initialize repository checkouts idempotently.
1928
+
1929
+ ### Acceptance and evidence
1930
+
1931
+ 1. `knodin init` preserves existing hook content, supports worktrees and
1932
+ `core.hooksPath`, and repeated initialization produces no duplicate knodin
1933
+ blocks.
1934
+ 2. Hook refresh selects paths through the engine's canonical policy, including
1935
+ TypeScript and supported Salesforce Apex, Aura, LWC, and metadata inputs.
1936
+ 3. `knodin repos init` discovers Git roots deterministically, deduplicates real
1937
+ paths/worktrees, supports dry-run and JSON output, and records individual
1938
+ failures without abandoning the remaining repositories. The historical
1939
+ `knodin fleet init` spelling remains a time-bounded compatibility alias.
1940
+ 4. `scripts/pack-install-smoke.ts` installs the exact tarball in a clean external
1941
+ repository and proves package-resolved hooks across post-commit,
1942
+ post-checkout, post-merge, post-rewrite, rename, and deletion lifecycle
1943
+ events.
1944
+
1945
+ Evidence: packaged hooks landed in `c3a10a5`, fleet initialization in `72211ab`,
1946
+ and the clean-consumer gate in `1a6dbca`. The 0.1.0 candidate gate completed 49
1947
+ sequential commands with 465 MiB peak aggregate RSS.
1948
+
1949
+ ### C53 — Make repair observable, cancellable, and resumable
1950
+
1951
+ - Status: implemented
1952
+ - Priority: P0
1953
+ - DependsOn: C47, C52
1954
+ - Touches: `src/engine/index.ts`, `bin/repair-progress.ts`, `bin/cli.ts`,
1955
+ `docs/specs/repair-progress-telemetry.md`, index-health and CLI progress tests,
1956
+ packed-install repair acceptance
1957
+ - What: expose truthful phase/file progress for long repairs, preserve completed
1958
+ transactions on cancellation, and provide composable human and machine CLI
1959
+ output.
1960
+
1961
+ ### Acceptance and evidence
1962
+
1963
+ 1. Engine events carry monotonic sequences and counters, repository-relative
1964
+ source-safe paths, Salesforce metadata families, and terminal completion,
1965
+ cancellation, or failure without allowing callback errors to fail repair.
1966
+ 2. `AbortSignal` boundaries leave committed files recoverable, avoid recording
1967
+ a successful reconciliation on cancellation, and let the next audit resume
1968
+ from `index_state`; concurrent callers receive their own event stream.
1969
+ 3. CLI progress supports automatic TTY, plain, JSON, JSONL, and silent modes;
1970
+ progress uses stderr except for JSONL envelopes, and final results retain
1971
+ deterministic stdout.
1972
+ 4. Unit/CLI tests and the installed-tarball smoke gate cover cancellation,
1973
+ resumption, throttling, heartbeat, stream separation, and argument
1974
+ validation.
1975
+
1976
+ Evidence: the engine event foundation landed in `e0dcb5f`, cancellation and
1977
+ resumption in `1dc74cb`, and CLI rendering/validation in `a05edcd` and
1978
+ `c0d6745`; the installed-tarball gate exercises all five CLI output modes.
1979
+
1980
+ ### C54 — Certify the local 0.1.0 release artifact
1981
+
1982
+ - Status: implemented
1983
+ - Priority: P0
1984
+ - DependsOn: C50, C52, C53
1985
+ - Touches: `package.json`, `bun.lock`, `README.md`,
1986
+ `docs/releases/0.1.0.md`, this roadmap,
1987
+ `docs/specs/repair-progress-telemetry.md`,
1988
+ `scripts/pack-install-smoke.ts`, compiled package/tarball evidence
1989
+ - What: version the first fleet-consumable package, prove it from a clean
1990
+ consumer under the 3 GiB development cap, reconcile roadmap/spec truth, and
1991
+ produce a local release commit and tag without claiming a remote or registry
1992
+ publication.
1993
+
1994
+ ### Acceptance and evidence
1995
+
1996
+ 1. Package metadata identifies `0.1.0`, and `bun.lock` is regenerated from that
1997
+ manifest (Bun's lock format does not duplicate the workspace version); the
1998
+ exact `knodin-0.1.0.tgz` passes the clean-consumer acceptance gate with
1999
+ an aggregate process-tree limit of 2,750 MiB.
2000
+ 2. Lint, typecheck, build, documentation checks, and the full suite pass
2001
+ sequentially under the memory cap; C50 is not closed through quarantine,
2002
+ skipped files, or forced exit. Fresh-worker shards use a bounded five-minute
2003
+ default test ceiling for loaded-host graph builds while replay subprocesses
2004
+ retain tighter explicit limits.
2005
+ 3. README and release notes distinguish source-checkout use, local packed
2006
+ installation, and actual npm/remote publication state.
2007
+ 4. Tasks 1–3 in the repair telemetry specification are recorded complete while
2008
+ the MCP progress bridge, `--plan`, embedding detail, large Salesforce replay,
2009
+ and incremental embedding work remain explicitly pending.
2010
+
2011
+ Evidence to date: the final exact `knodin-0.1.0.tgz` replay passed 49
2012
+ sequential clean-consumer commands in 43,917 ms with 478 MiB peak aggregate RSS
2013
+ against the 2,750 MiB fail-closed limit. Final integration lint, typecheck,
2014
+ build, 503-test suite, and merged coverage gates pass under 3 GiB (the highest
2015
+ observed gate RSS is 901,840,896 bytes). The local release commit is tagged
2016
+ `v0.1.0`; no remote or registry publication is implied.
2017
+
2018
+ ### C55 — Pin and audit the authorized Token Optimizer source
2019
+
2020
+ - Status: implemented
2021
+ - Priority: P0
2022
+ - DependsOn: C24
2023
+ - Touches:
2024
+ `benchmarks/evaluations/token-optimizer-20260730/comparison-manifest.json`,
2025
+ `scripts/token-optimizer-source-audit.ts`, immutable raw audit evidence, and
2026
+ manifest-consistency tests
2027
+ - What: replace the demo-only source assumption with an authorized,
2028
+ commit-bound GHES audit while preserving the distinction between source
2029
+ evidence and equivalent behavioral measurement.
2030
+
2031
+ ### Acceptance and evidence
2032
+
2033
+ 1. The comparison pins repository, revision, Git tree, file blobs, README hash,
2034
+ and demo hash; branch movement cannot alter the evaluated input.
2035
+ 2. An executable audit fetches blobs by immutable revision and fails closed on
2036
+ a different commit, truncated tree, missing file, contradicted finding, or
2037
+ failed behavioral probe.
2038
+ 3. The audit verifies parser strategy, checked-in test inventory, shell and
2039
+ containment posture, path scope, telemetry fields/defaults, token/baseline
2040
+ math, compression budget behavior, dependency range, update/provenance
2041
+ inventory, and installation steps.
2042
+ 4. The manifest remains `behavior-partial`: source verification must not be
2043
+ presented as production routing, savings, or end-to-end parity evidence.
2044
+
2045
+ Evidence:
2046
+ `benchmarks/evaluations/token-optimizer-20260730/raw-source-audit-20260730.json`
2047
+ pins commit `a2b9cef5efdfcf0393e9e7a5ded229ffd9f613d7`, 17 files, and 16 verified
2048
+ findings. Its behavioral probe returns 104 lines for 100 matching input lines
2049
+ despite `max_lines=50`.
2050
+
2051
+ ### C56 — Replay all seven Token Optimizer workflows
2052
+
2053
+ - Status: implemented
2054
+ - Priority: P0
2055
+ - DependsOn: C55
2056
+ - Evidence:
2057
+ `benchmarks/evaluations/token-optimizer-20260730/raw-behavior-replay-20260730.json`
2058
+
2059
+ The immutable replay executes 20 warm and five cold-process samples per
2060
+ operation, uses the real `o200k_base` tokenizer, and applies shared correctness
2061
+ oracles to all five structural workflows. Both products pass those oracles.
2062
+ The final C68 replay records knodin as smaller in real tokens and faster at warm
2063
+ p50/p95 for all five structural cases on the small fixture; its cold-process
2064
+ startup remains materially slower, and command execution remains gated. The competitor compressor
2065
+ preserves the labeled diagnostic strings but returns 29 lines for a requested
2066
+ maximum of 20. knodin's C57 compressor preserves 7/7 detected signals in 14
2067
+ lines with hard byte/line budgets, exact omission accounting, and retained
2068
+ local drill-down. Its selected content is smaller; its complete recovery
2069
+ envelope remains 18 tokens larger on this fixture.
2070
+
2071
+ ### C57 — Bounded recoverable diagnostic-output compression
2072
+
2073
+ - Status: implemented
2074
+ - Priority: P0
2075
+ - DependsOn: C3, C55
2076
+ - Evidence:
2077
+ `src/output-compression.ts`,
2078
+ `src/__tests__/unit/output-compression.spec.ts`,
2079
+ `src/__tests__/unit/output-compression-cli.spec.ts`,
2080
+ `benchmarks/evaluations/token-optimizer-20260730/raw-behavior-replay-20260730.json`,
2081
+ and `docs/COMMAND-OUTPUT-COMPRESSION.md`
2082
+
2083
+ knodin accepts already-produced combined text, ordered stdout/stderr events, or
2084
+ a repository-contained log artifact. It deterministically applies `smart`,
2085
+ `head-tail`, or `errors-only` selection with generic and ecosystem-specific
2086
+ adapters; enforces hard rendered-line and UTF-8 content-byte budgets; preserves
2087
+ exit/signal metadata; reports unpreserved signals rather than overstating
2088
+ fidelity; records exact omitted line/source-byte ranges; and retains bounded
2089
+ private drill-down by SHA-256 identity without rerunning a command.
2090
+
2091
+ The pinned adversarial replay measures 604 raw tokens. knodin returns 109
2092
+ content tokens and a 205-token recovery envelope while preserving 7/7 detected
2093
+ signals in 14/20 lines. The competitor returns 187 tokens and preserves every
2094
+ golden string, but emits 29 lines for a requested maximum of 20. This supports
2095
+ the narrow `verified-stronger-budget-and-recovery` classification, not a broad
2096
+ claim that every compression dimension is stronger.
2097
+
2098
+ Command execution is not part of C57. C58's portable-containment evaluation was
2099
+ rejected, and direct MCP `text` input does not prove context savings when raw
2100
+ text has already crossed the model boundary.
2101
+
2102
+ ### C58 — Gate command execution on reviewed portable containment
2103
+
2104
+ - Status: evaluated — rejected
2105
+ - Priority: P0
2106
+ - Disposition: evaluated — rejected
2107
+ - DependsOn: C57
2108
+ - Owner roles: security reviewer, release owner, and one platform certifier for
2109
+ each supported operating system
2110
+ - Touches only if approved: a disposable command-runner prototype, adversarial
2111
+ containment fixtures, platform evidence, threat model, and a separately
2112
+ reviewed production implementation item
2113
+ - Evidence: `benchmarks/evaluations/c58-containment/` contains the disposable
2114
+ prototype, adversarial fixture, native macOS raw result, methodology, threat
2115
+ model, tests, and independent `c58-security-review-r1` dated 2026-08-02.
2116
+ - Terminal disposition: `evaluated — rejected`. Native macOS evidence passed
2117
+ the bounded defense-in-depth cases, but repository cwd/path checks are not a
2118
+ filesystem or network sandbox, process polling retains escape/PID races, and
2119
+ Linux/Windows are explicitly unavailable. Production execution remains
2120
+ absent and no implementation item is created.
2121
+ - Post-evaluation gate hardening: the three-level fixture now starts its 100 ms
2122
+ containment window only after an exact bounded readiness handshake, with a
2123
+ separate startup deadline and silent-tree cleanup test. This removes a
2124
+ coverage-timing race without changing raw evidence, platform omissions,
2125
+ security limitations, or the rejected disposition.
2126
+
2127
+ The evaluation decides whether knodin can safely execute a caller-supplied
2128
+ command before compressing its output. C57's already-produced text and artifact
2129
+ inputs remain the production default. Merely matching a competitor workflow is
2130
+ not sufficient to authorize execution.
2131
+
2132
+ Evaluation gate:
2133
+
2134
+ 1. Build no production command surface. First produce a disposable prototype
2135
+ using no-shell argv execution, repository-root binding, a minimal sanitized
2136
+ environment, stdin closure, output/time/RSS/process-count limits, and
2137
+ process-tree termination on every proposed supported platform.
2138
+ 2. Check in adversarial fixtures for shell metacharacters, traversal, symlink
2139
+ escape, inherited secrets, descriptor inheritance, child/grandchild escape,
2140
+ detached processes, output floods, timeout races, signals, and cancellation.
2141
+ 3. Record native macOS, Linux, and Windows results independently. Unsupported
2142
+ containment primitives or an unavailable platform are explicit omissions,
2143
+ not inferred portability.
2144
+ 4. Require an independent security review of the threat model and residual
2145
+ risks. Retain its revision, date, findings, and disposition without including
2146
+ credentials or private source.
2147
+ 5. Close with exactly one disposition: `evaluated — rejected`, `evaluated —
2148
+ deferred`, or a newly numbered implementation item with its own acceptance
2149
+ criteria. C58 itself never authorizes shipping from prototype evidence.
2150
+ 6. If rejected or deferred, keep command execution absent and update the
2151
+ scorecard limitation. This is a valid terminal outcome and does not block
2152
+ C67 when the exclusion remains explicit.
2153
+
2154
+ ### C57–C67 — Token Optimizer parity and trusted-distribution program
2155
+
2156
+ - Status: parked-external-evidence — 2026-08-02
2157
+ - Previous status: active
2158
+ - Source objective: the checked-in comparison manifest and the DTS competitive
2159
+ goal accepted on 2026-07-30
2160
+ - Ordering: behavioral replay (C56) precedes claims; safe compression (C57)
2161
+ precedes failure-to-code intelligence (C59); command execution (C58) remains
2162
+ excluded after its rejected evaluation; signed trust (C62) precedes update
2163
+ mutation (C63), release attestation
2164
+ (C64), and adversarial supply-chain proof (C65); the scorecard and release are
2165
+ terminal gates.
2166
+
2167
+ The program must preserve these invariants:
2168
+
2169
+ 1. Every competitor capability is either verified equivalent/stronger, recorded
2170
+ weaker, or explicitly excluded by a reviewed safety gate.
2171
+ 2. Compression has a hard truthful budget, exact omission accounting, retained
2172
+ local drill-down data, and adversarial diagnostic-fidelity fixtures.
2173
+ 3. No command runner ships without no-shell argv execution, repository binding,
2174
+ sanitized environment, process-tree/resource/output containment, audit, and
2175
+ supported-platform evidence.
2176
+ 4. ROI telemetry is local and opt-in, uses a real tokenizer, separates measured
2177
+ and modeled baselines, and excludes source, raw output, command text,
2178
+ usernames, and absolute paths by default.
2179
+ 5. Update availability is non-blocking and privacy-preserving. Mutation requires
2180
+ expiring threshold-signed metadata, rollback/freeze/mix-and-match defenses,
2181
+ digest/length/provenance checks, quarantine, health validation, atomic
2182
+ activation, and rollback.
2183
+ 6. A compromised npm, Homebrew, Artifactory, GitHub SaaS, or GHES account alone
2184
+ cannot authorize installation.
2185
+ 7. The terminal scorecard links every cell to source, executable tests,
2186
+ benchmark output, and a limitation. Broad superiority language remains
2187
+ prohibited until C66 and C67 close.
2188
+
2189
+ ### C59 — Source-evidenced failure-to-code diagnosis
2190
+
2191
+ - Status: implemented
2192
+ - Priority: P1
2193
+ - DependsOn: C57
2194
+ - Evidence:
2195
+ `src/failure-diagnosis.ts`,
2196
+ `src/__tests__/unit/failure-diagnosis.spec.ts`,
2197
+ `src/__tests__/unit/output-compression-cli.spec.ts`,
2198
+ `src/__tests__/unit/knodin-tools.spec.ts`, and
2199
+ `docs/COMMAND-OUTPUT-COMPRESSION.md`
2200
+
2201
+ The retained compression artifact can be diagnosed through
2202
+ `knodin compress diagnose <artifact-id>` or
2203
+ `operation: "compress", compressionAction: "diagnose"` without rerunning the
2204
+ failed command or adding a second MCP tool. Direct already-produced text is
2205
+ also accepted behind the same input, redaction, and byte limits.
2206
+
2207
+ The resolver extracts common JavaScript/TypeScript, Python, Java, C#, Go, and
2208
+ Rust source locations; binds exact or unique-suffix paths to tracked files;
2209
+ refuses traversal, external symlinks, untracked paths, and ambiguous basenames;
2210
+ and returns source-evidenced owning symbols, stable identities, nearest package
2211
+ manifests, tests, upstream callers, downstream dependencies, recent changes,
2212
+ bounded snippets, and indexed/current commit freshness.
2213
+
2214
+ The aggregate source-context byte budget is hard. Incomplete input, exhausted
2215
+ context, unresolved locations, or non-current graph state makes the result
2216
+ `partial` rather than complete. Static relationships remain diagnostic
2217
+ candidates rather than runtime-causality proof; source maps, generated paths,
2218
+ dynamic dispatch, and framework wiring are recorded limitations.
2219
+
2220
+ ### C60 — Cross-client and GHES onboarding
2221
+
2222
+ - Status: parked-external-evidence — 2026-08-02
2223
+ - Previous status: active — the legacy GHES source and versioned 0.3.0 release path are
2224
+ live; local fail-closed mirror preparation is complete, while enterprise
2225
+ governance/workload identity, timed independent cross-platform onboarding,
2226
+ and Linux/Windows named-client certification remain external gates
2227
+ - Priority: P1
2228
+ - DependsOn: C52, C55
2229
+ - Evidence: `docs/PT-ACCESS-RECOMMENDATION.md`, `docs/INSTALLATION.md`,
2230
+ `docs/evidence/pt-ghes-onboarding-2026-07-31.md`,
2231
+ `docs/evidence/c60-macos-codex-native-2026-08-02.json`,
2232
+ `docs/evidence/c60-macos-native-rss-fix-2026-08-02.json`,
2233
+ `schemas/c60-named-client-evidence-v2.schema.json`,
2234
+ `scripts/c60-native-certify.ts`, `scripts/c60-native-platform.ts`,
2235
+ `scripts/c60-macos-codex-certify.ts`,
2236
+ `src/__tests__/unit/c60-native-platform.spec.ts`,
2237
+ `src/__tests__/unit/agent-integration.spec.ts`,
2238
+ `src/__tests__/unit/doctor.spec.ts`, `scripts/pack-install-smoke.ts`,
2239
+ `config/enterprise-mirror-policy.json`, and
2240
+ `docs/evidence/c60-autonomous-preparation-2026-08-02.md`
2241
+
2242
+ The `Enterprise-Apps/knodin` GHES repository is the DTS / Application Engineering
2243
+ read-only downstream mirror of the authoritative GitHub SaaS repository
2244
+ `DTS-Productivity-Engineering/knodin`. P&T engineers without GitHub SaaS
2245
+ access can download an exact versioned tarball from its
2246
+ release page, verify the recorded SHA-256, install with lifecycle scripts
2247
+ disabled, initialize an existing checkout, and diagnose the single MCP gateway.
2248
+ The observed 0.3.0 GHES asset is byte-identical to the retained CI artifact and
2249
+ public npm artifact. Docusign Artifactory currently stores the same bytes as a raw
2250
+ artifact, not npm registry metadata; the docs do not misrepresent that endpoint
2251
+ as an npm registry.
2252
+
2253
+ Named Claude, Codex, Gemini, and Antigravity configuration plus a real generic
2254
+ MCP handshake are executable tests. The immutable schema-v2 macOS arm64
2255
+ Codex-configuration plus generic-MCP record certifies the exact 0.6.0 tarball in
2256
+ 13.739 seconds using pre-populated offline npm/model caches: lifecycle-disabled
2257
+ install, existing-checkout init, project-local configuration without user-home
2258
+ mutation, bounded MCP initialize/list/status/query/doctor, and aggregate
2259
+ RSS/output evidence all pass.
2260
+ An isolated offline build and npm pack from the recorded source commit is
2261
+ byte-identical to the consumed tarball, with its prerequisite commands
2262
+ separately bounded and recorded. It does not prove cold-cache installation, an
2263
+ actual Codex process or active UI
2264
+ session, or the eventual distributed release artifact. Linux packed consumption
2265
+ remains release-gated. A first real Linux attempt falsified the original
2266
+ portability claim because zero-RSS kernel rows made RSS parsing fail before the
2267
+ first command completed; Windows PID 0 had the same defect. The fixed parser
2268
+ retains structural and integer validation, filters non-process/zero-resident
2269
+ rows, and rejects an all-filtered snapshot. A new provenance-bound macOS replay
2270
+ of the exact fix commit passed in 11.212 seconds with byte-identical package
2271
+ evidence, proving no macOS regression but not Linux/Windows readiness. The
2272
+ native Linux and Windows runs must restart from the fixed commit without source
2273
+ changes. GHES `main` is not protected, GHES releases are not
2274
+ immutable, and native Linux/Windows named-client runs, independent cross-platform review,
2275
+ and non-personal automated mirror synchronization are missing, so C60 remains
2276
+ active and no cross-platform-complete or universal two-minute claim is allowed.
2277
+
2278
+ ### C61 — Private real-token ROI telemetry and dashboard
2279
+
2280
+ - Status: implemented
2281
+ - Priority: P1
2282
+ - DependsOn: C16, C55
2283
+ - Evidence: `src/output-telemetry.ts`, `src/cli-model.ts`, `bin/cli.ts`,
2284
+ `src/tools/knodin-tools.ts`, `src/__tests__/unit/output-telemetry.spec.ts`,
2285
+ `src/__tests__/unit/output-telemetry-cli.spec.ts`,
2286
+ `src/__tests__/unit/knodin-tools.spec.ts`, and `docs/TELEMETRY.md`
2287
+
2288
+ Persistence remains disabled by default and local. Opt-in JSONL records use the
2289
+ pinned real tokenizer, executed full-file baselines only when bounded source is
2290
+ available, explicit unknowns otherwise, cold/warm and outcome dimensions,
2291
+ agent-round-trip, latency, RSS, freshness, compression-fidelity, indexing-time,
2292
+ and index-storage measurements. Repository grouping uses a one-way identifier;
2293
+ source, raw output, command text, usernames, and absolute paths are excluded.
2294
+
2295
+ The CLI and single MCP gateway expose status, static HTML report, sanitized JSON
2296
+ export, retention, and explicit clear operations. Inputs and outputs are
2297
+ repository-contained and symlink-refusing, persistence is atomic/private, and
2298
+ the self-contained dashboard reports summary, operation, time-series, private
2299
+ repository, latency p50/p95, indexing cost, freshness, and fidelity dimensions
2300
+ with light/dark rendering. Metrics with no defensible observation remain
2301
+ unavailable rather than modeled.
2302
+
2303
+ ### C62 — Signed update metadata trust
2304
+
2305
+ - Status: parked-external-evidence — 2026-08-02
2306
+ - Previous status: active — verification core and fail-closed local ceremony preparation
2307
+ implemented; production offline custodians, independent pin review, recovery
2308
+ evidence, and trust-pin activation remain external gates
2309
+ - Priority: P0
2310
+ - DependsOn: C55
2311
+ - Evidence: `src/update-trust.ts`,
2312
+ `src/__tests__/unit/update-trust.spec.ts`, `src/update-ceremony.ts`,
2313
+ `src/__tests__/unit/update-ceremony.spec.ts`,
2314
+ `scripts/root-ceremony-preflight.ts`,
2315
+ `schemas/root-ceremony-manifest-v1.schema.json`,
2316
+ `schemas/root-pin-review-receipt-v1.schema.json`,
2317
+ `docs/ROOT-CEREMONY-RUNBOOK.md`, and `docs/SIGNED-UPDATES.md`
2318
+
2319
+ The pure verifier implements deterministic signed payloads, content-addressed
2320
+ Ed25519 keys, disjoint root/targets/snapshot/timestamp roles, production
2321
+ thresholds for root and targets, bounded expiry, independent root pinning,
2322
+ dual-threshold one-step root rotation, version-and-digest rollback/equivocation
2323
+ state, timestamp→snapshot→targets length/hash/version binding, safe target
2324
+ paths, and signed artifact SHA-256/length verification. It performs no network
2325
+ or installation mutation.
2326
+
2327
+ Adversarial fixtures reject missing thresholds, field and artifact
2328
+ substitution, expired timestamp freezes, rollback, same-version equivocation,
2329
+ mix-and-match metadata, role-key reuse, root-pin substitution, future metadata,
2330
+ and rotations not signed by the old root threshold.
2331
+
2332
+ C62 is not closed merely because the verification code exists. No production
2333
+ private key is checked in or generated by tests. Release owners must complete
2334
+ the offline multi-custodian root ceremony, independently review and embed the
2335
+ initial root digest, and record rotation/recovery evidence. C63 now makes all
2336
+ doctor, CLI, and MCP update state consume only verified metadata; the former
2337
+ npm lookup has been removed. Production authorization still waits for the
2338
+ offline ceremony.
2339
+
2340
+ ### C63 — Signed-only update client and safe policy
2341
+
2342
+ - Status: parked-external-evidence — 2026-08-02
2343
+ - Previous status: active — implementation and adversarial unit/CLI evidence complete;
2344
+ closure requires a recorded fail-closed deferred-activation policy after C62
2345
+ - Priority: P0
2346
+ - DependsOn: C62
2347
+ - Evidence: `src/update-policy.ts`,
2348
+ `src/__tests__/unit/update-policy.spec.ts`,
2349
+ `src/__tests__/unit/update-policy-cli.spec.ts`,
2350
+ `docs/DOCTOR-AND-UPDATES.md`, and `docs/SIGNED-UPDATES.md`
2351
+
2352
+ The five `knodin update` actions work outside a repository. Periodic checks use
2353
+ an atomic background lease and never put network latency on ordinary commands.
2354
+ Fixed signed metadata paths are HTTPS-origin/path bound, redirect-free, capped,
2355
+ timed, privacy-preserving, and optionally byte-identical across mirrors.
2356
+ Verified state is atomic/private and retains monotonic evidence across failed
2357
+ checks.
2358
+
2359
+ Apply enforces notify/download/patch/minor policy, minimum age, channels,
2360
+ enterprise mirrors, maintenance windows, exact artifact length/digest, signed
2361
+ provenance digest, candidate/rollback quarantine, no-shell exact local npm
2362
+ artifact invocation, symlink-safe exclusive quarantine, signed-version path
2363
+ validation, allowlisted enterprise cross-check origins, health validation,
2364
+ health-verified automatic rollback, and offline re-verification for explicit
2365
+ rollback. Mutable registry tags are never used.
2366
+
2367
+ The production CLI remains fail-closed as `trust-unconfigured`: a root and its
2368
+ pin from one mutable user file are not accepted. No certified sandbox smoke
2369
+ adapter or Homebrew/Node-manager rollback adapter is enabled yet. C62's offline
2370
+ ceremony must embed the independent anchor. C63 closes by recording that
2371
+ automatic production activation remains disabled; it does not wait for C65.
2372
+ C65 then uses the terminal signed-only client to certify the smoke sandbox,
2373
+ multi-step rotation/recovery, compromised channels, and manager-specific
2374
+ activation before any later auto-apply enablement. This ordering removes a
2375
+ C63↔C65 closure cycle without weakening the fail-closed policy.
2376
+
2377
+ ### C64 — Provenance, SBOM, and cross-channel release attestation
2378
+
2379
+ - Status: parked-external-evidence — 2026-08-02
2380
+ - Previous status: active — implementation and adversarial fixtures complete; a live
2381
+ five-channel release attestation remains a C67 publication gate
2382
+ - Priority: P0
2383
+ - DependsOn: C62
2384
+ - Evidence: `src/release-attestation.ts`,
2385
+ `src/__tests__/unit/release-attestation.spec.ts`,
2386
+ `src/__tests__/unit/release-attestation-cli.spec.ts`,
2387
+ `src/__tests__/unit/release-workflow.spec.ts`,
2388
+ `src/__tests__/unit/release-candidate-workflow.spec.ts`,
2389
+ `scripts/release-attestation.ts`, `.github/workflows/publish.yml`,
2390
+ `.github/workflows/release-candidate.yml`, and `docs/RELEASING.md`
2391
+
2392
+ The release workflow installs dependencies with lifecycle scripts disabled,
2393
+ runs tests, and packs the artifact in an unprivileged job with read-only
2394
+ repository access. A separate job with OIDC and attestation authority downloads
2395
+ only the retained artifact, generates a production-dependency CycloneDX SBOM,
2396
+ and uses full-commit-pinned official GitHub actions to sign SLSA build provenance
2397
+ and the SBOM. It independently verifies both retained bundles against the exact
2398
+ artifact, source digest, tag ref, repository, workflow certificate identity,
2399
+ and GitHub OIDC issuer before npm publication. A manual retry must run at the
2400
+ tag ref; checked-out commit, workflow SHA, tag, and package version disagreement
2401
+ is a hard failure. The exact tarball, signed evidence, and verification receipts
2402
+ are retained for 90 days before that tarball is passed to npm and Artifactory.
2403
+
2404
+ The separate manual-only release-candidate workflow now structurally stops
2405
+ after the exact artifact and rollback bytes, provenance and SBOM bundles,
2406
+ verification and identity receipts, bounded input manifest, and incomplete
2407
+ five-channel attestation draft are retained for 90 days. Its build job is
2408
+ unprivileged and lifecycle-disabled; its OIDC job downloads the retained input
2409
+ set without a source checkout or dependency install and has no npm/JFrog/channel
2410
+ credential, publishing command, reusable publisher, or distribution mutation.
2411
+ This is local implementation evidence only: C64 remains active until a live OIDC
2412
+ run at an immutable tag produces the retained candidate bundle and independent
2413
+ verification. C67 alone authorizes distribution after that candidate is
2414
+ independently verified.
2415
+
2416
+ The canonical schema derives, rather than accepts, `verified`, `incomplete`, or
2417
+ `quarantined` status. It records the exact artifact SHA-256/SHA-512/npm
2418
+ integrity, source commit, build identity, provenance and SBOM digests,
2419
+ publication timestamps and locations, all five approved destinations, and an
2420
+ exact rollback target whose digest the CLI derives from downloaded bytes rather
2421
+ than accepting from the manifest. Missing evidence is incomplete; any artifact, version,
2422
+ or source disagreement is quarantined. The CLI binds inputs beneath one
2423
+ directory, refuses direct or nested symlinks/traversal/unbounded files/output replacement,
2424
+ writes visible failure evidence, and exits nonzero for an incomplete final or
2425
+ any mismatch. It ignores a manifest's claimed verification state and invokes
2426
+ an explicitly approved absolute `gh attestation verify` executable without a
2427
+ shell or inherited credentials/environment against both local bundles and the
2428
+ exact artifact/source/ref/workflow/certificate policy before deriving verified
2429
+ signature evidence.
2430
+
2431
+ Signed targets may bind the provenance, SBOM, SBOM-attestation, and release-
2432
+ attestation digests. Update quarantine and explicit rollback re-fetch or re-read
2433
+ all four objects, validate their exact bytes and attested channel state, and
2434
+ refuse missing or inconsistent evidence before a package manager is invoked.
2435
+ Apply cross-checks the candidate's attested rollback version, path, and digest
2436
+ against the separately quarantined rollback artifact.
2437
+
2438
+ C64 is not a claim that the current 0.3.0 channels have been retroactively
2439
+ attested. GitHub/Sigstore build signatures are independently useful evidence,
2440
+ but do not replace C62's offline-controlled threshold root. C64 closes when the
2441
+ exact release-candidate artifact has verified provenance and SBOM bundles plus
2442
+ a complete attestation draft containing the five intended destinations. The
2443
+ draft remains `incomplete` until distribution. C67 publishes those exact bytes,
2444
+ fills the destination receipts, derives the final `verified` record, and
2445
+ TUF-authorizes it. This split prevents C64 and C67 from depending on each
2446
+ other's terminal evidence.
2447
+
2448
+ The signature proves workflow identity and exact bytes, not that a compromised
2449
+ self-hosted runner built those bytes faithfully. Independent build reproduction
2450
+ and runner-compromise drills remain C65/C67 gates. The JFrog reusable workflow's
2451
+ `@master` reference also remains an explicit external constraint because its
2452
+ Vault OIDC policy rejects commit-pinned `job_workflow_ref` claims; this is not
2453
+ misreported as full workflow immutability.
2454
+
2455
+ ### C65 — Compromised-channel, rollback, freeze, and recovery defenses
2456
+
2457
+ - Status: parked-external-evidence — 2026-08-02
2458
+ - Previous status: active — bounded multi-step root recovery is implemented; production
2459
+ key ceremony, smoke sandbox, manager-specific activation, and live drills
2460
+ remain required
2461
+ - Priority: P0
2462
+ - DependsOn: C63, C64
2463
+ - Evidence: `src/update-trust.ts`, `src/update-policy.ts`,
2464
+ `src/__tests__/unit/update-trust.spec.ts`,
2465
+ `src/__tests__/unit/update-policy.spec.ts`, `docs/SIGNED-UPDATES.md`, and
2466
+ `docs/evidence/0.4.3-update-defense-drill.json`
2467
+
2468
+ The update client can recover from up to 32 missed root rotations per check by
2469
+ fetching every numbered intermediate root, cross-checking its exact bytes across
2470
+ configured mirrors, and verifying each version under both the previous and new
2471
+ root thresholds. It never accepts a direct version jump or promotes a root from
2472
+ mutable state. Missing, malformed, wrong-version, skipped, or mirror-divergent
2473
+ intermediates fail before new trust is persisted; the rotation bound prevents a
2474
+ malicious terminal version from causing unbounded requests.
2475
+
2476
+ This closes the implemented multi-step retrieval gap only. C65 still requires
2477
+ the offline multi-custodian production root/recovery rehearsal, certified smoke
2478
+ sandbox, npm/Homebrew/manager activation and rollback drills, compromised
2479
+ transport and runner exercises, and checked-in evidence from those executions.
2480
+ No automatic production activation or complete supply-chain claim is allowed
2481
+ until those gates and C67's live release are complete.
2482
+
2483
+ The checked-in 0.4.3 isolated drill records one combined execution of 52 trust,
2484
+ policy, CLI, and attestation scenarios, including compromised-fixture refusal,
2485
+ multi-step recovery, automatic rollback, and unhealthy-rollback refusal. It
2486
+ strengthens repeatable local evidence but deliberately leaves the production
2487
+ ceremony, certified adapters/sandbox, real manager mutation, and live
2488
+ transport/runner exercises open.
2489
+
2490
+ ### C66 — Evidence-linked Token Optimizer scorecard
2491
+
2492
+ - Status: parked-external-evidence — 2026-08-02
2493
+ - Previous status: active — scorecard published; terminal prerelease snapshot awaits
2494
+ active dependencies
2495
+ - Priority: P0
2496
+ - DependsOn: C56, C57, C59, C60, C61, C65, C68, C75
2497
+ - Evidence: `docs/TOKEN-OPTIMIZER-SCORECARD.md`,
2498
+ `benchmarks/evaluations/token-optimizer-20260730/scorecard.json`,
2499
+ `src/__tests__/unit/docs-integrity.spec.ts`, and the pinned audit/replay
2500
+ artifacts under `benchmarks/evaluations/token-optimizer-20260730/`
2501
+
2502
+ The human scorecard and machine manifest classify every advertised competitor
2503
+ workflow plus the cross-cutting product dimensions. Every row links checked-in
2504
+ implementation, verification source/test paths, benchmark or live evidence,
2505
+ scope, readiness, and a material limitation. The manifest pins the competitor,
2506
+ the raw replay digest, its declared base revision, every evaluated source hash,
2507
+ and the first merged knodin commit containing that exact composite. Its integrity test
2508
+ rejects dimension omissions/duplicates, revision or replay-digest drift,
2509
+ missing evidence paths, or empty scope/limitation fields.
2510
+
2511
+ The scorecard supports a qualified “more complete local code-intelligence
2512
+ system” statement. It explicitly prohibits “better across all dimensions”:
2513
+ command execution remains intentionally excluded, Windows/timed onboarding is
2514
+ not certified, cold startup is slower, and production trust ceremony, drills,
2515
+ and five-channel release evidence remain open. C66 publication does not close
2516
+ its active dependencies or C67. C66 becomes terminal when every listed
2517
+ dependency is terminal and the scorecard records the prerelease truth. C67
2518
+ appends the final release record and receipts as a non-gating amendment; C66
2519
+ does not depend on the release it gates.
2520
+
2521
+ ### C67 — Certify and distribute the completed competitive release
2522
+
2523
+ - Status: parked-external-evidence — 2026-08-02
2524
+ - Previous status: proposed
2525
+ - Priority: P0
2526
+ - Disposition: Must close
2527
+ - DependsOn: C60, C62, C63, C64, C65, C66, C75
2528
+ - Owner roles: release owner, offline-root custodians meeting the recorded
2529
+ threshold, security reviewer, five channel operators, and Windows/macOS/Linux
2530
+ certifiers
2531
+ - Evidence directory: `docs/evidence/releases/<version>/` with immutable raw
2532
+ receipts under `benchmarks/evaluations/releases/<version>/`
2533
+ - Preflight implementation: `src/release-preflight.ts`,
2534
+ `scripts/release-preflight.ts`, `schemas/release-plan-v1.schema.json`,
2535
+ `src/__tests__/unit/release-preflight.spec.ts`, and `docs/RELEASING.md`
2536
+
2537
+ C67 is the terminal release and roadmap gate. It certifies one exact version
2538
+ from one source commit and distributes one byte-identical package through all
2539
+ approved channels without relaxing any unresolved limitation.
2540
+
2541
+ The pure C67 preflight is implemented without release authority. It validates
2542
+ a bounded plan, exact clean annotated-tag identity, the terminal roadmap and
2543
+ ledger state of every dependency, tracked repository-contained nonsymlink
2544
+ evidence including an immutable terminal artifact per dependency, certified
2545
+ native targets, exact verification commands, an independently reviewed
2546
+ production root/ceremony/digest tuple, and a new version-bound evidence path. It performs no publishing,
2547
+ tagging, evidence-directory creation, credential access, or other external
2548
+ mutation and does not change C64's nonpublishing candidate workflow. C67 remains
2549
+ proposed and open: against the current authoritative ledger the preflight must
2550
+ fail until C60, C62, C63, C64, C65, and C66 close; a passing local unit suite is
2551
+ not a release authorization or distribution receipt.
2552
+
2553
+ Acceptance and execution contract:
2554
+
2555
+ 1. Preflight fails unless C60, C62, C63, C64, C65, C66, and C75 have terminal
2556
+ statuses and checked-in evidence. Record package version, tag, source commit,
2557
+ workflow commit, runtimes, target platforms, trust-root digest, and exact
2558
+ verification commands before mutation.
2559
+ 2. From a clean tag checkout, run `npm ci --ignore-scripts`, `npm run lint`,
2560
+ `npm run typecheck`, `npm run build`, `npm run test`,
2561
+ `npm run test:pack-install`, `npm run check:roadmap`, and the release-specific
2562
+ competitive, trust, defense, and attestation replays. Record exit status,
2563
+ elapsed time, peak RSS, and immutable artifact digests without overwriting an
2564
+ earlier run.
2565
+ 3. Build once in the unprivileged release job. Verify the retained package,
2566
+ provenance, SBOM, SBOM attestation, and release-attestation draft before any
2567
+ channel receives the artifact. A retry must use the same tag and bytes.
2568
+ 4. Publish SHA-256/length-identical bytes to npm, Homebrew, Artifactory, GitHub
2569
+ SaaS, and GHES. Record immutable channel identifiers, timestamps, downloaded
2570
+ digests, and independently fetched receipts. Any missing or divergent channel
2571
+ fails the release and triggers quarantine/recovery rather than success.
2572
+ 5. Produce threshold-signed targets/snapshot/timestamp metadata bound to the
2573
+ artifact and C64 evidence digests. Exercise update status, check, explain,
2574
+ apply, health validation, and rollback from clean supported consumers.
2575
+ 6. Re-run named-client onboarding on certified macOS, Linux, and Windows and
2576
+ record C60's two-minute measurement. Verify `knodin init`, lifecycle refresh,
2577
+ CLI query, and the single MCP gateway against an existing checkout.
2578
+ 7. Publish a release record listing every gate, receipt, degradation,
2579
+ unsupported case, rollback target, and authorized product claim. Append those
2580
+ receipts to C66's already-terminal prerelease scorecard without reopening its
2581
+ gate; preserve every limitation whose independent gate did not pass.
2582
+ 8. Mark C67 terminal and remove its ledger row in the same change, then require
2583
+ a fresh `npm run check:roadmap -- --complete` to succeed. If authority or
2584
+ infrastructure is unavailable, retain `blocked` with the exact missing
2585
+ authority and safe next action; never fabricate a receipt.
2586
+
2587
+ ### C68 — Close structural output and latency gaps
2588
+
2589
+ - Status: implemented
2590
+ - Priority: P0
2591
+ - DependsOn: C56
2592
+ - Evidence: the initial C56 replay identified compactness and warm-latency gaps;
2593
+ the final pinned replay proves the C68 compact mode passes its narrow shared
2594
+ oracles and beats competitor response tokens and warm p50/p95 for all five
2595
+ structural operations on the small fixture. Cold startup remains slower.
2596
+
2597
+ Acceptance:
2598
+
2599
+ 1. Add a purpose-built compact structural response mode that preserves stable
2600
+ identity, source location, ambiguity, freshness, and truthful budgets without
2601
+ returning unrelated graph fields or telemetry envelopes.
2602
+ 2. Measure gateway schema overhead separately from per-call output.
2603
+ 3. Avoid repeated freshness/index startup work inside one session and document
2604
+ the irreducible cold cost of a persistent graph separately from operation
2605
+ latency.
2606
+ 4. Re-run the exact C56 fixture. knodin must match or beat competitor response
2607
+ tokens for all five structural workflows and must not be slower in warm
2608
+ operation p50/p95 without an explicitly approved, evidence-backed tradeoff.
2609
+ 5. Correctness, ambiguity, and stale-index oracles must remain passing; size or
2610
+ latency cannot be improved by weakening evidence.
2611
+
2612
+ Implementation evidence: `detailLevel: "compact"` now selects a purpose-built
2613
+ compact structural contract for file summary, exact source, structural search,
2614
+ batch outline, and project overview. Stable `~` identity prefixes are accepted
2615
+ as ambiguity-safe selectors, returned/total counts remain explicit, and no
2616
+ generic telemetry or response-budget envelope is added. Generation-scoped
2617
+ structural caching plus a watcher-invalidated freshness lease removes repeated
2618
+ session work; a no-subprocess HEAD check prevents Git movement from reusing the
2619
+ lease. The pinned C56 replay records gateway schema cost separately and passes
2620
+ all five correctness, token, warm-p50, and warm-p95 gates. Cold-process cost
2621
+ remains explicitly reported as a persistent-graph tradeoff.
2622
+
2623
+ ### C69–C75 — Large-portfolio lifecycle hardening
2624
+
2625
+ - Status: certified with recorded degradations
2626
+ - Priority: P0/P1 as listed above
2627
+ - Source objective:
2628
+ `knodin-Portfolio-Integration-Issues-2026-07-30.md` and the tested candidate
2629
+ commit `c981480db89ffbbb646206a73660e9ab842686ba`
2630
+ - Scope: portfolio discovery, selection, initialization, lifecycle truth,
2631
+ diagnostics, and Salesforce metadata candidate quality.
2632
+ - C75 Evidence: `docs/evidence/portfolio-initialization-2026-07-31.md`
2633
+
2634
+ The candidate commit proves that a 256 MiB subprocess buffer, early selection,
2635
+ and sequential inventory can discover the current 61-repository portfolio, but
2636
+ that commit is evidence rather than the terminal design. It still materializes
2637
+ SalesforceCI's roughly 105 MB `git ls-files` response and portfolio dry-run
2638
+ peaked near the one-GB incremental-memory ceiling.
2639
+
2640
+ Acceptance:
2641
+
2642
+ 1. Inventory tracked files with a streaming or equivalently bounded design.
2643
+ SalesforceCI's 904,072 tracked files cannot abort the portfolio. A large,
2644
+ malformed, timed-out, or otherwise failed repository produces a
2645
+ machine-readable degraded record while other repositories continue.
2646
+ 2. Apply include/exclude selection before tracked-file inventory. An include
2647
+ selector that matches no discovered repository returns an explicit
2648
+ diagnostic rather than a successful empty result.
2649
+ 3. `repos discover ~/code --json` returns valid JSON for all 61 current
2650
+ repositories while remaining below one GB of incremental RSS. Selected
2651
+ searches do not inventory excluded repositories.
2652
+ 4. Make portfolio init and dry-run sequential, resource-releasing, resumable,
2653
+ and bounded. Dry-run performs no model or database work, estimates scale and
2654
+ memory risk, and reports init/update/repair actions accurately. A
2655
+ per-process memory ceiling degrades and continues rather than killing the
2656
+ portfolio run.
2657
+ 5. `knodin configure` is either clearly configuration-only or atomically
2658
+ produces a usable graph and hooks. `knodin init` drains compatible queued
2659
+ lifecycle events after graph health is established; incompatible events
2660
+ receive an exact remediation. A `knodin wait --fresh` deadline must not
2661
+ terminate a healthy lifecycle processor; dead or failed attempts are
2662
+ retried at a bounded interval. Immediate status and search must agree with
2663
+ the success message.
2664
+ 6. Running `knodin doctor` on a non-Git portfolio parent warns and directs the
2665
+ user to portfolio diagnostics. MCP diagnostics separately report on-disk
2666
+ configuration, subprocess handshake, active-client exposure being unknown,
2667
+ reload requirements, and the exact configuration path.
2668
+ 7. Architecture and roadmap candidates use basename/directory-aware evidence,
2669
+ bounded lists, reasons, and confidence. Salesforce object or metadata names
2670
+ containing words such as `Design`, `WizardDesign`, or `Backlog` are not
2671
+ promoted without qualifying path/content evidence.
2672
+ 8. Certification records the 61-repository result, peak RSS, elapsed time,
2673
+ degraded repositories, selector behavior, immediate queryability, and all
2674
+ regression/typecheck/build/package-install gates. C66 and C67 cannot make a
2675
+ portfolio-readiness claim until C75 closes.
2676
+
2677
+ C72 lifecycle evidence is recorded in
2678
+ `docs/evidence/lifecycle-event-drain-2026-07-31.md`. Generated events now bind
2679
+ to the actual worktree root and status separately exposes a pending event and
2680
+ `idle`, `running`, or `stale-lock` processor state. The checked-in wait path
2681
+ retries a stranded compatible processor. The evidence explicitly does not
2682
+ claim that the installed 0.3.0 Homebrew release contains those source changes;
2683
+ distribution remains a C67 gate.
2684
+
2685
+ C71 is implemented by `src/repository-init-process.ts` and the sequential
2686
+ orchestration in `src/repository-management.ts`. Every live repository runs in
2687
+ a disposable process with a 768 MiB RSS ceiling, independent parent-side RSS
2688
+ polling plus worker heartbeats, process-group termination on macOS/Linux,
2689
+ timeout escalation, repository-scoped degradation, aggregate resource evidence,
2690
+ and resumable manifests. `src/__tests__/unit/repository-init-process.spec.ts`
2691
+ and `src/__tests__/unit/repository-management.spec.ts` prove termination,
2692
+ continuation, and single-worker sequencing. That implementation evidence is
2693
+ separate from the following live C75 certification.
2694
+
2695
+ C75 live evidence is recorded in
2696
+ `docs/evidence/portfolio-initialization-2026-07-31.md`. The 69-record run stayed
2697
+ below one GiB, completed 48 selected repositories, degraded 16 workers at the
2698
+ memory ceiling, preserved tracked team integration in four repositories, and
2699
+ continued after every failure. C75 therefore certifies bounded portfolio
2700
+ behavior with named limitations; it does not claim that all repositories are
2701
+ healthy or initialized.
2702
+
2703
+ ### C76 — Responsive declarative CLI model
2704
+
2705
+ - Status: implemented
2706
+ - Priority: P1
2707
+ - DependsOn: C73
2708
+ - Evidence: PR #29 replaced the hand-authored parser/help template with the
2709
+ declarative Commander model and pseudo-terminal width acceptance coverage.
2710
+
2711
+ Acceptance:
2712
+
2713
+ 1. Immediately reflow help to the detected TTY width with deterministic
2714
+ non-TTY output and narrow/wide pseudo-terminal tests.
2715
+ 2. Migrate parsing, validation, global options, nested subcommands, and help to
2716
+ one declarative command model so documentation cannot drift from accepted
2717
+ arguments.
2718
+ 3. Preserve every checked-in legacy invocation and JSON stdout contract.
2719
+ 4. Prefer Commander unless an executable bakeoff disproves the choice:
2720
+ current Commander supplies width-aware help, nested commands, TypeScript
2721
+ support, a documented security policy, and zero runtime dependencies.
2722
+ 5. Pin the exact dependency and integrity through the normal lockfiles and
2723
+ release provenance. Do not adopt a broader dependency tree solely for
2724
+ cosmetic terminal rendering.
2725
+
2726
+ Implementation evidence: `src/cli-model.ts` is the single Commander-backed
2727
+ grammar for global options, root and nested commands, positional contracts,
2728
+ numeric parsing, unknown-option rejection, and root or command help.
2729
+ `bin/cli.ts` validates every non-help invocation through that model and reads
2730
+ shared option values from its parsed result before compatibility dispatch.
2731
+ Commander 15.0.0 is exact-pinned in both lockfiles and has no transitive runtime
2732
+ dependencies. `src/__tests__/unit/cli-model.spec.ts`,
2733
+ `src/__tests__/unit/cli-pty-help.spec.ts`, the existing CLI subprocess suite,
2734
+ and `docs/CLI.md` cover narrow/wide deterministic output, a real Unix
2735
+ pseudo-terminal when available, nested help, invalid syntax, legacy invocation
2736
+ compatibility, packaging, and the Windows-safe non-PTY path.
2737
+
2738
+ ### C77 — Replay CodeFlow edge-provenance and architecture-export claims
2739
+
2740
+ - Status: evaluated — retained knodin unchanged
2741
+ - Priority: P2
2742
+ - Disposition: Evaluate
2743
+ - DependsOn: C24, C27, C42
2744
+ - Motivation: CodeFlow is a browser-local, MIT architecture mapper with broad
2745
+ language coverage, blast-radius views, and raw JSON export. A secondary
2746
+ LinkedIn post says each dependency identifies its extraction mechanism, but
2747
+ that claim was not found in the current repository or README. Knodin already
2748
+ exposes provenance, confidence, exact-versus-heuristic labels, source evidence,
2749
+ and source lines; only a shared replay can establish whether CodeFlow presents
2750
+ uncertain edges or architecture exports more usefully.
2751
+ - In scope: add a pinned CodeFlow adapter or documented browser-local protocol;
2752
+ compare dependency precision/recall, unsupported edges, uncertainty labels,
2753
+ source evidence, blast-radius completeness, export fidelity, output size,
2754
+ latency, and peak RSS on the existing architecture fixtures.
2755
+ - Out of scope: adopting CodeFlow, copying its browser UI, adding a new service,
2756
+ or changing Knodin's extraction architecture solely from vendor or social-post
2757
+ claims.
2758
+ - Touches: `benchmarks/competitors/TRACKER.md`, a CodeFlow comparison report and
2759
+ immutable raw result, the shared competitive harness only if its existing
2760
+ adapter contract is insufficient, and this roadmap.
2761
+ - Collision risk: competitor tracker and shared harness mappings; schedule apart
2762
+ from other competitive-replay edits.
2763
+ - Source: https://github.com/braedonsaunders/codeflow
2764
+ - Source task: `6h8V3JGJMhV3m63h`
2765
+ - Evidence: `benchmarks/competitors/codeflow-vs-knodin.raw.json` records the
2766
+ exact commit and MIT license, local-only protocol, original/normalized edges,
2767
+ false-positive/false-negative sets, provenance coverage, export fidelity,
2768
+ response bytes, 3+20 warm and five-process cold samples, and independently
2769
+ measured RSS. Both products achieved precision/recall 1.0, unsupported-edge
2770
+ rate 0, and blast completeness 1.0. CodeFlow used 105,955,328 bytes peak RSS
2771
+ and had 122.788/186.788 ms cold p50/p95; knodin used 591,446,016 bytes and had
2772
+ 1,056.246/6,332.857 ms. Knodin warm p50/p95 was 0.532/0.736 ms versus
2773
+ CodeFlow's 3.101/8.202 ms. CodeFlow supplied none of the six evaluated
2774
+ file-relationship provenance categories; knodin supplied all six for all
2775
+ three edges. Architecture block exports remain explicitly incomparable.
2776
+ - Disposition: retain knodin unchanged. CodeFlow demonstrates a real cold-start
2777
+ and memory advantage, while knodin already supplies the evaluated provenance
2778
+ and warm-query behavior. Neither the incomparable architecture exports nor
2779
+ feature presence authorizes an architecture or presentation change, and this
2780
+ evaluation makes no aggregate superiority claim.
2781
+
2782
+ Acceptance:
2783
+
2784
+ 1. Pin the tested CodeFlow commit and record its MIT license, local execution
2785
+ path, setup, and limitations; do not require an account, hosted service,
2786
+ paid API, source egress, or new Knodin runtime dependency.
2787
+ 2. Run both products against identical checked-in fixtures and correctness
2788
+ oracles. Preserve false positives, false negatives, unresolved edges, and
2789
+ unavailable cases in immutable raw results.
2790
+ 3. For every exported relationship, record whether the tool supplies relation
2791
+ kind, extractor identity, exact/heuristic confidence, source file and line,
2792
+ and bounded source evidence. Treat the LinkedIn extractor-identity claim as
2793
+ unverified unless the pinned implementation emits it.
2794
+ 4. Report precision/recall, unsupported-edge rate, blast-radius completeness,
2795
+ export fidelity, response bytes, cold/warm p50 and p95, and independently
2796
+ measured peak RSS. Do not publish a win from incomparable or oracle-free rows.
2797
+ 5. Keep incremental spend at zero and incremental memory below 1 GB. Keep all
2798
+ fixture source local; if CodeFlow cannot run without GitHub egress, record a
2799
+ blocker rather than uploading private source.
2800
+ 6. Close with one evidence-based disposition: retain Knodin unchanged, improve
2801
+ provenance presentation/export, or propose a separately reviewed roadmap
2802
+ item. Do not infer an architecture change from feature presence alone.
2803
+
2804
+ ### C78 — Add bounded source-to-sink resource reachability
2805
+
2806
+ - Status: evaluated — retained C8 unchanged
2807
+ - Priority: P1
2808
+ - Disposition: Evaluate
2809
+ - DependsOn: C2, C3, C8, C24
2810
+ - Motivation: Code-Graph-RAG's opt-in `READS_FROM`, `WRITES_TO`, and `FLOWS_TO`
2811
+ edges answer a useful provenance question Knodin cannot currently answer:
2812
+ whether a value from an environment variable, file, database, socket, or
2813
+ network source can reach a log, file, database, or outbound-network sink.
2814
+ Knodin's C8 `flow_analysis` is deliberately limited to bounded, on-demand
2815
+ TS/JS facts inside one selected symbol; call-argument evidence is explicitly
2816
+ heuristic and does not compose source-to-sink resource reachability.
2817
+ - In scope: prototype synthetic resource identities plus source-evidenced read,
2818
+ write, assignment, argument, return, and kill facts; begin with environment or
2819
+ local configuration flowing to logging, network, and database sinks; expose a
2820
+ bounded query with exact coverage and omission metadata.
2821
+ - Out of scope: adopting Code-Graph-RAG, Memgraph, Qdrant, Docker, a hosted
2822
+ model, unrestricted graph queries, autonomous vulnerability verdicts, a full
2823
+ persisted PDG, SSA, alias-complete or path-sensitive analysis, or claims of
2824
+ runtime reachability.
2825
+ - Touches if approved: language extraction, optional persisted resource/flow
2826
+ schema or a cheaper on-demand representation, one compact query pattern,
2827
+ source/sink registry, documentation, and labeled multi-language fixtures.
2828
+ - Collision risk: `src/engine/index.ts`, graph schema/migrations, MCP/CLI query
2829
+ enums, and shared language extractors; schedule apart from other graph-schema
2830
+ or query-surface work.
2831
+ - Source: https://github.com/vitali87/code-graph-rag/blob/d0b257b402bb25b54cf8f1dee0af3a71623c8d9e/docs/architecture/data-flow-edges.md
2832
+ - Source task: `6hC5Cv934v8xj2Gh`
2833
+ - Evidence: `benchmarks/evaluations/c78-resource-reachability/` contains the
2834
+ labeled TS/JS oracle, disposable prototype, immutable raw result, tests, and
2835
+ methodology. It achieved 100% precision but 62.5% recall, missing argument,
2836
+ return, and recursion/cycle handoff. Three local repositories measured
2837
+ bounded on-demand analysis against temporary-file persisted JSON; peak RSS was
2838
+ 129,859,584 bytes with zero spend and no egress.
2839
+ - Disposition: retain C8's bounded statement flow and call-argument evidence
2840
+ unchanged. The precision gate passed, but recall failed the 80% minimum useful
2841
+ threshold, so neither persistence nor a production query is authorized and no
2842
+ separately numbered implementation item is proposed.
2843
+
2844
+ Acceptance:
2845
+
2846
+ 1. First build a disposable prototype and labeled oracle containing true flows,
2847
+ clean overwrites, shadowing, unrelated source/sink co-occurrence, dynamic
2848
+ resource names, argument handoff, return handoff, recursion/cycles, and
2849
+ deliberately unsupported constructs. Do not change the production schema
2850
+ until the prototype passes the gate.
2851
+ 2. Require at least 95% precision and report recall per supported language and
2852
+ source/sink class. False positives are a hard failure because the result may
2853
+ inform security review; unsupported or ambiguous paths must remain explicit.
2854
+ 3. Every result identifies the source and sink, relation kind, bounded path,
2855
+ source file and line evidence, extraction provenance, confidence, supported
2856
+ language/registry coverage, omissions, truncation, and freshness. Label every
2857
+ path static and heuristic; never call it proof of exploitability, secrecy, or
2858
+ runtime reachability.
2859
+ 4. Compare a bounded on-demand implementation with persisted resource/flow
2860
+ edges. Choose the smallest design meeting correctness and warm-latency goals;
2861
+ preserve stable identity, ambiguity safety, monotonic bounds, deterministic
2862
+ ordering, hard item/byte/token budgets, and fail-closed stale-index behavior.
2863
+ 5. Measure clean and incremental index time, database growth, cold/warm p50 and
2864
+ p95 query latency, response bytes, and peak RSS on at least three real local
2865
+ repositories. Incremental spend remains zero and incremental memory remains
2866
+ below 1 GB, with no hosted service, credentials, source egress, or required
2867
+ production model.
2868
+ 6. Start with a deliberately narrow language and source/sink matrix justified by
2869
+ fixtures. Adding a registry entry requires positive and negative acceptance
2870
+ cases; a language without a verified registry emits an explicit omission,
2871
+ not inferred coverage.
2872
+ 7. If the gate fails, retain C8's current bounded statement flow and call-
2873
+ argument evidence unchanged. Do not broaden or persist the feature merely to
2874
+ match a competitor's schema.
2875
+
2876
+ ### C79 — Bounded repository applicability signals
2877
+
2878
+ - Status: implemented
2879
+ - Priority: P1
2880
+ - Disposition: Evaluate
2881
+ - DependsOn: C52, C69
2882
+ - Touches: `repos discover` CLI option/schema, bounded repository inspection,
2883
+ Git-remote parsing, worktree-aware discovery fixtures, CLI help/reference, and
2884
+ compatibility evidence. It does not change `doctor` or the MCP surface.
2885
+ - What: decide whether repository marker and configuration detection belongs in
2886
+ knodin's repository-discovery charter. If it does, add an additive, opt-in
2887
+ `knodin repos discover <roots...> --json --signals` facet so consumers can use
2888
+ knodin's existing repository/worktree classification instead of duplicating a
2889
+ second traversal and a less reliable definition of a repository.
2890
+ - Scope: without `--signals`, output must remain byte-identical. With the flag,
2891
+ every repository gains `signals`, using `{}` for honest absence. Detection is
2892
+ read-only, offline, credential-free, bounded by existing bytes/tokens/items
2893
+ budgets, and limited to a documented allowlist of marker paths rather than a
2894
+ tree glob. Never read potentially secret file contents; the only content-read
2895
+ exception is allowlisted hook-manager configuration needed to report whether
2896
+ it references `aidev-track`. Linked worktrees inspect their own checkout and
2897
+ preserve existing skip/include behavior.
2898
+ - Evidence: either a checked-in charter decision closing the item unchanged, or
2899
+ a checked-in implementation with compatibility bytes, help/docs, unit
2900
+ fixtures, bounded-resource results, and verifier output.
2901
+
2902
+ Charter decision (2026-08-03): accepted. These path-derived, repository-scoped
2903
+ signals belong in discovery because they reuse knodin's existing bounded
2904
+ repository and linked-worktree classification. The facet remains opt-in and
2905
+ descriptive: it does not decide whether an agent skill applies or whether a
2906
+ detected service is configured correctly.
2907
+
2908
+ Acceptance:
2909
+
2910
+ 1. First record the charter decision. If marker/config policy is outside
2911
+ knodin's charter, close C79 unchanged with checked-in rationale; no feature
2912
+ implementation is required. Otherwise implement only the opt-in facet.
2913
+ 2. A golden compatibility test proves `repos discover <roots...> --json`
2914
+ remains byte-identical without the flag. With `--signals`, every repository
2915
+ entry has a deterministic `signals` object, including `{}` when the
2916
+ allowlisted inspection finds nothing; schema version changes only if the
2917
+ existing compatibility policy requires it.
2918
+ 3. When detected, the deterministic contract must return `hookManager`,
2919
+ allowlisted `markerFiles` (including Sonar, CI/PR-workflow, agent, and MCP
2920
+ client markers), sanitized `remotes` with name/host/owner, `ciProviders`,
2921
+ `agentConfigs`, and `aidevTrackReferenced: true|false` when an allowlisted
2922
+ hook-manager config is present. Detection performs
2923
+ no writes, network access, credential access, source egress, or secret-file
2924
+ content reads, except the narrow allowlisted hook-manager parse for the
2925
+ `aidev-track` reference. Incremental spend is zero.
2926
+ 4. Unit fixtures cover lefthook, husky, pre-commit, no hook manager,
2927
+ `aidev-track` reference detection, multiple remotes, multiple remote hosts,
2928
+ missing `.git`, every marker absent, and linked worktrees. Tests prove fixed
2929
+ allowlists and existing bytes/tokens/items budgets bound inspection and
2930
+ serialization.
2931
+ 5. `knodin doctor` behavior is unchanged. The repos CLI reference and
2932
+ `repos discover --help` document the flag, returned fields, allowlist,
2933
+ omissions, privacy constraints, and worktree behavior.
2934
+
2935
+ Implementation evidence: the deterministic verifier and bounded-resource
2936
+ result are checked in at
2937
+ [`docs/evidence/c79-repository-signals-2026-08-03.md`](../docs/evidence/c79-repository-signals-2026-08-03.md)
2938
+ and
2939
+ [`docs/evidence/c79-repository-signals-2026-08-03.json`](../docs/evidence/c79-repository-signals-2026-08-03.json).
2940
+ The unit acceptance fixtures are in
2941
+ `src/__tests__/unit/repository-management.spec.ts`; CLI help coverage is in
2942
+ `src/__tests__/unit/cli-model.spec.ts`, and the unchanged doctor regression is
2943
+ `src/__tests__/unit/doctor.spec.ts`.
2944
+
2945
+ Known limitations and decision gate: marker presence describes repository
2946
+ configuration; it does not prove that a skill applies or that a service is
2947
+ configured correctly. If the charter review rejects this application-policy
2948
+ facet, closing C79 unchanged is a valid terminal outcome and downstream tools
2949
+ may inspect only the repository paths knodin returns.
2950
+
2951
+ ### C80 — Structural-first retrieval routing
2952
+
2953
+ - Status: implemented
2954
+ - Priority: P0
2955
+ - Disposition: Must close
2956
+ - DependsOn: C3, C24, C42, C43
2957
+ - Touches: retrieval dispatch and telemetry, search/context tests, a checked-in
2958
+ cross-file oracle, and deterministic ablation replay artifacts.
2959
+ - What: route stable identity and exact name/signature/path evidence through
2960
+ lexical/FTS and bounded graph expansion before loading embeddings. Serialize
2961
+ the selected route and `embeddingsUsed` so behavior is explainable.
2962
+ - Scope: preserve the one-tool gateway, stable identities, ambiguity safety,
2963
+ freshness, deterministic ordering, and current hard response budgets.
2964
+ - Evidence: a local checked-in implementation, fixtures, raw three-arm replay,
2965
+ verifier, and report covering factual precision/recall, unsupported claims,
2966
+ latency, and independently measured peak RSS.
2967
+
2968
+ Acceptance:
2969
+
2970
+ 1. Exact symbol, path, and source-evidenced cross-file questions do not load the
2971
+ embedding model when structural evidence is sufficient; tests assert the
2972
+ route and `embeddingsUsed: false` through CLI and MCP-compatible responses.
2973
+ 2. A checked-in multi-hop oracle replays lexical/vector-only, bounded graph
2974
+ expansion, and the complete route deterministically. Graph expansion becomes
2975
+ default only if correctness improves without increasing unsupported,
2976
+ ambiguous, or stale claims; otherwise retain the narrower passing route.
2977
+ 3. Preserve raw results and report precision, recall, omissions/refusals,
2978
+ latency, response size, and peak RSS. Incremental spend is zero, source stays
2979
+ local, no network or account is required, and incremental memory stays below
2980
+ 1 GB.
2981
+
2982
+ Delivered evidence: structural-first dispatch and request-local retrieval
2983
+ telemetry are implemented in `src/engine/index.ts`; CLI/MCP parity, stable
2984
+ identity, exact path/name, lexical sufficiency, embedding fallback, bounded
2985
+ work, and duplicate-name safety are covered by
2986
+ `src/__tests__/unit/structural-routing.spec.ts`. The checked-in multi-hop and
2987
+ duplicate-name oracle, fixture, deterministic four-route raw replay, and
2988
+ verifier outputs are under
2989
+ `benchmarks/evaluations/c80-structural-routing/`; reproduce them with
2990
+ `npm run bench:c80 && npm run verify:c80`. The limitations-aware precision,
2991
+ recall, unsupported/ambiguous/stale, omission, latency, response-size, spend,
2992
+ locality, and independently isolated RSS report is
2993
+ `docs/evidence/c80-structural-routing-2026-08-03.md`.
2994
+
2995
+ Decision: bounded graph expansion is the default structural route because the
2996
+ checked-in gate improves mean oracle recall from 0.389 to 1.000 while preserving
2997
+ 1.000 precision, zero unsupported and stale claims, and the same one surfaced
2998
+ duplicate-name ambiguity. The complete route preserves those results and loads
2999
+ embeddings only when exact or all-token structural evidence is insufficient.
3000
+ The largest isolated incremental RSS sample was 160,284,672 bytes; incremental
3001
+ spend was zero and no network, account, credential, or hosted service was used.
3002
+
3003
+ Known limitations and decision gate: structural routing is not proof that
3004
+ semantic retrieval is unnecessary. Embeddings remain a bounded fallback, and
3005
+ token reduction alone cannot authorize a route that weakens correctness. Graph
3006
+ expansion follows resolved source relationships only, is capped at two hops and
3007
+ 200 selected nodes, and cannot recover runtime-only or unresolved relationships.
3008
+ The three-case deterministic oracle proves this decision only on its checked-in
3009
+ cross-file chains; sampled RSS can miss short-lived peaks between samples.
3010
+
3011
+ ### C81 — Powered end-to-end benchmark with a diagnose arm
3012
+
3013
+ - Status: evaluated — retained knodin unchanged
3014
+ - Priority: P0
3015
+ - Disposition: Must close
3016
+ - DependsOn: C42, C43, C59, C80
3017
+ - Touches: competitive benchmark corpus/runner, task labels, deterministic
3018
+ patch/test oracles, raw trajectories, statistics, and evidence report.
3019
+ - What: compare ordinary filesystem tools, knodin, and one pinned local
3020
+ competitor on retrieval and failing-build/test diagnosis using a powered,
3021
+ paired design.
3022
+ - Scope: at least 40 tasks, at least three repetitions, pinned repositories,
3023
+ commits, prompts, models, versions, seeds, and budgets; confidence intervals
3024
+ and pre-registered analysis are mandatory.
3025
+ - Evidence: checked-in corpus and labels, immutable raw trajectories, verifier
3026
+ output, confidence intervals, and a limitations-aware report.
3027
+
3028
+ Acceptance:
3029
+
3030
+ 1. Each arm runs the same tasks and deterministic correctness oracles. The
3031
+ diagnose subset scores turns to a correct patch and passing test, not an LLM
3032
+ judge alone; unavailable and incomparable rows remain explicit.
3033
+ 2. Cache behavior is controlled and recorded. Do not report billed-cost claims
3034
+ unless cache economics are comparable; this item must incur zero model/API
3035
+ spend and require no hosted service, credentials, or source egress.
3036
+ Incremental spend is zero.
3037
+ 3. Record patch-application rate, turns to correct edit, factual correctness,
3038
+ tokens where measured, cold/warm latency, and peak RSS under 1 GB. Preserve
3039
+ losing cases and raw data without overwriting prior evidence.
3040
+
3041
+ Known limitations and decision gate: the benchmark may justify retaining the
3042
+ product unchanged. It authorizes no superiority claim beyond oracle-qualified,
3043
+ comparable rows and no production feature outside a separately numbered item.
3044
+
3045
+ Delivered evidence: the pinned 40-task corpus, three-repetition paired runner,
3046
+ deterministic retrieval and patch/test oracles, immutable accepted and excluded
3047
+ pilot trajectories, integrity hashes, and verifier are under
3048
+ `benchmarks/evaluations/c81-powered-benchmark/`; reproduce validation and
3049
+ confidence intervals with `npm run verify:c81`. The limitations-aware report is
3050
+ [`C81 powered benchmark evidence`](../docs/evidence/c81-powered-benchmark-2026-08-03.md).
3051
+
3052
+ Decision: retain knodin unchanged. All three available arms achieved 1.000
3053
+ factual correctness and patch-plus-passing-test rate on this narrow synthetic
3054
+ corpus, with paired correctness differences of 0.000. The result supports no
3055
+ superiority, token, billed-cost, or production claim. The conservative
3056
+ worker-plus-descendant peak-RSS upper bound was 351,469,568 bytes; spend,
3057
+ hosted-service use, credentials, and source egress were zero.
3058
+
3059
+ ### C82 — Productize `compress diagnose`
3060
+
3061
+ - Status: implemented
3062
+ - Priority: P0
3063
+ - Disposition: Must close
3064
+ - DependsOn: C59, C81
3065
+ - Touches: existing compress diagnose CLI/MCP dispatch, discoverability and
3066
+ documentation, stable response schema, fixtures, telemetry, and replay.
3067
+ - What: make the retained-failure-to-graph workflow discoverable and robust so
3068
+ an artifact resolves to owning symbols, tests, callers, and bounded source
3069
+ without rerunning the failed command.
3070
+ - Scope: reuse the existing single MCP tool and retained C59 primitive; preserve
3071
+ stable artifact identity, freshness, ambiguity, traversal, and budget bounds.
3072
+ - Evidence: checked-in adversarial diagnostic fixtures, CLI/MCP parity tests,
3073
+ golden responses, raw replay, verifier, and local evidence report.
3074
+
3075
+ Acceptance:
3076
+
3077
+ 1. Representative compiler, test, stack-trace, multiline, truncated, malformed,
3078
+ stale, ambiguous, and missing-artifact cases return stable schema fields and
3079
+ exact owning source evidence or an explicit bounded refusal.
3080
+ 2. CLI and MCP produce equivalent source identities, callers/tests, omissions,
3081
+ freshness, confidence, and telemetry without a new MCP tool, command rerun,
3082
+ hosted service, credential, source egress, or model/API spend.
3083
+ Incremental spend is zero.
3084
+ 3. The C81 diagnose corpus demonstrates the workflow end to end with checked-in
3085
+ deterministic oracles, hard item/byte/token caps, recoverable continuation,
3086
+ and peak incremental memory below 1 GB.
3087
+
3088
+ Known limitations and decision gate: static graph evidence cannot prove runtime
3089
+ causality. Documentation must say that diagnoses are bounded source-evidenced
3090
+ candidates, not guaranteed root causes or automatic fixes.
3091
+
3092
+ Delivered evidence: the existing CLI and single-tool MCP dispatch now share an
3093
+ additive stable diagnosis envelope with confidence, exact diagnostic/relation
3094
+ omissions, retained-artifact continuation, and command-rerun/cap telemetry.
3095
+ Focused parity and refusal coverage is in `src/__tests__/unit/`; adversarial
3096
+ fixtures, golden responses, immutable raw replay, hard-cap verification, and
3097
+ resource evidence are under
3098
+ `benchmarks/evaluations/c82-compress-diagnose/`. Reproduce all ten adversarial
3099
+ cases and all ten C81 diagnose tasks with `npm run verify:c82`. The evidence
3100
+ report is [`C82 compress diagnose evidence`](../docs/evidence/c82-compress-diagnose-2026-08-03.md).
3101
+
3102
+ The conservative measured maximum RSS was 309,116,928 bytes, below 1 GB.
3103
+ Incremental spend, command reruns, network/hosted service use, credentials, and
3104
+ source egress were zero. Static candidates remain neither runtime root-cause
3105
+ proof nor automatic fixes; the synthetic replay does not cover every toolchain.
3106
+
3107
+ ### C83 — Progressive evidence delivery and hash handshake
3108
+
3109
+ - Status: implemented
3110
+ - Priority: P0
3111
+ - Disposition: Must close
3112
+ - DependsOn: C80, C82
3113
+ - Touches: gateway response contracts, locate/outline/evidence/expand dispatch,
3114
+ continuation/evidence handles, content hashing, CLI/MCP tests, and docs.
3115
+ - What: deliver recoverable evidence levels and allow delta/omission only when
3116
+ the client supplies an exact content hash for the baseline it holds.
3117
+ - Scope: every level reports returned and omitted material, reason, `more`, a
3118
+ stable continuation, freshness, confidence, and truthful hard budgets.
3119
+ - Evidence: checked-in golden protocol fixtures, tamper/staleness/rename tests,
3120
+ raw bounded replay, verifier, and compatibility report.
3121
+
3122
+ Delivered evidence: the additive `evidence` operation and CLI command share the
3123
+ source-preserving implementation in `src/progressive-evidence.ts`. Golden and
3124
+ raw 14-case protocol replays, fixture, resource record, compatibility report,
3125
+ and limitations are under `benchmarks/evaluations/c83-progressive-evidence/`;
3126
+ reproduce them with `npm run verify:c83`. Unit acceptance covers complete
3127
+ omission recovery, adversarial hashes/handles/paths, stale and renamed handles,
3128
+ and byte-identical CLI/MCP edit anchors. The evidence report is
3129
+ [`C83 progressive evidence evidence`](../docs/evidence/c83-progressive-evidence-2026-08-03.md).
3130
+
3131
+ Acceptance:
3132
+
3133
+ 1. Exact hash match permits the documented delta or omission path; missing,
3134
+ stale, truncated, renamed, or tampered baselines return full required source
3135
+ and state which path was taken. Client assertion alone is never trusted.
3136
+ 2. Locate, outline, evidence, and expand remain deterministic and recoverable
3137
+ under item/byte/token caps; tests prove omitted evidence can be requested and
3138
+ that edit anchors remain intact across CLI and MCP.
3139
+ 3. The protocol remains local, zero-spend, offline, under 1 GB incremental RSS,
3140
+ backward-compatible where promised, and fail-closed on stale or ambiguous
3141
+ evidence. Checked-in fixtures include adversarial hash and handle inputs.
3142
+
3143
+ Known limitations and decision gate: this is not lossy summarization and does
3144
+ not authorize token-savings or task-success claims until C81 measures them.
3145
+ The V1 surface addresses one exact regular repository file at a time, refuses
3146
+ symlinks and ambiguous/path-traversal input, and uses a lightweight top-level
3147
+ outline rather than claiming the engine's complete relationship model. The
3148
+ synthetic replay authorizes no production-readiness or cross-platform claim.
3149
+
3150
+ ### C84 — Optional SCIP import
3151
+
3152
+ - Status: implemented
3153
+ - Priority: P1
3154
+ - Disposition: Must close
3155
+ - DependsOn: C24, C80, C83
3156
+ - Touches: optional SCIP ingestion, stable identity/provenance mapping, conflict
3157
+ handling, LSIF/native fallbacks, fixtures, lifecycle bounds, and docs.
3158
+ - What: import compiler-produced SCIP facts as an opt-in precision tier while
3159
+ retaining LSIF compatibility and the native AST/structural/heuristic ladder.
3160
+ - Scope: SCIP files are local inputs; import is deterministic, bounded, and
3161
+ optional. Live LSP integration is explicitly not part of C84.
3162
+ - Evidence: checked-in SCIP fixtures and expected graph facts, malformed/path/
3163
+ oversize/conflict tests, fallback replay, verifier, and resource report.
3164
+
3165
+ Delivered evidence: `knodin index --scip <file>` explicitly imports a bounded
3166
+ local SCIP protobuf snapshot without discovering or running a producer. The
3167
+ npm-integrity-pinned `@sourcegraph/scip-typescript@0.4.0` fixture, binary hash,
3168
+ source, configuration, expected facts, duplicate document-local symbols, and
3169
+ adversarial cases are under `fixtures/c84-scip/` and `src/__tests__/unit/`.
3170
+ The raw offline replay, distinct relationship kinds, native/LSIF fallback,
3171
+ stable identity refresh, 64 MiB/10,000-file/250,000-fact/30-second bounds, and
3172
+ 502,431,744-byte incremental RSS record are under
3173
+ `benchmarks/evaluations/c84-scip-import/`; reproduce them with
3174
+ `npm run verify:c84`. The evidence report is
3175
+ [`C84 optional SCIP import evidence`](../docs/evidence/c84-optional-scip-import-2026-08-03.md).
3176
+
3177
+ Acceptance:
3178
+
3179
+ 1. Pinned fixtures produce deterministic stable identities, relationships,
3180
+ source ranges, and `scip` provenance. Conflicting tiers surface ambiguity and
3181
+ never silently replace higher-confidence or fresher evidence.
3182
+ 2. Missing, malformed, hostile-path, stale, and oversized inputs fail safely;
3183
+ repositories without SCIP retain current native/LSIF behavior with no daemon,
3184
+ account, network, source egress, or model/API spend.
3185
+ Incremental spend is zero.
3186
+ 3. Import and query replays remain within explicit file/fact/time limits and
3187
+ below 1 GB incremental RSS. Raw local results, verifier output, and known
3188
+ language/indexer coverage are checked in.
3189
+
3190
+ Known limitations and decision gate: SCIP precision is bounded by the producer
3191
+ and languages represented in the imported index. Any live LSP tier requires a
3192
+ separately approved roadmap item and must never become a baseline dependency.
3193
+
3194
+ ### C85 — Bounded git-history review signals
3195
+
3196
+ - Status: implemented
3197
+ - Priority: P1
3198
+ - Disposition: Must close
3199
+ - DependsOn: C24, C44, C81
3200
+ - Touches: review risk serialization, bounded Git history queries, unavailable/
3201
+ shallow states, synthetic repository fixtures, telemetry, and docs.
3202
+ - What: add churn, co-change, and coupling as separately itemized review signals
3203
+ beside graph impact, test gaps, and structural centrality; forbid opaque scores.
3204
+ - Scope: all Git operations have explicit commit/file/time bounds and preserve
3205
+ deterministic ordering, source evidence, and graceful unavailable states.
3206
+ - Evidence: checked-in synthetic Git DAG, exact signal oracles, bounded-command
3207
+ tests, raw review replay, verifier, and resource report.
3208
+ - Evidence delivered: `src/engine/git-history.ts` and
3209
+ `src/__tests__/unit/git-history-signals.spec.ts` cover rename-aware itemized
3210
+ facts, merge/unrelated separation, shallow deepen, unborn, missing Git,
3211
+ non-repository, arbitrary HEAD failure, timeout, path containment, hard bounds,
3212
+ deterministic truncation, cache validity, and CLI/MCP serialization.
3213
+ `benchmarks/evaluations/c85-git-history/` checks in the exact synthetic DAG,
3214
+ raw replay, latency/command/resource evidence, limitations, and verifier;
3215
+ `npm run verify:c85` reproduces churn 2, co-change 2, coupling 1, three bounded
3216
+ history commands, zero locality costs, and peak RSS below 1 GB.
3217
+
3218
+ Acceptance:
3219
+
3220
+ 1. Fixtures cover renames, merges, unrelated co-occurrence, shallow history,
3221
+ unborn repositories, missing Git, and bounded truncation. Each output exposes
3222
+ the contributing facts and omissions rather than only a composite number.
3223
+ 2. Review remains deterministic, fail-closed, and useful when history is absent;
3224
+ commands cannot escape the repository or exceed configured commit/file/time
3225
+ limits. CLI/MCP-compatible evidence retains freshness and confidence.
3226
+ 3. Checked-in replay reports correctness, latency, command counts, and peak RSS
3227
+ below 1 GB with zero spend, no network, no credentials, and no source egress.
3228
+
3229
+ Known limitations and decision gate: co-change is correlation, not causation.
3230
+ History signals may influence itemized review context but cannot independently
3231
+ assert defect likelihood, ownership, or mandatory remediation.
3232
+
3233
+ ### C86 — Static-embedding bake-off
3234
+
3235
+ - Status: evaluated — retained MiniLM unchanged
3236
+ - Priority: P2
3237
+ - Disposition: Evaluate
3238
+ - DependsOn: C46, C51, C80, C81
3239
+ - Touches: isolated Model2Vec/MiniLM evaluation runner, fixed corpus and labels,
3240
+ model provenance/digests, raw measurements, verifier, and decision record.
3241
+ - What: evaluate a pinned static-embedding candidate against retained MiniLM
3242
+ before authorizing any runtime, index-format, or default-search change.
3243
+ - Scope: measure cold/warm indexing and query performance, RSS, storage, recall,
3244
+ NDCG, and downstream task correctness with at least three repetitions.
3245
+ - Evidence: checked-in corpus/query labels, model identifiers and digests, raw
3246
+ repeated results, verifier output, confidence/variance report, and disposition.
3247
+ - Evidence delivered: `benchmarks/evaluations/c86-static-embeddings/` contains
3248
+ the pinned 20-document corpus, ten relevance/task labels, model revisions and
3249
+ ONNX digests, six fresh-process raw repetitions, cold/warm timing, RSS and
3250
+ storage measurements, descriptive confidence/variance, unsupported cases,
3251
+ limitations, and the retain-MiniLM decision. `npm run verify:c86` checks the
3252
+ identical-route/budget oracle, three repetitions per arm, local/offline and
3253
+ zero-spend contract, sub-1-GB incremental RSS, artifact pins, and evaluation
3254
+ isolation without requiring either model artifact.
3255
+
3256
+ Acceptance:
3257
+
3258
+ 1. Both candidates run on identical local fixtures, retrieval routes, budgets,
3259
+ and deterministic relevance/task oracles. Record setup, unsupported cases,
3260
+ cold/warm timing, index size, recall/NDCG, correctness, and peak RSS.
3261
+ 2. The replay is offline after documented local preparation, adds no production
3262
+ dependency, uses no hosted API/account/source egress, incurs zero spend, and
3263
+ stays below 1 GB incremental memory. Pin model identity and artifact digest.
3264
+ 3. Retain MiniLM unless the pre-registered decision gate shows a reproducible
3265
+ correctness-preserving win. Close by retaining MiniLM, deferring/rejecting
3266
+ Model2Vec, or proposing a separately reviewed implementation item; evaluation
3267
+ alone must not change production behavior or schema.
3268
+
3269
+ Known limitations and decision gate: corpus results may not generalize to every
3270
+ language or repository. Publish only measured local outcomes, never a universal
3271
+ embedding-quality or product-superiority claim.
3272
+
3273
+ Decision: retain MiniLM unchanged. On the checked-in fixture both models reached
3274
+ 1.000 recall@5, 1.000 NDCG@10, and 10/10 top-1 task correctness in every
3275
+ repetition. Model2Vec was repeatably faster and smaller locally, but the corpus
3276
+ does not resolve repository-scale or multilingual-query quality, so the
3277
+ pre-registered gate's no-material-unsupported-category condition did not pass.
3278
+ No production runtime, dependency, index format, schema, default, or MCP surface
3279
+ changed, and no separately numbered implementation item is authorized.
3280
+
3281
+ ### C87 — Structural cold-start fast path
3282
+
3283
+ - Status: implemented
3284
+ - Priority: P0
3285
+ - Disposition: Must close
3286
+ - DependsOn: C56, C68
3287
+ - Touches: packaged CLI launcher, structural snapshot publication, direct-file
3288
+ fallback, compression routing, benchmarks, and truth-bound response fields.
3289
+ - What: bypass graph/model initialization for structural-only cold requests
3290
+ while preserving the persistent graph route for relationship intelligence.
3291
+ - Scope: file outline, exact file-scoped extraction, batch outline, project
3292
+ overview, and pure output compression only.
3293
+ - Evidence: `src/structural-snapshot.ts`, `src/structural-fast-path.ts`,
3294
+ `src/pure-compression-cli.ts`, focused unit tests, and
3295
+ `benchmarks/evaluations/c87-structural-cold-start/raw-results.json`.
3296
+
3297
+ Indexing now publishes an atomic compact file/symbol snapshot. The packaged
3298
+ launcher routes file outlines, exact file-scoped extraction, batch outlines,
3299
+ project overview, and pure compression without importing the graph or embedding
3300
+ modules. A missing or fingerprint-stale requested file is parsed directly and
3301
+ labeled `lexical-structural-fallback`; graph-backed search, impact, review,
3302
+ relationships, and architecture retain their existing initialization and
3303
+ freshness contract.
3304
+
3305
+ Acceptance:
3306
+
3307
+ 1. File-outline, exact-symbol, and project-overview fresh-process p50 are
3308
+ at most 50 ms and p95 at most 75 ms after separating benchmark-parent spawn/
3309
+ scheduling overhead from measured child-process lifetime. Preserve both raw
3310
+ observations; do not hide an application or runtime-floor miss.
3311
+ 2. Warm p50 regresses by no more than 10%, response tokens by no more than 5%,
3312
+ and the existing shared correctness oracles remain 100% for both fast and
3313
+ graph routes.
3314
+ 3. Every fast answer includes path/fingerprint evidence, exact locations,
3315
+ evidence quality, `graphEnriched: false`, file-scoped uniqueness limits, and
3316
+ an upgrade handle. Dependency evidence proves no graph/model import.
3317
+ 4. Graph-backed operations retain stable identity, relationships, and freshness
3318
+ evidence. Snapshot absence/staleness never becomes a repository-wide
3319
+ freshness claim. Checked-in replay evidence requires zero hosted spend,
3320
+ network, account, credential, or source egress.
3321
+
3322
+ Known limitations and decision gate: the direct fallback is a bounded lexical
3323
+ structural parser, not the full Tree-sitter graph extractor. Its classification
3324
+ is explicit and a later graph operation is the supported enrichment path. Keep
3325
+ the fast route only while all checked-in acceptance gates pass.
3326
+
3327
+ ### C88 — Profile-based contained execution
3328
+
3329
+ - Status: implemented — macOS certified; Linux/Windows unavailable
3330
+ - Priority: P1
3331
+ - Disposition: Must close
3332
+ - DependsOn: C57, C58, C59
3333
+ - Touches: one-tool schema/dispatcher, profile configuration, macOS Seatbelt
3334
+ adapter, output compression/diagnosis, audit records, adversarial evidence,
3335
+ and security documentation.
3336
+ - What: permit only preapproved single-process profiles behind enforceable
3337
+ macOS filesystem/network/process boundaries and fail closed elsewhere.
3338
+ - Scope: test, lint, typecheck, and build-style profiles with immutable argv;
3339
+ arbitrary commands, child processes, and portable containment claims remain
3340
+ excluded.
3341
+ - Evidence: `src/execution-profile.ts`, `docs/CONTAINED-EXECUTION.md`, gateway/
3342
+ adversarial unit tests, and `benchmarks/evaluations/c88-contained-execution/`.
3343
+
3344
+ C58's portable polling prototype remains rejected. C88 authorizes a narrower
3345
+ production contract: independent global and repository enablement, immutable
3346
+ repository profile identity, absolute configured executable and fixed argv,
3347
+ no shell/caller suffix/cwd/environment, sanitized environment, repository cwd,
3348
+ hard timeout/output/fork limits, private profile-only audit, recoverable
3349
+ compression, and failure-to-code diagnosis.
3350
+
3351
+ Acceptance:
3352
+
3353
+ 1. macOS native evidence proves literal metacharacter argv, denied ambient
3354
+ credentials/user-data reads, denied network sockets, denied forks,
3355
+ repository-only authorization, closed stdin, output/timeout termination, and
3356
+ no raw command text in audit/telemetry.
3357
+ 2. `fully-contained` is reported only by the certified macOS Seatbelt adapter.
3358
+ Linux and Windows report `containment-unavailable` and spawn nothing until
3359
+ independent native adapters and evidence exist.
3360
+ 3. MCP accepts only `profile`; it cannot accept an executable, arguments, cwd,
3361
+ environment, or limits. Every result names the profile, actual containment,
3362
+ configured bounds, exit state, compressed artifact, and diagnosis state.
3363
+ 4. Preserve the one-tool schema reduction gate and local-only operation. The
3364
+ checked-in evidence requires zero hosted spend and zero source egress while
3365
+ preserving the original C58 limitations. Do not describe deprecated
3366
+ `sandbox-exec` as a supported Apple public security API.
3367
+
3368
+ Known limitations and decision gate: profiles are single-process
3369
+ (`maxProcesses: 1`), system runtime paths remain readable, approved repository
3370
+ code can modify its checkout, and Linux/Windows execution is unavailable. Keep
3371
+ the surface only while native evidence proves the stated boundary; any broader
3372
+ profile or platform requires a separately certified adapter. Preserve no unrestricted command surface.
3373
+
3374
+ ### C89 — MCP request reliability and recovery
3375
+
3376
+ - Status: implemented
3377
+ - Priority: P0
3378
+ - Disposition: Must close
3379
+ - DependsOn: C3, C16, C53
3380
+ - Touches: MCP server lifecycle, request dispatcher, cancellation/progress
3381
+ bridge, durable diagnostic log, error taxonomy, and adversarial transport tests.
3382
+ - What: prevent one expensive or stuck operation from ending the useful MCP
3383
+ session and make every failure attributable and recoverable.
3384
+ - Scope: per-operation deadlines, client cancellation, progress notifications,
3385
+ request and trace IDs, bounded exit-safe logs, stuck-request isolation, and
3386
+ distinct crash, lock, memory-pressure, timeout, and client-disconnect errors.
3387
+ - Evidence: `src/mcp-worker-supervisor.ts`, `src/mcp-graph-worker.ts`,
3388
+ `src/mcp-reliability.ts`, `src/server.ts`, `docs/MCP.md`, and the checked-in
3389
+ supervisor/lifecycle/CLI replays prove versioned IPC, real held-lock handling,
3390
+ abrupt worker exit, the observed `Transport closed` disconnect cancellation
3391
+ race, hard termination, bounded restart, predecessor-to-next-success
3392
+ correlation, warm reuse, source-safe
3393
+ bounded journals, and linked-worktree isolation without hosted dependencies.
3394
+
3395
+ Acceptance:
3396
+
3397
+ 1. Every request receives stable request/trace correlation, an operation-aware
3398
+ deadline, and cancellation that releases its resources without corrupting the
3399
+ graph; expensive operations emit bounded truthful progress when supported.
3400
+ 2. A killed worker, held graph lock, memory-pressure sentinel, deadline, and
3401
+ disconnected client each produce a distinct actionable error. The server
3402
+ remains usable or restarts automatically without retry loops or lost logs.
3403
+ 3. Durable logs survive abrupt process exit, rotate within documented byte/time
3404
+ bounds, omit repository source and secrets, and correlate progress, failure,
3405
+ cleanup, restart, and the next successful request.
3406
+ 4. Checked-in CLI/MCP parity, concurrency, cancellation-race, forced-exit, and
3407
+ recovery replays pass locally with zero hosted spend, credentials, network,
3408
+ or source egress.
3409
+
3410
+ Known limitations and decision gate: deadlines and memory-pressure attribution
3411
+ are bounded diagnoses, not proof of an operating-system root cause. Do not claim
3412
+ transparent recovery unless the next independent request succeeds in replay.
3413
+
3414
+ ### C90 — Privacy-safe diagnostic support bundle
3415
+
3416
+ - Status: implemented
3417
+ - Priority: P0
3418
+ - Disposition: Must close
3419
+ - DependsOn: C16, C59, C89
3420
+ - Touches: diagnostics CLI, MCP trace store, lifecycle/graph health, redaction,
3421
+ archive manifest, preview/report rendering, and privacy adversarial tests.
3422
+ - What: turn local diagnostics into one support-ready bundle whose exact shared
3423
+ contents users can inspect before export.
3424
+ - Scope: recent MCP traces/failures, lifecycle and graph health, versions and
3425
+ runtime, repository scale without source, manifest, human report, and preview.
3426
+ - Evidence: `src/diagnostics.ts`, `src/diagnostics-write-helper.ts`,
3427
+ `schemas/support-bundle-v2.schema.json`, the C90 adversarial fixture and unit
3428
+ replays, and packed-install preview/archive/inspect smoke prove a fail-closed
3429
+ allowlist, exact content-addressed preview parity, bounded C89 trace recovery,
3430
+ honest unavailable fields, linked-worktree isolation, bounded retention, and
3431
+ repository-bound no-follow writes without hosted dependencies or source
3432
+ egress.
3433
+
3434
+ Acceptance:
3435
+
3436
+ 1. Preview and final manifest enumerate every file and field, byte size,
3437
+ retention window, redaction applied, and unavailable section before sharing;
3438
+ the concise report includes request/trace IDs and actionable recovery steps.
3439
+ 2. Adversarial fixtures containing source, diffs, paths outside the repository,
3440
+ credentials, environment values, usernames, and command output prove none can
3441
+ enter the default bundle. Repository scale is aggregate metadata only.
3442
+ 3. The bundle includes exit-surviving C89 traces and recent classified failures,
3443
+ lifecycle/graph health, knodin/Node/platform versions, and honest missing or
3444
+ stale states without initiating network activity.
3445
+ 4. Checked-in deterministic CLI/MCP tests prove preview/archive parity, bounded
3446
+ size and retention, zero hosted spend, no credentials, and no source egress.
3447
+
3448
+ Known limitations and decision gate: automated redaction cannot certify arbitrary
3449
+ future fields. New diagnostic fields fail closed until included in the explicit
3450
+ allowlist and privacy oracle; sending a bundle remains a user action.
3451
+
3452
+ ### C91 — Boring macOS installation and runtime handoff
3453
+
3454
+ - Status: parked-external-evidence — 2026-08-05
3455
+ - Previous status: owner-approved adoption priority
3456
+ - Priority: P0
3457
+ - Disposition: Must close
3458
+ - DependsOn: C52, C73, C89, C90
3459
+ - Touches: package/install CI, runtime resolver, doctor repairs, client setup
3460
+ guides, macOS acceptance fixtures, and package compatibility checks.
3461
+ - What: make install through first useful MCP call repeatable on environments the
3462
+ project can actually validate.
3463
+ - Scope: macOS; Node 20 project handoff to an available Node 24 runtime; mise,
3464
+ nvm, fnm, Volta, asdf, Homebrew, npm, and Artifactory; Copilot, Claude, Gemini,
3465
+ and Codex; safe doctor repairs.
3466
+ - Evidence: checked-in automated install matrix, package-content/compatibility
3467
+ gates, isolated runtime-manager fixtures, and timed acceptance receipts.
3468
+
3469
+ Acceptance:
3470
+
3471
+ 1. Each named installation path is exercised from a clean macOS fixture through
3472
+ install, `knodin init`, `knodin status`, one useful MCP call, and uninstall or
3473
+ rollback; package compatibility is checked automatically in CI.
3474
+ 2. A Node 20 project reliably hands execution to an already available compatible
3475
+ Node 24 runtime without modifying the project's declared runtime. Missing or
3476
+ conflicting runtimes produce one actionable diagnostic.
3477
+ 3. Copilot, Claude, Gemini, and Codex each complete the same acceptance flow in
3478
+ no more than two minutes under the documented warm/cold boundary, with exact
3479
+ client/version/configuration evidence and no universal client claim.
3480
+ 4. Doctor applies only enumerated idempotent repairs with preview and rollback;
3481
+ checked-in tests require zero hosted spend for local fixtures and no source
3482
+ egress. Credentialed Artifactory evidence is recorded only when available.
3483
+
3484
+ Known limitations and decision gate: Linux and Windows certification are not a
3485
+ goal and unavailable platform evidence must stay unavailable. Runtime-manager
3486
+ fixtures prove supported configurations, not every shell customization.
3487
+
3488
+ Code-readiness on 2026-08-05 added a merge-preserving GitHub Copilot/VS Code
3489
+ MCP adapter, doctor visibility, idempotent CLI-only reversal, and fail-closed
3490
+ malformed/symlink handling. Certification remains parked until native macOS
3491
+ install-to-uninstall receipts exist for nvm, fnm, Volta, and asdf; an exact
3492
+ GitHub Copilot/VS Code timed receipt joins the available Claude, Gemini, and
3493
+ Codex versions; a published Homebrew formula and credentialed Artifactory path
3494
+ are exercised where authorized; and the same four clients complete the bounded
3495
+ install, `init`, `status`, useful MCP call, and rollback flow. Doctor also still
3496
+ needs an enumerated transactional preview/apply/rollback implementation and
3497
+ acceptance replay. No fixture result may be promoted to this missing native
3498
+ evidence, and C94 remains dependency-blocked.
3499
+
3500
+ ### C92 — Unified five-channel release orchestration
3501
+
3502
+ - Status: parked-external-evidence — 2026-08-06
3503
+ - Previous status: owner-approved release priority
3504
+ - Priority: P0
3505
+ - Disposition: Must close
3506
+ - DependsOn: C64, C91
3507
+ - Touches: release workflows, preflight, immutable package staging, npm,
3508
+ Artifactory, Homebrew, GitHub SaaS/GHES synchronization, receipts, and retry tests.
3509
+ - What: replace fragmented publication paths with one fail-closed orchestrator
3510
+ that promotes one immutable tarball everywhere.
3511
+ - Scope: npm, Artifactory, Homebrew, GitHub SaaS, and the read-only GHES mirror;
3512
+ availability preflight, private-repository attestations, retries, and receipts.
3513
+ - Evidence: checked-in dry-run channel adapters, fault matrix, immutable digest
3514
+ oracle, retry replay, receipt schema, and private-repository attestation fixture.
3515
+
3516
+ Acceptance:
3517
+
3518
+ 1. Before any publication, preflight proves every required channel, credential,
3519
+ permission, target version, and attestation mode is available; one failure
3520
+ publishes nothing and leaves a machine-readable failure receipt.
3521
+ 2. One staged tarball digest and length are promoted to all five channels.
3522
+ Retries reconcile idempotently without moving or recreating source tags and
3523
+ refuse any channel whose existing bytes differ.
3524
+ 3. Private-repository attestations follow the supported trust path rather than
3525
+ assuming public-repository identity behavior. The final receipt records each
3526
+ channel's artifact identity, digest, status, attempt, and authoritative URL.
3527
+ 4. Checked-in fault-injection tests cover every preflight/publish boundary with
3528
+ zero real publication or spend. Live channel receipts remain Track B evidence
3529
+ and cannot be simulated by this local implementation item.
3530
+
3531
+ Known limitations and decision gate: dry-run adapters do not certify production
3532
+ availability. GitHub SaaS remains authoritative; GHES is a read-only mirror and
3533
+ must never be forced. Stop on divergence.
3534
+
3535
+ Parking decision: C92 remains open behind C64 and C91. Ordinary 0.x publishing
3536
+ continues through the existing fail-closed tag workflow; this does not claim the
3537
+ unimplemented unified five-channel orchestrator or manufacture channel receipts.
3538
+
3539
+ ### C93 — Paired engineering-outcome measurement
3540
+
3541
+ - Status: implemented
3542
+ - Priority: P1
3543
+ - Disposition: Must close
3544
+ - DependsOn: C24, C81
3545
+ - Touches: competitive runner, task corpus, patch evaluator, dependency oracle,
3546
+ agent telemetry import, cost accounting, and outcome report.
3547
+ - What: measure whether knodin improves completed engineering work, not whether
3548
+ it merely emits fewer tokens.
3549
+ - Scope: identical with/without-knodin tasks measuring task success,
3550
+ missed-dependency rate, corrective round trips, time to first correct edit,
3551
+ patch-application success, total billed cost, and performance.
3552
+ - Evidence: preregistered checked-in corpus/oracles, paired raw runs, variance and
3553
+ failure accounting, verifier, and a limitations-bound report.
3554
+ - Evidence delivered: `benchmarks/evaluations/c93-engineering-outcomes/`
3555
+ contains the preregistered two-arm corpus, executable dependency/patch oracle,
3556
+ source-free raw event trajectories, deterministic local runner, paired
3557
+ variance/uncertainty report, and visible limitations. `npm run verify:c93`
3558
+ recreates all 12 runs from the identical fixture commit on isolated linked
3559
+ worktrees and non-default branches, verifies independent arm commits and a
3560
+ clean base checkout, and records zero hosted spend, observed network requests,
3561
+ or source egress plus explicit process-isolation bounds and unavailable—not
3562
+ inferred—token and billed-cost data.
3563
+
3564
+ Acceptance:
3565
+
3566
+ 1. Both arms use identical repositories, commits, tasks, model/client versions,
3567
+ prompts, permissions, time limits, repetitions, and correctness oracles; the
3568
+ only intended treatment difference is knodin availability/instructions.
3569
+ 2. Every named outcome is reported per task and in aggregate, including failures,
3570
+ censored time, unavailable token/cost data, and confidence/variance. Token
3571
+ counts may be diagnostic but are never the success metric.
3572
+ 3. Dependency and patch oracles inspect the resulting edit and tests rather than
3573
+ grading prose. Raw events support auditable round-trip and first-correct-edit
3574
+ reconstruction without storing repository source in telemetry.
3575
+ 4. The local harness and synthetic fixtures are checked in and run with zero hosted
3576
+ spend; any paid model replay records actual total billed cost and is
3577
+ never inferred from tokens alone.
3578
+ Checked-in evidence requires zero hosted spend.
3579
+
3580
+ Known limitations and decision gate: a bounded corpus cannot prove universal
3581
+ product superiority. Publish only measured effect sizes and uncertainty; a null
3582
+ or negative outcome must remain visible and may require product rollback.
3583
+
3584
+ ### C94 — First-use clarity and five-minute demonstration
3585
+
3586
+ - Status: parked-external-evidence — 2026-08-06
3587
+ - Previous status: owner-approved clarity priority
3588
+ - Blocked by: C91 external certification
3589
+ - Priority: P1
3590
+ - Disposition: Must close
3591
+ - DependsOn: C89, C90, C91
3592
+ - Touches: MCP/CLI help, status responses, client guides, examples, demo fixture,
3593
+ freshness/error copy, and documentation integrity tests.
3594
+ - What: make the first session explain the tool identity, next action, evidence
3595
+ limits, freshness, and recovery without requiring prior product knowledge.
3596
+ - Scope: `knodin.knodin` naming, operation discovery/examples, next-useful-action
3597
+ recommendations, concise limitations/recovery, and a real-repository demo.
3598
+ - Evidence: checked-in golden help/status/error output, four-client guide checks,
3599
+ timed demo transcript, and source-evidence links.
3600
+
3601
+ Acceptance:
3602
+
3603
+ 1. Setup docs explain that clients display `knodin.knodin` as server name plus
3604
+ tool name. Schema/help gives copyable examples and makes operation discovery
3605
+ possible without adding MCP tools.
3606
+ 2. Status recommends a context-sensitive next useful operation and presents
3607
+ freshness, availability, known bounds, and repair/reindex actions concisely;
3608
+ no healthy, stale, ambiguous, or repair-needed state is conflated.
3609
+ 3. A checked-in five-minute demonstration runs `knodin init`, status, orientation,
3610
+ source-evidenced impact/review, and bounded context on a real repository and
3611
+ explicitly demonstrates value, ROI, tomorrow's action, and secret sauce.
3612
+ 4. Golden CLI/MCP and docs-integrity tests run locally with zero hosted spend,
3613
+ no credentials, and no source egress.
3614
+
3615
+ Known limitations and decision gate: a scripted demonstration is onboarding
3616
+ evidence, not an outcome or superiority benchmark. Client UI labels may vary by
3617
+ version and must name the version actually checked.
3618
+
3619
+ Parking decision: C94 remains open behind C91. The existing local help and
3620
+ documentation remain available, but no fixture is promoted into the missing
3621
+ native client demonstration or certification evidence.
3622
+
3623
+ ### C95 — Defensible behavioral-contract replay
3624
+
3625
+ - Status: implemented
3626
+ - Priority: P1
3627
+ - Disposition: Must close
3628
+ - DependsOn: C3, C24, C83
3629
+ - Touches: contract manifest, cross-operation fixtures, freshness/identity/
3630
+ evidence/budget/compression/diagnosis/privacy tests, and product positioning.
3631
+ - What: make the integrated behavior—not individual graph feature count—the
3632
+ product's checked and release-gated contract.
3633
+ - Scope: truthful budgets, explicit freshness/availability, source evidence,
3634
+ ambiguity-safe stable identity, recoverable compression, failure-to-symbol
3635
+ diagnosis, and local operation without source egress.
3636
+ - Evidence: `contracts/behavior-contract-v1.json`,
3637
+ `benchmarks/evaluations/c95-behavior-contract/runner.ts`,
3638
+ `scripts/verify-c95.ts`, `docs/BEHAVIORAL-CONTRACT.md`, and the release
3639
+ preflight gate map and dispatch all 22 production operations plus every
3640
+ schema-declared query/action discriminator against linked-worktree fixtures.
3641
+ The replays cover hard-budget continuation, stable evidence,
3642
+ recoverable compression, freshness and ambiguity qualification, diagnosis,
3643
+ and observed local privacy across fetch, HTTP(S), TCP/TLS, DNS, native-addon,
3644
+ and subprocess boundaries. All 458 mapped fixture/clause rows have isolated
3645
+ pre-serialization production-response mutations that fail the same release
3646
+ oracle; omitted or mutated clauses fail closed.
3647
+
3648
+ Acceptance:
3649
+
3650
+ 1. Every production operation is mapped to the applicable contract clauses and
3651
+ adversarial fixtures prove honest truncation, stale/unavailable states,
3652
+ evidence provenance, ambiguity, recovery handles, diagnosis confidence, and
3653
+ privacy behavior.
3654
+ 2. The replay verifies composition: a bounded response can be expanded by a
3655
+ stable handle without changing identity/evidence, and stale or ambiguous
3656
+ evidence cannot become a confident diagnosis or impact claim.
3657
+ 3. Release preflight fails on a contract regression. Product documentation
3658
+ answers value, ROI, tomorrow, and secret sauce with checked-in evidence and
3659
+ preserves every roadmap limitation.
3660
+ 4. Tests and replays run locally with zero hosted spend, credentials, network,
3661
+ or source egress; no universal superiority claim is inferred.
3662
+
3663
+ Known limitations and decision gate: the contract proves the checked operation
3664
+ and action fixtures, not all repositories or clients. Network construction is
3665
+ denied before request-body streaming, and opaque transport inside an
3666
+ already-loaded native library remains unobservable. A new operation or action
3667
+ must declare its clauses, positive response semantics, and isolated production
3668
+ mutation before release.
3669
+
3670
+ ### C96 — Replay-gated capability discipline
3671
+
3672
+ - Status: evaluated — deferred
3673
+ - Priority: P2
3674
+ - Disposition: Evaluate
3675
+ - DependsOn: C24, C93, C95
3676
+ - Touches: roadmap policy, proposal template, outcome replay, one narrow
3677
+ cross-substrate fixture, verifier, and decision record.
3678
+ - What: require evidence of correctness or workflow improvement before expanding
3679
+ production capability, beginning with one bounded cross-substrate impact case.
3680
+ - Scope: one relationship crossing two explicitly named substrates; no broad
3681
+ platform, language, repository, or superiority promise.
3682
+ - Evidence: checked-in baseline/treatment fixture, exact dependency and task
3683
+ oracles, cost/resource report, limitations, and retain/defer/implement decision.
3684
+
3685
+ Acceptance:
3686
+
3687
+ 1. Preregister one concrete missed-dependency or workflow failure and the exact
3688
+ improvement threshold before implementation. Compare unchanged knodin with a
3689
+ bounded prototype on the same task using C93 outcomes and C95 contracts.
3690
+ 2. The fixture names supported substrates, edge provenance, ambiguity behavior,
3691
+ freshness, budgets, and false-positive/negative oracles; unsupported crossings
3692
+ remain explicit rather than generalized.
3693
+ 3. Close by retaining knodin unchanged, deferring/rejecting the capability, or
3694
+ creating a separately numbered implementation item only when the gate passes.
3695
+ 4. Evaluation artifacts and verifier are checked in, require zero hosted spend,
3696
+ credentials, network, or source egress, and preserve losing results.
3697
+
3698
+ Known limitations and decision gate: one passing cross-substrate fixture does
3699
+ not authorize a broad platform claim. Implementation requires a separately
3700
+ reviewed item and must preserve the behavioral contract.
3701
+
3702
+ Decision: defer production expansion. C93 found no correctness advantage on its
3703
+ narrow paired fixture, and C95 already requires every new operation or action to
3704
+ declare and replay its behavioral clauses. No preregistered cross-substrate
3705
+ missed-dependency case currently justifies a prototype, so knodin remains
3706
+ unchanged. Reopen under a separately numbered item only when a concrete failure,
3707
+ supported substrates, exact oracle, and improvement threshold are preregistered.
3708
+ Evidence: `benchmarks/evaluations/c93-engineering-outcomes/verification.json`,
3709
+ `contracts/behavior-contract-v1.json`, and
3710
+ `benchmarks/evaluations/c95-behavior-contract/runner.ts`.
3711
+
3712
+ ## Explicit non-goals from the GitNexus comparison
3713
+
3714
+ - Do not split knodin's gateway into one MCP schema per capability; the one-tool
3715
+ surface is a measured differentiator.
3716
+ - Do not copy unrestricted Cypher directly into the MCP surface.
3717
+ - Do not require Milvus, Qdrant, Ollama, credentials, or any hosted service;
3718
+ Claude Context and mcp-codebase-index validate this differentiator.
3719
+ - Do not add one MCP tool per capability; Serena's 29-tool and
3720
+ code-review-graph's 30-tool surfaces reinforce the one-gateway advantage.
3721
+ - Do not make vector visualization, agent memory, or skill generation core graph
3722
+ requirements. C35 is limited to deterministic local architecture/call-flow
3723
+ artifacts with a concrete developer workflow; it does not authorize a hosted
3724
+ visualization service or a generalized vector-visualization feature.
3725
+ - Do not enable network embeddings, hosted indexing, telemetry, or source-code
3726
+ egress to chase competitor features.
3727
+ - Do not claim API-shape or cross-repository superiority until dedicated local
3728
+ fixtures produce evidence.
3729
+ - Do not interpret every competitor option as a product requirement; parameters
3730
+ that merely select nonexistent groups, branches, or services need real
3731
+ fixtures before roadmap promotion.
3732
+
3733
+ ## Verification bar
3734
+
3735
+ Every implementation item must run the repository's discovered gates and add a
3736
+ targeted acceptance fixture. At minimum:
3737
+
3738
+ ```bash
3739
+ bun run typecheck
3740
+ bun run lint
3741
+ bun run test
3742
+ ```
3743
+
3744
+ Competitive claims must also re-run the relevant harness under
3745
+ `benchmarks/competitors/` and update its result file without deleting losing
3746
+ cases.
3747
+
3748
+ ## Competitive replay — 2026-07-24 (Phase 2)
3749
+
3750
+ Fresh, provenance-bound replays of every dimension were run against the real
3751
+ competitor CLIs and saved as **new** immutable artifacts (never overwriting
3752
+ prior results): search `raw-results-20260724T182246Z.json`, review + architecture
3753
+ `…T182448Z.json`, and symbols/editing/visualization/context/lifecycle
3754
+ `…T182707Z.json` under `benchmarks/evaluations/competitive-*/`.
3755
+
3756
+ Applying the measurement contract (a WIN needs equivalent operation/fixture/mode/
3757
+ scope/budget/oracle **and** an oracle-qualified row; unavailable/unverified/
3758
+ incomparable rows are excluded from win claims):
3759
+
3760
+ ### Where knodin is ahead or level (oracle-qualified, comparable)
3761
+
3762
+ - **Diff review — WIN.** traversal recall 1.0 and 733 B vs gitnexus (recall 1.0,
3763
+ 1458 B), code-review-graph (recall 0, 2269 B), graphify (recall 0.667).
3764
+ - **Architecture — WIN.** full package coverage + hubs + bridges; every completed
3765
+ competitor is at best equal on coverage, none better.
3766
+ - **Search — WIN vs the only comparable competitor.** nDCG 0.855 vs gitnexus
3767
+ 0.677; dead-code precision 1.0 vs 0.667. codegraph (probe-only, unscored),
3768
+ grepai (MCP probe blocked), claude-context (needs Ollama+Milvus) are
3769
+ contract-excluded.
3770
+ - **Symbols — WIN/level.** precision 1 / recall 1: ties gitnexus (1/1), beats
3771
+ codegraph (0/0), serena (0/0), codebase-memory (1/0.5).
3772
+ - **Lifecycle capability — WIN.** indexed/queryable/closed/stale-detection/
3773
+ per-query telemetry all qualified; codebase-memory and grepai advertise
3774
+ neither stale detection nor per-query telemetry.
3775
+
3776
+ ### Where the comparison is not oracle-qualified (excluded, not losses)
3777
+
3778
+ - **Editing / Visualization / Context.** knodin passes its golden/export oracles;
3779
+ the competitors either did not run (serena, code-review-graph, aider
3780
+ unavailable/blocked) or ran without a shared correctness oracle (graphify,
3781
+ repomix, code2prompt "complete" but unscored). No comparable head-to-head.
3782
+ - **Project memory.** knodin has no memory feature (deferred, ADR 004); Serena's
3783
+ own value-beyond-docs gate also failed. No contest.
3784
+
3785
+ ### Where knodin still costs more — and why
3786
+
3787
+ - **Lifecycle peak RSS: 83.8 MB vs codebase-memory 31.3 MB, grepai 22.9 MB.**
3788
+ This is the one axis where knodin's raw number is worse. It is **not** an
3789
+ oracle-qualified head-to-head: knodin is measured as a dedicated, long-lived
3790
+ product process holding the real semantic model + graph resident (which is
3791
+ precisely what wins the search-quality contest above), while the competitors
3792
+ are measured at spawn-per-command peak. The RSS cost and the search-quality
3793
+ lead are two sides of the same resident-model design; knodin cannot undercut
3794
+ them on RSS without abandoning the model that beats them on search. C51
3795
+ (native training-free quantization) targets the *stored-embedding* portion of
3796
+ that footprint; the model-runtime portion is inherent to local semantic search.
3797
+
3798
+ Bottom line: on every dimension with an oracle-qualified head-to-head against an
3799
+ available competitor, knodin is **as good or better** (wins or ties). The only
3800
+ place a competitor's raw number beats knodin is lifecycle RSS, which is an
3801
+ incomparable isolation model and the direct cost of knodin's search-quality lead.