knodin 0.7.6 → 0.8.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (134) hide show
  1. package/README.md +41 -12
  2. package/benchmarks/competitors/SYNTHESIS.md +66 -0
  3. package/dist/bin/cli.js +2181 -108
  4. package/dist/bin/launcher.js +25 -3
  5. package/dist/src/agent-integration.js +304 -0
  6. package/dist/src/artifact-refresh.js +82 -0
  7. package/dist/src/cli-args.js +292 -0
  8. package/dist/src/cli-model.js +384 -0
  9. package/dist/src/codeflow-replay.js +81 -0
  10. package/dist/src/compact-structural.js +96 -0
  11. package/dist/src/compare.js +39 -0
  12. package/dist/src/competitive-cold-mcp.js +40 -0
  13. package/dist/src/competitive-constraints.js +21 -0
  14. package/dist/src/competitive-manifest.js +411 -0
  15. package/dist/src/competitive-measurement.js +183 -0
  16. package/dist/src/competitive-runner.js +487 -0
  17. package/dist/src/competitive-sandbox.js +108 -0
  18. package/dist/src/context-export.js +423 -0
  19. package/dist/src/context.js +102 -0
  20. package/dist/src/deterministic-random.js +34 -0
  21. package/dist/src/diagnostics-write-helper.js +473 -0
  22. package/dist/src/diagnostics.js +1476 -0
  23. package/dist/src/docs-sections.js +142 -0
  24. package/dist/src/doctor.js +382 -0
  25. package/dist/src/engine/ann-hnsw.js +261 -0
  26. package/dist/src/engine/embeddings.js +193 -0
  27. package/dist/src/engine/file-walker.js +49 -0
  28. package/dist/src/engine/git-history.js +289 -0
  29. package/dist/src/engine/index.js +15094 -0
  30. package/dist/src/engine/perf.js +115 -0
  31. package/dist/src/engine/prune.js +112 -0
  32. package/dist/src/engine/sarif-import.js +341 -0
  33. package/dist/src/engine/scip-import.js +423 -0
  34. package/dist/src/engine/source-policy.js +85 -0
  35. package/dist/src/engine/sqlite.js +71 -0
  36. package/dist/src/engine/state-paths.js +175 -0
  37. package/dist/src/engine/symbol-delete.js +58 -0
  38. package/dist/src/execution-profile.js +208 -0
  39. package/dist/src/failure-diagnosis.js +655 -0
  40. package/dist/src/fleet.js +7 -0
  41. package/dist/src/git-executable.js +31 -0
  42. package/dist/src/graph-layout.js +173 -0
  43. package/dist/src/graph-query-health.js +115 -0
  44. package/dist/src/hook-manager-integration.js +156 -0
  45. package/dist/src/index-activity.js +126 -0
  46. package/dist/src/init-progress-worker.js +70 -2
  47. package/dist/src/init-progress.js +155 -0
  48. package/dist/src/init.js +1295 -0
  49. package/dist/src/lifecycle-health.js +282 -0
  50. package/dist/src/lsp-readonly.js +217 -0
  51. package/dist/src/mcp-graph-worker.js +69 -0
  52. package/dist/src/mcp-reliability.js +154 -0
  53. package/dist/src/mcp-worker-supervisor.js +350 -0
  54. package/dist/src/mirror.js +290 -0
  55. package/dist/src/node-runtime.js +157 -0
  56. package/dist/src/output-compression.js +630 -0
  57. package/dist/src/output-telemetry.js +368 -0
  58. package/dist/src/pr-triage.js +638 -0
  59. package/dist/src/progress-worker-runtime.js +46 -0
  60. package/dist/src/progressive-evidence.js +477 -0
  61. package/dist/src/pure-compression-cli.js +102 -0
  62. package/dist/src/relationship-adapters.js +377 -0
  63. package/dist/src/release-attestation.js +533 -0
  64. package/dist/src/release-preflight.js +513 -0
  65. package/dist/src/repair-lease.js +85 -0
  66. package/dist/src/repair-progress-worker.js +83 -2
  67. package/dist/src/repair-progress.js +262 -0
  68. package/dist/src/repository-init-process.js +177 -0
  69. package/dist/src/repository-management.js +1261 -0
  70. package/dist/src/resource-reachability.js +456 -0
  71. package/dist/src/response-budget.js +200 -0
  72. package/dist/src/server.js +217 -0
  73. package/dist/src/structural-fast-path.js +344 -0
  74. package/dist/src/structural-snapshot.js +37 -0
  75. package/dist/src/system-config.js +638 -0
  76. package/dist/src/terminal-help.js +83 -0
  77. package/dist/src/tools/knodin-tools.js +1645 -0
  78. package/dist/src/update-ceremony.js +162 -0
  79. package/dist/src/update-policy.js +944 -0
  80. package/dist/src/update-trust.js +504 -0
  81. package/dist/src/version.js +13 -0
  82. package/dist/src/visualization.js +515 -0
  83. package/dist/src/wait-for-fresh.js +98 -0
  84. package/dist/src/worktree-lifecycle.js +234 -0
  85. package/docs/BEHAVIORAL-CONTRACT.md +114 -0
  86. package/docs/CLI.md +30 -1
  87. package/docs/COMPARISON.md +413 -0
  88. package/docs/COMPETITIVE-LANDSCAPE-2026-08.md +267 -0
  89. package/docs/CONTAINED-EXECUTION.md +77 -0
  90. package/docs/DEMO.md +49 -0
  91. package/docs/DIAGNOSTICS.md +80 -0
  92. package/docs/GIT-HISTORY-REVIEW.md +39 -0
  93. package/docs/HANDOFF.md +180 -0
  94. package/docs/INSTALLATION.md +21 -18
  95. package/docs/MCP.md +64 -8
  96. package/docs/PROGRESSIVE-EVIDENCE.md +37 -0
  97. package/docs/PT-ACCESS-RECOMMENDATION.md +89 -0
  98. package/docs/RELEASE-0.3-EVIDENCE.md +73 -0
  99. package/docs/REPOSITORIES-AND-WORKTREES.md +18 -6
  100. package/docs/SCIP-IMPORT.md +62 -0
  101. package/docs/SIGNED-UPDATES.md +151 -0
  102. package/docs/TELEMETRY.md +46 -0
  103. package/docs/TOKEN-OPTIMIZER-SCORECARD.md +79 -0
  104. package/docs/assets/knodin-favicon.svg +4 -0
  105. package/docs/releases/0.3.0.md +46 -0
  106. package/docs/releases/0.4.0.md +68 -0
  107. package/docs/releases/0.4.1.md +28 -0
  108. package/docs/releases/0.4.2.md +27 -0
  109. package/docs/releases/0.4.3.md +23 -0
  110. package/docs/releases/0.5.0.md +29 -0
  111. package/docs/releases/0.5.1.md +17 -0
  112. package/docs/releases/0.6.0.md +18 -0
  113. package/docs/releases/0.7.0.md +24 -0
  114. package/docs/releases/0.7.1.md +21 -0
  115. package/docs/releases/0.7.2.md +21 -0
  116. package/docs/releases/0.7.3.md +23 -0
  117. package/docs/releases/0.7.4.md +17 -0
  118. package/docs/releases/0.7.5.md +20 -0
  119. package/docs/releases/0.8.0.md +74 -0
  120. package/docs/releases/0.8.2.md +34 -0
  121. package/docs/releases/0.8.3.md +47 -0
  122. package/package.json +139 -4
  123. package/roadmap/competitive-roadmap.md +3896 -0
  124. package/schemas/release-attestation-v1.schema.json +210 -0
  125. package/schemas/support-bundle-v2.schema.json +212 -0
  126. package/dist/chunks/chunk-DMQAGX77.js +0 -654
  127. package/dist/chunks/chunk-F4Z3Z766.js +0 -4
  128. package/dist/chunks/chunk-SIJAQVSX.js +0 -3
  129. package/dist/chunks/chunk-X6M4HUUE.js +0 -2
  130. package/dist/chunks/chunk-YPRMY2LP.js +0 -8
  131. package/dist/chunks/pure-compression-cli-4TA2TQD5.js +0 -5
  132. package/dist/chunks/server-7EDF4CBY.js +0 -14
  133. package/dist/chunks/structural-fast-path-KD5KQSPX.js +0 -4
  134. package/docs/releases/0.7.6.md +0 -25
@@ -0,0 +1,3896 @@
1
+ # knodin — Competitive Roadmap
2
+
3
+ This is the authoritative active work queue after the original parity program
4
+ closed on 2026-07-21. The completed R1–R48 history, including measured no-go and
5
+ deferred performance decisions, is preserved in
6
+ [`archive/parity-roadmap-2026-07-21.md`](archive/parity-roadmap-2026-07-21.md).
7
+
8
+ ## Product constraint
9
+
10
+ knodin wins by being local, low-latency, and compact: one MCP tool, no hosted
11
+ service, no required credentials, and no source-code egress. Competitor features
12
+ belong here only when live evidence shows that they improve correctness or a
13
+ core developer loop without undermining that constraint.
14
+
15
+ ## Evidence policy
16
+
17
+ Every competitive item must link to a reproducible result under
18
+ `benchmarks/competitors/`. A competitor having a feature is not sufficient by
19
+ itself. Items use four dispositions:
20
+
21
+ - **Must close** — correctness, safety, or a prerequisite for the stated product
22
+ position.
23
+ - **Should close** — meaningful capability or efficiency improvement with a
24
+ bounded implementation path.
25
+ - **Evaluate** — evidence is incomplete or implementation cost may outweigh the
26
+ benefit; do not implement before its gate passes.
27
+ - **Won't pursue** — conflicts with the product constraint or duplicates a
28
+ better knodin mechanism.
29
+
30
+ The first evidence set is
31
+ [`GitNexus vs knodin`](../benchmarks/competitors/gitnexus-vs-knodin.md), backed
32
+ by 49 local cases and their raw MCP results. Later competitor bake-offs append
33
+ new items or strengthen existing ones; they do not reopen the archived roadmap.
34
+
35
+ ## Execution order
36
+
37
+ | ID | Priority | Disposition | Item | Depends on |
38
+ |---|---:|---|---|---|
39
+ | C1 | P0 | Must close | Declare a permissive license | — |
40
+ | C2 | P0 | Must close | Stable symbol identity and ambiguity-safe resolution | — |
41
+ | C3 | P0 | Must close | Enforce truthful response budgets and minimal modes | — |
42
+ | C4 | P1 | Must close | Symbol-level directional impact analysis | C2, C3 |
43
+ | C5 | P1 | Should close | Explicit git diff scopes | C3 |
44
+ | C6 | P1 | Should close | Deterministic import-cycle analysis | C2 |
45
+ | C7 | P2 | Should close | MCP tool and handler mapping | C2 |
46
+ | C8 | P2 | Evaluate | Local statement-level PDG | C2, C3 |
47
+ | C9 | P2 | Evaluate | API contract and response-shape analysis | C2, C3 |
48
+ | C10 | P3 | Evaluate | Bounded expert graph-query DSL | C2, C3 |
49
+ | C11 | P0 | Must close | Call/reference precision and structural identity | C2 |
50
+ | C12 | P1 | Should close | Uniform filters, pagination, and compact drill-downs | C2, C3 |
51
+ | C13 | P1 | Should close | Typed traversal and architecture facets | C2, C3, C11 |
52
+ | C14 | P2 | Should close | Portable budgeted context export | C3, C5 |
53
+ | C15 | P2 | Evaluate | LSP diagnostics and guarded symbol editing | C2, C11 |
54
+ | C16 | P2 | Should close | Index health, repair, and output telemetry | C3 |
55
+ | C17 | P3 | Evaluate | DFS and feature-path navigation | C2, C3, C11 |
56
+ | C18 | P1 | Should close | Unified competitive regression harness | — |
57
+ | C19 | P2 | Evaluate | Salesforce LWC bundles and Apex/data bridges | C2, C3, C11 |
58
+ | C20 | P2 | Evaluate | Salesforce metadata and Flow/automation model | C2, C3, C11 |
59
+ | C21 | P3 | Evaluate | Aura bundles and Visualforce/Aura/LWC interop | C19 |
60
+ | C22 | P3 | Evaluate | LWR and Experience Cloud topology | C19 |
61
+ | C23 | P2 | Evaluate | Apex platform-entry and asynchronous semantics | C2, C3, C11 |
62
+ | C24 | P0 | Must close | Competitive correctness contract and replay baseline | C18 |
63
+ | C25 | P0 | Must close | Prove symbol identity and reference precision leadership | C24, C2, C11 |
64
+ | C26 | P1 | Must close | Prove diff-review and directional traversal leadership | C24, C4, C5, C13 |
65
+ | C27 | P1 | Must close | Prove architecture and community-analysis leadership | C24, C3, C12, C13 |
66
+ | C28 | P1 | Must close | Prove dead-code and semantic-search quality | C24, C3, C11, C12 |
67
+ | C29 | P1 | Must close | Prove context-packing and export leadership | C24, C3, C5, C14 |
68
+ | C30 | P2 | Evaluate | Prove guarded editing and diagnostics leadership | C24, C2, C11, C15 |
69
+ | C31 | P1 | Must close | Prove index lifecycle and output-telemetry leadership | C24, C3, C16, C18 |
70
+ | C32 | P2 | Evaluate | Prove local API and statement-flow analysis leadership | C24, C2, C3, C8, C9 |
71
+ | C33 | P0 | Must close | Preserve local privacy and one-tool schema economy | C24 |
72
+ | C34 | P1 | Must close | Keep local graph indexes current across repository lifecycle events | C16, C24, C33 |
73
+ | C35 | P2 | implemented | Produce local interactive architecture and call-flow visualizations | C3, C13, C24, C33 |
74
+ | C36 | P0 | implemented | Instrument the current warm-operation performance paths | C18, C24 |
75
+ | C37 | P0 | implemented | Reuse generation-scoped graph analytics and traversal snapshots | C36 |
76
+ | C38 | P1 | implemented | Compose orientation context from one shared snapshot | C37 |
77
+ | C39 | P1 | implemented | Separate fast truthful status from explicit deep audit | C36 |
78
+ | C40 | P1 | implemented | Make context packing and artifact access linear and cache-safe | C36 |
79
+ | C41 | P0 | evaluated — deferred | Replay like-for-like performance after the implementation batch | C37, C38, C39, C40 |
80
+ | C42 | P0 | implemented | Enforce comparable competitive cases and correctness oracles | C41 |
81
+ | C43 | P0 | implemented | Measure cold/warm p50, p95, and independent process RSS | C42 |
82
+ | C44 | P0 | implemented | Restore low-latency diff review without weakening evidence | C43 |
83
+ | C45 | P1 | implemented | Add facet-selective architecture fast paths | C43 |
84
+ | C46 | P1 | implemented | Cache search state and hydrate only the requested page | C43 |
85
+ | C47 | P1 | implemented | Reduce status, repair, and measured lifecycle resource cost | C43 |
86
+ | C48 | P2 | evaluated — deferred | Decide an optional local project-memory boundary | C42 |
87
+ | C49 | P1 | implemented | Keep immutable runs lint-safe and make the full suite terminate | — |
88
+ | C50 | P0 | implemented | Make the single-worker full-suite gate start reliably | C49 |
89
+ | C51 | P2 | evaluated — rejected for production | Native training-free embedding quantization | C46 |
90
+ | C52 | P0 | implemented | Package repository lifecycle initialization and refresh | C16, C34, C50 |
91
+ | C53 | P0 | implemented | Make repair observable, cancellable, and resumable | C47, C52 |
92
+ | C54 | P0 | implemented | Certify the local 0.1.0 release artifact | C50, C52, C53 |
93
+ | C55 | P0 | implemented | Pin and audit the authorized Token Optimizer source | C24 |
94
+ | C56 | P0 | implemented | Replay all seven Token Optimizer workflows on shared fixtures | C55 |
95
+ | C57 | P0 | implemented | Add bounded recoverable diagnostic-output compression | C3, C55 |
96
+ | C58 | P0 | evaluated — rejected | Gate command execution on reviewed portable containment | C57 |
97
+ | C59 | P1 | implemented | Connect diagnostic failures to source-evidenced graph context | C57 |
98
+ | C60 | P1 | parked-external-evidence — 2026-08-02 (previous: active — legacy GHES 0.3.0 live; external mirror governance, workload identity, and native certification pending) | Complete two-minute cross-client and GHES onboarding | C52, C55 |
99
+ | C61 | P1 | implemented | Deliver private real-token ROI telemetry and dashboard | C16, C55 |
100
+ | C62 | P0 | parked-external-evidence — 2026-08-02 (previous: active — verifier implemented; production ceremony pending) | Define and implement signed update metadata trust | C55 |
101
+ | C63 | P0 | parked-external-evidence — 2026-08-02 (previous: active — signed-only client implemented; production activation gated) | Add safe update status, check, explain, apply, and rollback policy | C62 |
102
+ | C64 | P0 | parked-external-evidence — 2026-08-02 (previous: active — implementation and adversarial fixtures complete; live release evidence pending) | Produce provenance, SBOM, and cross-channel release attestation | C62 |
103
+ | C65 | P0 | parked-external-evidence — 2026-08-02 (previous: active — bounded multi-step root recovery implemented; production drills pending) | Prove compromised-channel, rollback, freeze, and recovery defenses | C63, C64 |
104
+ | C66 | P0 | parked-external-evidence — 2026-08-02 (previous: active — scorecard published; terminal refresh awaits active dependencies) | Publish an evidence-linked Token Optimizer scorecard | C56, C57, C59, C60, C61, C65, C68, C75 |
105
+ | C67 | P0 | parked-external-evidence — 2026-08-02 (previous: Must close) | Certify and distribute the completed competitive release | C60, C62, C63, C64, C65, C66, C75 |
106
+ | C68 | P0 | implemented | Close Token Optimizer structural output and latency gaps | C56 |
107
+ | C69 | P0 | implemented | Stream large repository inventories with per-repository degradation | C52 |
108
+ | C70 | P0 | implemented | Select portfolio repositories before inventory and reject unknown selectors | C69 |
109
+ | C71 | P0 | implemented | Make portfolio init bounded, resumable, and truthful in dry-run mode | C69, C70 |
110
+ | C72 | P0 | implemented | Make configure/init lifecycle claims atomic and drain compatible queued events | C52 |
111
+ | C73 | P1 | implemented | Distinguish portfolio doctor and installed-versus-active MCP states | C60, C69 |
112
+ | C74 | P1 | implemented | Bound and classify Salesforce metadata architecture candidates | C69 |
113
+ | C75 | P0 | certified with recorded degradations | Certify the 61-repository portfolio under the one-GB memory ceiling | C69, C70, C71, C72, C73, C74 |
114
+ | C76 | P1 | implemented | Replace duplicated CLI parsing/help with one responsive declarative command model | C73 |
115
+ | C77 | P2 | evaluated — retained knodin unchanged | Replay CodeFlow edge-provenance and architecture-export claims | C24, C27, C42 |
116
+ | C78 | P1 | evaluated — retained C8 unchanged | Add bounded source-to-sink resource reachability | C2, C3, C8, C24 |
117
+ | C79 | P1 | implemented | Evaluate or add bounded repository applicability signals | C52, C69 |
118
+ | C80 | P0 | implemented | Route retrieval through structural evidence before embeddings | C3, C24, C42, C43 |
119
+ | C81 | P0 | evaluated — retained knodin unchanged | Run a powered end-to-end benchmark with a diagnose arm | C42, C43, C59, C80 |
120
+ | C82 | P0 | implemented | Productize the retained compress diagnose workflow | C59, C81 |
121
+ | C83 | P0 | implemented | Deliver progressive evidence with a verified hash handshake | C80, C82 |
122
+ | C84 | P1 | implemented | Import optional SCIP evidence without a live-LSP dependency | C24, C80, C83 |
123
+ | C85 | P1 | implemented | Add bounded, itemized git-history review signals | C24, C44, C81 |
124
+ | C86 | P2 | evaluated — retained MiniLM unchanged | Evaluate static embeddings against retained MiniLM | C46, C51, C80, C81 |
125
+ | C87 | P0 | implemented | Add a truthful structural cold-start fast path | C56, C68 |
126
+ | C88 | P1 | implemented — macOS certified; Linux/Windows unavailable | Add profile-based contained execution | C57, C58, C59 |
127
+ | C89 | P0 | implemented | Harden MCP request reliability and recovery | C3, C16, C53 |
128
+ | C90 | P0 | implemented | Produce previewable privacy-safe support bundles | C16, C59, C89 |
129
+ | C91 | P0 | parked-external-evidence — 2026-08-05 (previous: owner-approved adoption priority) | Certify boring macOS installation and Node runtime handoff | C52, C73, C89, C90 |
130
+ | C92 | P0 | parked-external-evidence — 2026-08-06 (previous: proposed) | Unify fail-closed five-channel release orchestration | C64, C91 |
131
+ | C93 | P1 | implemented | Measure engineering task outcomes with and without knodin | C24, C81 |
132
+ | C94 | P1 | parked-external-evidence — 2026-08-06 (previous: blocked) | Make first use self-explanatory and demonstrable | C89, C90, C91 |
133
+ | C95 | P1 | implemented | Codify and replay the defensible behavioral contract | C3, C24, C83 |
134
+ | C96 | P2 | evaluated — deferred | Gate every new capability on a narrow outcome replay | C24, C93, C95 |
135
+ | C97 | P1 | evaluated — retained existing design | Reclaim idle semantic memory without changing one-process architecture | C46, C51, C80, C87 |
136
+ | C98 | P1 | implemented | Add bounded source-to-sink resource reachability | C2, C3, C8, C24, C78, C95 |
137
+ | C99 | P1 | implemented | Resolve bounded TypeScript dependency-injection wiring | C2, C3, C24, C95 |
138
+ | C100 | P1 | implemented | Prove bounded Salesforce Flow-to-Apex paths | C2, C3, C20, C23, C95, C96 |
139
+ | C101 | P1 | implemented | Strengthen contract-led positioning | C95, C97, C98, C99, C100 |
140
+
141
+ ### C98 — Bounded source-to-sink resource reachability
142
+
143
+ - Status: implemented
144
+ - Evidence: `benchmarks/evaluations/c98-resource-reachability/` freezes the strengthened eight-path C78 oracle and records 100% precision, 100% recall, deterministic ordering, and historical cold/warm/RSS measurements across three local repositories. The verifier independently replays correctness and its current sub-1-GB memory gate; portfolio timings remain descriptive rather than acceptance gates.
145
+ - Implementation: `query resource_reachability` performs cached, on-demand TS/JS analysis without persisted resource-flow tables. It recognizes literal `process.env` and `fs.readFileSync` sources, `console.log`, `fetch`, and `db.query` sinks, plus assignment, argument, return, and bounded recursive handoff.
146
+ - Bounds: maximum depth 6 and 100 paths. Results are static heuristics, never runtime reachability or exploitability proof; dynamic names, reflection, computed aliases, unsupported languages, cycles, truncation, continuation, coverage, and freshness remain explicit.
147
+ - Locality: no hosted service, authentication, credentials, source egress, network access, or production model is required.
148
+
149
+ ### C99 — Bounded TypeScript dependency-injection wiring
150
+
151
+ - Status: implemented
152
+ - Evidence: `benchmarks/evaluations/c99-typescript-di/` records 100% precision and recall on its declared Inversify/tsyringe oracle, including an explicit ambiguous-token omission; `npm run verify:c99` rejects stale evidence or any missed/extra edge.
153
+ - Implementation: exact source-only `di_binding` and `di_resolution` references cover `bind().to`, `toSelf`, `register(useClass)`, `registerSingleton`, `@inject`, `get`, and `resolve`. Existing explain, impact, traversal, shortest-path, review, and architecture consumers reuse the persisted relationships.
154
+ - Known bounds: only source-proven framework imports, receivers, token declarations, and uniquely resolved implementations qualify. Multiple bindings, computed tokens, aliases, factories, runtime container modules, conditionals, and reflection remain explicit omissions rather than inferred edges.
155
+
156
+ ### C100 — Bounded Salesforce Flow-to-Apex paths
157
+
158
+ - Status: implemented
159
+ - Evidence: `benchmarks/evaluations/c100-flow-apex-path/` records 100% precision and recall for one preregistered Salesforce DX Flow `actionCalls` to uniquely declared Apex `@InvocableMethod` boundary, including unannotated, missing, and overloaded negative cases. The paired missed-dependency replay improves from an imprecise file answer to a stable method identity with exact XML and Apex evidence.
160
+ - Implementation: Flow indexing persists an exact `flow_apex_action` method reference in addition to its compatible file topology. `query cross_substrate_path <from> <to>` requires both endpoints and returns only that supported substrate transition with stable identities, per-step substrates, extracted provenance, confidence, freshness, exact source evidence, truthful truncation/continuation, and explicit omissions.
161
+ - Known bounds: this is not a general path engine and does not support Terraform, dbt, arbitrary Salesforce metadata, runtime dispatch, namespaced packages, factories, aliases, or ambiguous/multiple Apex methods. One passing boundary and paired task do not establish broad cross-substrate or competitor superiority.
162
+ - Locality: the replay requires no Salesforce org, hosted service, authentication, credentials, network access, or source egress.
163
+
164
+ ### C101 — Contract-led positioning closure
165
+
166
+ - Status: implemented
167
+ - Evidence: `benchmarks/evaluations/c101-contract-positioning/` and
168
+ `scripts/verify-c101.ts` source-bind the C97–C100 decisions, the behavioral
169
+ contract, public CLI/MCP declarations, positioning documents, and repository
170
+ stewardship instructions.
171
+ - Contract: eight adversarial mutations of actual production responses reject
172
+ resource/path stale evidence, continuation-free truncation, a truncated
173
+ response presented as complete, heuristic-to-exact promotion, DI source
174
+ removal, DI ambiguity, and a path omission retained beside a supported step.
175
+ Fixture-backed CLI and one-tool MCP calls return the same bounded resource and
176
+ Flow-to-Apex envelopes; DI reuses
177
+ existing explain/query/review/map surfaces rather than adding a discriminator.
178
+ - Positioning: value is less manual context assembly and fewer missed
179
+ dependencies; ROI's intended mechanism is fewer corrective round trips
180
+ without hosted-service, authentication, or source-egress overhead, while
181
+ C93's null two-task result authorizes no measured productivity claim;
182
+ tomorrow is `knodin init` plus the local CLI or single gateway; the secret
183
+ sauce is stable identity, source evidence, freshness, and truthful budgets
184
+ composed under one contract.
185
+ - Known bounds: C98 is a static TS/JS heuristic, C99 covers only its declared
186
+ Inversify/tsyringe patterns, C100 proves one Flow-to-Apex boundary, and C97
187
+ does not reclaim idle semantic RSS. The closure proves checked fixtures, not
188
+ runtime causality, arbitrary cross-substrate paths, broad framework coverage,
189
+ client/platform certification, or competitor superiority.
190
+ - Stewardship: pull requests invoke the full test matrix and Sonar analysis and
191
+ are therefore long and resource intensive. Related dependency-compatible
192
+ commits normally form one reasonably sized shared PR; assistants do not
193
+ auto-open per-item PRs, and unrelated changes are not bundled for size.
194
+
195
+ ## Roadmap completion contract
196
+
197
+ This section is the control plane for completing this roadmap with a persistent
198
+ goal prompt. The execution-order table remains the full historical ledger. The
199
+ owner split remaining work on 2026-08-02 into Track A, active feature work whose
200
+ acceptance evidence is local and checked in, and Track B, production-release and
201
+ trusted-distribution work parked until its external evidence arrives. Parking
202
+ does not close, weaken, or simulate any gate.
203
+
204
+ Run the roadmap loop with:
205
+
206
+ ```text
207
+ Complete Track A in roadmap/competitive-roadmap.md in dependency order. Preserve its
208
+ product constraint, evidence policy, limitations, user-owned changes, and
209
+ checked-in evidence. Work autonomously on every safe local step. For each item,
210
+ run its acceptance gates and the repository verification bar, update its Status
211
+ and evidence truthfully, and remove it from Track A only when its
212
+ terminal condition is met. Never turn an evaluation into an implementation or
213
+ a superiority claim unless its gate authorizes that outcome. When an external
214
+ credential, production channel, independent custodian, or platform is required,
215
+ route the item to Track B only by an explicit owner decision; do not simulate
216
+ production evidence. Continue until `npm run check:roadmap -- --complete`
217
+ passes. That command proves Track A completion only and reports Track B parked;
218
+ it is not a production-release-readiness claim.
219
+ ```
220
+
221
+ ### Status and closure rules
222
+
223
+ - `implemented`, `done`, `closed`, `certified`, `evaluated — deferred`,
224
+ `evaluated — rejected`, and `won't pursue` are terminal only when the item
225
+ links checked-in evidence satisfying its acceptance or evaluation gate.
226
+ - `proposed`, `active`, `permission-gated`, and `blocked` are non-terminal.
227
+ - An `Evaluate` item closes by recording one allowed disposition: retain the
228
+ current product unchanged, defer/reject the idea with evidence, or create a
229
+ separately numbered implementation item. Evaluation never silently expands
230
+ production scope.
231
+ - P0/P1/P2/P3 affect ordering, not whether an active item must be resolved.
232
+ Track A completion does not close Track B or the entire roadmap.
233
+ - External work is fail-closed. A missing credential, custodian, production
234
+ channel, Windows host, or approval may remain parked with exact evidence; it
235
+ may not produce a synthetic pass. Parked rows remain in Track B.
236
+ - Proposals in ADRs, landscape documents, task systems, or competitor reports
237
+ are outside this roadmap's completion boundary until assigned a C-number in
238
+ the execution-order table. Repository applicability signals are C79 and ADR
239
+ 006 items 1–7 are promoted intact as C80–C86;
240
+ later ADR 006 proposals remain outside the boundary until separately approved.
241
+ - Closing an item requires updating its item-level `Status`, evidence links,
242
+ limitations, the execution-order disposition if needed, and its ledger in
243
+ the same change. Do not rewrite immutable raw benchmark artifacts.
244
+ - `npm run check:roadmap -- --publish-ready` is the ordinary 0.x publication
245
+ guard. It requires Track A to be empty while retaining and reporting Track B.
246
+ The owner clarified on 2026-08-03 that early Docusign usage is a rollout
247
+ cohort, not a separate artifact channel; ordinary releases may publish while
248
+ Track B remains parked. `npm run check:roadmap -- --release-ready` remains the
249
+ stricter GA/trusted-distribution guard and fails unless both tracks are empty.
250
+ - While Track B remains open, documentation and positioning must not claim
251
+ cross-platform certification, trusted distribution, or production-release
252
+ readiness. All existing qualifications and recorded limitations remain in
253
+ force.
254
+
255
+ ### Track A — Active local-work ledger
256
+
257
+ Track A contains only feature or local-evidence work whose acceptance can be
258
+ proved by checked-in tests and replays. Rows are ordered by dependency and then
259
+ priority. The owner approved repository applicability signals as the top item
260
+ on 2026-08-03, followed by ADR 006 items 1–7 intact as C80–C86;
261
+ later ADR proposals remain outside the completion boundary until separately
262
+ approved and numbered.
263
+
264
+ | ID | State | Previous state | Depends on | Exit evidence |
265
+ |---|---|---|---|---|
266
+
267
+ ### Track B — Parked external-evidence ledger
268
+
269
+ Track B parked (10 items, external). The owner parked C60–C67 on 2026-08-02,
270
+ parked C91 after its code-readiness increment on 2026-08-05, and parked C92
271
+ and C94 on 2026-08-06 because their acceptance depends on C91 evidence.
272
+ Every former state, dependency, exit condition, acceptance criterion, and
273
+ limitation remains binding. Re-enter this track when either the documented C60
274
+ packet or C62 packet arrives; completing one packet does not waive the other.
275
+ The durable handoff index is
276
+ `docs/evidence/handoff/README.md`, the C60 packet is
277
+ `docs/evidence/c60-autonomous-preparation-2026-08-02.md`, and the C62 packet is
278
+ `docs/ROOT-CEREMONY-RUNBOOK.md`.
279
+
280
+ | ID | State | Previous state | Depends on | Exit evidence |
281
+ |---|---|---|---|---|
282
+ | C60 | parked-external-evidence — 2026-08-02 | active | C52, C55 | Mirror automation, timed onboarding, and Linux/Windows certification evidence |
283
+ | C62 | parked-external-evidence — 2026-08-02 | active | C55 | Offline root ceremony, independent pin review, and recovery receipt |
284
+ | C63 | parked-external-evidence — 2026-08-02 | active | C62 | Production activation/rollback evidence or explicit safe deferred policy |
285
+ | C64 | parked-external-evidence — 2026-08-02 | active | C62 | Verified release-candidate provenance/SBOM bundles and attestation draft |
286
+ | C65 | parked-external-evidence — 2026-08-02 | active | C63, C64 | Production ceremony and channel/runner/manager defense drills |
287
+ | C66 | parked-external-evidence — 2026-08-02 | active | C56, C57, C59, C60, C61, C65, C68, C75 | Terminal scorecard refresh linked to closed dependencies |
288
+ | C67 | parked-external-evidence — 2026-08-02 | proposed | C60, C62, C63, C64, C65, C66, C75 | Certified release record and distribution receipts |
289
+ | C91 | parked-external-evidence — 2026-08-05 | owner-approved adoption priority | C52, C73, C89, C90 | Native macOS manager/client/Homebrew/Artifactory receipts plus transactional doctor repairs |
290
+ | C92 | parked-external-evidence — 2026-08-06 | owner-approved release priority | C64, C91 | Dry-run preflight, retry replay, immutable artifact proof, and receipt schema |
291
+ | C94 | parked-external-evidence — 2026-08-06 | owner-approved clarity priority | C89, C90, C91 | First-use copy tests and a checked-in five-minute real-repository demonstration |
292
+
293
+ ### Goal-loop ordering
294
+
295
+ 1. Keep Track A active work independent of release authority and prove every
296
+ item with local, checked-in acceptance evidence.
297
+ 2. Re-enter Track B when either C60 or C62's documented evidence packet arrives.
298
+ Execute C60 and C62 in parallel where people and platforms permit.
299
+ 3. After C62, produce C64's release-candidate artifacts. Close C63 with its
300
+ fail-closed deferred-activation policy, then use that terminal client as the
301
+ subject of C65's adversarial drills. C65, not C63, authorizes any later
302
+ production activation.
303
+ 4. Produce C66's terminal prerelease snapshot only from terminal dependency
304
+ evidence. C67 may append release receipts without reopening C66.
305
+ 5. Execute C67 last. GA and trusted-distribution claims remain prohibited until
306
+ Track B is empty and `npm run check:roadmap -- --release-ready` passes.
307
+ Ordinary 0.x publication requires `npm run check:roadmap -- --publish-ready`.
308
+ 6. Keep C92 and C94 parked behind C91. Re-enter C91 only when the named native macOS
309
+ environments and credentialed channels are available and the transactional
310
+ doctor repair contract is implemented; deterministic fixtures alone do not
311
+ satisfy certification.
312
+
313
+ ---
314
+
315
+ ## C1 — Declare a permissive license
316
+
317
+ - Status: implemented
318
+ - Priority: P0
319
+ - Disposition: Must close
320
+ - Evidence: the GitNexus bake-off found GitNexus explicitly licensed under
321
+ PolyForm Noncommercial 1.0.0 while knodin has neither a `LICENSE` file nor a
322
+ `package.json` license field. “Local and free” is not a legally supportable
323
+ marketing claim until this is resolved.
324
+ - Touches: `LICENSE`, `package.json`, `README.md`, `docs/COMPARISON.md`
325
+
326
+ ### Acceptance criteria
327
+
328
+ 1. The repository owner selects and adds an OSI-approved permissive license;
329
+ implementation must not guess the owner's legal choice.
330
+ 2. `package.json` declares the matching SPDX identifier.
331
+ 3. README and comparison copy use only claims permitted by that license.
332
+ 4. Package contents include the license after `bun pm pack --dry-run` or the
333
+ repository's equivalent packaging check.
334
+
335
+ ---
336
+
337
+ ## C2 — Stable symbol identity and ambiguity-safe resolution
338
+
339
+ - Status: implemented
340
+ - Priority: P0
341
+ - Disposition: Must close
342
+ - Evidence: GitNexus's `trace(main, createEngine)` returned three ranked `main`
343
+ candidates and required disambiguation. knodin silently returned
344
+ `main → createEngine`, with no indication of which definition it selected.
345
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
346
+ symbol/query/explain/rename/path tests, persisted schema if IDs must survive
347
+ rebuilds
348
+
349
+ ### Required behavior
350
+
351
+ 1. Every symbol-facing result exposes a stable repository-scoped identity.
352
+ 2. `explain`, structured queries, rename preview/apply, and shortest path accept
353
+ optional identity, file, and kind selectors.
354
+ 3. A bare-name request with multiple viable definitions returns an explicit
355
+ ambiguity result with ranked candidates; it never silently merges or chooses
356
+ definitions when that can change the answer.
357
+ 4. Existing unambiguous bare-name calls remain source-compatible.
358
+
359
+ ### Acceptance criteria
360
+
361
+ 1. A fixture with at least three same-named functions reproduces the GitNexus
362
+ `main` case and proves no conflated path is returned.
363
+ 2. Selecting each candidate by identity returns only that candidate's real
364
+ callers, callees, impact, and paths.
365
+ 3. Rename by identity edits only the selected definition and its resolved
366
+ references.
367
+ 4. Migration/rebuild behavior for persisted IDs is deterministic and tested.
368
+
369
+ ---
370
+
371
+ ## C3 — Enforce truthful response budgets and minimal modes
372
+
373
+ - Status: implemented
374
+ - Priority: P0
375
+ - Disposition: Must close
376
+ - Evidence: knodin's minimal symbol context still returned 22,843 bytes because
377
+ source was always present, and minimal `map` returned 254,102 bytes. GitNexus
378
+ was usually more compact when source was not requested. Conversely,
379
+ GitNexus's source-inclusive query returned 606,235 bytes, demonstrating why a
380
+ hard cap—not convention—is required.
381
+ - Reinforced by Graphify's 417-byte top-10 hubs versus knodin's 258,628-byte
382
+ mapped response, Aider's 425-byte mean at a 128-token budget versus knodin's
383
+ roughly 190 KB mapped mean, and Repomix/code2prompt's explicit token metrics.
384
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, response-size and
385
+ operation tests, docs
386
+
387
+ ### Required behavior
388
+
389
+ 1. Every operation has a tested serialized-byte ceiling and reports truncation
390
+ plus total counts when the ceiling binds.
391
+ 2. Minimal `explain` omits source unless explicitly requested; standard explain
392
+ retains edit-ready source within its cap.
393
+ 3. Minimal `map` returns counts, aggregates, and bounded top-N summaries only;
394
+ it never leaks full community member arrays or full edge lists.
395
+ 4. Caps are applied after serialization-aware accounting so nested content
396
+ cannot bypass them.
397
+ 5. Callers may choose token, byte, and item budgets; all three produce honest
398
+ totals, truncation state, and a continuation/drill-down route.
399
+ 6. Hub, bridge, community, traversal, search, source, and export operations honor
400
+ their requested top-N or budget without constructing a full map response.
401
+
402
+ ### Acceptance criteria
403
+
404
+ 1. The GitNexus benchmark's `context-name-min` and `cypher-communities`
405
+ counterparts remain below documented budgets.
406
+ 2. Large-symbol and large-map fixtures assert both byte ceilings and honest
407
+ `truncated`/total-count metadata.
408
+ 3. No standard response silently loses edit-critical source without a clear
409
+ continuation or drill-down route.
410
+
411
+ ---
412
+
413
+ ## C4 — Symbol-level directional impact analysis
414
+
415
+ - Status: implemented
416
+ - Priority: P1
417
+ - Disposition: Must close
418
+ - DependsOn: C2, C3
419
+ - Evidence: GitNexus distinguished upstream/downstream/both, depths 1/3/6,
420
+ tests, relation confidence, processes, and modules for `createEngine`.
421
+ knodin's closest public impact query was file-level and returned the same
422
+ broad 31-row set for every directional comparison.
423
+ - Reinforced by codebase-memory-mcp's directional typed trace and argument-level
424
+ data-flow evidence, and Graphify's relation-filtered direct neighbors.
425
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
426
+ impact/query tests and docs
427
+
428
+ ### Required behavior
429
+
430
+ 1. Impact accepts a stable symbol selector or an explicit file-level mode.
431
+ 2. Symbol mode supports direction, maximum depth, relation kinds, confidence
432
+ threshold, test inclusion, result limit, argument/data-flow evidence, and
433
+ compact summary.
434
+ 3. Results preserve depth and edge evidence and identify affected flows and
435
+ communities without presenting heuristic reach as exact.
436
+ 4. File-level impact remains available and is labeled distinctly.
437
+
438
+ ### Acceptance criteria
439
+
440
+ 1. The `createEngine` fixture produces meaningfully different upstream and
441
+ downstream results.
442
+ 2. Depth limits are monotonic and enforced; confidence/relation filters have
443
+ targeted fixtures.
444
+ 3. Ambiguous symbols use C2's candidate response rather than conflated reach.
445
+ 4. Compact output satisfies C3's byte budget.
446
+
447
+ ---
448
+
449
+ ## C5 — Explicit git diff scopes
450
+
451
+ - Status: implemented
452
+ - Priority: P1
453
+ - Disposition: Should close
454
+ - DependsOn: C3
455
+ - Evidence: GitNexus supports `unstaged`, `staged`, `all`, and `compare`.
456
+ knodin cannot request staged-only review, and its result can omit changed
457
+ files when no indexed symbol maps to the diff.
458
+ - Reinforced by code-review-graph's caller-supplied changed-file review and
459
+ code2prompt's working-tree diff, branch diff, and branch log exports.
460
+ - Touches: git diff parsing, engine review contract, MCP/CLI schemas, review
461
+ tests and docs
462
+
463
+ ### Acceptance criteria
464
+
465
+ 1. Review accepts `unstaged|staged|all|compare`; compare requires or defaults a
466
+ documented base ref.
467
+ 2. Each scope has a fixture containing different staged and unstaged edits.
468
+ 3. Results always report changed files, including non-code/unindexed files,
469
+ separately from mapped changed symbols.
470
+ 4. Existing `base` behavior has a documented compatibility mapping.
471
+ 5. Review accepts an explicit file list or revision pair without requiring the
472
+ checkout to contain an equivalent live diff.
473
+
474
+ ---
475
+
476
+ ## C6 — Deterministic import-cycle analysis
477
+
478
+ - Status: implemented
479
+ - Priority: P1
480
+ - Disposition: Should close
481
+ - DependsOn: C2
482
+ - Evidence: GitNexus exposes a bounded read-only cycle check; knodin has no
483
+ direct equivalent despite already storing import dependencies.
484
+ - Touches: `src/engine/index.ts`, query schema/docs, cycle fixtures
485
+
486
+ ### Acceptance criteria
487
+
488
+ 1. A structured query returns canonical deterministic file-import cycles.
489
+ 2. Equivalent rotations of one cycle are deduplicated.
490
+ 3. Limits and truncation metadata prevent exponential output.
491
+ 4. Fixtures cover no cycle, one cycle, overlapping cycles, and type-only/import
492
+ edge policy.
493
+
494
+ ---
495
+
496
+ ## C7 — MCP tool and handler mapping
497
+
498
+ - Status: implemented
499
+ - Priority: P2
500
+ - Disposition: Should close
501
+ - DependsOn: C2
502
+ - Evidence: GitNexus correctly detected knodin's `knodin` MCP tool and its
503
+ source file. knodin has endpoint patterns but no MCP registration-to-handler
504
+ map.
505
+ - Touches: language extraction, schema, structured queries, MCP fixtures/docs
506
+
507
+ ### Acceptance criteria
508
+
509
+ 1. Index common TypeScript MCP SDK registration forms and associate tool name,
510
+ description, schema declaration, handler symbol, and file.
511
+ 2. A query supports all tools and lookup by tool name.
512
+ 3. Dynamic/unresolved registrations are labeled heuristic rather than exact.
513
+ 4. The repository's own `knodin` registration is found without matching prose
514
+ examples or benchmark fixture strings.
515
+
516
+ ---
517
+
518
+ ## C8 — On-demand local statement-level flow analysis
519
+
520
+ - Status: implemented (bounded on-demand analysis; no persisted PDG)
521
+ - Priority: P2
522
+ - Disposition: Evaluate
523
+ - DependsOn: C2, C3
524
+ - Evidence: GitNexus's opt-in local PDG returned concrete control-dependence and
525
+ reaching-definition rows, including variable-filtered `db` flows. knodin has
526
+ no statement-level answer. GitNexus also produced large 14–42 KB PDG
527
+ responses, so copying its output contract would conflict with knodin's
528
+ compactness goal.
529
+ - Touches if approved: parser/extractor architecture, persisted schema,
530
+ performance harness, query surface, multi-language fixtures
531
+ - Evaluation evidence: [`C8 summary and methodology`](../benchmarks/evaluations/c8-pdg/summary-methodology.md)
532
+ - Decision: the disposable prototype passed the mechanical gate at 100% labeled
533
+ fixture precision, with three real repositories measured and responses below
534
+ the C3 ceiling. The supported implementation is the narrower on-demand
535
+ `flow_analysis` query for one selected TS/JS symbol, with optional variable
536
+ filtering, compact source evidence, C3 budgets, and explicit unresolved cases;
537
+ it does not persist a broad PDG.
538
+
539
+ ### Evaluation gate
540
+
541
+ 1. Prototype TS/JS control-flow and reaching-definitions on disposable code,
542
+ not the production schema.
543
+ 2. Measure clean-index time, incremental-index time, database growth, peak RSS,
544
+ query latency, precision, and response size against at least three real
545
+ repositories.
546
+ 3. Require ≥90% precision on a labeled control/data-flow fixture and a bounded
547
+ response below C3's ceiling.
548
+ 4. Approve implementation only if the value cannot be delivered more cheaply
549
+ through compiler/LSP facts or a narrower on-demand analysis.
550
+
551
+ ---
552
+
553
+ ## C9 — Evaluate API contract and response-shape analysis
554
+
555
+ - Status: implemented
556
+ - Priority: P2
557
+ - Disposition: Evaluate
558
+ - DependsOn: C2, C3
559
+ - Evidence: GitNexus exposes route maps, response-shape checking, and API impact,
560
+ but this repository had zero detected HTTP routes. The bake-off proves the
561
+ surface exists, not that it is correct or useful. knodin now persists
562
+ source-evidenced literal Express/Fastify route and fetch/Axios client facts,
563
+ with a bounded `api_contract_mismatches` query.
564
+ - Touches: route extraction, schema/shape model, cross-file/client usage
565
+ resolution, fixtures and query surface
566
+ - Evaluation evidence: [`C9 summary`](../benchmarks/evaluations/c9-api/summary.md)
567
+ - Decision: the evaluation gate passed. GitNexus found the actionable `profile`
568
+ response mismatch that current handler, endpoint, and impact queries missed.
569
+ The production implementation shipped after the checked-in fixture produced
570
+ 3 true positives, 0 false positives, and 0 false negatives (precision 1.00,
571
+ recall 1.00); dynamic routes remain unresolved and absolute-origin clients do
572
+ not link to local endpoints. See [`C9 production acceptance`](../benchmarks/evaluations/c9-api/production-acceptance.md).
573
+
574
+ ### Evaluation gate
575
+
576
+ 1. Build a local fixture with multiple verbs on one route, middleware, typed and
577
+ inferred response bodies, internal clients, mismatches, and dynamic routes.
578
+ 2. Run GitNexus and knodin endpoint primitives against the same fixture and
579
+ manually label precision/recall.
580
+ 3. Approve implementation only if shape analysis finds actionable breakage that
581
+ current handler/endpoint/impact queries miss.
582
+
583
+ ---
584
+
585
+ ## C10 — Evaluate a bounded expert graph-query DSL
586
+
587
+ - Status: closed; general DSL will not be pursued
588
+ - Priority: P3
589
+ - Disposition: Evaluate
590
+ - DependsOn: C2, C3
591
+ - Evidence: GitNexus's parameterized Cypher expressed caller, community
592
+ aggregate, and custom dead-code queries outside a fixed enum. knodin's fixed
593
+ patterns are easier to route and safer but cannot answer arbitrary structural
594
+ questions. GitNexus's naive custom dead-code query also returned many false
595
+ positives, illustrating the cost of unconstrained power.
596
+ - Touches if reopened: parser/validator, query planner, response budget,
597
+ read-only safety tests and docs
598
+ - Evaluation evidence: [`C10 summary`](../benchmarks/evaluations/c10-dsl/summary.md)
599
+ - Decision: no DSL is justified in this roadmap slice. Three bounded fixed-pattern
600
+ candidates covered 12 of 13 provenance-backed questions (92.3%), so the DSL
601
+ prototype gate is a no-go.
602
+
603
+ ### Evaluation gate
604
+
605
+ 1. Collect at least ten real questions that existing patterns cannot answer.
606
+ 2. Determine whether two or three new fixed patterns cover most demand.
607
+ 3. If not, design a read-only bounded DSL with explicit node/edge allowlists,
608
+ mandatory limits, traversal-depth caps, and no raw SQL/Cypher execution.
609
+ 4. Approve only with deterministic cost bounds and C3-compliant output.
610
+
611
+ ---
612
+
613
+ ## C11 — Call/reference precision and structural identity
614
+
615
+ - Status: implemented
616
+ - Priority: P0
617
+ - Disposition: Must close
618
+ - DependsOn: C2
619
+ - Evidence: Graphify found four credible call neighbors for `createEngine` while
620
+ knodin included a known name-collision false positive; Serena returned precise
621
+ LSP references and structural `KnodinEngine` implementations knodin's nominal
622
+ inheritance query missed. grepai and codebase-memory-mcp also demonstrated
623
+ that imports/files must not be mislabeled as callers.
624
+
625
+ ### Acceptance criteria
626
+
627
+ 1. Labeled fixtures distinguish call, import, usage, containment, inheritance,
628
+ and TypeScript structural implementation edges.
629
+ 2. Callers/callees reach at least 95% precision and recall on those fixtures;
630
+ built-ins and same-name symbols never create cross-definition edges.
631
+ 3. Structural implementations are queryable separately from nominal inheritors.
632
+ 4. Tests cover top-level calls, aliases, barrels, object-literal interface
633
+ implementations, scripts, tests, and ambiguous names.
634
+
635
+ ---
636
+
637
+ ## C12 — Uniform filters, pagination, and compact drill-downs
638
+
639
+ - Status: implemented
640
+ - Priority: P1
641
+ - Disposition: Should close
642
+ - DependsOn: C2, C3
643
+ - Evidence: code-review-graph filters semantic kind, large-code threshold/kind,
644
+ flow/community sort and detail; codebase-memory-mcp adds label, qualified name,
645
+ file, relationship, degree, total/offset pagination; Claude Context filters
646
+ extensions. Graphify honors top-N hubs and relation filters.
647
+
648
+ ### Acceptance criteria
649
+
650
+ 1. Search accepts language/extension, kind, path, test/production, and source
651
+ inclusion filters plus offset/cursor pagination with total and `hasMore`.
652
+ 2. Large-code queries accept explicit line/complexity thresholds, symbol kind,
653
+ and path; no requested parameter is silently ignored.
654
+ 3. Hub, bridge, community, flow, and traversal drill-downs honor top-N, sort,
655
+ relation, and detail controls under C3.
656
+ 4. Parameter-combination tests prove filters compose rather than overwrite one
657
+ another.
658
+
659
+ ---
660
+
661
+ ## C13 — Typed traversal and architecture facets
662
+
663
+ - Status: implemented
664
+ - Priority: P1
665
+ - Disposition: Should close
666
+ - DependsOn: C2, C3, C11
667
+ - Evidence: codebase-memory-mcp beat knodin on inbound/outbound/both traces,
668
+ edge types, argument expressions, tests, risk labels, packages, layers,
669
+ boundaries, hotspots, and path scoping. Graphify adds direct relation-filtered
670
+ neighbors.
671
+
672
+ ### Acceptance criteria
673
+
674
+ 1. Traversal supports direction and explicit edge types with per-hop evidence.
675
+ 2. Call/data-flow results preserve argument expressions where extraction is
676
+ compiler-grounded and label heuristic facts honestly.
677
+ 3. Architecture exposes independently selectable packages, layers, boundaries,
678
+ hotspots, entry points, languages, and path scope.
679
+ 4. Results satisfy C11 precision fixtures and C3 budgets.
680
+
681
+ ---
682
+
683
+ ## C14 — Portable budgeted context export
684
+
685
+ - Status: implemented
686
+ - Priority: P2
687
+ - Disposition: Should close
688
+ - DependsOn: C3, C5
689
+ - Evidence: Repomix, code2prompt, and Aider provide portable source-shaped
690
+ context, deterministic formatting, token accounting, file policies, ranged
691
+ retrieval, mention/chat awareness, and Git context. These are complementary
692
+ to knodin's graph rather than replacements for it.
693
+
694
+ ### Acceptance criteria
695
+
696
+ 1. Export supports Markdown, JSON, and XML with relative paths, optional line
697
+ numbers/tree, deterministic ordering, and explicit tokenizer cost estimate.
698
+ 2. Include/exclude rules and per-file `full|summary|structure-only` policies are
699
+ composable with a hard token/byte budget.
700
+ 3. Already-present/chat files can be excluded from duplicated source context.
701
+ 4. Optional diff/log sections use C5 scopes; packed artifacts support bounded
702
+ ranged reads and exact regex grep.
703
+
704
+ ---
705
+
706
+ ## C15 — Evaluate LSP diagnostics and guarded symbol editing
707
+
708
+ - Status: implemented (read-only TypeScript language-service queries)
709
+ - Priority: P2
710
+ - Disposition: Evaluate
711
+ - DependsOn: C2, C11
712
+ - Evidence: Serena successfully exposed diagnostics, declarations,
713
+ implementations, symbol-body replacement, before/after insertion, safe delete,
714
+ and guarded multi-file replacement. knodin's verified rename is safer but much
715
+ narrower. The 2026-07-22 disposable evaluation exercised TypeScript Language
716
+ Server, Pyright, JDTLS, Java/Javac, .NET compilation, and the official Roslyn
717
+ C# Language Server with guarded TypeScript previews and rollback. The final
718
+ raw result records a C# `textDocument/diagnostic` response and
719
+ `textDocument/implementation` navigation to the fixture's line-10
720
+ implementation; the passing compiler check remains corroboration rather than a
721
+ substitute for LSP evidence. knodin now ships optional local TypeScript
722
+ diagnostics, definition, declaration, and implementation queries with no
723
+ daemon, subprocess, workspace mutation, or index initialization. Missing or
724
+ unsupported adapters return explicit unavailable results; guarded editing and
725
+ cross-language mutation remain unapproved.
726
+
727
+ ### Evaluation gate
728
+
729
+ 1. Prototype a local TypeScript/LSP adapter for diagnostics, implementation, and
730
+ edit previews without adding a required daemon.
731
+ 2. Require all edits to produce a preview and reuse rename's compiler check and
732
+ rollback guarantees; never expose unguarded filesystem mutation.
733
+ 3. Measure startup/RSS/latency and verify behavior across TS, Python, Java, and
734
+ C# fixtures before approving a general editing surface.
735
+
736
+ ---
737
+
738
+ ## C16 — Index health, repair, and output telemetry
739
+
740
+ - Status: implemented
741
+ - Priority: P2
742
+ - Disposition: Should close
743
+ - DependsOn: C3
744
+ - Evidence: codebase-memory-mcp exposes portable lifecycle modes and schema;
745
+ Claude Context exposes clear/status lifecycle; mcp-codebase-index statically
746
+ defines health/repair; grepai records output/token savings. knodin has strong
747
+ freshness internals but limited public health and budget telemetry.
748
+
749
+ ### Acceptance criteria
750
+
751
+ 1. Status reports schema/model/version, file and symbol coverage, orphaned or
752
+ missing records, last successful reconciliation, and actionable repair steps.
753
+ 2. A local repair command can rebuild only damaged/missing state and verify the
754
+ result without deleting healthy data.
755
+ 3. Benchmark telemetry records operation, latency, serialized bytes, estimated
756
+ tokens, truncation, and detail mode without recording source content.
757
+ 4. All telemetry is local, opt-in for persistence, and documented as such.
758
+
759
+ ---
760
+
761
+ ## C17 — Evaluate DFS and feature-path navigation
762
+
763
+ - Status: implemented
764
+ - Priority: P3
765
+ - Disposition: Evaluate
766
+ - DependsOn: C2, C3, C11
767
+ - Evidence: Graphify and code-review-graph expose DFS; grepai adds deterministic
768
+ functional feature paths. `query feature_path <symbol>` now supplies a compact,
769
+ deterministic downstream DFS over C2/C11-resolved references. It includes
770
+ file/line edge evidence, cycle guards, depth/node caps, and explicit
771
+ truncation without adding persistent identity state. The focused fixture-led
772
+ regression in `feature-path.spec.ts` proves a branching result observably
773
+ differs from BFS while remaining stable after harmless line shifts.
774
+
775
+ ### Delivered behavior
776
+
777
+ 1. `feature_path` follows only resolved downstream references in stable DFS
778
+ order; `traverse` remains the existing selectable-direction BFS neighborhood.
779
+ 2. Standard-detail non-root nodes include C11-resolved identity; every hop
780
+ includes source file/line, kind, provenance, and confidence evidence.
781
+ 3. Cycles are visited once, depth is clamped to 1–6, the item cap is enforced,
782
+ and a larger reachable graph reports `truncated: true`.
783
+ 4. The feature is source-only: dynamic dispatch, reflection, unresolved calls,
784
+ and runtime execution order remain explicitly outside the contract.
785
+
786
+ ---
787
+
788
+ ## C18 — Unified competitive regression harness
789
+
790
+ - Status: implemented
791
+ - Priority: P1
792
+ - Disposition: Should close
793
+ - Evidence: the first program produced 12 independent protocols and 713 paired
794
+ cases. Without one version-pinned entry point, improvements to C1-C17 could
795
+ silently lose coverage or be compared against a different tool installation.
796
+ - Touches: `scripts/competitive-bakeoff.ts`, `src/competitive-manifest.ts`,
797
+ `src/competitive-runner.ts`, `benchmarks/competitors/README.md`
798
+
799
+ ### Acceptance criteria
800
+
801
+ 1. One manifest selects competitors directly or by C1-C17 roadmap coverage and
802
+ records exact versions, prerequisites, preparation, commands, and blockers.
803
+ 2. Runs refuse dirty checkouts by default, validate installed versions, isolate
804
+ fixture mutations, and preserve immutable timestamp-and-commit artifacts.
805
+ 3. Heterogeneous raw results normalize into one summary and compare against the
806
+ checked-in baseline with a configurable latency tolerance and nonzero
807
+ regression exit status.
808
+ 4. List, prerequisite-check, and dry-run modes are non-mutating; commands use
809
+ argument arrays rather than shell interpolation.
810
+ 5. Unit tests cover selection, immutable history, normalization, blockers, and
811
+ material-regression detection.
812
+
813
+ ---
814
+
815
+ ## Salesforce scope boundary and remaining limitations
816
+
817
+ The baseline parses Apex and Visualforce and now includes the bounded C19–C23
818
+ implementations for LWC, selected metadata/Flow automation, Aura interop, LWR
819
+ topology, and Apex platform entry/asynchronous semantics. Their individual
820
+ sections and fixtures define the supported static evidence. This remains
821
+ partial Salesforce coverage, not proof that every deployed-org dependency or
822
+ runtime dispatch is resolved. Static dead-code candidates still require
823
+ corroboration against deployed-org and platform dependency data before deletion.
824
+
825
+ OmniStudio and industry-cloud packages are intentionally out of scope: promote
826
+ them only after a concrete, legally usable Salesforce repository demonstrates a
827
+ workflow that C19–C23 do not cover.
828
+
829
+ Their proposed shared `src/engine/index.ts` footprint is a collision, so they
830
+ must be scheduled sequentially even where their logical dependencies are
831
+ independent.
832
+
833
+ ---
834
+
835
+ ## C19 — Evaluate Salesforce LWC bundles and Apex/data bridges
836
+
837
+ - Status: implemented
838
+ - Priority: P2
839
+ - Disposition: Evaluate
840
+ - DependsOn: C2, C3, C11
841
+ - Evidence: the version-pinned, source-only Salesforce DX fixture now records
842
+ 13/13 exact bundle, component, Apex, schema, wire-adapter, and Lightning Data
843
+ Service labels (100% precision, 100% recall) in
844
+ `benchmarks/evaluations/c19-salesforce-lwc/raw-results-20260722T132842614Z.json`.
845
+ The replay copies the fixture to a temporary directory, needs no org login,
846
+ credentials, network access, or source egress, and records one-file LWC/Apex
847
+ reindex plus C3-bounded response measurements. Dynamic and namespaced forms
848
+ remain omitted rather than exact claims. The raw result hashes the evaluated
849
+ engine, runner, and label revision in addition to the fixture. This authorizes no broader Salesforce
850
+ product commitment beyond the evaluated static source edges.
851
+ - Touches if approved: `src/engine/index.ts`, `src/__tests__/unit/salesforce.spec.ts`,
852
+ `src/__tests__/unit/lwc.spec.ts` (new),
853
+ `benchmarks/evaluations/c19-salesforce-lwc/` (new),
854
+ `src/competitive-manifest.ts`
855
+
856
+ ### Evaluation gate
857
+
858
+ 1. Build a self-contained Salesforce DX fixture with LWC JavaScript, HTML, CSS,
859
+ and `*.js-meta.xml` bundle members; include component-to-component imports,
860
+ `@salesforce/apex`, `@salesforce/schema`, wire adapters, and Lightning Data
861
+ Service usage alongside resolvable Apex methods.
862
+ 2. Label bundle containment plus LWC-to-Apex, LWC-to-object/field, and
863
+ component-to-component edges. Require at least 95% precision and 90% recall;
864
+ unresolved, dynamic, or namespaced references must be omitted or labeled
865
+ heuristic rather than reported as exact.
866
+ 3. Re-index a one-file LWC and one-file Apex change, measure index/response
867
+ cost against C3, and prove the analysis needs neither an org login nor source
868
+ egress.
869
+
870
+ ---
871
+
872
+ ## C20 — Evaluate Salesforce metadata and Flow/automation model
873
+
874
+ - Status: implemented
875
+ - Priority: P2
876
+ - Disposition: Evaluate
877
+ - DependsOn: C2, C3, C11
878
+ - Evidence: the version-pinned source-only Salesforce DX fixture records 15/15
879
+ exact custom-object/field/record-type, permission-set, Flow, workflow, and
880
+ approval-process labels (100% precision and recall) in
881
+ `benchmarks/evaluations/c20-salesforce-metadata/raw-results-20260722T134732728Z.json`.
882
+ The replay copies the fixture to a temporary directory, runs a metadata-only
883
+ one-file update, stays within C3's response budget, and needs no org login,
884
+ deployment, credentials, network access, or source egress. Formula and
885
+ runtime-selected targets remain explicit unknowns rather than exact edges.
886
+ This authorizes no broader Salesforce product commitment beyond the evaluated
887
+ static source relationships.
888
+ - Touches if approved: `src/engine/index.ts`,
889
+ `src/__tests__/unit/salesforce-metadata.spec.ts` (new),
890
+ `benchmarks/evaluations/c20-salesforce-metadata/` (new),
891
+ `src/competitive-manifest.ts`
892
+
893
+ ### Evaluation gate
894
+
895
+ 1. Build a version-pinned Salesforce DX fixture containing custom objects and
896
+ fields, record types, permission sets, a record-triggered Flow, a screen or
897
+ autolaunched Flow, workflow/approval metadata, and a Flow-invoked Apex
898
+ action.
899
+ 2. Label only statically declared metadata edges: Flow/automation to Apex,
900
+ object, field, permission, and referenced UI component. Require at least
901
+ 95% precision and 90% recall; dynamic formulas and runtime-selected targets
902
+ must remain explicit unknowns.
903
+ 3. Demonstrate that a metadata-only change updates the affected topology without
904
+ requiring deployment to an org, and keep output within C3's budget.
905
+
906
+ ---
907
+
908
+ ## C21 — Evaluate Aura bundles and Visualforce/Aura/LWC interop
909
+
910
+ - Status: implemented
911
+ - Priority: P3
912
+ - Disposition: Evaluate
913
+ - DependsOn: C19
914
+ - Evidence: the version-pinned, source-only Salesforce DX fixture records 10/10
915
+ exact Aura containment, Aura-to-Apex, Aura-to-LWC, and Visualforce controller
916
+ labels (100% precision and recall) in
917
+ `benchmarks/evaluations/c21-salesforce-interop/raw-results-20260722T161801709Z.json`.
918
+ The replay copies its fixture to a temporary directory, needs no org login,
919
+ credentials, network access, or source egress, and verifies source evidence
920
+ for every claimed cross-stack edge. knodin supports these static source
921
+ relationships; `$A` runtime lookup strings remain omitted rather than exact
922
+ claims.
923
+ - Touches: `src/engine/index.ts`,
924
+ `src/__tests__/unit/visualforce.spec.ts`,
925
+ `src/__tests__/unit/salesforce-interop.spec.ts` (new),
926
+ `benchmarks/evaluations/c21-salesforce-interop/` (new),
927
+ `src/competitive-manifest.ts`
928
+
929
+ ### Evaluation gate
930
+
931
+ 1. Build a fixture with Aura component, application, interface, event, and
932
+ design members; Visualforce page/component/controller bindings; and a
933
+ supported Aura-to-LWC or Visualforce-to-Lightning bridge.
934
+ 2. Label containment and cross-stack edges, including Aura-to-Apex and the
935
+ declared bridge target. Require at least 95% precision and 90% recall, while
936
+ `$A` runtime lookup strings and unresolvable expressions remain heuristic.
937
+ 3. Confirm that unchanged Apex/Visualforce behavior retains its current test
938
+ results and that cross-stack results expose the source evidence for every
939
+ claimed edge.
940
+
941
+ ---
942
+
943
+ ## C22 — Evaluate LWR and Experience Cloud topology
944
+
945
+ - Status: implemented
946
+ - Priority: P3
947
+ - Disposition: Evaluate
948
+ - DependsOn: C19
949
+ - Evidence: the version-pinned, source-only LWR and Experience Cloud fixture
950
+ records 9/9 exact site, route, view, LWC/Aura, theme, and SVG asset labels
951
+ (100% precision and recall) in
952
+ `benchmarks/evaluations/c22-salesforce-experience/raw-results-20260722T162757428Z.json`.
953
+ The replay copies its fixture to a temporary directory, measures a route-only
954
+ reindex, stays within C3's 65,536-byte map budget, and needs no org login,
955
+ credentials, network access, source egress, server, or deployment. knodin
956
+ supports this static source topology; runtime route expressions remain
957
+ unresolved rather than fabricated.
958
+ - Touches: `src/engine/index.ts`,
959
+ `src/__tests__/unit/salesforce-experience.spec.ts` (new),
960
+ `benchmarks/evaluations/c22-salesforce-experience/` (new),
961
+ `src/competitive-manifest.ts`
962
+
963
+ ### Evaluation gate
964
+
965
+ 1. Build a source-only fixture containing LWR configuration plus the supported
966
+ Experience Cloud metadata for one site, its routes, views, theme/assets, and
967
+ declared LWC or Aura targets.
968
+ 2. Label site-to-route-to-view-to-component and static-asset relationships.
969
+ Require at least 95% precision and 90% recall; route values assembled at
970
+ runtime must be labeled unresolved rather than fabricated.
971
+ 3. Measure a route/view-only incremental update and prove that topology output
972
+ is compact, locally computed, and does not require serving or deploying the
973
+ site.
974
+
975
+ ---
976
+
977
+ ## C23 — Evaluate Apex platform-entry and asynchronous semantics
978
+
979
+ - Status: implemented
980
+ - Priority: P2
981
+ - Disposition: Evaluate
982
+ - DependsOn: C2, C3, C11
983
+ - Evidence: the version-pinned, source-only Salesforce DX fixture records 20/20
984
+ exact REST/SOAP/invocable/trigger, async callback, enqueue, subscription, and
985
+ Salesforce Function labels (100% precision and recall) in
986
+ `benchmarks/evaluations/c23-apex-platform-semantics/raw-results-20260722T140720946Z.json`.
987
+ The replay copies the fixture to a temporary directory, reindexes one Apex
988
+ file, keeps map responses under C3's 65,536-byte budget, and requires no org
989
+ login, credentials, network access, or source egress. Transaction ordering,
990
+ bulk behavior, retries, and runtime payload shape remain unknown.
991
+ - Touches if approved: `src/engine/index.ts`,
992
+ `src/__tests__/unit/salesforce.spec.ts`,
993
+ `src/__tests__/unit/apex-platform-semantics.spec.ts` (new),
994
+ `benchmarks/evaluations/c23-apex-platform-semantics/` (new),
995
+ `src/competitive-manifest.ts`
996
+
997
+ ### Evaluation gate
998
+
999
+ 1. Build a version-pinned fixture covering `@RestResource` and verb handlers,
1000
+ `webService`/SOAP, `@InvocableMethod`, trigger object/event declarations,
1001
+ `@future`, Queueable, Batchable, Schedulable, Platform Event or change-event
1002
+ subscribers, and statically resolvable Salesforce Function invocations.
1003
+ 2. Label entry-point, enqueue, callback, subscription, and invocation edges;
1004
+ require at least 95% precision and 90% recall. Never infer transaction order,
1005
+ bulk behavior, retries, or runtime payload shape from source alone.
1006
+ 3. Validate that each result carries an exact-versus-heuristic confidence label,
1007
+ stays within C3's response budget, and is produced with no org credentials or
1008
+ network access.
1009
+
1010
+ ---
1011
+
1012
+ ## Competitive leadership program
1013
+
1014
+ The competitive audit is evidence, not a second roadmap. C24-C33 turn every
1015
+ row of [`COMPETITIVE-AUDIT.md`](../benchmarks/competitors/COMPETITIVE-AUDIT.md)
1016
+ into a measurable delivery contract. A row is complete only when knodin meets
1017
+ or exceeds the stated criterion against the pinned competitor and shared
1018
+ fixture. A missing local prerequisite, incompatible license, hosted-service
1019
+ requirement, or resource failure is a reproducible blocker, never a pass.
1020
+
1021
+ All replays must record the knodin commit, competitor version, command,
1022
+ fixture revision, warm/cold mode, requested scope, response budget, and oracle
1023
+ result. Correctness wins require a labeled oracle; latency, response size, or
1024
+ feature presence alone cannot establish a win. A replay that proves a loss
1025
+ must add a follow-on implementation item before the parent may close.
1026
+
1027
+ ### Product-claim gate
1028
+
1029
+ Every roadmap item that changes a user-facing claim must state and verify all
1030
+ four parts of the adoption case:
1031
+
1032
+ 1. **Value:** the named developer or agent workflow and the concrete failure or
1033
+ delay it removes.
1034
+ 2. **ROI:** a labeled correctness, latency, size, safety, or effort measure
1035
+ with its baseline and measurement window; do not substitute a vendor claim
1036
+ or feature count.
1037
+ 3. **Tomorrow:** a local, credential-free adoption path that is compatible with
1038
+ existing Git, editor, and test workflows, including explicit unavailable or
1039
+ degraded states.
1040
+ 4. **Defensibility:** the evidence, integration, or outcome that remains hard
1041
+ to substitute after the underlying model capability becomes commonplace.
1042
+
1043
+ “Uses AI” and an unmeasured capability are never roadmap completion signals.
1044
+ An item without a measurable adoption case remains evaluation-only; a replay
1045
+ that finds a quality gap updates the roadmap and prioritizes the smallest
1046
+ workflow correction before any new surface area is added.
1047
+
1048
+ | Audit row | Leadership item | Win target |
1049
+ |---|---|---|
1050
+ | Exact symbol identity and ambiguity | C25 | Correct selected identity or explicit ambiguity for every labeled duplicate-name case |
1051
+ | Caller/callee/reference accuracy | C25 | ≥95% precision and recall; zero cross-definition false positives |
1052
+ | Diff review and change scope | C26 | Exact changed-file/symbol oracle across each supported scope |
1053
+ | Blast radius and traversal | C26 | Oracle-checked typed edges, depth, filters, and truncation |
1054
+ | Architecture and communities | C27 | Oracle-checked topology plus deterministic bounded drill-downs |
1055
+ | Dead code | C28 | Reported precision/recall and no unlabelled heuristic claim |
1056
+ | Semantic orientation/search | C28 | Frozen relevance judgments with reported nDCG and recall@k |
1057
+ | Code/context packing | C29 | Exact inclusion/range/token oracle under equivalent budgets |
1058
+ | Source editing and diagnostics | C30 | Preview, compile/test, and rollback oracle on disposable fixtures |
1059
+ | Resource/index lifecycle | C31 | Comparable cold/warm/repair/RSS/disk evidence with successful recovery oracle |
1060
+ | API/PDG/specialized analysis | C32 | Precision/recall and bounded source-evidence oracle |
1061
+ | Local privacy and schema economy | C33 | Zero required credentials/egress/hosted service and one top-level MCP tool |
1062
+
1063
+ ### C24 — Competitive correctness contract and replay baseline
1064
+
1065
+ - Status: implemented
1066
+ - Priority: P0
1067
+ - Disposition: Must close
1068
+ - DependsOn: C18
1069
+ - Evidence: the audit records 713 paired cases but explicitly says that most
1070
+ current comparisons lack a labeled correctness oracle. Existing label sets
1071
+ under `benchmarks/evaluations/` are fragmented and are not represented in
1072
+ the normalized competitive-run summary.
1073
+ - Competitors: all locally runnable entries in `COMPETITIVE_MANIFEST`; keep
1074
+ `mcp-codebase-index` blocked while it requires hosted Gemini embeddings.
1075
+ - Touches: `src/competitive-runner.ts`, `src/competitive-manifest.ts`,
1076
+ `scripts/competitive-bakeoff.ts`,
1077
+ `src/__tests__/unit/competitive-runner.spec.ts`,
1078
+ `benchmarks/evaluations/`, `benchmarks/competitors/COMPETITIVE-AUDIT.md`,
1079
+ `benchmarks/competitors/SYNTHESIS.md`, this roadmap
1080
+ - What: normalize competitor cases into a version-pinned, oracle-aware replay
1081
+ contract and reconcile every historical audit claim with the current code.
1082
+
1083
+ ### Acceptance criteria
1084
+
1085
+ 1. The normalized summary separately records correctness-oracle status,
1086
+ exact-match/precision/recall/false-positive results where applicable, and
1087
+ performance/resource measurements; unavailable fields remain explicit nulls.
1088
+ 2. Every audit row has a fixture inventory that either reuses a checked-in
1089
+ labeled fixture or identifies the missing fixture as a blocked prerequisite.
1090
+ 3. Every competitor run records pinned version, command, fixture revision,
1091
+ mode, budgets, and reproducible blockers; a partial run cannot be reported
1092
+ as a competitive win.
1093
+ 4. The audit labels all existing rows `implemented—needs replay`, `unverified`,
1094
+ `demonstrated gap`, or `non-goal` instead of treating historical results as
1095
+ current product state.
1096
+
1097
+ ### C25 — Prove symbol identity and reference precision leadership
1098
+
1099
+ - Status: implemented
1100
+ - Priority: P0
1101
+ - Disposition: Must close
1102
+ - DependsOn: C24, C2, C11
1103
+ - Competitors: GitNexus, Serena, codebase-memory-mcp, CodeGraph
1104
+ - Touches: `benchmarks/evaluations/competitive-symbols/`,
1105
+ `src/__tests__/unit/symbol-identity.spec.ts`,
1106
+ `src/__tests__/unit/reference-precision.spec.ts`,
1107
+ `src/competitive-manifest.ts`, relevant competitor bake-off scripts,
1108
+ `src/engine/index.ts` and `src/tools/knodin-tools.ts` only if replay fails
1109
+ - What: make exact target selection, ambiguity safety, and callers/callees/
1110
+ references objectively comparable on shared duplicate-name and structural
1111
+ implementation fixtures.
1112
+ - Replay evidence:
1113
+ `benchmarks/evaluations/competitive-symbols/raw-results-20260723.json`.
1114
+ knodin achieved 100% precision and recall with zero cross-definition false
1115
+ positives under the shared response budget; every pinned competitor completed
1116
+ against the same fixture and its oracle score is preserved.
1117
+
1118
+ ### Acceptance criteria
1119
+
1120
+ 1. A labeled fixture covers duplicate names, aliases, barrels, imports versus
1121
+ calls, structural implementations, tests, and at least one unresolved case.
1122
+ 2. knodin returns the selected stable identity or an explicit ambiguity result;
1123
+ it never silently conflates viable candidates.
1124
+ 3. knodin meets at least 95% precision and recall for labeled resolved
1125
+ call/reference edges, with zero cross-definition false positives.
1126
+ 4. The pinned competitors and knodin run against the same fixture and budget;
1127
+ a win requires oracle parity or superiority, not a smaller response.
1128
+
1129
+ ### C26 — Prove diff-review and directional traversal leadership
1130
+
1131
+ - Status: implemented
1132
+ - Priority: P1
1133
+ - Disposition: Must close
1134
+ - DependsOn: C24, C4, C5, C13
1135
+ - Competitors: GitNexus, code-review-graph, Graphify, codebase-memory-mcp
1136
+ - Touches: `benchmarks/evaluations/competitive-review/`,
1137
+ `src/__tests__/unit/review-scopes.spec.ts`,
1138
+ `src/__tests__/unit/impact.spec.ts`,
1139
+ `src/__tests__/unit/typed-traversal.spec.ts`, relevant bake-off scripts,
1140
+ engine/tool/CLI code only if replay fails
1141
+ - What: compare staged, unstaged, untracked, and revision-pair review results
1142
+ plus typed directional traversal using source-evidenced expected edges.
1143
+ - Replay evidence:
1144
+ `benchmarks/evaluations/competitive-review/raw-results-20260723.json`.
1145
+ knodin matched every diff-scope, unindexed-file, directional traversal, and
1146
+ edge-evidence oracle under the frozen budget. GitNexus reached traversal
1147
+ parity; Graphify and code-review-graph did not, and the unavailable
1148
+ codebase-memory query adapter remains an explicit blocker rather than a win.
1149
+
1150
+ ### Acceptance criteria
1151
+
1152
+ 1. The fixture labels changed files, mapped symbols, unindexed files, direct
1153
+ and transitive typed edges, depth, and expected truncation.
1154
+ 2. knodin supports every documented diff scope and reports all changed files
1155
+ independently of index coverage.
1156
+ 3. Upstream/downstream/both traversal obeys selected relation, confidence,
1157
+ depth, test, and budget controls with oracle-checked edge evidence.
1158
+ 4. A replay declares success only when knodin meets the oracle and normalized
1159
+ budget; any competitor prerequisite failure is preserved as a blocker.
1160
+
1161
+ ### C27 — Prove architecture and community-analysis leadership
1162
+
1163
+ - Status: implemented
1164
+ - Priority: P1
1165
+ - Disposition: Must close
1166
+ - DependsOn: C24, C3, C12, C13
1167
+ - Competitors: GitNexus, Graphify, codebase-memory-mcp, code-review-graph
1168
+ - Touches: `benchmarks/evaluations/competitive-architecture/`,
1169
+ `src/__tests__/unit/map.spec.ts`, `src/__tests__/unit/uniform-filters.spec.ts`,
1170
+ `src/__tests__/unit/typed-traversal.spec.ts`, relevant bake-off scripts,
1171
+ engine/tool code only if replay fails
1172
+ - What: test boundaries, layers, packages, hubs, bridges, communities,
1173
+ relation filters, pagination, and compact drill-downs against a labeled
1174
+ module-topology fixture.
1175
+ - Replay evidence:
1176
+ `benchmarks/evaluations/competitive-architecture/raw-results-20260723.json`.
1177
+ knodin met the labeled topology oracle, deterministic pagination, and C3
1178
+ budget. Pinned GitNexus, Graphify, and code-review-graph replays completed
1179
+ against the same fixture; codebase-memory remains an explicit policy blocker.
1180
+
1181
+ ### Acceptance criteria
1182
+
1183
+ 1. The fixture labels module boundaries, allowed cross-boundary edges,
1184
+ packages/layers, hub/bridge expectations, and known absent relationships.
1185
+ 2. knodin produces deterministic, filterable, paginated results with explicit
1186
+ totals and C3-compliant byte/token/item budgets.
1187
+ 3. Architecture claims retain exact-versus-heuristic confidence and never turn
1188
+ an inferred grouping into a precise boundary claim.
1189
+ 4. A pinned replay proves oracle quality and budget parity or better.
1190
+
1191
+ ### C28 — Prove dead-code and semantic-search quality
1192
+
1193
+ - Status: implemented
1194
+ - Priority: P1
1195
+ - Disposition: Must close
1196
+ - DependsOn: C24, C3, C11, C12
1197
+ - Competitors: GitNexus, CodeGraph, grepai, Claude Context
1198
+ - Touches: `benchmarks/evaluations/competitive-search/`,
1199
+ `src/__tests__/unit/query.spec.ts`, `src/__tests__/unit/search.spec.ts`,
1200
+ `src/__tests__/unit/response-budget.spec.ts`, relevant bake-off scripts,
1201
+ engine/tool code only if replay fails
1202
+ - What: measure dead-code precision/recall and semantic-orientation relevance
1203
+ using known-live/known-dead and judged-query fixtures.
1204
+ - Replay evidence:
1205
+ `benchmarks/evaluations/competitive-search/raw-results-20260723T214300000Z.json`
1206
+ preserves the frozen local oracle and exact version checks. The pinned
1207
+ GitNexus 1.6.9 adapter normalizes ranked function definitions and its
1208
+ read-only no-incoming-`CALLS` Cypher heuristic to the same `file::symbol`
1209
+ labels. knodin records dead-code precision/recall of 1.0/1.0 versus
1210
+ GitNexus's 0.667/1.0, and mean semantic nDCG/recall@3 of 0.855/1.0 versus
1211
+ 0.677/1.0. Both stay within the 32 KiB response budget with no source egress.
1212
+ The remaining CodeGraph, grepai, and Claude Context capability blockers are
1213
+ retained honestly but no longer prevent a comparable pinned competitor replay.
1214
+
1215
+ ### Acceptance criteria
1216
+
1217
+ 1. Dead-code labels cover exports, imports, reflection-like unresolved uses,
1218
+ tests, barrels, and known live/dead symbols; results report precision,
1219
+ recall, and false positives.
1220
+ 2. Search has frozen relevance judgments and reports nDCG and recall@k under a
1221
+ fixed source and response budget.
1222
+ 3. knodin labels heuristic dead-code or semantic evidence honestly and offers a
1223
+ bounded drill-down for excluded evidence.
1224
+ 4. The replay establishes a win only with oracle parity or better and no
1225
+ regression to C3 budgets.
1226
+
1227
+ ### C29 — Prove context-packing and export leadership
1228
+
1229
+ - Status: implemented
1230
+ - Priority: P1
1231
+ - Disposition: Must close
1232
+ - DependsOn: C24, C3, C5, C14
1233
+ - Competitors: Repomix, Aider repo-map, code2prompt
1234
+ - Touches: `benchmarks/evaluations/competitive-context/`,
1235
+ `src/__tests__/unit/context-export.spec.ts`,
1236
+ `src/__tests__/unit/response-budget.spec.ts`, relevant bake-off scripts,
1237
+ `src/context-export.ts` and tool/CLI code only if replay fails
1238
+ - What: compare deterministic packing, exact token accounting, file policy,
1239
+ incremental/ranged retrieval, and optional Git context on a shared corpus.
1240
+ - Replay evidence:
1241
+ `benchmarks/evaluations/competitive-context/raw-results-20260723.json`.
1242
+ knodin met the shared inclusion, exclusion, range, Git, token, and byte
1243
+ oracle for deterministic Markdown, JSON, and XML exports. Repomix and
1244
+ code2prompt completed the bounded corpus replay with their unavailable
1245
+ controls recorded; Aider is explicitly blocked because it has no comparable
1246
+ artifact surface.
1247
+
1248
+ ### Acceptance criteria
1249
+
1250
+ 1. The fixture labels required/excluded source, tree, line ranges, Git changes,
1251
+ and expected token/byte bounds for each request.
1252
+ 2. knodin produces deterministic Markdown, JSON, and XML exports with relative
1253
+ paths and a declared tokenizer estimate.
1254
+ 3. Include/exclude and per-file policies compose with hard budgets; ranged
1255
+ reads and exact regex grep return verifiable source positions.
1256
+ 4. The competitor replay compares equivalent scope and budget and preserves
1257
+ any unavailable command or format as an explicit blocker.
1258
+
1259
+ ### C30 — Prove guarded editing and diagnostics leadership
1260
+
1261
+ - Status: implemented (evaluation found no production-surface gap)
1262
+ - Priority: P2
1263
+ - Disposition: Evaluate
1264
+ - DependsOn: C24, C2, C11, C15
1265
+ - Competitors: Serena, code-review-graph
1266
+ - Touches: `benchmarks/evaluations/competitive-editing/`,
1267
+ `src/__tests__/unit/lsp-readonly.spec.ts`,
1268
+ `src/__tests__/unit/rename-apply.spec.ts`, relevant bake-off scripts,
1269
+ `src/lsp-readonly.ts` and editing surfaces only if evaluation approves them
1270
+ - What: establish whether additional local, guarded edit previews deliver
1271
+ enough value beyond compiler-verified rename and read-only TypeScript
1272
+ language-service queries.
1273
+ - Replay evidence:
1274
+ `benchmarks/evaluations/competitive-editing/raw-results-20260723.json`.
1275
+ The disposable fixture verified diagnostic and navigation goldens,
1276
+ preview-first compiler/test validation, rejected-edit rollback, and explicit
1277
+ unavailable competitor adapters. It found no actionable gap requiring a new
1278
+ production editing surface.
1279
+
1280
+ ### Acceptance criteria
1281
+
1282
+ 1. Fixtures include diagnostic goldens, definition/implementation navigation,
1283
+ valid edits, rejected edits, rollback, and compile/test verification.
1284
+ 2. Any proposed mutation is preview-first, bounded to a disposable copy, and
1285
+ produces a compiler/test result before it can be judged successful.
1286
+ 3. No daemon, credential, network access, or unguarded workspace mutation is
1287
+ required; unsupported language adapters return explicit unavailable results.
1288
+ 4. Promote an implementation only if the competitor comparison finds an
1289
+ actionable, locally deliverable gap that rename and read-only queries cannot
1290
+ cover.
1291
+
1292
+ ### C31 — Prove index lifecycle and output-telemetry leadership
1293
+
1294
+ - Status: done
1295
+ - Priority: P1
1296
+ - Disposition: Must close
1297
+ - DependsOn: C24, C3, C16, C18
1298
+ - Competitors: codebase-memory-mcp, grepai, Claude Context
1299
+ - Touches: `benchmarks/evaluations/competitive-lifecycle/`,
1300
+ `src/__tests__/unit/index-health.spec.ts`,
1301
+ `src/__tests__/unit/perf.spec.ts`, `src/competitive-runner.ts`, relevant
1302
+ bake-off scripts, engine/tool/CLI code only if replay fails
1303
+ - What: compare cold index, incremental index, corruption repair, stale-index
1304
+ recovery, output telemetry, peak RSS, disk, and failure behavior locally.
1305
+ - Replay evidence:
1306
+ `benchmarks/evaluations/competitive-lifecycle/raw-results-20260723T220325Z.json`
1307
+ records disposable four-phase replays for pinned codebase-memory-mcp and
1308
+ grepai with verified versions and behavioral lifecycle oracles. Both expose
1309
+ stale indexed state after a unique symbol replacement, recover through their
1310
+ documented refresh/reindex mechanism, make the replacement queryable, remove
1311
+ the old symbol, and preserve the healthy symbol. Native telemetry inspection
1312
+ deterministically establishes the remaining capability gaps:
1313
+ codebase-memory reports graph/index counts and repository state but lacks
1314
+ detailed per-query output telemetry; grepai reports aggregate query/token
1315
+ savings but lacks per-query latency, truncation, and detail mode. knodin
1316
+ supplies those opt-in local fields. Claude Context remains a separately
1317
+ classified Ollama/Milvus setup-resource blocker rather than an inferred win.
1318
+
1319
+ ### Acceptance criteria
1320
+
1321
+ 1. The fixture records cold and warm timing, peak RSS, database/disk growth,
1322
+ failure injection, repair outcome, and telemetry fields without source
1323
+ content.
1324
+ 2. knodin detects stale/damaged state, provides an actionable repair, and
1325
+ verifies the repaired result without deleting healthy state.
1326
+ 3. Persisted telemetry is opt-in, local, and includes operation, latency,
1327
+ serialized bytes, estimated tokens, truncation, and detail mode.
1328
+ 4. A replay compares equivalent lifecycle phases; setup or resource failures
1329
+ remain distinct from correctness and performance outcomes.
1330
+
1331
+ ### C32 — Prove local API and statement-flow analysis leadership
1332
+
1333
+ - Status: implemented (evaluation found no production-surface gap)
1334
+ - Priority: P2
1335
+ - Disposition: Evaluate
1336
+ - DependsOn: C24, C2, C3, C8, C9
1337
+ - Competitors: GitNexus
1338
+ - Touches: `benchmarks/evaluations/competitive-api-flow/`,
1339
+ `src/__tests__/unit/flow-analysis.spec.ts`,
1340
+ `src/__tests__/unit/api-contracts.spec.ts`, relevant bake-off scripts,
1341
+ engine/tool code only if evaluation proves a gap
1342
+ - What: replay bounded API contract mismatch and statement-level flow results
1343
+ against a labeled fixture, retaining knodin's compact local-only contract.
1344
+ - Replay evidence:
1345
+ `benchmarks/evaluations/competitive-api-flow/raw-results-20260723.json`.
1346
+ The labeled local replay produced 1.00 precision and recall with zero false
1347
+ positives for both bounded API mismatches and variable-filtered statement
1348
+ flow. It verifies source evidence, an explicit C3 continuation at the flow
1349
+ budget, and the one-gateway/local-only constraints. No broader persisted PDG
1350
+ or API surface is justified by this evaluation.
1351
+
1352
+ ### Acceptance criteria
1353
+
1354
+ 1. The fixture labels true API mismatches, true control/data-flow rows,
1355
+ variable filtering, unresolved cases, and expected source evidence.
1356
+ 2. knodin reports precision, recall, false positives, unresolved cases,
1357
+ truncation, and continuation routes under C3 budgets.
1358
+ 3. The comparison never treats broad persisted PDG output, raw Cypher, or
1359
+ hosted analysis as required parity; it evaluates only proven local workflows.
1360
+ 4. Promote broader implementation only when it finds actionable breakage that
1361
+ existing bounded API/flow queries miss.
1362
+
1363
+ ### C33 — Preserve local privacy and one-tool schema economy
1364
+
1365
+ - Status: implemented
1366
+ - Priority: P0
1367
+ - Disposition: Must close
1368
+ - DependsOn: C24
1369
+ - Competitors: all; `mcp-codebase-index` and Claude Context supply negative
1370
+ evidence where hosted services or local service dependencies are required.
1371
+ - Touches: `src/tools/knodin-tools.ts`, `src/server.ts`,
1372
+ `src/competitive-manifest.ts`, `src/__tests__/unit/mcp-tools.spec.ts`,
1373
+ `benchmarks/evaluations/competitive-constraints/`, this roadmap
1374
+ - What: make the local/no-auth/no-egress and one-gateway constraints measurable
1375
+ release gates for every competitive claim.
1376
+
1377
+ ### Acceptance criteria
1378
+
1379
+ 1. A clean local replay requires no credentials, hosted service, network
1380
+ embeddings, source-code egress, or mandatory external daemon.
1381
+ 2. The MCP surface remains one `knodin` gateway; a capability addition must
1382
+ document schema-token cost and cannot create a top-level tool.
1383
+ 3. Any optional local service is reported as optional and unavailable states
1384
+ are explicit rather than silently degraded success.
1385
+ 4. No item C24-C32 can close unless its replay records this constraint check.
1386
+
1387
+ ### C34 — Keep local graph indexes current across repository lifecycle events
1388
+
1389
+ - Status: implemented
1390
+ - Priority: P1
1391
+ - Disposition: Must close
1392
+ - DependsOn: C16, C24, C33
1393
+ - Competitors: GitNexus, Graphify, codebase-memory-mcp
1394
+ - Touches: `bin/cli.ts`, `src/engine/index.ts`, `lefthook.yml`,
1395
+ `templates/hooks/`, `src/__tests__/unit/freshness.spec.ts`,
1396
+ `src/__tests__/unit/watcher-lifecycle.spec.ts`,
1397
+ `benchmarks/evaluations/competitive-lifecycle/`, docs
1398
+ - What: make fresh, reconciled, stale, and unavailable graph states explicit
1399
+ across file edits, branch switches, pulls, rebases, merges, and a restarted
1400
+ local process; provide safe, local refresh commands for every supported graph
1401
+ artifact without pretending an external index is current.
1402
+ - Evidence: `knodin refresh-artifacts [checkout|merge|code-change]` is an
1403
+ explicit opt-in, 30-second-bounded refresh for locally installed GitNexus and
1404
+ Graphify. It prefers the repository's GitNexus runner when present, records
1405
+ success, failure, or skipped state without source content, and is never wired
1406
+ into a commit hook. The lifecycle fixtures prove restart, replacement,
1407
+ branch-switch, merge, watcher-loss, and external-tool states; the replay
1408
+ records freshness latency, RSS, disk, and verified state.
1409
+
1410
+ ### Acceptance criteria
1411
+
1412
+ 1. knodin continues to reconcile its persisted index at cold start and before
1413
+ answers after offline drift, reporting `fresh`, `reconciled`, or `unknown`.
1414
+ 2. The repository exposes a documented local command or opt-in hook path to
1415
+ refresh GitNexus after checkout/merge and Graphify after code changes; it
1416
+ records success, failure, or skipped state and never blocks a commit on an
1417
+ unbounded rebuild.
1418
+ 3. A fixture proves correct behavior for edit, deletion, branch switch, merge,
1419
+ pull/rebase-equivalent replacement, watcher loss, and restart; no answer may
1420
+ silently claim an unverified index is fresh.
1421
+ 4. The lifecycle replay reports freshness latency and resource cost separately
1422
+ from query quality and preserves any unavailable external tool as a blocker.
1423
+
1424
+ ### C35 — Produce local interactive architecture and call-flow visualizations
1425
+
1426
+ - Status: implemented
1427
+ - Priority: P2
1428
+ - Disposition: Should close
1429
+ - DependsOn: C3, C13, C24, C33
1430
+ - Competitors: Graphify, GitNexus
1431
+ - Touches: `bin/cli.ts`, `src/visualization.ts`,
1432
+ `src/__tests__/unit/visualization.spec.ts`,
1433
+ `benchmarks/evaluations/competitive-visualization/`, docs
1434
+ - What: evaluate and, if the evidence justifies it, generate deterministic
1435
+ local HTML architecture and call-flow artifacts from knodin's bounded graph
1436
+ facts without adding a new MCP tool, hosted renderer, source egress, or a
1437
+ mandatory long-lived service.
1438
+ - Evaluation evidence:
1439
+ `benchmarks/evaluations/competitive-visualization/raw-results-20260723.json`
1440
+ records Graphify's pinned community-graph, tree, and call-flow artifacts and
1441
+ GitNexus's explicit lack of a deterministic static-artifact command. knodin's
1442
+ checked-in evaluation artifact is self-contained, deterministic, bounded,
1443
+ source-commit/index-freshness tagged, and source-evidenced. `knodin visualize
1444
+ <entry> --output <path.html>` now exposes that bounded artifact as a local
1445
+ CLI-only export with 1–6 depth, a 4–64 KiB generation ceiling, safe in-repo
1446
+ output paths, and stable identity/file/kind entry disambiguation. It adds no
1447
+ MCP operation, hosted renderer, credential, telemetry, or source egress.
1448
+
1449
+ ### Acceptance criteria
1450
+
1451
+ 1. The evaluation compares Graphify's community graph, tree, and call-flow
1452
+ artifacts with a knodin artifact generated from the same pinned fixture.
1453
+ 2. Generated HTML is self-contained or uses only local static assets, records
1454
+ its source commit/index freshness, and links every displayed edge to bounded
1455
+ source evidence or an explicit heuristic label.
1456
+ 3. The artifact supports at least subsystem drill-down, hub/bridge inspection,
1457
+ and bounded call-flow navigation while respecting C3 output and generation
1458
+ budgets.
1459
+ 4. It remains an optional CLI/export capability behind the existing `knodin`
1460
+ gateway: no new top-level MCP tool, network service, credential, or source
1461
+ egress is introduced.
1462
+
1463
+ ### C36 — Instrument the current warm-operation performance paths
1464
+
1465
+ - Status: implemented
1466
+ - Priority: P0
1467
+ - Disposition: Must close
1468
+ - DependsOn: C18, C24
1469
+ - Evidence: opt-in, source-free phase accounting now attributes freshness,
1470
+ graph analytics, traversal snapshots, architecture facets, context
1471
+ composition, status audit, pack walking/serialization, and artifact
1472
+ read/grep. `perf.spec.ts` verifies the complete phase record without latency
1473
+ thresholds.
1474
+ - Touches: `src/engine/perf.ts`, `src/engine/index.ts`, `src/context.ts`,
1475
+ `src/context-export.ts`, targeted unit tests
1476
+ - What: add opt-in phase accounting around the current production paths without
1477
+ changing response contracts or enabling a competitor replay.
1478
+
1479
+ ### Acceptance criteria
1480
+
1481
+ 1. Phase names separately cover freshness, graph analytics, traversal snapshot,
1482
+ architecture facets, context composition, status audit, pack walk,
1483
+ serialization, and artifact read/grep.
1484
+ 2. Instrumentation remains inert unless a performance session is active and
1485
+ records no source content.
1486
+ 3. Targeted tests verify phase accounting without wall-clock thresholds.
1487
+ 4. No competitor or before/after comparison is run until C37-C40 are complete
1488
+ and the repository owner explicitly approves the replay.
1489
+
1490
+ ### C37 — Reuse generation-scoped graph analytics and traversal snapshots
1491
+
1492
+ - Status: implemented
1493
+ - Priority: P0
1494
+ - Disposition: Must close
1495
+ - DependsOn: C36
1496
+ - Evidence: minimal and standard map now share one generation-keyed analytics
1497
+ snapshot; traversal uses a compact signed adjacency plus cached definition
1498
+ metadata; architecture membership/cohesion/coupling aggregation is one-pass.
1499
+ Option-keyed analytics/minimal-map caches and per-repository traversal/status
1500
+ caches have explicit entry ceilings rather than unbounded process growth.
1501
+ `typed-traversal.spec.ts` proves both map and traversal snapshots invalidate
1502
+ after an incremental index generation.
1503
+ - Touches: `src/engine/index.ts`, `src/__tests__/unit/map.spec.ts`,
1504
+ `src/__tests__/unit/typed-traversal.spec.ts`,
1505
+ `src/__tests__/unit/uniform-filters.spec.ts`
1506
+ - What: cache immutable graph analytics, bidirectional typed adjacency, and
1507
+ symbol metadata by repository/federation configuration plus index generation.
1508
+
1509
+ ### Acceptance criteria
1510
+
1511
+ 1. Minimal and standard map modes share one generation-scoped analytics source
1512
+ while retaining their existing response shapes and truthful totals.
1513
+ 2. Traversal performs bounded BFS over cached adjacency with no per-result
1514
+ symbol lookup and preserves identity, direction, filters, evidence, and
1515
+ deterministic ordering.
1516
+ 3. Architecture facets and coupling use one-pass identity-to-community indexes
1517
+ rather than repeated community-by-edge searches.
1518
+ 4. Index, reconcile, watcher flush, repair, and close invalidate every affected
1519
+ cache; tests prove no stale result survives a generation change.
1520
+
1521
+ ### C38 — Compose orientation context from one shared snapshot
1522
+
1523
+ - Status: implemented
1524
+ - Priority: P1
1525
+ - Disposition: Should close
1526
+ - DependsOn: C37
1527
+ - Evidence: `buildKnodinContext` establishes the standard map snapshot once,
1528
+ then composes stats, flows, and optional review over that initialized
1529
+ generation. `context-composition.spec.ts` verifies at most one probe and the
1530
+ existing compact caps/contracts without timing assertions.
1531
+ - Touches: `src/context.ts`, `src/engine/index.ts`,
1532
+ `src/__tests__/unit/knodin-tools.spec.ts`,
1533
+ `src/__tests__/unit/freshness.spec.ts`
1534
+ - What: acquire one freshness-qualified engine snapshot, reuse C37 analytics,
1535
+ and compose independent context fields without repeating graph construction.
1536
+
1537
+ ### Acceptance criteria
1538
+
1539
+ 1. One context request performs at most one freshness probe and one graph
1540
+ analytics construction for its repository generation.
1541
+ 2. Stats, communities, hubs, flows, and optional risk retain their current
1542
+ contract, sorting, caps, and evidence.
1543
+ 3. Independent work may run concurrently only after initialization is complete;
1544
+ no SQLite writer concurrency or duplicate engine singleton is introduced.
1545
+ 4. Tests use counters and cache state, not timing thresholds.
1546
+
1547
+ ### C39 — Separate fast truthful status from explicit deep audit
1548
+
1549
+ - Status: implemented
1550
+ - Priority: P1
1551
+ - Disposition: Should close
1552
+ - DependsOn: C36
1553
+ - Evidence: warm status reuses a generation/database-fingerprint keyed deep
1554
+ audit only after the existing bounded drift probe verifies the source state.
1555
+ Responses identify `deep-audit` versus
1556
+ `cached-after-freshness-probe` and include `verifiedAt`; `--deep` and MCP
1557
+ `statusAudit: deep` force the complete audit. `index-health.spec.ts` proves
1558
+ warm reuse, explicit deep mode, and file-change invalidation.
1559
+ - Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
1560
+ `src/__tests__/unit/index-health.spec.ts`
1561
+ - What: return an invalidation-safe cached health snapshot for warm status and
1562
+ expose the existing filesystem/database verification as an explicit deep
1563
+ audit, with verification age and mode reported truthfully.
1564
+
1565
+ ### Acceptance criteria
1566
+
1567
+ 1. Default warm status never claims a fresh audit when it is returning cached
1568
+ evidence; it reports verification mode and timestamp.
1569
+ 2. Deep audit preserves current missing, damaged, orphaned, coverage, and repair
1570
+ behavior.
1571
+ 3. Index generation, watcher changes, repair, schema change, and relevant file
1572
+ events invalidate the cached snapshot.
1573
+ 4. Existing callers remain source-compatible and a caller can explicitly
1574
+ request the deep audit.
1575
+
1576
+ ### C40 — Make context packing and artifact access linear and cache-safe
1577
+
1578
+ - Status: implemented
1579
+ - Priority: P1
1580
+ - Disposition: Should close
1581
+ - DependsOn: C36
1582
+ - Evidence: glob/policy matchers are bounded and compiled once, walking uses one
1583
+ accumulator, exact per-file serialization deltas replace accumulated
1584
+ reserialization, and artifact lines use a two-entry/four-MiB-per-file
1585
+ path/mtime/size cache. Context-export and response-budget tests preserve
1586
+ formats, policies, ordering, hard budgets, safe paths, reads, grep, and cache
1587
+ invalidation.
1588
+ - Touches: `src/context-export.ts`,
1589
+ `src/__tests__/unit/context-export.spec.ts`,
1590
+ `src/__tests__/unit/response-budget.spec.ts`
1591
+ - What: precompile selection policy, prune excluded directories, account bytes
1592
+ incrementally, serialize once, and reuse a bounded artifact line index keyed
1593
+ by path/mtime/size.
1594
+
1595
+ ### Acceptance criteria
1596
+
1597
+ 1. Candidate walking and pattern matching are linear in visited paths plus
1598
+ patterns; patterns are compiled once per request.
1599
+ 2. Packing does not rebuild the entire accumulated artifact per candidate and
1600
+ still enforces exact response byte/token ceilings.
1601
+ 3. Artifact read/grep reuse a bounded cache only when path, size, and mtime
1602
+ match; mutation and replacement invalidate it.
1603
+ 4. Markdown, JSON, XML, policies, tree/Git sections, deterministic ordering,
1604
+ ranged reads, regex grep, and telemetry remain byte-for-byte compatible in
1605
+ golden fixtures.
1606
+
1607
+ ### C41 — Replay like-for-like performance after the implementation batch
1608
+
1609
+ - Status: evaluated — deferred
1610
+ - Priority: P0
1611
+ - Disposition: Evaluate
1612
+ - DependsOn: C37, C38, C39, C40
1613
+ - Evidence: the terminal bounded replay and strengthened-oracle verdict are
1614
+ preserved at
1615
+ `benchmarks/competitors/runs/20260802T203822521Z-3e27547d095e/`. Its
1616
+ `c41-verdict.json` verifies all five competitors, explicit warm/cold p50/p95,
1617
+ independently attributed two-sided RSS, and grepai cached/deep lifecycle
1618
+ rows. Under the strengthened labels, 19 mapped rows passed, 10 failed, four
1619
+ grepai trace rows were unavailable, and six intentionally unmatched Repomix
1620
+ rows remained incomparable. Graphify passed its direct-call rows but not hub
1621
+ rows; codebase-memory passed six of ten search/traversal/package rows;
1622
+ Repomix passed seven pack/read/grep rows; code2prompt passed both pack rows;
1623
+ and grepai produced no oracle-qualified latency row because its local index
1624
+ remained empty. The bounded local-Ollama indexing attempts at
1625
+ `20260802T205026646Z-18c1107af915/`,
1626
+ `20260802T205625288Z-e97db4c32eb6/` preserve that setup limitation. Earlier
1627
+ runs retain the sandbox-launch, timeout, and uninitialized-index failures
1628
+ rather than overwriting them.
1629
+ - Evaluation disposition: retain the current product unchanged and defer any
1630
+ aggregate performance claim or optimization response. Only oracle-passing
1631
+ exact rows may support operation-specific observations; failures,
1632
+ unavailable rows, and setup limits remain separate dimensions and cannot
1633
+ become wins.
1634
+ - Touches if approved: competitor harness mappings, immutable run artifacts,
1635
+ `benchmarks/competitors/COMPETITIVE-AUDIT.md`, this roadmap
1636
+ - What: after explicit repository-owner permission, run warm/cold p50/p95/RSS
1637
+ measurements using equivalent operations, budgets, fixtures, and correctness
1638
+ oracles. A user instruction to complete this roadmap constitutes permission
1639
+ for the checked-in, zero-spend, local replay only; network access, installation
1640
+ of unpinned software, publication, or external mutation still requires its own
1641
+ authority.
1642
+
1643
+ ### Evaluation gate
1644
+
1645
+ 1. Graphify traversal/map rows use the same direction, relation, depth, item,
1646
+ token, and evidence scope.
1647
+ 2. codebase-memory architecture/search rows use the same facets, pagination,
1648
+ source inclusion, and correctness labels.
1649
+ 3. code2prompt and Repomix compare with knodin `pack`, packed read, and packed
1650
+ grep—not semantic `context`, `search`, or `file_summary`.
1651
+ 4. grepai status compares cached summary with cached summary and deep audit with
1652
+ an equivalent verified lifecycle phase.
1653
+ 5. Results record correctness before latency and treat RSS/setup failures as
1654
+ separate dimensions.
1655
+
1656
+ ### C42 — Enforce comparable competitive cases and correctness oracles
1657
+
1658
+ - Status: implemented
1659
+ - Priority: P0
1660
+ - DependsOn: C41
1661
+ - Evidence: `src/competitive-contract.ts` and the competitor mappings enforce
1662
+ typed equivalence/oracle dispositions; the Repomix replay qualifies 18
1663
+ comparable rows and excludes wrong-fast, partial, unverified, and
1664
+ incomparable rows from wins.
1665
+ - Touches: `src/competitive-manifest.ts`, `src/competitive-runner.ts`,
1666
+ `scripts/competitive-bakeoff.ts`, competitor adapters, labeled competitive
1667
+ fixtures and runner tests
1668
+ - What: define one typed case contract for equivalent operation scope, budgets,
1669
+ fixtures, and correctness. Incomparable or wrong-fast rows must never count as
1670
+ latency wins.
1671
+
1672
+ ### Acceptance criteria
1673
+
1674
+ 1. Every timed pair records use case, fixture revision, mode, scope, budgets,
1675
+ and a comparable/incomparable disposition with a reason.
1676
+ 2. Graphify, codebase-memory, Repomix, code2prompt, and grepai mappings satisfy
1677
+ C41's exact equivalence rules; unmatched rows are excluded explicitly.
1678
+ 3. Comparable rows carry checked-in correctness oracles, and summaries
1679
+ distinguish oracle pass, oracle failure, unavailable, and incomparable.
1680
+ 4. Partial runs and rows without passing oracles cannot produce a competitive
1681
+ win or aggregate performance claim.
1682
+
1683
+ ### C43 — Measure cold/warm p50, p95, and independent process RSS
1684
+
1685
+ - Status: implemented
1686
+ - Priority: P0
1687
+ - DependsOn: C42
1688
+ - Evidence: `benchmarks/competitors/MEASUREMENT-CONTRACT.md` and shared sampling
1689
+ record three warmups, at least 20 warm samples, raw values, p50/p95/range,
1690
+ independently attributed RSS, descendant cleanup, and atomic partial
1691
+ summaries across all 11 adapters.
1692
+ - Touches: shared competitive measurement helper, competitor adapters,
1693
+ `scripts/competitive-bakeoff.ts`, lifecycle runner, normalized summaries and
1694
+ tests
1695
+ - What: replace copied three-sample loops and aggregate process-tree ceilings
1696
+ with a shared statistical and resource-measurement contract.
1697
+
1698
+ ### Acceptance criteria
1699
+
1700
+ 1. Warm and cold samples are separate and retain raw samples, warm-up count,
1701
+ p50, p95, minimum, and maximum; the documented warm sample minimum supports a
1702
+ meaningful p95.
1703
+ 2. knodin and competitor RSS are attributed independently by process and phase;
1704
+ combined harness RSS is never assigned to either product.
1705
+ 3. Per-call and per-competitor timeouts reap descendants, preserve partial
1706
+ evidence, identify the failed phase/case, and still write a terminal summary.
1707
+ 4. Regression checks compare matching modes and enforce both p50 and p95
1708
+ tolerances after correctness passes.
1709
+
1710
+ ### C44 — Restore low-latency diff review without weakening evidence
1711
+
1712
+ - Status: implemented
1713
+ - Priority: P0
1714
+ - DependsOn: C43
1715
+ - Evidence: `benchmarks/evaluations/competitive-review/` records an
1716
+ oracle-qualified bounded-gateway replay at 53.018 ms p50, 89.449 ms p95, and
1717
+ 733 bytes; review tests preserve path-scoped churn/risk, diff scopes, flows,
1718
+ cache invalidation, and immutable evidence provenance.
1719
+ - Touches: `src/engine/index.ts`, `src/engine/perf.ts`, review/flow tests and
1720
+ competitive-review evidence
1721
+ - What: remove per-file Git history processes and whole-repository repeated work
1722
+ from warm review while preserving changed-file, risk, flow, and source
1723
+ evidence.
1724
+
1725
+ ### Acceptance criteria
1726
+
1727
+ 1. Review phase telemetry separates diff discovery, churn, parsing, flow lookup,
1728
+ and serialization.
1729
+ 2. Churn/history is collected in one bounded pass or a repository/HEAD cache,
1730
+ and affected flows use a generation-scoped file-to-flow lookup.
1731
+ 3. All diff scopes and existing correctness fixtures remain unchanged.
1732
+ 4. The oracle-qualified warm minimal review meets the recorded C43 regression
1733
+ threshold on both p50 and p95.
1734
+
1735
+ ### C45 — Add facet-selective architecture fast paths
1736
+
1737
+ - Status: implemented
1738
+ - Priority: P1
1739
+ - DependsOn: C43
1740
+ - Evidence: `benchmarks/evaluations/competitive-architecture/` records an
1741
+ oracle-qualified bounded-gateway replay at 0.082 ms p50 and 0.097 ms p95;
1742
+ architecture tests prove facet-selective construction, zero minimal edge
1743
+ materialization, mutation isolation, budgets, and generation/close
1744
+ invalidation.
1745
+ - Touches: `src/engine/index.ts`, architecture/map tests and
1746
+ competitive-architecture evidence
1747
+ - What: compute only requested architecture facets from reusable
1748
+ generation-scoped metadata instead of building a complete standard map.
1749
+
1750
+ ### Acceptance criteria
1751
+
1752
+ 1. Language, package, layer, entry-point, boundary, coupling, and hotspot
1753
+ requests construct only their required inputs.
1754
+ 2. Minimal facet requests do not build full edge lists or hydrate unrelated
1755
+ source/symbol detail.
1756
+ 3. Identity, filters, pagination, totals, budgets, and cache invalidation remain
1757
+ correct.
1758
+ 4. Equivalent architecture rows pass their oracle and C43 thresholds.
1759
+
1760
+ ### C46 — Cache search state and hydrate only the requested page
1761
+
1762
+ - Status: implemented
1763
+ - Priority: P1
1764
+ - DependsOn: C43
1765
+ - Evidence: `benchmarks/evaluations/competitive-search/` records a
1766
+ same-oracle 500-symbol offset-page improvement from 5.309/5.865 ms to
1767
+ 0.310/0.481 ms p50/p95; tests preserve exact-default search, global ordering,
1768
+ totals, pagination, filters, ANN gating, federation, corruption, and budgets.
1769
+ - Touches: `src/engine/index.ts`, search/tool contracts, search and ANN tests,
1770
+ competitive-search evidence
1771
+ - What: reuse decoded generation-scoped search metadata, rank before hydration,
1772
+ and perform source/caller work only for the requested page.
1773
+
1774
+ ### Acceptance criteria
1775
+
1776
+ 1. Cheap filters run before vector scoring and source hydration; community
1777
+ lookup does not build a full map.
1778
+ 2. Exact search remains the default unless the existing ANN recall gate passes.
1779
+ 3. Global ordering, offset pagination, totals, corruption handling, federation,
1780
+ and response budgets remain deterministic.
1781
+ 4. A separately specified opaque cursor may be added only if it binds query,
1782
+ filters, order, and generation and truthfully rejects stale cursors.
1783
+
1784
+ ### C47 — Reduce status, repair, and measured lifecycle resource cost
1785
+
1786
+ - Status: implemented
1787
+ - Priority: P1
1788
+ - DependsOn: C43
1789
+ - Evidence: `benchmarks/evaluations/competitive-lifecycle/` records isolated,
1790
+ ownership-labeled lifecycle stages and immutable provenance; health tests
1791
+ prove lease expiry guards, single-flight deduplicated batch repair, one
1792
+ invalidation, and final forced deep verification.
1793
+ - Touches: `src/engine/index.ts`, `src/engine/perf.ts`, lifecycle runner,
1794
+ freshness/index-health/watcher tests
1795
+ - What: eliminate duplicate freshness/audit work, batch repair indexing, and
1796
+ reduce only resource retention demonstrated by isolated C43 measurements.
1797
+
1798
+ ### Acceptance criteria
1799
+
1800
+ 1. Cached status reuses a valid freshness lease and deep-audit snapshot without
1801
+ claiming unverified freshness; lease expiry still performs the missed-watcher
1802
+ guard.
1803
+ 2. Repair reuses valid audit evidence, deduplicates paths, batches indexing and
1804
+ orphan cleanup, invalidates once, and completes one final deep verification.
1805
+ 3. The lifecycle runner records isolated baseline, post-index, post-query,
1806
+ post-close, peak RSS, and disk measurements for each product.
1807
+ 4. Cache/parser changes are accepted only when attribution proves a retained
1808
+ owner and lifecycle correctness remains intact.
1809
+
1810
+ ### C48 — Decide an optional local project-memory boundary
1811
+
1812
+ - Status: evaluated — deferred
1813
+ - Priority: P2
1814
+ - DependsOn: C42
1815
+ - Decision: defer production work. The evaluated workflows do not yet
1816
+ demonstrate value beyond repository-owned documentation, while a memory
1817
+ surface would add a second source of truth, retention/privacy obligations,
1818
+ and an output-budget surface. See
1819
+ [`docs/adr/004-optional-project-memory.md`](../docs/adr/004-optional-project-memory.md)
1820
+ and the
1821
+ [C48 evaluation](../benchmarks/evaluations/c48-project-memory/evaluation.md).
1822
+ - Touches: roadmap/ADR and Serena memory evaluation evidence; no production
1823
+ code because the evaluation gate did not pass
1824
+ - What: evaluate concrete durable-note workflows without turning generalized
1825
+ agent memory into a core graph/index requirement.
1826
+
1827
+ ### Evaluation gate
1828
+
1829
+ 1. Record user workflows, path containment, retention/backup, privacy, output
1830
+ budgets, and RSS/disk cost against Serena's memory operations.
1831
+ 2. Prefer an optional local `.reckon/notes/` adapter that does not affect graph
1832
+ identity, indexing, freshness, or the one-tool gateway count.
1833
+ 3. Approve production work only when a labeled workflow demonstrates value
1834
+ beyond ordinary repository documentation.
1835
+ 4. If approved, require atomic writes, traversal protection, deterministic
1836
+ list/read/edit/rename/delete behavior, and explicit size budgets.
1837
+
1838
+ ### C49 — Keep immutable runs lint-safe and make the full suite terminate
1839
+
1840
+ - Status: implemented
1841
+ - Priority: P1
1842
+ - DependsOn: —
1843
+ - Evidence: Biome excludes only immutable `benchmarks/competitors/runs/**`;
1844
+ handle diagnostics and teardown coverage close watchers, databases, queues,
1845
+ and child processes. The full coverage suite terminates naturally within the
1846
+ documented six-minute ceiling (474/474 in the final C47 gate).
1847
+ - Touches: `biome.json`, `package.json`, Vitest configuration/setup,
1848
+ watcher/freshness teardown tests, competitive-run documentation
1849
+ - What: exclude immutable generated evidence from formatting gates while
1850
+ diagnosing and fixing the open handle that prevents the full suite from
1851
+ exiting.
1852
+
1853
+ ### Acceptance criteria
1854
+
1855
+ 1. Biome ignores `benchmarks/competitors/runs/**` but continues checking
1856
+ generators, baselines, labels, schemas, and source.
1857
+ 2. A hanging-process diagnostic command identifies active handles without
1858
+ masking them through force-exit.
1859
+ 3. Watchers, databases, queues, and spawned test children are asserted closed in
1860
+ teardown, including injected-failure paths.
1861
+ 4. `bun run lint` and the full test suite terminate naturally within the
1862
+ repository's documented ceiling.
1863
+
1864
+ ### C50 — Make the single-worker full-suite gate start reliably
1865
+
1866
+ - Status: implemented
1867
+ - Priority: P0
1868
+ - DependsOn: C49
1869
+ - Touches: Vitest configuration/setup, embedding test doubles and teardown,
1870
+ gate documentation, and a worker-start regression fixture
1871
+ - What: eliminate repeated native ONNX initialization or worker churn that can
1872
+ exhaust Vitest's 600-second pool ceiling before all test files start, without
1873
+ skipping files, forcing exit, weakening isolation blindly, or changing
1874
+ production embedding behavior.
1875
+
1876
+ ### Acceptance criteria
1877
+
1878
+ 1. A diagnostic reproduces and attributes the failed fork-worker startup for
1879
+ `knodin.test.ts`, `map.spec.ts`, and the lifecycle runner instead of
1880
+ increasing the timeout.
1881
+ 2. Tests use an explicit deterministic embedder where semantic model quality is
1882
+ not under test; dedicated embedding tests retain production-path coverage.
1883
+ 3. All test files start, the full suite passes, and each fresh-worker shard
1884
+ exits naturally under its bounded five-minute test ceiling with immutable
1885
+ run artifacts present in the checkout.
1886
+ 4. Lint, typecheck, build, coverage thresholds, native-resource teardown, and
1887
+ production local-embedding behavior remain unchanged.
1888
+
1889
+ ### Evidence and corrected diagnosis
1890
+
1891
+ - Root cause (corrected): the handoff attributed the startup failures to repeated
1892
+ ONNX init, but a diagnostic reproduced them with the three ONNX-free canary
1893
+ files alone. The real cause is `isolate: true` (Vitest default): the pool
1894
+ re-forks a **fresh process per test file** — the config's stated "keep one
1895
+ long-lived worker" intent was never achieved — and under host load a per-file
1896
+ fork spawn intermittently exceeds Vitest's 60 s worker-start handshake
1897
+ (`START_TIMEOUT`), failing a neighbouring file. Actual test work is ~200 s
1898
+ (well under the ceiling); the overage was pure per-file fork churn.
1899
+ - Fix: `isolate: false` + `singleFork` (one reused process — the documented
1900
+ intent, not a blind isolation weakening), so native modules init once and only
1901
+ one fork is ever spawned. Cross-file safety is preserved by explicit guards:
1902
+ the shared deterministic lexical embedder is installed by setup and restored by
1903
+ every teardown; a per-test env/cwd snapshot-restore plus `restoreMocks`/
1904
+ `unstubEnvs`; the real ONNX model is isolated to `embeddings.spec` (loaded once,
1905
+ disposed after) and guarded by a `getRealEmbedderLoadCount()` regression; the
1906
+ six competitive benchmark `runner.spec.ts` files spawn their sub-benchmarks
1907
+ asynchronously (no synchronous `execFileSync` blocking the worker); and removed
1908
+ per-test repo state (SQLite handles + FSWatcher/FSEvent descriptors) is evicted
1909
+ each test.
1910
+ - Verified: the gate uses 16 strictly sequential, fresh-worker shards so
1911
+ native SQLite/watcher state cannot accumulate across the whole suite. The
1912
+ two semantic-ranking specs pass on the deterministic lexical embedder; lint,
1913
+ typecheck, and build pass. Final integration certification passed all 87
1914
+ files/503 tests in 1,471.79 seconds with 597,917,696 bytes peak RSS, and its
1915
+ merged 503-test coverage run passed at 89.5% statements, 79.25% branches,
1916
+ 90.62% functions, and 91.8% lines (901,840,896 bytes peak RSS). No tests or
1917
+ files are quarantined or skipped to obtain these results.
1918
+ - Release certification still reruns the fully accumulated integration checkout
1919
+ under C54; this distinction prevents cumulative-branch evidence from being
1920
+ mislabeled as a final integration-branch run.
1921
+
1922
+ ### C51 — Native training-free embedding quantization (memory-efficient vector search)
1923
+
1924
+ - Status: evaluated — rejected for production
1925
+ - Priority: P2
1926
+ - DependsOn: C46
1927
+ - Touches: `src/engine/embeddings.ts`, `src/engine/index.ts` (exact scan + ANN
1928
+ path), the `symbol_embeddings` storage, and a recall-gate fixture
1929
+ - What: add an opt-in, in-process, data-oblivious quantization of the 384-dim
1930
+ symbol embeddings — normalize → fixed deterministic rotation → precomputed
1931
+ bucketing (the TurboQuant concept, implemented natively in TypeScript) —
1932
+ shrinking stored vectors ~8× (4-bit) and speeding the cosine scan, with float32
1933
+ remaining authoritative and the safe default. This is the memory-efficient
1934
+ vector-search lever that keeps the zero-auth / local / no-egress / one-process
1935
+ guarantees intact, versus competitors that reach for a hosted vector DB. It is
1936
+ explicitly **not** a sidecar, external index, Rust dependency, or hosted store.
1937
+ See [ADR 005](../docs/adr/005-native-embedding-quantization.md).
1938
+ - Evidence: `benchmarks/evaluations/c51-quantization/raw-results.json` binds a
1939
+ 2,048-vector real MiniLM fixture and preserves per-query top-5/top-10 overlap,
1940
+ rank correlation, estimated database/resident payload bytes, conversion time,
1941
+ 3+20 warm and five cold samples, peak RSS, constraints, gates, and disposition. The fixed-seed
1942
+ block-Hadamard 4-bit prototype used 20 unique queries per tier and achieved
1943
+ mean top-10 recall 96%, 93.5%, and 92% at 128, 512, and 2,048 vectors. It
1944
+ therefore failed the >=95% gate at two engagement sizes.
1945
+ Payloads were 8× smaller, but warm p50 was not materially better and
1946
+ conversion-inclusive cold p50 was worse at every size; 99,876,864-byte peak
1947
+ process RSS did not establish an attributable end-to-end RSS win.
1948
+ - Disposition: reject production quantization from this design. Float32 remains
1949
+ authoritative and unchanged. Representation shrink without the recall gate
1950
+ and a material end-to-end latency or memory win does not authorize production
1951
+ work, so no separately numbered implementation item is proposed.
1952
+
1953
+ ### Acceptance criteria
1954
+
1955
+ Evaluation outcome: criteria 1 and 4 hold for the prototype; criterion 3 fails
1956
+ at 512 and 2,048 vectors. Criterion 2 is deliberately not implemented because
1957
+ the failed evaluation does not authorize a production path or environment gate.
1958
+
1959
+ 1. Encode/decode is deterministic (hard-coded rotation seed) and reproducible
1960
+ across machines, with no trained or per-repo codebook and no calibration pass.
1961
+ 2. Quantized codes are opt-in behind an env gate (default off); float32 stays the
1962
+ source of truth and is rebuildable/invalidated by the existing generation
1963
+ counter.
1964
+ 3. A committed real-vector fixture proves at least 95% mean top-10 recall against
1965
+ the float32 exact scan at every size where the path could engage, mirroring
1966
+ the R18/R48 gate. Report top-5/top-10 overlap, rank correlation, database and
1967
+ resident bytes, conversion time, cold/warm p50/p95, and peak RSS. Eligibility
1968
+ additionally requires a material end-to-end latency or memory win; an inner-
1969
+ loop-only improvement is insufficient.
1970
+ 4. No external index, sidecar, Rust/native dependency, hosted DB, or source/vector
1971
+ egress; the embedding model, its dimensionality, and its normalization are
1972
+ unchanged.
1973
+
1974
+ ### C52 — Package repository lifecycle initialization and refresh
1975
+
1976
+ - Status: implemented
1977
+ - Priority: P0
1978
+ - DependsOn: C16, C34, C50
1979
+ - Touches: `src/init.ts`, `src/repository-management.ts`, the deprecated
1980
+ `src/fleet.ts` compatibility adapter, `src/indexable-paths.ts`,
1981
+ `bin/cli.ts`, `src/engine/index.ts`, `scripts/pack-install-smoke.ts`,
1982
+ repository-init/CLI acceptance tests, package build output
1983
+ - What: make the compiled package usable across existing repositories without a
1984
+ knodin source checkout. Resolve the installed executable in hooks, share one
1985
+ canonical indexable-path policy, refresh after commit/checkout/merge/rewrite,
1986
+ and initialize repository checkouts idempotently.
1987
+
1988
+ ### Acceptance and evidence
1989
+
1990
+ 1. `knodin init` preserves existing hook content, supports worktrees and
1991
+ `core.hooksPath`, and repeated initialization produces no duplicate knodin
1992
+ blocks.
1993
+ 2. Hook refresh selects paths through the engine's canonical policy, including
1994
+ TypeScript and supported Salesforce Apex, Aura, LWC, and metadata inputs.
1995
+ 3. `knodin repos init` discovers Git roots deterministically, deduplicates real
1996
+ paths/worktrees, supports dry-run and JSON output, and records individual
1997
+ failures without abandoning the remaining repositories. The historical
1998
+ `knodin fleet init` spelling remains a time-bounded compatibility alias.
1999
+ 4. `scripts/pack-install-smoke.ts` installs the exact tarball in a clean external
2000
+ repository and proves package-resolved hooks across post-commit,
2001
+ post-checkout, post-merge, post-rewrite, rename, and deletion lifecycle
2002
+ events.
2003
+
2004
+ Evidence: packaged hooks landed in `c3a10a5`, fleet initialization in `72211ab`,
2005
+ and the clean-consumer gate in `1a6dbca`. The 0.1.0 candidate gate completed 49
2006
+ sequential commands with 465 MiB peak aggregate RSS.
2007
+
2008
+ ### C53 — Make repair observable, cancellable, and resumable
2009
+
2010
+ - Status: implemented
2011
+ - Priority: P0
2012
+ - DependsOn: C47, C52
2013
+ - Touches: `src/engine/index.ts`, `bin/repair-progress.ts`, `bin/cli.ts`,
2014
+ `docs/specs/repair-progress-telemetry.md`, index-health and CLI progress tests,
2015
+ packed-install repair acceptance
2016
+ - What: expose truthful phase/file progress for long repairs, preserve completed
2017
+ transactions on cancellation, and provide composable human and machine CLI
2018
+ output.
2019
+
2020
+ ### Acceptance and evidence
2021
+
2022
+ 1. Engine events carry monotonic sequences and counters, repository-relative
2023
+ source-safe paths, Salesforce metadata families, and terminal completion,
2024
+ cancellation, or failure without allowing callback errors to fail repair.
2025
+ 2. `AbortSignal` boundaries leave committed files recoverable, avoid recording
2026
+ a successful reconciliation on cancellation, and let the next audit resume
2027
+ from `index_state`; concurrent callers receive their own event stream.
2028
+ 3. CLI progress supports automatic TTY, plain, JSON, JSONL, and silent modes;
2029
+ progress uses stderr except for JSONL envelopes, and final results retain
2030
+ deterministic stdout.
2031
+ 4. Unit/CLI tests and the installed-tarball smoke gate cover cancellation,
2032
+ resumption, throttling, heartbeat, stream separation, and argument
2033
+ validation.
2034
+
2035
+ Evidence: the engine event foundation landed in `e0dcb5f`, cancellation and
2036
+ resumption in `1dc74cb`, and CLI rendering/validation in `a05edcd` and
2037
+ `c0d6745`; the installed-tarball gate exercises all five CLI output modes.
2038
+
2039
+ ### C54 — Certify the local 0.1.0 release artifact
2040
+
2041
+ - Status: implemented
2042
+ - Priority: P0
2043
+ - DependsOn: C50, C52, C53
2044
+ - Touches: `package.json`, `bun.lock`, `README.md`,
2045
+ `docs/releases/0.1.0.md`, this roadmap,
2046
+ `docs/specs/repair-progress-telemetry.md`,
2047
+ `scripts/pack-install-smoke.ts`, compiled package/tarball evidence
2048
+ - What: version the first fleet-consumable package, prove it from a clean
2049
+ consumer under the 3 GiB development cap, reconcile roadmap/spec truth, and
2050
+ produce a local release commit and tag without claiming a remote or registry
2051
+ publication.
2052
+
2053
+ ### Acceptance and evidence
2054
+
2055
+ 1. Package metadata identifies `0.1.0`, and `bun.lock` is regenerated from that
2056
+ manifest (Bun's lock format does not duplicate the workspace version); the
2057
+ exact `knodin-0.1.0.tgz` passes the clean-consumer acceptance gate with
2058
+ an aggregate process-tree limit of 2,750 MiB.
2059
+ 2. Lint, typecheck, build, documentation checks, and the full suite pass
2060
+ sequentially under the memory cap; C50 is not closed through quarantine,
2061
+ skipped files, or forced exit. Fresh-worker shards use a bounded five-minute
2062
+ default test ceiling for loaded-host graph builds while replay subprocesses
2063
+ retain tighter explicit limits.
2064
+ 3. README and release notes distinguish source-checkout use, local packed
2065
+ installation, and actual npm/remote publication state.
2066
+ 4. Tasks 1–3 in the repair telemetry specification are recorded complete while
2067
+ the MCP progress bridge, `--plan`, embedding detail, large Salesforce replay,
2068
+ and incremental embedding work remain explicitly pending.
2069
+
2070
+ Evidence to date: the final exact `knodin-0.1.0.tgz` replay passed 49
2071
+ sequential clean-consumer commands in 43,917 ms with 478 MiB peak aggregate RSS
2072
+ against the 2,750 MiB fail-closed limit. Final integration lint, typecheck,
2073
+ build, 503-test suite, and merged coverage gates pass under 3 GiB (the highest
2074
+ observed gate RSS is 901,840,896 bytes). The local release commit is tagged
2075
+ `v0.1.0`; no remote or registry publication is implied.
2076
+
2077
+ ### C55 — Pin and audit the authorized Token Optimizer source
2078
+
2079
+ - Status: implemented
2080
+ - Priority: P0
2081
+ - DependsOn: C24
2082
+ - Touches:
2083
+ `benchmarks/evaluations/token-optimizer-20260730/comparison-manifest.json`,
2084
+ `scripts/token-optimizer-source-audit.ts`, immutable raw audit evidence, and
2085
+ manifest-consistency tests
2086
+ - What: replace the demo-only source assumption with an authorized,
2087
+ commit-bound GHES audit while preserving the distinction between source
2088
+ evidence and equivalent behavioral measurement.
2089
+
2090
+ ### Acceptance and evidence
2091
+
2092
+ 1. The comparison pins repository, revision, Git tree, file blobs, README hash,
2093
+ and demo hash; branch movement cannot alter the evaluated input.
2094
+ 2. An executable audit fetches blobs by immutable revision and fails closed on
2095
+ a different commit, truncated tree, missing file, contradicted finding, or
2096
+ failed behavioral probe.
2097
+ 3. The audit verifies parser strategy, checked-in test inventory, shell and
2098
+ containment posture, path scope, telemetry fields/defaults, token/baseline
2099
+ math, compression budget behavior, dependency range, update/provenance
2100
+ inventory, and installation steps.
2101
+ 4. The manifest remains `behavior-partial`: source verification must not be
2102
+ presented as production routing, savings, or end-to-end parity evidence.
2103
+
2104
+ Evidence:
2105
+ `benchmarks/evaluations/token-optimizer-20260730/raw-source-audit-20260730.json`
2106
+ pins commit `a2b9cef5efdfcf0393e9e7a5ded229ffd9f613d7`, 17 files, and 16 verified
2107
+ findings. Its behavioral probe returns 104 lines for 100 matching input lines
2108
+ despite `max_lines=50`.
2109
+
2110
+ ### C56 — Replay all seven Token Optimizer workflows
2111
+
2112
+ - Status: implemented
2113
+ - Priority: P0
2114
+ - DependsOn: C55
2115
+ - Evidence:
2116
+ `benchmarks/evaluations/token-optimizer-20260730/raw-behavior-replay-20260730.json`
2117
+
2118
+ The immutable replay executes 20 warm and five cold-process samples per
2119
+ operation, uses the real `o200k_base` tokenizer, and applies shared correctness
2120
+ oracles to all five structural workflows. Both products pass those oracles.
2121
+ The final C68 replay records knodin as smaller in real tokens and faster at warm
2122
+ p50/p95 for all five structural cases on the small fixture; its cold-process
2123
+ startup remains materially slower, and command execution remains gated. The competitor compressor
2124
+ preserves the labeled diagnostic strings but returns 29 lines for a requested
2125
+ maximum of 20. knodin's C57 compressor preserves 7/7 detected signals in 14
2126
+ lines with hard byte/line budgets, exact omission accounting, and retained
2127
+ local drill-down. Its selected content is smaller; its complete recovery
2128
+ envelope remains 18 tokens larger on this fixture.
2129
+
2130
+ ### C57 — Bounded recoverable diagnostic-output compression
2131
+
2132
+ - Status: implemented
2133
+ - Priority: P0
2134
+ - DependsOn: C3, C55
2135
+ - Evidence:
2136
+ `src/output-compression.ts`,
2137
+ `src/__tests__/unit/output-compression.spec.ts`,
2138
+ `src/__tests__/unit/output-compression-cli.spec.ts`,
2139
+ `benchmarks/evaluations/token-optimizer-20260730/raw-behavior-replay-20260730.json`,
2140
+ and `docs/COMMAND-OUTPUT-COMPRESSION.md`
2141
+
2142
+ knodin accepts already-produced combined text, ordered stdout/stderr events, or
2143
+ a repository-contained log artifact. It deterministically applies `smart`,
2144
+ `head-tail`, or `errors-only` selection with generic and ecosystem-specific
2145
+ adapters; enforces hard rendered-line and UTF-8 content-byte budgets; preserves
2146
+ exit/signal metadata; reports unpreserved signals rather than overstating
2147
+ fidelity; records exact omitted line/source-byte ranges; and retains bounded
2148
+ private drill-down by SHA-256 identity without rerunning a command.
2149
+
2150
+ The pinned adversarial replay measures 604 raw tokens. knodin returns 109
2151
+ content tokens and a 205-token recovery envelope while preserving 7/7 detected
2152
+ signals in 14/20 lines. The competitor returns 187 tokens and preserves every
2153
+ golden string, but emits 29 lines for a requested maximum of 20. This supports
2154
+ the narrow `verified-stronger-budget-and-recovery` classification, not a broad
2155
+ claim that every compression dimension is stronger.
2156
+
2157
+ Command execution is not part of C57. C58's portable-containment evaluation was
2158
+ rejected, and direct MCP `text` input does not prove context savings when raw
2159
+ text has already crossed the model boundary.
2160
+
2161
+ ### C58 — Gate command execution on reviewed portable containment
2162
+
2163
+ - Status: evaluated — rejected
2164
+ - Priority: P0
2165
+ - Disposition: evaluated — rejected
2166
+ - DependsOn: C57
2167
+ - Owner roles: security reviewer, release owner, and one platform certifier for
2168
+ each supported operating system
2169
+ - Touches only if approved: a disposable command-runner prototype, adversarial
2170
+ containment fixtures, platform evidence, threat model, and a separately
2171
+ reviewed production implementation item
2172
+ - Evidence: `benchmarks/evaluations/c58-containment/` contains the disposable
2173
+ prototype, adversarial fixture, native macOS raw result, methodology, threat
2174
+ model, tests, and independent `c58-security-review-r1` dated 2026-08-02.
2175
+ - Terminal disposition: `evaluated — rejected`. Native macOS evidence passed
2176
+ the bounded defense-in-depth cases, but repository cwd/path checks are not a
2177
+ filesystem or network sandbox, process polling retains escape/PID races, and
2178
+ Linux/Windows are explicitly unavailable. Production execution remains
2179
+ absent and no implementation item is created.
2180
+ - Post-evaluation gate hardening: the three-level fixture now starts its 100 ms
2181
+ containment window only after an exact bounded readiness handshake, with a
2182
+ separate startup deadline and silent-tree cleanup test. This removes a
2183
+ coverage-timing race without changing raw evidence, platform omissions,
2184
+ security limitations, or the rejected disposition.
2185
+
2186
+ The evaluation decides whether knodin can safely execute a caller-supplied
2187
+ command before compressing its output. C57's already-produced text and artifact
2188
+ inputs remain the production default. Merely matching a competitor workflow is
2189
+ not sufficient to authorize execution.
2190
+
2191
+ Evaluation gate:
2192
+
2193
+ 1. Build no production command surface. First produce a disposable prototype
2194
+ using no-shell argv execution, repository-root binding, a minimal sanitized
2195
+ environment, stdin closure, output/time/RSS/process-count limits, and
2196
+ process-tree termination on every proposed supported platform.
2197
+ 2. Check in adversarial fixtures for shell metacharacters, traversal, symlink
2198
+ escape, inherited secrets, descriptor inheritance, child/grandchild escape,
2199
+ detached processes, output floods, timeout races, signals, and cancellation.
2200
+ 3. Record native macOS, Linux, and Windows results independently. Unsupported
2201
+ containment primitives or an unavailable platform are explicit omissions,
2202
+ not inferred portability.
2203
+ 4. Require an independent security review of the threat model and residual
2204
+ risks. Retain its revision, date, findings, and disposition without including
2205
+ credentials or private source.
2206
+ 5. Close with exactly one disposition: `evaluated — rejected`, `evaluated —
2207
+ deferred`, or a newly numbered implementation item with its own acceptance
2208
+ criteria. C58 itself never authorizes shipping from prototype evidence.
2209
+ 6. If rejected or deferred, keep command execution absent and update the
2210
+ scorecard limitation. This is a valid terminal outcome and does not block
2211
+ C67 when the exclusion remains explicit.
2212
+
2213
+ ### C57–C67 — Token Optimizer parity and trusted-distribution program
2214
+
2215
+ - Status: parked-external-evidence — 2026-08-02
2216
+ - Previous status: active
2217
+ - Source objective: the checked-in comparison manifest and the DTS competitive
2218
+ goal accepted on 2026-07-30
2219
+ - Ordering: behavioral replay (C56) precedes claims; safe compression (C57)
2220
+ precedes failure-to-code intelligence (C59); command execution (C58) remains
2221
+ excluded after its rejected evaluation; signed trust (C62) precedes update
2222
+ mutation (C63), release attestation
2223
+ (C64), and adversarial supply-chain proof (C65); the scorecard and release are
2224
+ terminal gates.
2225
+
2226
+ The program must preserve these invariants:
2227
+
2228
+ 1. Every competitor capability is either verified equivalent/stronger, recorded
2229
+ weaker, or explicitly excluded by a reviewed safety gate.
2230
+ 2. Compression has a hard truthful budget, exact omission accounting, retained
2231
+ local drill-down data, and adversarial diagnostic-fidelity fixtures.
2232
+ 3. No command runner ships without no-shell argv execution, repository binding,
2233
+ sanitized environment, process-tree/resource/output containment, audit, and
2234
+ supported-platform evidence.
2235
+ 4. ROI telemetry is local and opt-in, uses a real tokenizer, separates measured
2236
+ and modeled baselines, and excludes source, raw output, command text,
2237
+ usernames, and absolute paths by default.
2238
+ 5. Update availability is non-blocking and privacy-preserving. Mutation requires
2239
+ expiring threshold-signed metadata, rollback/freeze/mix-and-match defenses,
2240
+ digest/length/provenance checks, quarantine, health validation, atomic
2241
+ activation, and rollback.
2242
+ 6. A compromised npm, Homebrew, Artifactory, GitHub SaaS, or GHES account alone
2243
+ cannot authorize installation.
2244
+ 7. The terminal scorecard links every cell to source, executable tests,
2245
+ benchmark output, and a limitation. Broad superiority language remains
2246
+ prohibited until C66 and C67 close.
2247
+
2248
+ ### C59 — Source-evidenced failure-to-code diagnosis
2249
+
2250
+ - Status: implemented
2251
+ - Priority: P1
2252
+ - DependsOn: C57
2253
+ - Evidence:
2254
+ `src/failure-diagnosis.ts`,
2255
+ `src/__tests__/unit/failure-diagnosis.spec.ts`,
2256
+ `src/__tests__/unit/output-compression-cli.spec.ts`,
2257
+ `src/__tests__/unit/knodin-tools.spec.ts`, and
2258
+ `docs/COMMAND-OUTPUT-COMPRESSION.md`
2259
+
2260
+ The retained compression artifact can be diagnosed through
2261
+ `knodin compress diagnose <artifact-id>` or
2262
+ `operation: "compress", compressionAction: "diagnose"` without rerunning the
2263
+ failed command or adding a second MCP tool. Direct already-produced text is
2264
+ also accepted behind the same input, redaction, and byte limits.
2265
+
2266
+ The resolver extracts common JavaScript/TypeScript, Python, Java, C#, Go, and
2267
+ Rust source locations; binds exact or unique-suffix paths to tracked files;
2268
+ refuses traversal, external symlinks, untracked paths, and ambiguous basenames;
2269
+ and returns source-evidenced owning symbols, stable identities, nearest package
2270
+ manifests, tests, upstream callers, downstream dependencies, recent changes,
2271
+ bounded snippets, and indexed/current commit freshness.
2272
+
2273
+ The aggregate source-context byte budget is hard. Incomplete input, exhausted
2274
+ context, unresolved locations, or non-current graph state makes the result
2275
+ `partial` rather than complete. Static relationships remain diagnostic
2276
+ candidates rather than runtime-causality proof; source maps, generated paths,
2277
+ dynamic dispatch, and framework wiring are recorded limitations.
2278
+
2279
+ ### C60 — Cross-client and GHES onboarding
2280
+
2281
+ - Status: parked-external-evidence — 2026-08-02
2282
+ - Previous status: active — the legacy GHES source and versioned 0.3.0 release path are
2283
+ live; local fail-closed mirror preparation is complete, while enterprise
2284
+ governance/workload identity, timed independent cross-platform onboarding,
2285
+ and Linux/Windows named-client certification remain external gates
2286
+ - Priority: P1
2287
+ - DependsOn: C52, C55
2288
+ - Evidence: `docs/PT-ACCESS-RECOMMENDATION.md`, `docs/INSTALLATION.md`,
2289
+ `docs/evidence/pt-ghes-onboarding-2026-07-31.md`,
2290
+ `docs/evidence/c60-macos-codex-native-2026-08-02.json`,
2291
+ `docs/evidence/c60-macos-native-rss-fix-2026-08-02.json`,
2292
+ `schemas/c60-named-client-evidence-v2.schema.json`,
2293
+ `scripts/c60-native-certify.ts`, `scripts/c60-native-platform.ts`,
2294
+ `scripts/c60-macos-codex-certify.ts`,
2295
+ `src/__tests__/unit/c60-native-platform.spec.ts`,
2296
+ `src/__tests__/unit/agent-integration.spec.ts`,
2297
+ `src/__tests__/unit/doctor.spec.ts`, `scripts/pack-install-smoke.ts`,
2298
+ `config/enterprise-mirror-policy.json`, and
2299
+ `docs/evidence/c60-autonomous-preparation-2026-08-02.md`
2300
+
2301
+ The `Enterprise-Apps/knodin` GHES repository is the DTS / Application Engineering
2302
+ read-only downstream mirror of the authoritative GitHub SaaS repository
2303
+ `DTS-Productivity-Engineering/knodin`. P&T engineers without GitHub SaaS
2304
+ access can download an exact versioned tarball from its
2305
+ release page, verify the recorded SHA-256, install with lifecycle scripts
2306
+ disabled, initialize an existing checkout, and diagnose the single MCP gateway.
2307
+ The observed 0.3.0 GHES asset is byte-identical to the retained CI artifact and
2308
+ public npm artifact. Docusign Artifactory currently stores the same bytes as a raw
2309
+ artifact, not npm registry metadata; the docs do not misrepresent that endpoint
2310
+ as an npm registry.
2311
+
2312
+ Named Claude, Codex, Gemini, and Antigravity configuration plus a real generic
2313
+ MCP handshake are executable tests. The immutable schema-v2 macOS arm64
2314
+ Codex-configuration plus generic-MCP record certifies the exact 0.6.0 tarball in
2315
+ 13.739 seconds using pre-populated offline npm/model caches: lifecycle-disabled
2316
+ install, existing-checkout init, project-local configuration without user-home
2317
+ mutation, bounded MCP initialize/list/status/query/doctor, and aggregate
2318
+ RSS/output evidence all pass.
2319
+ An isolated offline build and npm pack from the recorded source commit is
2320
+ byte-identical to the consumed tarball, with its prerequisite commands
2321
+ separately bounded and recorded. It does not prove cold-cache installation, an
2322
+ actual Codex process or active UI
2323
+ session, or the eventual distributed release artifact. Linux packed consumption
2324
+ remains release-gated. A first real Linux attempt falsified the original
2325
+ portability claim because zero-RSS kernel rows made RSS parsing fail before the
2326
+ first command completed; Windows PID 0 had the same defect. The fixed parser
2327
+ retains structural and integer validation, filters non-process/zero-resident
2328
+ rows, and rejects an all-filtered snapshot. A new provenance-bound macOS replay
2329
+ of the exact fix commit passed in 11.212 seconds with byte-identical package
2330
+ evidence, proving no macOS regression but not Linux/Windows readiness. The
2331
+ native Linux and Windows runs must restart from the fixed commit without source
2332
+ changes. GHES `main` is not protected, GHES releases are not
2333
+ immutable, and native Linux/Windows named-client runs, independent cross-platform review,
2334
+ and non-personal automated mirror synchronization are missing, so C60 remains
2335
+ active and no cross-platform-complete or universal two-minute claim is allowed.
2336
+
2337
+ ### C61 — Private real-token ROI telemetry and dashboard
2338
+
2339
+ - Status: implemented
2340
+ - Priority: P1
2341
+ - DependsOn: C16, C55
2342
+ - Evidence: `src/output-telemetry.ts`, `src/cli-model.ts`, `bin/cli.ts`,
2343
+ `src/tools/knodin-tools.ts`, `src/__tests__/unit/output-telemetry.spec.ts`,
2344
+ `src/__tests__/unit/output-telemetry-cli.spec.ts`,
2345
+ `src/__tests__/unit/knodin-tools.spec.ts`, and `docs/TELEMETRY.md`
2346
+
2347
+ Persistence remains disabled by default and local. Opt-in JSONL records use the
2348
+ pinned real tokenizer, executed full-file baselines only when bounded source is
2349
+ available, explicit unknowns otherwise, cold/warm and outcome dimensions,
2350
+ agent-round-trip, latency, RSS, freshness, compression-fidelity, indexing-time,
2351
+ and index-storage measurements. Repository grouping uses a one-way identifier;
2352
+ source, raw output, command text, usernames, and absolute paths are excluded.
2353
+
2354
+ The CLI and single MCP gateway expose status, static HTML report, sanitized JSON
2355
+ export, retention, and explicit clear operations. Inputs and outputs are
2356
+ repository-contained and symlink-refusing, persistence is atomic/private, and
2357
+ the self-contained dashboard reports summary, operation, time-series, private
2358
+ repository, latency p50/p95, indexing cost, freshness, and fidelity dimensions
2359
+ with light/dark rendering. Metrics with no defensible observation remain
2360
+ unavailable rather than modeled.
2361
+
2362
+ ### C62 — Signed update metadata trust
2363
+
2364
+ - Status: parked-external-evidence — 2026-08-02
2365
+ - Previous status: active — verification core and fail-closed local ceremony preparation
2366
+ implemented; production offline custodians, independent pin review, recovery
2367
+ evidence, and trust-pin activation remain external gates
2368
+ - Priority: P0
2369
+ - DependsOn: C55
2370
+ - Evidence: `src/update-trust.ts`,
2371
+ `src/__tests__/unit/update-trust.spec.ts`, `src/update-ceremony.ts`,
2372
+ `src/__tests__/unit/update-ceremony.spec.ts`,
2373
+ `scripts/root-ceremony-preflight.ts`,
2374
+ `schemas/root-ceremony-manifest-v1.schema.json`,
2375
+ `schemas/root-pin-review-receipt-v1.schema.json`,
2376
+ `docs/ROOT-CEREMONY-RUNBOOK.md`, and `docs/SIGNED-UPDATES.md`
2377
+
2378
+ The pure verifier implements deterministic signed payloads, content-addressed
2379
+ Ed25519 keys, disjoint root/targets/snapshot/timestamp roles, production
2380
+ thresholds for root and targets, bounded expiry, independent root pinning,
2381
+ dual-threshold one-step root rotation, version-and-digest rollback/equivocation
2382
+ state, timestamp→snapshot→targets length/hash/version binding, safe target
2383
+ paths, and signed artifact SHA-256/length verification. It performs no network
2384
+ or installation mutation.
2385
+
2386
+ Adversarial fixtures reject missing thresholds, field and artifact
2387
+ substitution, expired timestamp freezes, rollback, same-version equivocation,
2388
+ mix-and-match metadata, role-key reuse, root-pin substitution, future metadata,
2389
+ and rotations not signed by the old root threshold.
2390
+
2391
+ C62 is not closed merely because the verification code exists. No production
2392
+ private key is checked in or generated by tests. Release owners must complete
2393
+ the offline multi-custodian root ceremony, independently review and embed the
2394
+ initial root digest, and record rotation/recovery evidence. C63 now makes all
2395
+ doctor, CLI, and MCP update state consume only verified metadata; the former
2396
+ npm lookup has been removed. Production authorization still waits for the
2397
+ offline ceremony.
2398
+
2399
+ ### C63 — Signed-only update client and safe policy
2400
+
2401
+ - Status: parked-external-evidence — 2026-08-02
2402
+ - Previous status: active — implementation and adversarial unit/CLI evidence complete;
2403
+ closure requires a recorded fail-closed deferred-activation policy after C62
2404
+ - Priority: P0
2405
+ - DependsOn: C62
2406
+ - Evidence: `src/update-policy.ts`,
2407
+ `src/__tests__/unit/update-policy.spec.ts`,
2408
+ `src/__tests__/unit/update-policy-cli.spec.ts`,
2409
+ `docs/DOCTOR-AND-UPDATES.md`, and `docs/SIGNED-UPDATES.md`
2410
+
2411
+ The five `knodin update` actions work outside a repository. Periodic checks use
2412
+ an atomic background lease and never put network latency on ordinary commands.
2413
+ Fixed signed metadata paths are HTTPS-origin/path bound, redirect-free, capped,
2414
+ timed, privacy-preserving, and optionally byte-identical across mirrors.
2415
+ Verified state is atomic/private and retains monotonic evidence across failed
2416
+ checks.
2417
+
2418
+ Apply enforces notify/download/patch/minor policy, minimum age, channels,
2419
+ enterprise mirrors, maintenance windows, exact artifact length/digest, signed
2420
+ provenance digest, candidate/rollback quarantine, no-shell exact local npm
2421
+ artifact invocation, symlink-safe exclusive quarantine, signed-version path
2422
+ validation, allowlisted enterprise cross-check origins, health validation,
2423
+ health-verified automatic rollback, and offline re-verification for explicit
2424
+ rollback. Mutable registry tags are never used.
2425
+
2426
+ The production CLI remains fail-closed as `trust-unconfigured`: a root and its
2427
+ pin from one mutable user file are not accepted. No certified sandbox smoke
2428
+ adapter or Homebrew/Node-manager rollback adapter is enabled yet. C62's offline
2429
+ ceremony must embed the independent anchor. C63 closes by recording that
2430
+ automatic production activation remains disabled; it does not wait for C65.
2431
+ C65 then uses the terminal signed-only client to certify the smoke sandbox,
2432
+ multi-step rotation/recovery, compromised channels, and manager-specific
2433
+ activation before any later auto-apply enablement. This ordering removes a
2434
+ C63↔C65 closure cycle without weakening the fail-closed policy.
2435
+
2436
+ ### C64 — Provenance, SBOM, and cross-channel release attestation
2437
+
2438
+ - Status: parked-external-evidence — 2026-08-02
2439
+ - Previous status: active — implementation and adversarial fixtures complete; a live
2440
+ five-channel release attestation remains a C67 publication gate
2441
+ - Priority: P0
2442
+ - DependsOn: C62
2443
+ - Evidence: `src/release-attestation.ts`,
2444
+ `src/__tests__/unit/release-attestation.spec.ts`,
2445
+ `src/__tests__/unit/release-attestation-cli.spec.ts`,
2446
+ `src/__tests__/unit/release-workflow.spec.ts`,
2447
+ `src/__tests__/unit/release-candidate-workflow.spec.ts`,
2448
+ `scripts/release-attestation.ts`, `.github/workflows/publish.yml`,
2449
+ `.github/workflows/release-candidate.yml`, and `docs/RELEASING.md`
2450
+
2451
+ The release workflow installs dependencies with lifecycle scripts disabled,
2452
+ runs tests, and packs the artifact in an unprivileged job with read-only
2453
+ repository access. A separate job with OIDC and attestation authority downloads
2454
+ only the retained artifact, generates a production-dependency CycloneDX SBOM,
2455
+ and uses full-commit-pinned official GitHub actions to sign SLSA build provenance
2456
+ and the SBOM. It independently verifies both retained bundles against the exact
2457
+ artifact, source digest, tag ref, repository, workflow certificate identity,
2458
+ and GitHub OIDC issuer before npm publication. A manual retry must run at the
2459
+ tag ref; checked-out commit, workflow SHA, tag, and package version disagreement
2460
+ is a hard failure. The exact tarball, signed evidence, and verification receipts
2461
+ are retained for 90 days before that tarball is passed to npm and Artifactory.
2462
+
2463
+ The separate manual-only release-candidate workflow now structurally stops
2464
+ after the exact artifact and rollback bytes, provenance and SBOM bundles,
2465
+ verification and identity receipts, bounded input manifest, and incomplete
2466
+ five-channel attestation draft are retained for 90 days. Its build job is
2467
+ unprivileged and lifecycle-disabled; its OIDC job downloads the retained input
2468
+ set without a source checkout or dependency install and has no npm/JFrog/channel
2469
+ credential, publishing command, reusable publisher, or distribution mutation.
2470
+ This is local implementation evidence only: C64 remains active until a live OIDC
2471
+ run at an immutable tag produces the retained candidate bundle and independent
2472
+ verification. C67 alone authorizes distribution after that candidate is
2473
+ independently verified.
2474
+
2475
+ The canonical schema derives, rather than accepts, `verified`, `incomplete`, or
2476
+ `quarantined` status. It records the exact artifact SHA-256/SHA-512/npm
2477
+ integrity, source commit, build identity, provenance and SBOM digests,
2478
+ publication timestamps and locations, all five approved destinations, and an
2479
+ exact rollback target whose digest the CLI derives from downloaded bytes rather
2480
+ than accepting from the manifest. Missing evidence is incomplete; any artifact, version,
2481
+ or source disagreement is quarantined. The CLI binds inputs beneath one
2482
+ directory, refuses direct or nested symlinks/traversal/unbounded files/output replacement,
2483
+ writes visible failure evidence, and exits nonzero for an incomplete final or
2484
+ any mismatch. It ignores a manifest's claimed verification state and invokes
2485
+ an explicitly approved absolute `gh attestation verify` executable without a
2486
+ shell or inherited credentials/environment against both local bundles and the
2487
+ exact artifact/source/ref/workflow/certificate policy before deriving verified
2488
+ signature evidence.
2489
+
2490
+ Signed targets may bind the provenance, SBOM, SBOM-attestation, and release-
2491
+ attestation digests. Update quarantine and explicit rollback re-fetch or re-read
2492
+ all four objects, validate their exact bytes and attested channel state, and
2493
+ refuse missing or inconsistent evidence before a package manager is invoked.
2494
+ Apply cross-checks the candidate's attested rollback version, path, and digest
2495
+ against the separately quarantined rollback artifact.
2496
+
2497
+ C64 is not a claim that the current 0.3.0 channels have been retroactively
2498
+ attested. GitHub/Sigstore build signatures are independently useful evidence,
2499
+ but do not replace C62's offline-controlled threshold root. C64 closes when the
2500
+ exact release-candidate artifact has verified provenance and SBOM bundles plus
2501
+ a complete attestation draft containing the five intended destinations. The
2502
+ draft remains `incomplete` until distribution. C67 publishes those exact bytes,
2503
+ fills the destination receipts, derives the final `verified` record, and
2504
+ TUF-authorizes it. This split prevents C64 and C67 from depending on each
2505
+ other's terminal evidence.
2506
+
2507
+ The signature proves workflow identity and exact bytes, not that a compromised
2508
+ self-hosted runner built those bytes faithfully. Independent build reproduction
2509
+ and runner-compromise drills remain C65/C67 gates. The JFrog reusable workflow's
2510
+ `@master` reference also remains an explicit external constraint because its
2511
+ Vault OIDC policy rejects commit-pinned `job_workflow_ref` claims; this is not
2512
+ misreported as full workflow immutability.
2513
+
2514
+ ### C65 — Compromised-channel, rollback, freeze, and recovery defenses
2515
+
2516
+ - Status: parked-external-evidence — 2026-08-02
2517
+ - Previous status: active — bounded multi-step root recovery is implemented; production
2518
+ key ceremony, smoke sandbox, manager-specific activation, and live drills
2519
+ remain required
2520
+ - Priority: P0
2521
+ - DependsOn: C63, C64
2522
+ - Evidence: `src/update-trust.ts`, `src/update-policy.ts`,
2523
+ `src/__tests__/unit/update-trust.spec.ts`,
2524
+ `src/__tests__/unit/update-policy.spec.ts`, `docs/SIGNED-UPDATES.md`, and
2525
+ `docs/evidence/0.4.3-update-defense-drill.json`
2526
+
2527
+ The update client can recover from up to 32 missed root rotations per check by
2528
+ fetching every numbered intermediate root, cross-checking its exact bytes across
2529
+ configured mirrors, and verifying each version under both the previous and new
2530
+ root thresholds. It never accepts a direct version jump or promotes a root from
2531
+ mutable state. Missing, malformed, wrong-version, skipped, or mirror-divergent
2532
+ intermediates fail before new trust is persisted; the rotation bound prevents a
2533
+ malicious terminal version from causing unbounded requests.
2534
+
2535
+ This closes the implemented multi-step retrieval gap only. C65 still requires
2536
+ the offline multi-custodian production root/recovery rehearsal, certified smoke
2537
+ sandbox, npm/Homebrew/manager activation and rollback drills, compromised
2538
+ transport and runner exercises, and checked-in evidence from those executions.
2539
+ No automatic production activation or complete supply-chain claim is allowed
2540
+ until those gates and C67's live release are complete.
2541
+
2542
+ The checked-in 0.4.3 isolated drill records one combined execution of 52 trust,
2543
+ policy, CLI, and attestation scenarios, including compromised-fixture refusal,
2544
+ multi-step recovery, automatic rollback, and unhealthy-rollback refusal. It
2545
+ strengthens repeatable local evidence but deliberately leaves the production
2546
+ ceremony, certified adapters/sandbox, real manager mutation, and live
2547
+ transport/runner exercises open.
2548
+
2549
+ ### C66 — Evidence-linked Token Optimizer scorecard
2550
+
2551
+ - Status: parked-external-evidence — 2026-08-02
2552
+ - Previous status: active — scorecard published; terminal prerelease snapshot awaits
2553
+ active dependencies
2554
+ - Priority: P0
2555
+ - DependsOn: C56, C57, C59, C60, C61, C65, C68, C75
2556
+ - Evidence: `docs/TOKEN-OPTIMIZER-SCORECARD.md`,
2557
+ `benchmarks/evaluations/token-optimizer-20260730/scorecard.json`,
2558
+ `src/__tests__/unit/docs-integrity.spec.ts`, and the pinned audit/replay
2559
+ artifacts under `benchmarks/evaluations/token-optimizer-20260730/`
2560
+
2561
+ The human scorecard and machine manifest classify every advertised competitor
2562
+ workflow plus the cross-cutting product dimensions. Every row links checked-in
2563
+ implementation, verification source/test paths, benchmark or live evidence,
2564
+ scope, readiness, and a material limitation. The manifest pins the competitor,
2565
+ the raw replay digest, its declared base revision, every evaluated source hash,
2566
+ and the first merged knodin commit containing that exact composite. Its integrity test
2567
+ rejects dimension omissions/duplicates, revision or replay-digest drift,
2568
+ missing evidence paths, or empty scope/limitation fields.
2569
+
2570
+ The scorecard supports a qualified “more complete local code-intelligence
2571
+ system” statement. It explicitly prohibits “better across all dimensions”:
2572
+ command execution remains intentionally excluded, Windows/timed onboarding is
2573
+ not certified, cold startup is slower, and production trust ceremony, drills,
2574
+ and five-channel release evidence remain open. C66 publication does not close
2575
+ its active dependencies or C67. C66 becomes terminal when every listed
2576
+ dependency is terminal and the scorecard records the prerelease truth. C67
2577
+ appends the final release record and receipts as a non-gating amendment; C66
2578
+ does not depend on the release it gates.
2579
+
2580
+ ### C67 — Certify and distribute the completed competitive release
2581
+
2582
+ - Status: parked-external-evidence — 2026-08-02
2583
+ - Previous status: proposed
2584
+ - Priority: P0
2585
+ - Disposition: Must close
2586
+ - DependsOn: C60, C62, C63, C64, C65, C66, C75
2587
+ - Owner roles: release owner, offline-root custodians meeting the recorded
2588
+ threshold, security reviewer, five channel operators, and Windows/macOS/Linux
2589
+ certifiers
2590
+ - Evidence directory: `docs/evidence/releases/<version>/` with immutable raw
2591
+ receipts under `benchmarks/evaluations/releases/<version>/`
2592
+ - Preflight implementation: `src/release-preflight.ts`,
2593
+ `scripts/release-preflight.ts`, `schemas/release-plan-v1.schema.json`,
2594
+ `src/__tests__/unit/release-preflight.spec.ts`, and `docs/RELEASING.md`
2595
+
2596
+ C67 is the terminal release and roadmap gate. It certifies one exact version
2597
+ from one source commit and distributes one byte-identical package through all
2598
+ approved channels without relaxing any unresolved limitation.
2599
+
2600
+ The pure C67 preflight is implemented without release authority. It validates
2601
+ a bounded plan, exact clean annotated-tag identity, the terminal roadmap and
2602
+ ledger state of every dependency, tracked repository-contained nonsymlink
2603
+ evidence including an immutable terminal artifact per dependency, certified
2604
+ native targets, exact verification commands, an independently reviewed
2605
+ production root/ceremony/digest tuple, and a new version-bound evidence path. It performs no publishing,
2606
+ tagging, evidence-directory creation, credential access, or other external
2607
+ mutation and does not change C64's nonpublishing candidate workflow. C67 remains
2608
+ proposed and open: against the current authoritative ledger the preflight must
2609
+ fail until C60, C62, C63, C64, C65, and C66 close; a passing local unit suite is
2610
+ not a release authorization or distribution receipt.
2611
+
2612
+ Acceptance and execution contract:
2613
+
2614
+ 1. Preflight fails unless C60, C62, C63, C64, C65, C66, and C75 have terminal
2615
+ statuses and checked-in evidence. Record package version, tag, source commit,
2616
+ workflow commit, runtimes, target platforms, trust-root digest, and exact
2617
+ verification commands before mutation.
2618
+ 2. From a clean tag checkout, run `npm ci --ignore-scripts`, `npm run lint`,
2619
+ `npm run typecheck`, `npm run build`, `npm run test`,
2620
+ `npm run test:pack-install`, `npm run check:roadmap`, and the release-specific
2621
+ competitive, trust, defense, and attestation replays. Record exit status,
2622
+ elapsed time, peak RSS, and immutable artifact digests without overwriting an
2623
+ earlier run.
2624
+ 3. Build once in the unprivileged release job. Verify the retained package,
2625
+ provenance, SBOM, SBOM attestation, and release-attestation draft before any
2626
+ channel receives the artifact. A retry must use the same tag and bytes.
2627
+ 4. Publish SHA-256/length-identical bytes to npm, Homebrew, Artifactory, GitHub
2628
+ SaaS, and GHES. Record immutable channel identifiers, timestamps, downloaded
2629
+ digests, and independently fetched receipts. Any missing or divergent channel
2630
+ fails the release and triggers quarantine/recovery rather than success.
2631
+ 5. Produce threshold-signed targets/snapshot/timestamp metadata bound to the
2632
+ artifact and C64 evidence digests. Exercise update status, check, explain,
2633
+ apply, health validation, and rollback from clean supported consumers.
2634
+ 6. Re-run named-client onboarding on certified macOS, Linux, and Windows and
2635
+ record C60's two-minute measurement. Verify `knodin init`, lifecycle refresh,
2636
+ CLI query, and the single MCP gateway against an existing checkout.
2637
+ 7. Publish a release record listing every gate, receipt, degradation,
2638
+ unsupported case, rollback target, and authorized product claim. Append those
2639
+ receipts to C66's already-terminal prerelease scorecard without reopening its
2640
+ gate; preserve every limitation whose independent gate did not pass.
2641
+ 8. Mark C67 terminal and remove its ledger row in the same change, then require
2642
+ a fresh `npm run check:roadmap -- --complete` to succeed. If authority or
2643
+ infrastructure is unavailable, retain `blocked` with the exact missing
2644
+ authority and safe next action; never fabricate a receipt.
2645
+
2646
+ ### C68 — Close structural output and latency gaps
2647
+
2648
+ - Status: implemented
2649
+ - Priority: P0
2650
+ - DependsOn: C56
2651
+ - Evidence: the initial C56 replay identified compactness and warm-latency gaps;
2652
+ the final pinned replay proves the C68 compact mode passes its narrow shared
2653
+ oracles and beats competitor response tokens and warm p50/p95 for all five
2654
+ structural operations on the small fixture. Cold startup remains slower.
2655
+
2656
+ Acceptance:
2657
+
2658
+ 1. Add a purpose-built compact structural response mode that preserves stable
2659
+ identity, source location, ambiguity, freshness, and truthful budgets without
2660
+ returning unrelated graph fields or telemetry envelopes.
2661
+ 2. Measure gateway schema overhead separately from per-call output.
2662
+ 3. Avoid repeated freshness/index startup work inside one session and document
2663
+ the irreducible cold cost of a persistent graph separately from operation
2664
+ latency.
2665
+ 4. Re-run the exact C56 fixture. knodin must match or beat competitor response
2666
+ tokens for all five structural workflows and must not be slower in warm
2667
+ operation p50/p95 without an explicitly approved, evidence-backed tradeoff.
2668
+ 5. Correctness, ambiguity, and stale-index oracles must remain passing; size or
2669
+ latency cannot be improved by weakening evidence.
2670
+
2671
+ Implementation evidence: `detailLevel: "compact"` now selects a purpose-built
2672
+ compact structural contract for file summary, exact source, structural search,
2673
+ batch outline, and project overview. Stable `~` identity prefixes are accepted
2674
+ as ambiguity-safe selectors, returned/total counts remain explicit, and no
2675
+ generic telemetry or response-budget envelope is added. Generation-scoped
2676
+ structural caching plus a watcher-invalidated freshness lease removes repeated
2677
+ session work; a no-subprocess HEAD check prevents Git movement from reusing the
2678
+ lease. The pinned C56 replay records gateway schema cost separately and passes
2679
+ all five correctness, token, warm-p50, and warm-p95 gates. Cold-process cost
2680
+ remains explicitly reported as a persistent-graph tradeoff.
2681
+
2682
+ ### C69–C75 — Large-portfolio lifecycle hardening
2683
+
2684
+ - Status: certified with recorded degradations
2685
+ - Priority: P0/P1 as listed above
2686
+ - Source objective:
2687
+ `knodin-Portfolio-Integration-Issues-2026-07-30.md` and the tested candidate
2688
+ commit `c981480db89ffbbb646206a73660e9ab842686ba`
2689
+ - Scope: portfolio discovery, selection, initialization, lifecycle truth,
2690
+ diagnostics, and Salesforce metadata candidate quality.
2691
+ - C75 Evidence: `docs/evidence/portfolio-initialization-2026-07-31.md`
2692
+
2693
+ The candidate commit proves that a 256 MiB subprocess buffer, early selection,
2694
+ and sequential inventory can discover the current 61-repository portfolio, but
2695
+ that commit is evidence rather than the terminal design. It still materializes
2696
+ SalesforceCI's roughly 105 MB `git ls-files` response and portfolio dry-run
2697
+ peaked near the one-GB incremental-memory ceiling.
2698
+
2699
+ Acceptance:
2700
+
2701
+ 1. Inventory tracked files with a streaming or equivalently bounded design.
2702
+ SalesforceCI's 904,072 tracked files cannot abort the portfolio. A large,
2703
+ malformed, timed-out, or otherwise failed repository produces a
2704
+ machine-readable degraded record while other repositories continue.
2705
+ 2. Apply include/exclude selection before tracked-file inventory. An include
2706
+ selector that matches no discovered repository returns an explicit
2707
+ diagnostic rather than a successful empty result.
2708
+ 3. `repos discover ~/code --json` returns valid JSON for all 61 current
2709
+ repositories while remaining below one GB of incremental RSS. Selected
2710
+ searches do not inventory excluded repositories.
2711
+ 4. Make portfolio init and dry-run sequential, resource-releasing, resumable,
2712
+ and bounded. Dry-run performs no model or database work, estimates scale and
2713
+ memory risk, and reports init/update/repair actions accurately. A
2714
+ per-process memory ceiling degrades and continues rather than killing the
2715
+ portfolio run.
2716
+ 5. `knodin configure` is either clearly configuration-only or atomically
2717
+ produces a usable graph and hooks. `knodin init` drains compatible queued
2718
+ lifecycle events after graph health is established; incompatible events
2719
+ receive an exact remediation. A `knodin wait --fresh` deadline must not
2720
+ terminate a healthy lifecycle processor; dead or failed attempts are
2721
+ retried at a bounded interval. Immediate status and search must agree with
2722
+ the success message.
2723
+ 6. Running `knodin doctor` on a non-Git portfolio parent warns and directs the
2724
+ user to portfolio diagnostics. MCP diagnostics separately report on-disk
2725
+ configuration, subprocess handshake, active-client exposure being unknown,
2726
+ reload requirements, and the exact configuration path.
2727
+ 7. Architecture and roadmap candidates use basename/directory-aware evidence,
2728
+ bounded lists, reasons, and confidence. Salesforce object or metadata names
2729
+ containing words such as `Design`, `WizardDesign`, or `Backlog` are not
2730
+ promoted without qualifying path/content evidence.
2731
+ 8. Certification records the 61-repository result, peak RSS, elapsed time,
2732
+ degraded repositories, selector behavior, immediate queryability, and all
2733
+ regression/typecheck/build/package-install gates. C66 and C67 cannot make a
2734
+ portfolio-readiness claim until C75 closes.
2735
+
2736
+ C72 lifecycle evidence is recorded in
2737
+ `docs/evidence/lifecycle-event-drain-2026-07-31.md`. Generated events now bind
2738
+ to the actual worktree root and status separately exposes a pending event and
2739
+ `idle`, `running`, or `stale-lock` processor state. The checked-in wait path
2740
+ retries a stranded compatible processor. The evidence explicitly does not
2741
+ claim that the installed 0.3.0 Homebrew release contains those source changes;
2742
+ distribution remains a C67 gate.
2743
+
2744
+ C71 is implemented by `src/repository-init-process.ts` and the sequential
2745
+ orchestration in `src/repository-management.ts`. Every live repository runs in
2746
+ a disposable process with a 768 MiB RSS ceiling, independent parent-side RSS
2747
+ polling plus worker heartbeats, process-group termination on macOS/Linux,
2748
+ timeout escalation, repository-scoped degradation, aggregate resource evidence,
2749
+ and resumable manifests. `src/__tests__/unit/repository-init-process.spec.ts`
2750
+ and `src/__tests__/unit/repository-management.spec.ts` prove termination,
2751
+ continuation, and single-worker sequencing. That implementation evidence is
2752
+ separate from the following live C75 certification.
2753
+
2754
+ C75 live evidence is recorded in
2755
+ `docs/evidence/portfolio-initialization-2026-07-31.md`. The 69-record run stayed
2756
+ below one GiB, completed 48 selected repositories, degraded 16 workers at the
2757
+ memory ceiling, preserved tracked team integration in four repositories, and
2758
+ continued after every failure. C75 therefore certifies bounded portfolio
2759
+ behavior with named limitations; it does not claim that all repositories are
2760
+ healthy or initialized.
2761
+
2762
+ ### C76 — Responsive declarative CLI model
2763
+
2764
+ - Status: implemented
2765
+ - Priority: P1
2766
+ - DependsOn: C73
2767
+ - Evidence: PR #29 replaced the hand-authored parser/help template with the
2768
+ declarative Commander model and pseudo-terminal width acceptance coverage.
2769
+
2770
+ Acceptance:
2771
+
2772
+ 1. Immediately reflow help to the detected TTY width with deterministic
2773
+ non-TTY output and narrow/wide pseudo-terminal tests.
2774
+ 2. Migrate parsing, validation, global options, nested subcommands, and help to
2775
+ one declarative command model so documentation cannot drift from accepted
2776
+ arguments.
2777
+ 3. Preserve every checked-in legacy invocation and JSON stdout contract.
2778
+ 4. Prefer Commander unless an executable bakeoff disproves the choice:
2779
+ current Commander supplies width-aware help, nested commands, TypeScript
2780
+ support, a documented security policy, and zero runtime dependencies.
2781
+ 5. Pin the exact dependency and integrity through the normal lockfiles and
2782
+ release provenance. Do not adopt a broader dependency tree solely for
2783
+ cosmetic terminal rendering.
2784
+
2785
+ Implementation evidence: `src/cli-model.ts` is the single Commander-backed
2786
+ grammar for global options, root and nested commands, positional contracts,
2787
+ numeric parsing, unknown-option rejection, and root or command help.
2788
+ `bin/cli.ts` validates every non-help invocation through that model and reads
2789
+ shared option values from its parsed result before compatibility dispatch.
2790
+ Commander 15.0.0 is exact-pinned in both lockfiles and has no transitive runtime
2791
+ dependencies. `src/__tests__/unit/cli-model.spec.ts`,
2792
+ `src/__tests__/unit/cli-pty-help.spec.ts`, the existing CLI subprocess suite,
2793
+ and `docs/CLI.md` cover narrow/wide deterministic output, a real Unix
2794
+ pseudo-terminal when available, nested help, invalid syntax, legacy invocation
2795
+ compatibility, packaging, and the Windows-safe non-PTY path.
2796
+
2797
+ ### C77 — Replay CodeFlow edge-provenance and architecture-export claims
2798
+
2799
+ - Status: evaluated — retained knodin unchanged
2800
+ - Priority: P2
2801
+ - Disposition: Evaluate
2802
+ - DependsOn: C24, C27, C42
2803
+ - Motivation: CodeFlow is a browser-local, MIT architecture mapper with broad
2804
+ language coverage, blast-radius views, and raw JSON export. A secondary
2805
+ LinkedIn post says each dependency identifies its extraction mechanism, but
2806
+ that claim was not found in the current repository or README. Knodin already
2807
+ exposes provenance, confidence, exact-versus-heuristic labels, source evidence,
2808
+ and source lines; only a shared replay can establish whether CodeFlow presents
2809
+ uncertain edges or architecture exports more usefully.
2810
+ - In scope: add a pinned CodeFlow adapter or documented browser-local protocol;
2811
+ compare dependency precision/recall, unsupported edges, uncertainty labels,
2812
+ source evidence, blast-radius completeness, export fidelity, output size,
2813
+ latency, and peak RSS on the existing architecture fixtures.
2814
+ - Out of scope: adopting CodeFlow, copying its browser UI, adding a new service,
2815
+ or changing Knodin's extraction architecture solely from vendor or social-post
2816
+ claims.
2817
+ - Touches: `benchmarks/competitors/TRACKER.md`, a CodeFlow comparison report and
2818
+ immutable raw result, the shared competitive harness only if its existing
2819
+ adapter contract is insufficient, and this roadmap.
2820
+ - Collision risk: competitor tracker and shared harness mappings; schedule apart
2821
+ from other competitive-replay edits.
2822
+ - Source: https://github.com/braedonsaunders/codeflow
2823
+ - Source task: `6h8V3JGJMhV3m63h`
2824
+ - Evidence: `benchmarks/competitors/codeflow-vs-knodin.raw.json` records the
2825
+ exact commit and MIT license, local-only protocol, original/normalized edges,
2826
+ false-positive/false-negative sets, provenance coverage, export fidelity,
2827
+ response bytes, 3+20 warm and five-process cold samples, and independently
2828
+ measured RSS. Both products achieved precision/recall 1.0, unsupported-edge
2829
+ rate 0, and blast completeness 1.0. CodeFlow used 105,955,328 bytes peak RSS
2830
+ and had 122.788/186.788 ms cold p50/p95; knodin used 591,446,016 bytes and had
2831
+ 1,056.246/6,332.857 ms. Knodin warm p50/p95 was 0.532/0.736 ms versus
2832
+ CodeFlow's 3.101/8.202 ms. CodeFlow supplied none of the six evaluated
2833
+ file-relationship provenance categories; knodin supplied all six for all
2834
+ three edges. Architecture block exports remain explicitly incomparable.
2835
+ - Disposition: retain knodin unchanged. CodeFlow demonstrates a real cold-start
2836
+ and memory advantage, while knodin already supplies the evaluated provenance
2837
+ and warm-query behavior. Neither the incomparable architecture exports nor
2838
+ feature presence authorizes an architecture or presentation change, and this
2839
+ evaluation makes no aggregate superiority claim.
2840
+
2841
+ Acceptance:
2842
+
2843
+ 1. Pin the tested CodeFlow commit and record its MIT license, local execution
2844
+ path, setup, and limitations; do not require an account, hosted service,
2845
+ paid API, source egress, or new Knodin runtime dependency.
2846
+ 2. Run both products against identical checked-in fixtures and correctness
2847
+ oracles. Preserve false positives, false negatives, unresolved edges, and
2848
+ unavailable cases in immutable raw results.
2849
+ 3. For every exported relationship, record whether the tool supplies relation
2850
+ kind, extractor identity, exact/heuristic confidence, source file and line,
2851
+ and bounded source evidence. Treat the LinkedIn extractor-identity claim as
2852
+ unverified unless the pinned implementation emits it.
2853
+ 4. Report precision/recall, unsupported-edge rate, blast-radius completeness,
2854
+ export fidelity, response bytes, cold/warm p50 and p95, and independently
2855
+ measured peak RSS. Do not publish a win from incomparable or oracle-free rows.
2856
+ 5. Keep incremental spend at zero and incremental memory below 1 GB. Keep all
2857
+ fixture source local; if CodeFlow cannot run without GitHub egress, record a
2858
+ blocker rather than uploading private source.
2859
+ 6. Close with one evidence-based disposition: retain Knodin unchanged, improve
2860
+ provenance presentation/export, or propose a separately reviewed roadmap
2861
+ item. Do not infer an architecture change from feature presence alone.
2862
+
2863
+ ### C78 — Add bounded source-to-sink resource reachability
2864
+
2865
+ - Status: evaluated — retained C8 unchanged
2866
+ - Priority: P1
2867
+ - Disposition: Evaluate
2868
+ - DependsOn: C2, C3, C8, C24
2869
+ - Motivation: Code-Graph-RAG's opt-in `READS_FROM`, `WRITES_TO`, and `FLOWS_TO`
2870
+ edges answer a useful provenance question Knodin cannot currently answer:
2871
+ whether a value from an environment variable, file, database, socket, or
2872
+ network source can reach a log, file, database, or outbound-network sink.
2873
+ Knodin's C8 `flow_analysis` is deliberately limited to bounded, on-demand
2874
+ TS/JS facts inside one selected symbol; call-argument evidence is explicitly
2875
+ heuristic and does not compose source-to-sink resource reachability.
2876
+ - In scope: prototype synthetic resource identities plus source-evidenced read,
2877
+ write, assignment, argument, return, and kill facts; begin with environment or
2878
+ local configuration flowing to logging, network, and database sinks; expose a
2879
+ bounded query with exact coverage and omission metadata.
2880
+ - Out of scope: adopting Code-Graph-RAG, Memgraph, Qdrant, Docker, a hosted
2881
+ model, unrestricted graph queries, autonomous vulnerability verdicts, a full
2882
+ persisted PDG, SSA, alias-complete or path-sensitive analysis, or claims of
2883
+ runtime reachability.
2884
+ - Touches if approved: language extraction, optional persisted resource/flow
2885
+ schema or a cheaper on-demand representation, one compact query pattern,
2886
+ source/sink registry, documentation, and labeled multi-language fixtures.
2887
+ - Collision risk: `src/engine/index.ts`, graph schema/migrations, MCP/CLI query
2888
+ enums, and shared language extractors; schedule apart from other graph-schema
2889
+ or query-surface work.
2890
+ - Source: https://github.com/vitali87/code-graph-rag/blob/d0b257b402bb25b54cf8f1dee0af3a71623c8d9e/docs/architecture/data-flow-edges.md
2891
+ - Source task: `6hC5Cv934v8xj2Gh`
2892
+ - Evidence: `benchmarks/evaluations/c78-resource-reachability/` contains the
2893
+ labeled TS/JS oracle, disposable prototype, immutable raw result, tests, and
2894
+ methodology. It achieved 100% precision but 62.5% recall, missing argument,
2895
+ return, and recursion/cycle handoff. Three local repositories measured
2896
+ bounded on-demand analysis against temporary-file persisted JSON; peak RSS was
2897
+ 129,859,584 bytes with zero spend and no egress.
2898
+ - Disposition: retain C8's bounded statement flow and call-argument evidence
2899
+ unchanged. The precision gate passed, but recall failed the 80% minimum useful
2900
+ threshold, so neither persistence nor a production query is authorized and no
2901
+ separately numbered implementation item is proposed.
2902
+
2903
+ Acceptance:
2904
+
2905
+ 1. First build a disposable prototype and labeled oracle containing true flows,
2906
+ clean overwrites, shadowing, unrelated source/sink co-occurrence, dynamic
2907
+ resource names, argument handoff, return handoff, recursion/cycles, and
2908
+ deliberately unsupported constructs. Do not change the production schema
2909
+ until the prototype passes the gate.
2910
+ 2. Require at least 95% precision and report recall per supported language and
2911
+ source/sink class. False positives are a hard failure because the result may
2912
+ inform security review; unsupported or ambiguous paths must remain explicit.
2913
+ 3. Every result identifies the source and sink, relation kind, bounded path,
2914
+ source file and line evidence, extraction provenance, confidence, supported
2915
+ language/registry coverage, omissions, truncation, and freshness. Label every
2916
+ path static and heuristic; never call it proof of exploitability, secrecy, or
2917
+ runtime reachability.
2918
+ 4. Compare a bounded on-demand implementation with persisted resource/flow
2919
+ edges. Choose the smallest design meeting correctness and warm-latency goals;
2920
+ preserve stable identity, ambiguity safety, monotonic bounds, deterministic
2921
+ ordering, hard item/byte/token budgets, and fail-closed stale-index behavior.
2922
+ 5. Measure clean and incremental index time, database growth, cold/warm p50 and
2923
+ p95 query latency, response bytes, and peak RSS on at least three real local
2924
+ repositories. Incremental spend remains zero and incremental memory remains
2925
+ below 1 GB, with no hosted service, credentials, source egress, or required
2926
+ production model.
2927
+ 6. Start with a deliberately narrow language and source/sink matrix justified by
2928
+ fixtures. Adding a registry entry requires positive and negative acceptance
2929
+ cases; a language without a verified registry emits an explicit omission,
2930
+ not inferred coverage.
2931
+ 7. If the gate fails, retain C8's current bounded statement flow and call-
2932
+ argument evidence unchanged. Do not broaden or persist the feature merely to
2933
+ match a competitor's schema.
2934
+
2935
+ ### C79 — Bounded repository applicability signals
2936
+
2937
+ - Status: implemented
2938
+ - Priority: P1
2939
+ - Disposition: Evaluate
2940
+ - DependsOn: C52, C69
2941
+ - Touches: `repos discover` CLI option/schema, bounded repository inspection,
2942
+ Git-remote parsing, worktree-aware discovery fixtures, CLI help/reference, and
2943
+ compatibility evidence. It does not change `doctor` or the MCP surface.
2944
+ - What: decide whether repository marker and configuration detection belongs in
2945
+ knodin's repository-discovery charter. If it does, add an additive, opt-in
2946
+ `knodin repos discover <roots...> --json --signals` facet so consumers can use
2947
+ knodin's existing repository/worktree classification instead of duplicating a
2948
+ second traversal and a less reliable definition of a repository.
2949
+ - Scope: without `--signals`, output must remain byte-identical. With the flag,
2950
+ every repository gains `signals`, using `{}` for honest absence. Detection is
2951
+ read-only, offline, credential-free, bounded by existing bytes/tokens/items
2952
+ budgets, and limited to a documented allowlist of marker paths rather than a
2953
+ tree glob. Never read potentially secret file contents; the only content-read
2954
+ exception is allowlisted hook-manager configuration needed to report whether
2955
+ it references `aidev-track`. Linked worktrees inspect their own checkout and
2956
+ preserve existing skip/include behavior.
2957
+ - Evidence: either a checked-in charter decision closing the item unchanged, or
2958
+ a checked-in implementation with compatibility bytes, help/docs, unit
2959
+ fixtures, bounded-resource results, and verifier output.
2960
+
2961
+ Charter decision (2026-08-03): accepted. These path-derived, repository-scoped
2962
+ signals belong in discovery because they reuse knodin's existing bounded
2963
+ repository and linked-worktree classification. The facet remains opt-in and
2964
+ descriptive: it does not decide whether an agent skill applies or whether a
2965
+ detected service is configured correctly.
2966
+
2967
+ Acceptance:
2968
+
2969
+ 1. First record the charter decision. If marker/config policy is outside
2970
+ knodin's charter, close C79 unchanged with checked-in rationale; no feature
2971
+ implementation is required. Otherwise implement only the opt-in facet.
2972
+ 2. A golden compatibility test proves `repos discover <roots...> --json`
2973
+ remains byte-identical without the flag. With `--signals`, every repository
2974
+ entry has a deterministic `signals` object, including `{}` when the
2975
+ allowlisted inspection finds nothing; schema version changes only if the
2976
+ existing compatibility policy requires it.
2977
+ 3. When detected, the deterministic contract must return `hookManager`,
2978
+ allowlisted `markerFiles` (including Sonar, CI/PR-workflow, agent, and MCP
2979
+ client markers), sanitized `remotes` with name/host/owner, `ciProviders`,
2980
+ `agentConfigs`, and `aidevTrackReferenced: true|false` when an allowlisted
2981
+ hook-manager config is present. Detection performs
2982
+ no writes, network access, credential access, source egress, or secret-file
2983
+ content reads, except the narrow allowlisted hook-manager parse for the
2984
+ `aidev-track` reference. Incremental spend is zero.
2985
+ 4. Unit fixtures cover lefthook, husky, pre-commit, no hook manager,
2986
+ `aidev-track` reference detection, multiple remotes, multiple remote hosts,
2987
+ missing `.git`, every marker absent, and linked worktrees. Tests prove fixed
2988
+ allowlists and existing bytes/tokens/items budgets bound inspection and
2989
+ serialization.
2990
+ 5. `knodin doctor` behavior is unchanged. The repos CLI reference and
2991
+ `repos discover --help` document the flag, returned fields, allowlist,
2992
+ omissions, privacy constraints, and worktree behavior.
2993
+
2994
+ Implementation evidence: the deterministic verifier and bounded-resource
2995
+ result are checked in at
2996
+ [`docs/evidence/c79-repository-signals-2026-08-03.md`](../docs/evidence/c79-repository-signals-2026-08-03.md)
2997
+ and
2998
+ [`docs/evidence/c79-repository-signals-2026-08-03.json`](../docs/evidence/c79-repository-signals-2026-08-03.json).
2999
+ The unit acceptance fixtures are in
3000
+ `src/__tests__/unit/repository-management.spec.ts`; CLI help coverage is in
3001
+ `src/__tests__/unit/cli-model.spec.ts`, and the unchanged doctor regression is
3002
+ `src/__tests__/unit/doctor.spec.ts`.
3003
+
3004
+ Known limitations and decision gate: marker presence describes repository
3005
+ configuration; it does not prove that a skill applies or that a service is
3006
+ configured correctly. If the charter review rejects this application-policy
3007
+ facet, closing C79 unchanged is a valid terminal outcome and downstream tools
3008
+ may inspect only the repository paths knodin returns.
3009
+
3010
+ ### C80 — Structural-first retrieval routing
3011
+
3012
+ - Status: implemented
3013
+ - Priority: P0
3014
+ - Disposition: Must close
3015
+ - DependsOn: C3, C24, C42, C43
3016
+ - Touches: retrieval dispatch and telemetry, search/context tests, a checked-in
3017
+ cross-file oracle, and deterministic ablation replay artifacts.
3018
+ - What: route stable identity and exact name/signature/path evidence through
3019
+ lexical/FTS and bounded graph expansion before loading embeddings. Serialize
3020
+ the selected route and `embeddingsUsed` so behavior is explainable.
3021
+ - Scope: preserve the one-tool gateway, stable identities, ambiguity safety,
3022
+ freshness, deterministic ordering, and current hard response budgets.
3023
+ - Evidence: a local checked-in implementation, fixtures, raw three-arm replay,
3024
+ verifier, and report covering factual precision/recall, unsupported claims,
3025
+ latency, and independently measured peak RSS.
3026
+
3027
+ Acceptance:
3028
+
3029
+ 1. Exact symbol, path, and source-evidenced cross-file questions do not load the
3030
+ embedding model when structural evidence is sufficient; tests assert the
3031
+ route and `embeddingsUsed: false` through CLI and MCP-compatible responses.
3032
+ 2. A checked-in multi-hop oracle replays lexical/vector-only, bounded graph
3033
+ expansion, and the complete route deterministically. Graph expansion becomes
3034
+ default only if correctness improves without increasing unsupported,
3035
+ ambiguous, or stale claims; otherwise retain the narrower passing route.
3036
+ 3. Preserve raw results and report precision, recall, omissions/refusals,
3037
+ latency, response size, and peak RSS. Incremental spend is zero, source stays
3038
+ local, no network or account is required, and incremental memory stays below
3039
+ 1 GB.
3040
+
3041
+ Delivered evidence: structural-first dispatch and request-local retrieval
3042
+ telemetry are implemented in `src/engine/index.ts`; CLI/MCP parity, stable
3043
+ identity, exact path/name, lexical sufficiency, embedding fallback, bounded
3044
+ work, and duplicate-name safety are covered by
3045
+ `src/__tests__/unit/structural-routing.spec.ts`. The checked-in multi-hop and
3046
+ duplicate-name oracle, fixture, deterministic four-route raw replay, and
3047
+ verifier outputs are under
3048
+ `benchmarks/evaluations/c80-structural-routing/`; reproduce them with
3049
+ `npm run bench:c80 && npm run verify:c80`. The limitations-aware precision,
3050
+ recall, unsupported/ambiguous/stale, omission, latency, response-size, spend,
3051
+ locality, and independently isolated RSS report is
3052
+ `docs/evidence/c80-structural-routing-2026-08-03.md`.
3053
+
3054
+ Decision: bounded graph expansion is the default structural route because the
3055
+ checked-in gate improves mean oracle recall from 0.389 to 1.000 while preserving
3056
+ 1.000 precision, zero unsupported and stale claims, and the same one surfaced
3057
+ duplicate-name ambiguity. The complete route preserves those results and loads
3058
+ embeddings only when exact or all-token structural evidence is insufficient.
3059
+ The largest isolated incremental RSS sample was 160,284,672 bytes; incremental
3060
+ spend was zero and no network, account, credential, or hosted service was used.
3061
+
3062
+ Known limitations and decision gate: structural routing is not proof that
3063
+ semantic retrieval is unnecessary. Embeddings remain a bounded fallback, and
3064
+ token reduction alone cannot authorize a route that weakens correctness. Graph
3065
+ expansion follows resolved source relationships only, is capped at two hops and
3066
+ 200 selected nodes, and cannot recover runtime-only or unresolved relationships.
3067
+ The three-case deterministic oracle proves this decision only on its checked-in
3068
+ cross-file chains; sampled RSS can miss short-lived peaks between samples.
3069
+
3070
+ ### C81 — Powered end-to-end benchmark with a diagnose arm
3071
+
3072
+ - Status: evaluated — retained knodin unchanged
3073
+ - Priority: P0
3074
+ - Disposition: Must close
3075
+ - DependsOn: C42, C43, C59, C80
3076
+ - Touches: competitive benchmark corpus/runner, task labels, deterministic
3077
+ patch/test oracles, raw trajectories, statistics, and evidence report.
3078
+ - What: compare ordinary filesystem tools, knodin, and one pinned local
3079
+ competitor on retrieval and failing-build/test diagnosis using a powered,
3080
+ paired design.
3081
+ - Scope: at least 40 tasks, at least three repetitions, pinned repositories,
3082
+ commits, prompts, models, versions, seeds, and budgets; confidence intervals
3083
+ and pre-registered analysis are mandatory.
3084
+ - Evidence: checked-in corpus and labels, immutable raw trajectories, verifier
3085
+ output, confidence intervals, and a limitations-aware report.
3086
+
3087
+ Acceptance:
3088
+
3089
+ 1. Each arm runs the same tasks and deterministic correctness oracles. The
3090
+ diagnose subset scores turns to a correct patch and passing test, not an LLM
3091
+ judge alone; unavailable and incomparable rows remain explicit.
3092
+ 2. Cache behavior is controlled and recorded. Do not report billed-cost claims
3093
+ unless cache economics are comparable; this item must incur zero model/API
3094
+ spend and require no hosted service, credentials, or source egress.
3095
+ Incremental spend is zero.
3096
+ 3. Record patch-application rate, turns to correct edit, factual correctness,
3097
+ tokens where measured, cold/warm latency, and peak RSS under 1 GB. Preserve
3098
+ losing cases and raw data without overwriting prior evidence.
3099
+
3100
+ Known limitations and decision gate: the benchmark may justify retaining the
3101
+ product unchanged. It authorizes no superiority claim beyond oracle-qualified,
3102
+ comparable rows and no production feature outside a separately numbered item.
3103
+
3104
+ Delivered evidence: the pinned 40-task corpus, three-repetition paired runner,
3105
+ deterministic retrieval and patch/test oracles, immutable accepted and excluded
3106
+ pilot trajectories, integrity hashes, and verifier are under
3107
+ `benchmarks/evaluations/c81-powered-benchmark/`; reproduce validation and
3108
+ confidence intervals with `npm run verify:c81`. The limitations-aware report is
3109
+ [`C81 powered benchmark evidence`](../docs/evidence/c81-powered-benchmark-2026-08-03.md).
3110
+
3111
+ Decision: retain knodin unchanged. All three available arms achieved 1.000
3112
+ factual correctness and patch-plus-passing-test rate on this narrow synthetic
3113
+ corpus, with paired correctness differences of 0.000. The result supports no
3114
+ superiority, token, billed-cost, or production claim. The conservative
3115
+ worker-plus-descendant peak-RSS upper bound was 351,469,568 bytes; spend,
3116
+ hosted-service use, credentials, and source egress were zero.
3117
+
3118
+ ### C82 — Productize `compress diagnose`
3119
+
3120
+ - Status: implemented
3121
+ - Priority: P0
3122
+ - Disposition: Must close
3123
+ - DependsOn: C59, C81
3124
+ - Touches: existing compress diagnose CLI/MCP dispatch, discoverability and
3125
+ documentation, stable response schema, fixtures, telemetry, and replay.
3126
+ - What: make the retained-failure-to-graph workflow discoverable and robust so
3127
+ an artifact resolves to owning symbols, tests, callers, and bounded source
3128
+ without rerunning the failed command.
3129
+ - Scope: reuse the existing single MCP tool and retained C59 primitive; preserve
3130
+ stable artifact identity, freshness, ambiguity, traversal, and budget bounds.
3131
+ - Evidence: checked-in adversarial diagnostic fixtures, CLI/MCP parity tests,
3132
+ golden responses, raw replay, verifier, and local evidence report.
3133
+
3134
+ Acceptance:
3135
+
3136
+ 1. Representative compiler, test, stack-trace, multiline, truncated, malformed,
3137
+ stale, ambiguous, and missing-artifact cases return stable schema fields and
3138
+ exact owning source evidence or an explicit bounded refusal.
3139
+ 2. CLI and MCP produce equivalent source identities, callers/tests, omissions,
3140
+ freshness, confidence, and telemetry without a new MCP tool, command rerun,
3141
+ hosted service, credential, source egress, or model/API spend.
3142
+ Incremental spend is zero.
3143
+ 3. The C81 diagnose corpus demonstrates the workflow end to end with checked-in
3144
+ deterministic oracles, hard item/byte/token caps, recoverable continuation,
3145
+ and peak incremental memory below 1 GB.
3146
+
3147
+ Known limitations and decision gate: static graph evidence cannot prove runtime
3148
+ causality. Documentation must say that diagnoses are bounded source-evidenced
3149
+ candidates, not guaranteed root causes or automatic fixes.
3150
+
3151
+ Delivered evidence: the existing CLI and single-tool MCP dispatch now share an
3152
+ additive stable diagnosis envelope with confidence, exact diagnostic/relation
3153
+ omissions, retained-artifact continuation, and command-rerun/cap telemetry.
3154
+ Focused parity and refusal coverage is in `src/__tests__/unit/`; adversarial
3155
+ fixtures, golden responses, immutable raw replay, hard-cap verification, and
3156
+ resource evidence are under
3157
+ `benchmarks/evaluations/c82-compress-diagnose/`. Reproduce all ten adversarial
3158
+ cases and all ten C81 diagnose tasks with `npm run verify:c82`. The evidence
3159
+ report is [`C82 compress diagnose evidence`](../docs/evidence/c82-compress-diagnose-2026-08-03.md).
3160
+
3161
+ The conservative measured maximum RSS was 309,116,928 bytes, below 1 GB.
3162
+ Incremental spend, command reruns, network/hosted service use, credentials, and
3163
+ source egress were zero. Static candidates remain neither runtime root-cause
3164
+ proof nor automatic fixes; the synthetic replay does not cover every toolchain.
3165
+
3166
+ ### C83 — Progressive evidence delivery and hash handshake
3167
+
3168
+ - Status: implemented
3169
+ - Priority: P0
3170
+ - Disposition: Must close
3171
+ - DependsOn: C80, C82
3172
+ - Touches: gateway response contracts, locate/outline/evidence/expand dispatch,
3173
+ continuation/evidence handles, content hashing, CLI/MCP tests, and docs.
3174
+ - What: deliver recoverable evidence levels and allow delta/omission only when
3175
+ the client supplies an exact content hash for the baseline it holds.
3176
+ - Scope: every level reports returned and omitted material, reason, `more`, a
3177
+ stable continuation, freshness, confidence, and truthful hard budgets.
3178
+ - Evidence: checked-in golden protocol fixtures, tamper/staleness/rename tests,
3179
+ raw bounded replay, verifier, and compatibility report.
3180
+
3181
+ Delivered evidence: the additive `evidence` operation and CLI command share the
3182
+ source-preserving implementation in `src/progressive-evidence.ts`. Golden and
3183
+ raw 14-case protocol replays, fixture, resource record, compatibility report,
3184
+ and limitations are under `benchmarks/evaluations/c83-progressive-evidence/`;
3185
+ reproduce them with `npm run verify:c83`. Unit acceptance covers complete
3186
+ omission recovery, adversarial hashes/handles/paths, stale and renamed handles,
3187
+ and byte-identical CLI/MCP edit anchors. The evidence report is
3188
+ [`C83 progressive evidence evidence`](../docs/evidence/c83-progressive-evidence-2026-08-03.md).
3189
+
3190
+ Acceptance:
3191
+
3192
+ 1. Exact hash match permits the documented delta or omission path; missing,
3193
+ stale, truncated, renamed, or tampered baselines return full required source
3194
+ and state which path was taken. Client assertion alone is never trusted.
3195
+ 2. Locate, outline, evidence, and expand remain deterministic and recoverable
3196
+ under item/byte/token caps; tests prove omitted evidence can be requested and
3197
+ that edit anchors remain intact across CLI and MCP.
3198
+ 3. The protocol remains local, zero-spend, offline, under 1 GB incremental RSS,
3199
+ backward-compatible where promised, and fail-closed on stale or ambiguous
3200
+ evidence. Checked-in fixtures include adversarial hash and handle inputs.
3201
+
3202
+ Known limitations and decision gate: this is not lossy summarization and does
3203
+ not authorize token-savings or task-success claims until C81 measures them.
3204
+ The V1 surface addresses one exact regular repository file at a time, refuses
3205
+ symlinks and ambiguous/path-traversal input, and uses a lightweight top-level
3206
+ outline rather than claiming the engine's complete relationship model. The
3207
+ synthetic replay authorizes no production-readiness or cross-platform claim.
3208
+
3209
+ ### C84 — Optional SCIP import
3210
+
3211
+ - Status: implemented
3212
+ - Priority: P1
3213
+ - Disposition: Must close
3214
+ - DependsOn: C24, C80, C83
3215
+ - Touches: optional SCIP ingestion, stable identity/provenance mapping, conflict
3216
+ handling, LSIF/native fallbacks, fixtures, lifecycle bounds, and docs.
3217
+ - What: import compiler-produced SCIP facts as an opt-in precision tier while
3218
+ retaining LSIF compatibility and the native AST/structural/heuristic ladder.
3219
+ - Scope: SCIP files are local inputs; import is deterministic, bounded, and
3220
+ optional. Live LSP integration is explicitly not part of C84.
3221
+ - Evidence: checked-in SCIP fixtures and expected graph facts, malformed/path/
3222
+ oversize/conflict tests, fallback replay, verifier, and resource report.
3223
+
3224
+ Delivered evidence: `knodin index --scip <file>` explicitly imports a bounded
3225
+ local SCIP protobuf snapshot without discovering or running a producer. The
3226
+ npm-integrity-pinned `@sourcegraph/scip-typescript@0.4.0` fixture, binary hash,
3227
+ source, configuration, expected facts, duplicate document-local symbols, and
3228
+ adversarial cases are under `fixtures/c84-scip/` and `src/__tests__/unit/`.
3229
+ The raw offline replay, distinct relationship kinds, native/LSIF fallback,
3230
+ stable identity refresh, 64 MiB/10,000-file/250,000-fact/30-second bounds, and
3231
+ 502,431,744-byte incremental RSS record are under
3232
+ `benchmarks/evaluations/c84-scip-import/`; reproduce them with
3233
+ `npm run verify:c84`. The evidence report is
3234
+ [`C84 optional SCIP import evidence`](../docs/evidence/c84-optional-scip-import-2026-08-03.md).
3235
+
3236
+ Acceptance:
3237
+
3238
+ 1. Pinned fixtures produce deterministic stable identities, relationships,
3239
+ source ranges, and `scip` provenance. Conflicting tiers surface ambiguity and
3240
+ never silently replace higher-confidence or fresher evidence.
3241
+ 2. Missing, malformed, hostile-path, stale, and oversized inputs fail safely;
3242
+ repositories without SCIP retain current native/LSIF behavior with no daemon,
3243
+ account, network, source egress, or model/API spend.
3244
+ Incremental spend is zero.
3245
+ 3. Import and query replays remain within explicit file/fact/time limits and
3246
+ below 1 GB incremental RSS. Raw local results, verifier output, and known
3247
+ language/indexer coverage are checked in.
3248
+
3249
+ Known limitations and decision gate: SCIP precision is bounded by the producer
3250
+ and languages represented in the imported index. Any live LSP tier requires a
3251
+ separately approved roadmap item and must never become a baseline dependency.
3252
+
3253
+ ### C85 — Bounded git-history review signals
3254
+
3255
+ - Status: implemented
3256
+ - Priority: P1
3257
+ - Disposition: Must close
3258
+ - DependsOn: C24, C44, C81
3259
+ - Touches: review risk serialization, bounded Git history queries, unavailable/
3260
+ shallow states, synthetic repository fixtures, telemetry, and docs.
3261
+ - What: add churn, co-change, and coupling as separately itemized review signals
3262
+ beside graph impact, test gaps, and structural centrality; forbid opaque scores.
3263
+ - Scope: all Git operations have explicit commit/file/time bounds and preserve
3264
+ deterministic ordering, source evidence, and graceful unavailable states.
3265
+ - Evidence: checked-in synthetic Git DAG, exact signal oracles, bounded-command
3266
+ tests, raw review replay, verifier, and resource report.
3267
+ - Evidence delivered: `src/engine/git-history.ts` and
3268
+ `src/__tests__/unit/git-history-signals.spec.ts` cover rename-aware itemized
3269
+ facts, merge/unrelated separation, shallow deepen, unborn, missing Git,
3270
+ non-repository, arbitrary HEAD failure, timeout, path containment, hard bounds,
3271
+ deterministic truncation, cache validity, and CLI/MCP serialization.
3272
+ `benchmarks/evaluations/c85-git-history/` checks in the exact synthetic DAG,
3273
+ raw replay, latency/command/resource evidence, limitations, and verifier;
3274
+ `npm run verify:c85` reproduces churn 2, co-change 2, coupling 1, three bounded
3275
+ history commands, zero locality costs, and peak RSS below 1 GB.
3276
+
3277
+ Acceptance:
3278
+
3279
+ 1. Fixtures cover renames, merges, unrelated co-occurrence, shallow history,
3280
+ unborn repositories, missing Git, and bounded truncation. Each output exposes
3281
+ the contributing facts and omissions rather than only a composite number.
3282
+ 2. Review remains deterministic, fail-closed, and useful when history is absent;
3283
+ commands cannot escape the repository or exceed configured commit/file/time
3284
+ limits. CLI/MCP-compatible evidence retains freshness and confidence.
3285
+ 3. Checked-in replay reports correctness, latency, command counts, and peak RSS
3286
+ below 1 GB with zero spend, no network, no credentials, and no source egress.
3287
+
3288
+ Known limitations and decision gate: co-change is correlation, not causation.
3289
+ History signals may influence itemized review context but cannot independently
3290
+ assert defect likelihood, ownership, or mandatory remediation.
3291
+
3292
+ ### C86 — Static-embedding bake-off
3293
+
3294
+ - Status: evaluated — retained MiniLM unchanged
3295
+ - Priority: P2
3296
+ - Disposition: Evaluate
3297
+ - DependsOn: C46, C51, C80, C81
3298
+ - Touches: isolated Model2Vec/MiniLM evaluation runner, fixed corpus and labels,
3299
+ model provenance/digests, raw measurements, verifier, and decision record.
3300
+ - What: evaluate a pinned static-embedding candidate against retained MiniLM
3301
+ before authorizing any runtime, index-format, or default-search change.
3302
+ - Scope: measure cold/warm indexing and query performance, RSS, storage, recall,
3303
+ NDCG, and downstream task correctness with at least three repetitions.
3304
+ - Evidence: checked-in corpus/query labels, model identifiers and digests, raw
3305
+ repeated results, verifier output, confidence/variance report, and disposition.
3306
+ - Evidence delivered: `benchmarks/evaluations/c86-static-embeddings/` contains
3307
+ the pinned 20-document corpus, ten relevance/task labels, model revisions and
3308
+ ONNX digests, six fresh-process raw repetitions, cold/warm timing, RSS and
3309
+ storage measurements, descriptive confidence/variance, unsupported cases,
3310
+ limitations, and the retain-MiniLM decision. `npm run verify:c86` checks the
3311
+ identical-route/budget oracle, three repetitions per arm, local/offline and
3312
+ zero-spend contract, sub-1-GB incremental RSS, artifact pins, and evaluation
3313
+ isolation without requiring either model artifact.
3314
+
3315
+ Acceptance:
3316
+
3317
+ 1. Both candidates run on identical local fixtures, retrieval routes, budgets,
3318
+ and deterministic relevance/task oracles. Record setup, unsupported cases,
3319
+ cold/warm timing, index size, recall/NDCG, correctness, and peak RSS.
3320
+ 2. The replay is offline after documented local preparation, adds no production
3321
+ dependency, uses no hosted API/account/source egress, incurs zero spend, and
3322
+ stays below 1 GB incremental memory. Pin model identity and artifact digest.
3323
+ 3. Retain MiniLM unless the pre-registered decision gate shows a reproducible
3324
+ correctness-preserving win. Close by retaining MiniLM, deferring/rejecting
3325
+ Model2Vec, or proposing a separately reviewed implementation item; evaluation
3326
+ alone must not change production behavior or schema.
3327
+
3328
+ Known limitations and decision gate: corpus results may not generalize to every
3329
+ language or repository. Publish only measured local outcomes, never a universal
3330
+ embedding-quality or product-superiority claim.
3331
+
3332
+ Decision: retain MiniLM unchanged. On the checked-in fixture both models reached
3333
+ 1.000 recall@5, 1.000 NDCG@10, and 10/10 top-1 task correctness in every
3334
+ repetition. Model2Vec was repeatably faster and smaller locally, but the corpus
3335
+ does not resolve repository-scale or multilingual-query quality, so the
3336
+ pre-registered gate's no-material-unsupported-category condition did not pass.
3337
+ No production runtime, dependency, index format, schema, default, or MCP surface
3338
+ changed, and no separately numbered implementation item is authorized.
3339
+
3340
+ ### C87 — Structural cold-start fast path
3341
+
3342
+ - Status: implemented
3343
+ - Priority: P0
3344
+ - Disposition: Must close
3345
+ - DependsOn: C56, C68
3346
+ - Touches: packaged CLI launcher, structural snapshot publication, direct-file
3347
+ fallback, compression routing, benchmarks, and truth-bound response fields.
3348
+ - What: bypass graph/model initialization for structural-only cold requests
3349
+ while preserving the persistent graph route for relationship intelligence.
3350
+ - Scope: file outline, exact file-scoped extraction, batch outline, project
3351
+ overview, and pure output compression only.
3352
+ - Evidence: `src/structural-snapshot.ts`, `src/structural-fast-path.ts`,
3353
+ `src/pure-compression-cli.ts`, focused unit tests, and
3354
+ `benchmarks/evaluations/c87-structural-cold-start/raw-results.json`.
3355
+
3356
+ Indexing now publishes an atomic compact file/symbol snapshot. The packaged
3357
+ launcher routes file outlines, exact file-scoped extraction, batch outlines,
3358
+ project overview, and pure compression without importing the graph or embedding
3359
+ modules. A missing or fingerprint-stale requested file is parsed directly and
3360
+ labeled `lexical-structural-fallback`; graph-backed search, impact, review,
3361
+ relationships, and architecture retain their existing initialization and
3362
+ freshness contract.
3363
+
3364
+ Acceptance:
3365
+
3366
+ 1. File-outline, exact-symbol, and project-overview fresh-process p50 are
3367
+ at most 50 ms and p95 at most 75 ms after separating benchmark-parent spawn/
3368
+ scheduling overhead from measured child-process lifetime. Preserve both raw
3369
+ observations; do not hide an application or runtime-floor miss.
3370
+ 2. Warm p50 regresses by no more than 10%, response tokens by no more than 5%,
3371
+ and the existing shared correctness oracles remain 100% for both fast and
3372
+ graph routes.
3373
+ 3. Every fast answer includes path/fingerprint evidence, exact locations,
3374
+ evidence quality, `graphEnriched: false`, file-scoped uniqueness limits, and
3375
+ an upgrade handle. Dependency evidence proves no graph/model import.
3376
+ 4. Graph-backed operations retain stable identity, relationships, and freshness
3377
+ evidence. Snapshot absence/staleness never becomes a repository-wide
3378
+ freshness claim. Checked-in replay evidence requires zero hosted spend,
3379
+ network, account, credential, or source egress.
3380
+
3381
+ Known limitations and decision gate: the direct fallback is a bounded lexical
3382
+ structural parser, not the full Tree-sitter graph extractor. Its classification
3383
+ is explicit and a later graph operation is the supported enrichment path. Keep
3384
+ the fast route only while all checked-in acceptance gates pass.
3385
+
3386
+ ### C88 — Profile-based contained execution
3387
+
3388
+ - Status: implemented — macOS certified; Linux/Windows unavailable
3389
+ - Priority: P1
3390
+ - Disposition: Must close
3391
+ - DependsOn: C57, C58, C59
3392
+ - Touches: one-tool schema/dispatcher, profile configuration, macOS Seatbelt
3393
+ adapter, output compression/diagnosis, audit records, adversarial evidence,
3394
+ and security documentation.
3395
+ - What: permit only preapproved single-process profiles behind enforceable
3396
+ macOS filesystem/network/process boundaries and fail closed elsewhere.
3397
+ - Scope: test, lint, typecheck, and build-style profiles with immutable argv;
3398
+ arbitrary commands, child processes, and portable containment claims remain
3399
+ excluded.
3400
+ - Evidence: `src/execution-profile.ts`, `docs/CONTAINED-EXECUTION.md`, gateway/
3401
+ adversarial unit tests, and `benchmarks/evaluations/c88-contained-execution/`.
3402
+
3403
+ C58's portable polling prototype remains rejected. C88 authorizes a narrower
3404
+ production contract: independent global and repository enablement, immutable
3405
+ repository profile identity, absolute configured executable and fixed argv,
3406
+ no shell/caller suffix/cwd/environment, sanitized environment, repository cwd,
3407
+ hard timeout/output/fork limits, private profile-only audit, recoverable
3408
+ compression, and failure-to-code diagnosis.
3409
+
3410
+ Acceptance:
3411
+
3412
+ 1. macOS native evidence proves literal metacharacter argv, denied ambient
3413
+ credentials/user-data reads, denied network sockets, denied forks,
3414
+ repository-only authorization, closed stdin, output/timeout termination, and
3415
+ no raw command text in audit/telemetry.
3416
+ 2. `fully-contained` is reported only by the certified macOS Seatbelt adapter.
3417
+ Linux and Windows report `containment-unavailable` and spawn nothing until
3418
+ independent native adapters and evidence exist.
3419
+ 3. MCP accepts only `profile`; it cannot accept an executable, arguments, cwd,
3420
+ environment, or limits. Every result names the profile, actual containment,
3421
+ configured bounds, exit state, compressed artifact, and diagnosis state.
3422
+ 4. Preserve the one-tool schema reduction gate and local-only operation. The
3423
+ checked-in evidence requires zero hosted spend and zero source egress while
3424
+ preserving the original C58 limitations. Do not describe deprecated
3425
+ `sandbox-exec` as a supported Apple public security API.
3426
+
3427
+ Known limitations and decision gate: profiles are single-process
3428
+ (`maxProcesses: 1`), system runtime paths remain readable, approved repository
3429
+ code can modify its checkout, and Linux/Windows execution is unavailable. Keep
3430
+ the surface only while native evidence proves the stated boundary; any broader
3431
+ profile or platform requires a separately certified adapter. Preserve no unrestricted command surface.
3432
+
3433
+ ### C89 — MCP request reliability and recovery
3434
+
3435
+ - Status: implemented
3436
+ - Priority: P0
3437
+ - Disposition: Must close
3438
+ - DependsOn: C3, C16, C53
3439
+ - Touches: MCP server lifecycle, request dispatcher, cancellation/progress
3440
+ bridge, durable diagnostic log, error taxonomy, and adversarial transport tests.
3441
+ - What: prevent one expensive or stuck operation from ending the useful MCP
3442
+ session and make every failure attributable and recoverable.
3443
+ - Scope: per-operation deadlines, client cancellation, progress notifications,
3444
+ request and trace IDs, bounded exit-safe logs, stuck-request isolation, and
3445
+ distinct crash, lock, memory-pressure, timeout, and client-disconnect errors.
3446
+ - Evidence: `src/mcp-worker-supervisor.ts`, `src/mcp-graph-worker.ts`,
3447
+ `src/mcp-reliability.ts`, `src/server.ts`, `docs/MCP.md`, and the checked-in
3448
+ supervisor/lifecycle/CLI replays prove versioned IPC, real held-lock handling,
3449
+ abrupt worker exit, the observed `Transport closed` disconnect cancellation
3450
+ race, hard termination, bounded restart, predecessor-to-next-success
3451
+ correlation, warm reuse, source-safe
3452
+ bounded journals, and linked-worktree isolation without hosted dependencies.
3453
+
3454
+ Acceptance:
3455
+
3456
+ 1. Every request receives stable request/trace correlation, an operation-aware
3457
+ deadline, and cancellation that releases its resources without corrupting the
3458
+ graph; expensive operations emit bounded truthful progress when supported.
3459
+ 2. A killed worker, held graph lock, memory-pressure sentinel, deadline, and
3460
+ disconnected client each produce a distinct actionable error. The server
3461
+ remains usable or restarts automatically without retry loops or lost logs.
3462
+ 3. Durable logs survive abrupt process exit, rotate within documented byte/time
3463
+ bounds, omit repository source and secrets, and correlate progress, failure,
3464
+ cleanup, restart, and the next successful request.
3465
+ 4. Checked-in CLI/MCP parity, concurrency, cancellation-race, forced-exit, and
3466
+ recovery replays pass locally with zero hosted spend, credentials, network,
3467
+ or source egress.
3468
+
3469
+ Known limitations and decision gate: deadlines and memory-pressure attribution
3470
+ are bounded diagnoses, not proof of an operating-system root cause. Do not claim
3471
+ transparent recovery unless the next independent request succeeds in replay.
3472
+
3473
+ ### C90 — Privacy-safe diagnostic support bundle
3474
+
3475
+ - Status: implemented
3476
+ - Priority: P0
3477
+ - Disposition: Must close
3478
+ - DependsOn: C16, C59, C89
3479
+ - Touches: diagnostics CLI, MCP trace store, lifecycle/graph health, redaction,
3480
+ archive manifest, preview/report rendering, and privacy adversarial tests.
3481
+ - What: turn local diagnostics into one support-ready bundle whose exact shared
3482
+ contents users can inspect before export.
3483
+ - Scope: recent MCP traces/failures, lifecycle and graph health, versions and
3484
+ runtime, repository scale without source, manifest, human report, and preview.
3485
+ - Evidence: `src/diagnostics.ts`, `src/diagnostics-write-helper.ts`,
3486
+ `schemas/support-bundle-v2.schema.json`, the C90 adversarial fixture and unit
3487
+ replays, and packed-install preview/archive/inspect smoke prove a fail-closed
3488
+ allowlist, exact content-addressed preview parity, bounded C89 trace recovery,
3489
+ honest unavailable fields, linked-worktree isolation, bounded retention, and
3490
+ repository-bound no-follow writes without hosted dependencies or source
3491
+ egress.
3492
+
3493
+ Acceptance:
3494
+
3495
+ 1. Preview and final manifest enumerate every file and field, byte size,
3496
+ retention window, redaction applied, and unavailable section before sharing;
3497
+ the concise report includes request/trace IDs and actionable recovery steps.
3498
+ 2. Adversarial fixtures containing source, diffs, paths outside the repository,
3499
+ credentials, environment values, usernames, and command output prove none can
3500
+ enter the default bundle. Repository scale is aggregate metadata only.
3501
+ 3. The bundle includes exit-surviving C89 traces and recent classified failures,
3502
+ lifecycle/graph health, knodin/Node/platform versions, and honest missing or
3503
+ stale states without initiating network activity.
3504
+ 4. Checked-in deterministic CLI/MCP tests prove preview/archive parity, bounded
3505
+ size and retention, zero hosted spend, no credentials, and no source egress.
3506
+
3507
+ Known limitations and decision gate: automated redaction cannot certify arbitrary
3508
+ future fields. New diagnostic fields fail closed until included in the explicit
3509
+ allowlist and privacy oracle; sending a bundle remains a user action.
3510
+
3511
+ ### C91 — Boring macOS installation and runtime handoff
3512
+
3513
+ - Status: parked-external-evidence — 2026-08-05
3514
+ - Previous status: owner-approved adoption priority
3515
+ - Priority: P0
3516
+ - Disposition: Must close
3517
+ - DependsOn: C52, C73, C89, C90
3518
+ - Touches: package/install CI, runtime resolver, doctor repairs, client setup
3519
+ guides, macOS acceptance fixtures, and package compatibility checks.
3520
+ - What: make install through first useful MCP call repeatable on environments the
3521
+ project can actually validate.
3522
+ - Scope: macOS; Node 20 project handoff to an available Node 24 runtime; mise,
3523
+ nvm, fnm, Volta, asdf, Homebrew, npm, and Artifactory; Copilot, Claude, Gemini,
3524
+ and Codex; safe doctor repairs.
3525
+ - Evidence: checked-in automated install matrix, package-content/compatibility
3526
+ gates, isolated runtime-manager fixtures, and timed acceptance receipts.
3527
+
3528
+ Acceptance:
3529
+
3530
+ 1. Each named installation path is exercised from a clean macOS fixture through
3531
+ install, `knodin init`, `knodin status`, one useful MCP call, and uninstall or
3532
+ rollback; package compatibility is checked automatically in CI.
3533
+ 2. A Node 20 project reliably hands execution to an already available compatible
3534
+ Node 24 runtime without modifying the project's declared runtime. Missing or
3535
+ conflicting runtimes produce one actionable diagnostic.
3536
+ 3. Copilot, Claude, Gemini, and Codex each complete the same acceptance flow in
3537
+ no more than two minutes under the documented warm/cold boundary, with exact
3538
+ client/version/configuration evidence and no universal client claim.
3539
+ 4. Doctor applies only enumerated idempotent repairs with preview and rollback;
3540
+ checked-in tests require zero hosted spend for local fixtures and no source
3541
+ egress. Credentialed Artifactory evidence is recorded only when available.
3542
+
3543
+ Known limitations and decision gate: Linux and Windows certification are not a
3544
+ goal and unavailable platform evidence must stay unavailable. Runtime-manager
3545
+ fixtures prove supported configurations, not every shell customization.
3546
+
3547
+ Code-readiness on 2026-08-05 added a merge-preserving GitHub Copilot/VS Code
3548
+ MCP adapter, doctor visibility, idempotent CLI-only reversal, and fail-closed
3549
+ malformed/symlink handling. Certification remains parked until native macOS
3550
+ install-to-uninstall receipts exist for nvm, fnm, Volta, and asdf; an exact
3551
+ GitHub Copilot/VS Code timed receipt joins the available Claude, Gemini, and
3552
+ Codex versions; a published Homebrew formula and credentialed Artifactory path
3553
+ are exercised where authorized; and the same four clients complete the bounded
3554
+ install, `init`, `status`, useful MCP call, and rollback flow. Doctor also still
3555
+ needs an enumerated transactional preview/apply/rollback implementation and
3556
+ acceptance replay. No fixture result may be promoted to this missing native
3557
+ evidence, and C94 remains dependency-blocked.
3558
+
3559
+ ### C92 — Unified five-channel release orchestration
3560
+
3561
+ - Status: parked-external-evidence — 2026-08-06
3562
+ - Previous status: owner-approved release priority
3563
+ - Priority: P0
3564
+ - Disposition: Must close
3565
+ - DependsOn: C64, C91
3566
+ - Touches: release workflows, preflight, immutable package staging, npm,
3567
+ Artifactory, Homebrew, GitHub SaaS/GHES synchronization, receipts, and retry tests.
3568
+ - What: replace fragmented publication paths with one fail-closed orchestrator
3569
+ that promotes one immutable tarball everywhere.
3570
+ - Scope: npm, Artifactory, Homebrew, GitHub SaaS, and the read-only GHES mirror;
3571
+ availability preflight, private-repository attestations, retries, and receipts.
3572
+ - Evidence: checked-in dry-run channel adapters, fault matrix, immutable digest
3573
+ oracle, retry replay, receipt schema, and private-repository attestation fixture.
3574
+
3575
+ Acceptance:
3576
+
3577
+ 1. Before any publication, preflight proves every required channel, credential,
3578
+ permission, target version, and attestation mode is available; one failure
3579
+ publishes nothing and leaves a machine-readable failure receipt.
3580
+ 2. One staged tarball digest and length are promoted to all five channels.
3581
+ Retries reconcile idempotently without moving or recreating source tags and
3582
+ refuse any channel whose existing bytes differ.
3583
+ 3. Private-repository attestations follow the supported trust path rather than
3584
+ assuming public-repository identity behavior. The final receipt records each
3585
+ channel's artifact identity, digest, status, attempt, and authoritative URL.
3586
+ 4. Checked-in fault-injection tests cover every preflight/publish boundary with
3587
+ zero real publication or spend. Live channel receipts remain Track B evidence
3588
+ and cannot be simulated by this local implementation item.
3589
+
3590
+ Known limitations and decision gate: dry-run adapters do not certify production
3591
+ availability. GitHub SaaS remains authoritative; GHES is a read-only mirror and
3592
+ must never be forced. Stop on divergence.
3593
+
3594
+ Parking decision: C92 remains open behind C64 and C91. Ordinary 0.x publishing
3595
+ continues through the existing fail-closed tag workflow; this does not claim the
3596
+ unimplemented unified five-channel orchestrator or manufacture channel receipts.
3597
+
3598
+ ### C93 — Paired engineering-outcome measurement
3599
+
3600
+ - Status: implemented
3601
+ - Priority: P1
3602
+ - Disposition: Must close
3603
+ - DependsOn: C24, C81
3604
+ - Touches: competitive runner, task corpus, patch evaluator, dependency oracle,
3605
+ agent telemetry import, cost accounting, and outcome report.
3606
+ - What: measure whether knodin improves completed engineering work, not whether
3607
+ it merely emits fewer tokens.
3608
+ - Scope: identical with/without-knodin tasks measuring task success,
3609
+ missed-dependency rate, corrective round trips, time to first correct edit,
3610
+ patch-application success, total billed cost, and performance.
3611
+ - Evidence: preregistered checked-in corpus/oracles, paired raw runs, variance and
3612
+ failure accounting, verifier, and a limitations-bound report.
3613
+ - Evidence delivered: `benchmarks/evaluations/c93-engineering-outcomes/`
3614
+ contains the preregistered two-arm corpus, executable dependency/patch oracle,
3615
+ source-free raw event trajectories, deterministic local runner, paired
3616
+ variance/uncertainty report, and visible limitations. `npm run verify:c93`
3617
+ recreates all 12 runs from the identical fixture commit on isolated linked
3618
+ worktrees and non-default branches, verifies independent arm commits and a
3619
+ clean base checkout, and records zero hosted spend, observed network requests,
3620
+ or source egress plus explicit process-isolation bounds and unavailable—not
3621
+ inferred—token and billed-cost data.
3622
+
3623
+ Acceptance:
3624
+
3625
+ 1. Both arms use identical repositories, commits, tasks, model/client versions,
3626
+ prompts, permissions, time limits, repetitions, and correctness oracles; the
3627
+ only intended treatment difference is knodin availability/instructions.
3628
+ 2. Every named outcome is reported per task and in aggregate, including failures,
3629
+ censored time, unavailable token/cost data, and confidence/variance. Token
3630
+ counts may be diagnostic but are never the success metric.
3631
+ 3. Dependency and patch oracles inspect the resulting edit and tests rather than
3632
+ grading prose. Raw events support auditable round-trip and first-correct-edit
3633
+ reconstruction without storing repository source in telemetry.
3634
+ 4. The local harness and synthetic fixtures are checked in and run with zero hosted
3635
+ spend; any paid model replay records actual total billed cost and is
3636
+ never inferred from tokens alone.
3637
+ Checked-in evidence requires zero hosted spend.
3638
+
3639
+ Known limitations and decision gate: a bounded corpus cannot prove universal
3640
+ product superiority. Publish only measured effect sizes and uncertainty; a null
3641
+ or negative outcome must remain visible and may require product rollback.
3642
+
3643
+ ### C94 — First-use clarity and five-minute demonstration
3644
+
3645
+ - Status: parked-external-evidence — 2026-08-06
3646
+ - Previous status: owner-approved clarity priority
3647
+ - Blocked by: C91 external certification
3648
+ - Priority: P1
3649
+ - Disposition: Must close
3650
+ - DependsOn: C89, C90, C91
3651
+ - Touches: MCP/CLI help, status responses, client guides, examples, demo fixture,
3652
+ freshness/error copy, and documentation integrity tests.
3653
+ - What: make the first session explain the tool identity, next action, evidence
3654
+ limits, freshness, and recovery without requiring prior product knowledge.
3655
+ - Scope: `knodin.knodin` naming, operation discovery/examples, next-useful-action
3656
+ recommendations, concise limitations/recovery, and a real-repository demo.
3657
+ - Evidence: checked-in golden help/status/error output, four-client guide checks,
3658
+ timed demo transcript, and source-evidence links.
3659
+
3660
+ Acceptance:
3661
+
3662
+ 1. Setup docs explain that clients display `knodin.knodin` as server name plus
3663
+ tool name. Schema/help gives copyable examples and makes operation discovery
3664
+ possible without adding MCP tools.
3665
+ 2. Status recommends a context-sensitive next useful operation and presents
3666
+ freshness, availability, known bounds, and repair/reindex actions concisely;
3667
+ no healthy, stale, ambiguous, or repair-needed state is conflated.
3668
+ 3. A checked-in five-minute demonstration runs `knodin init`, status, orientation,
3669
+ source-evidenced impact/review, and bounded context on a real repository and
3670
+ explicitly demonstrates value, ROI, tomorrow's action, and secret sauce.
3671
+ 4. Golden CLI/MCP and docs-integrity tests run locally with zero hosted spend,
3672
+ no credentials, and no source egress.
3673
+
3674
+ Known limitations and decision gate: a scripted demonstration is onboarding
3675
+ evidence, not an outcome or superiority benchmark. Client UI labels may vary by
3676
+ version and must name the version actually checked.
3677
+
3678
+ Parking decision: C94 remains open behind C91. The existing local help and
3679
+ documentation remain available, but no fixture is promoted into the missing
3680
+ native client demonstration or certification evidence.
3681
+
3682
+ ### C95 — Defensible behavioral-contract replay
3683
+
3684
+ - Status: implemented
3685
+ - Priority: P1
3686
+ - Disposition: Must close
3687
+ - DependsOn: C3, C24, C83
3688
+ - Touches: contract manifest, cross-operation fixtures, freshness/identity/
3689
+ evidence/budget/compression/diagnosis/privacy tests, and product positioning.
3690
+ - What: make the integrated behavior—not individual graph feature count—the
3691
+ product's checked and release-gated contract.
3692
+ - Scope: truthful budgets, explicit freshness/availability, source evidence,
3693
+ ambiguity-safe stable identity, recoverable compression, failure-to-symbol
3694
+ diagnosis, and local operation without source egress.
3695
+ - Evidence: `contracts/behavior-contract-v1.json`,
3696
+ `benchmarks/evaluations/c95-behavior-contract/runner.ts`,
3697
+ `scripts/verify-c95.ts`, `docs/BEHAVIORAL-CONTRACT.md`, and the release
3698
+ preflight gate map and dispatch all 22 production operations plus every
3699
+ schema-declared query/action discriminator against linked-worktree fixtures.
3700
+ The replays cover hard-budget continuation, stable evidence,
3701
+ recoverable compression, freshness and ambiguity qualification, diagnosis,
3702
+ and observed local privacy across fetch, HTTP(S), TCP/TLS, DNS, native-addon,
3703
+ and subprocess boundaries. All 458 mapped fixture/clause rows have isolated
3704
+ pre-serialization production-response mutations that fail the same release
3705
+ oracle; omitted or mutated clauses fail closed.
3706
+
3707
+ Acceptance:
3708
+
3709
+ 1. Every production operation is mapped to the applicable contract clauses and
3710
+ adversarial fixtures prove honest truncation, stale/unavailable states,
3711
+ evidence provenance, ambiguity, recovery handles, diagnosis confidence, and
3712
+ privacy behavior.
3713
+ 2. The replay verifies composition: a bounded response can be expanded by a
3714
+ stable handle without changing identity/evidence, and stale or ambiguous
3715
+ evidence cannot become a confident diagnosis or impact claim.
3716
+ 3. Release preflight fails on a contract regression. Product documentation
3717
+ answers value, ROI, tomorrow, and secret sauce with checked-in evidence and
3718
+ preserves every roadmap limitation.
3719
+ 4. Tests and replays run locally with zero hosted spend, credentials, network,
3720
+ or source egress; no universal superiority claim is inferred.
3721
+
3722
+ Known limitations and decision gate: the contract proves the checked operation
3723
+ and action fixtures, not all repositories or clients. Network construction is
3724
+ denied before request-body streaming, and opaque transport inside an
3725
+ already-loaded native library remains unobservable. A new operation or action
3726
+ must declare its clauses, positive response semantics, and isolated production
3727
+ mutation before release.
3728
+
3729
+ ### C96 — Replay-gated capability discipline
3730
+
3731
+ - Status: evaluated — deferred
3732
+ - Priority: P2
3733
+ - Disposition: Evaluate
3734
+ - DependsOn: C24, C93, C95
3735
+ - Touches: roadmap policy, proposal template, outcome replay, one narrow
3736
+ cross-substrate fixture, verifier, and decision record.
3737
+ - What: require evidence of correctness or workflow improvement before expanding
3738
+ production capability, beginning with one bounded cross-substrate impact case.
3739
+ - Scope: one relationship crossing two explicitly named substrates; no broad
3740
+ platform, language, repository, or superiority promise.
3741
+ - Evidence: checked-in baseline/treatment fixture, exact dependency and task
3742
+ oracles, cost/resource report, limitations, and retain/defer/implement decision.
3743
+
3744
+ Acceptance:
3745
+
3746
+ 1. Preregister one concrete missed-dependency or workflow failure and the exact
3747
+ improvement threshold before implementation. Compare unchanged knodin with a
3748
+ bounded prototype on the same task using C93 outcomes and C95 contracts.
3749
+ 2. The fixture names supported substrates, edge provenance, ambiguity behavior,
3750
+ freshness, budgets, and false-positive/negative oracles; unsupported crossings
3751
+ remain explicit rather than generalized.
3752
+ 3. Close by retaining knodin unchanged, deferring/rejecting the capability, or
3753
+ creating a separately numbered implementation item only when the gate passes.
3754
+ 4. Evaluation artifacts and verifier are checked in, require zero hosted spend,
3755
+ credentials, network, or source egress, and preserve losing results.
3756
+
3757
+ Known limitations and decision gate: one passing cross-substrate fixture does
3758
+ not authorize a broad platform claim. Implementation requires a separately
3759
+ reviewed item and must preserve the behavioral contract.
3760
+
3761
+ Decision: defer production expansion. C93 found no correctness advantage on its
3762
+ narrow paired fixture, and C95 already requires every new operation or action to
3763
+ declare and replay its behavioral clauses. No preregistered cross-substrate
3764
+ missed-dependency case currently justifies a prototype, so knodin remains
3765
+ unchanged. Reopen under a separately numbered item only when a concrete failure,
3766
+ supported substrates, exact oracle, and improvement threshold are preregistered.
3767
+ Evidence: `benchmarks/evaluations/c93-engineering-outcomes/verification.json`,
3768
+ `contracts/behavior-contract-v1.json`, and
3769
+ `benchmarks/evaluations/c95-behavior-contract/runner.ts`.
3770
+
3771
+ ### C97 — One-process semantic residency reclamation
3772
+
3773
+ - Status: evaluated — retained existing design
3774
+ - Priority: P1
3775
+ - Disposition: Evaluate
3776
+ - DependsOn: C46, C51, C80, C87
3777
+ - Touches: evaluation and decision evidence only; no runtime, status, API, or
3778
+ storage change.
3779
+ - What: test whether idle model/vector/HNSW eviction can make a repeatable OS-RSS
3780
+ guarantee while preserving Knodin's one-local-process architecture.
3781
+ - Evidence: `benchmarks/evaluations/c97-one-process-semantic-residency/` and
3782
+ `scripts/verify-c97.ts` bind 10-trial fresh-process measurements to the runner,
3783
+ engine, package, and embedding-test sources.
3784
+
3785
+ Decision: retain the current lazy local model and bounded generation-scoped
3786
+ caches. No in-process candidate repeatedly reclaimed both 32 MiB and 15% of its
3787
+ attributable semantic increment across two corpus tiers. ONNX disposal and
3788
+ allocator behavior were nondeterministic; `worker_threads` retained the shared
3789
+ native address space; q8 lacks the required C28 retrieval-equivalence replay;
3790
+ fp16 failed to initialize. Child-process isolation is explicitly rejected
3791
+ because it would violate the one-local-process product contract. Do not claim
3792
+ idle RSS reclamation from model disposal or cache clearing on this evidence.
3793
+
3794
+ Known limitations: resident semantic search retains a measurable one-process
3795
+ RSS cost after first use. Structural-first routing still avoids loading the model
3796
+ when evidence is sufficient, and decoded-vector/HNSW caches remain bounded and
3797
+ generation-invalidated. Reopen only with a portable allocator/runtime mechanism
3798
+ that passes the unchanged absolute, percentage, p95, retrieval, and 20-cycle
3799
+ gates.
3800
+
3801
+ Final-stack refresh: C101 reran the full C97 measurement after C98–C100 changed
3802
+ the shared engine. Some individual cache/ONNX trials crossed the memory
3803
+ threshold, but no production candidate passed all 10 trials; the retained
3804
+ one-process decision therefore remains unchanged. The refreshed raw artifact
3805
+ replaces the mutable C97 measurement file and `verify:c97` recomputes this gate.
3806
+
3807
+ ## Explicit non-goals from the GitNexus comparison
3808
+
3809
+ - Do not split knodin's gateway into one MCP schema per capability; the one-tool
3810
+ surface is a measured differentiator.
3811
+ - Do not copy unrestricted Cypher directly into the MCP surface.
3812
+ - Do not require Milvus, Qdrant, Ollama, credentials, or any hosted service;
3813
+ Claude Context and mcp-codebase-index validate this differentiator.
3814
+ - Do not add one MCP tool per capability; Serena's 29-tool and
3815
+ code-review-graph's 30-tool surfaces reinforce the one-gateway advantage.
3816
+ - Do not make vector visualization, agent memory, or skill generation core graph
3817
+ requirements. C35 is limited to deterministic local architecture/call-flow
3818
+ artifacts with a concrete developer workflow; it does not authorize a hosted
3819
+ visualization service or a generalized vector-visualization feature.
3820
+ - Do not enable network embeddings, hosted indexing, telemetry, or source-code
3821
+ egress to chase competitor features.
3822
+ - Do not claim API-shape or cross-repository superiority until dedicated local
3823
+ fixtures produce evidence.
3824
+ - Do not interpret every competitor option as a product requirement; parameters
3825
+ that merely select nonexistent groups, branches, or services need real
3826
+ fixtures before roadmap promotion.
3827
+
3828
+ ## Verification bar
3829
+
3830
+ Every implementation item must run the repository's discovered gates and add a
3831
+ targeted acceptance fixture. At minimum:
3832
+
3833
+ ```bash
3834
+ bun run typecheck
3835
+ bun run lint
3836
+ bun run test
3837
+ ```
3838
+
3839
+ Competitive claims must also re-run the relevant harness under
3840
+ `benchmarks/competitors/` and update its result file without deleting losing
3841
+ cases.
3842
+
3843
+ ## Competitive replay — 2026-07-24 (Phase 2)
3844
+
3845
+ Fresh, provenance-bound replays of every dimension were run against the real
3846
+ competitor CLIs and saved as **new** immutable artifacts (never overwriting
3847
+ prior results): search `raw-results-20260724T182246Z.json`, review + architecture
3848
+ `…T182448Z.json`, and symbols/editing/visualization/context/lifecycle
3849
+ `…T182707Z.json` under `benchmarks/evaluations/competitive-*/`.
3850
+
3851
+ Applying the measurement contract (a WIN needs equivalent operation/fixture/mode/
3852
+ scope/budget/oracle **and** an oracle-qualified row; unavailable/unverified/
3853
+ incomparable rows are excluded from win claims):
3854
+
3855
+ ### Where knodin is ahead or level (oracle-qualified, comparable)
3856
+
3857
+ - **Diff review — WIN.** traversal recall 1.0 and 733 B vs gitnexus (recall 1.0,
3858
+ 1458 B), code-review-graph (recall 0, 2269 B), graphify (recall 0.667).
3859
+ - **Architecture — WIN.** full package coverage + hubs + bridges; every completed
3860
+ competitor is at best equal on coverage, none better.
3861
+ - **Search — WIN vs the only comparable competitor.** nDCG 0.855 vs gitnexus
3862
+ 0.677; dead-code precision 1.0 vs 0.667. codegraph (probe-only, unscored),
3863
+ grepai (MCP probe blocked), claude-context (needs Ollama+Milvus) are
3864
+ contract-excluded.
3865
+ - **Symbols — WIN/level.** precision 1 / recall 1: ties gitnexus (1/1), beats
3866
+ codegraph (0/0), serena (0/0), codebase-memory (1/0.5).
3867
+ - **Lifecycle capability — WIN.** indexed/queryable/closed/stale-detection/
3868
+ per-query telemetry all qualified; codebase-memory and grepai advertise
3869
+ neither stale detection nor per-query telemetry.
3870
+
3871
+ ### Where the comparison is not oracle-qualified (excluded, not losses)
3872
+
3873
+ - **Editing / Visualization / Context.** knodin passes its golden/export oracles;
3874
+ the competitors either did not run (serena, code-review-graph, aider
3875
+ unavailable/blocked) or ran without a shared correctness oracle (graphify,
3876
+ repomix, code2prompt "complete" but unscored). No comparable head-to-head.
3877
+ - **Project memory.** knodin has no memory feature (deferred, ADR 004); Serena's
3878
+ own value-beyond-docs gate also failed. No contest.
3879
+
3880
+ ### Where knodin still costs more — and why
3881
+
3882
+ - **Lifecycle peak RSS: 83.8 MB vs codebase-memory 31.3 MB, grepai 22.9 MB.**
3883
+ This is the one axis where knodin's raw number is worse. It is **not** an
3884
+ oracle-qualified head-to-head: knodin is measured as a dedicated, long-lived
3885
+ product process holding the real semantic model + graph resident (which is
3886
+ precisely what wins the search-quality contest above), while the competitors
3887
+ are measured at spawn-per-command peak. The RSS cost and the search-quality
3888
+ lead are two sides of the same resident-model design; knodin cannot undercut
3889
+ them on RSS without abandoning the model that beats them on search. C51
3890
+ (native training-free quantization) targets the *stored-embedding* portion of
3891
+ that footprint; the model-runtime portion is inherent to local semantic search.
3892
+
3893
+ Bottom line: on every dimension with an oracle-qualified head-to-head against an
3894
+ available competitor, knodin is **as good or better** (wins or ties). The only
3895
+ place a competitor's raw number beats knodin is lifecycle RSS, which is an
3896
+ incomparable isolation model and the direct cost of knodin's search-quality lead.