knodin 0.7.6 → 0.8.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +19 -7
- package/benchmarks/competitors/SYNTHESIS.md +66 -0
- package/dist/bin/cli.js +2164 -108
- package/dist/bin/launcher.js +25 -3
- package/dist/src/agent-integration.js +304 -0
- package/dist/src/artifact-refresh.js +82 -0
- package/dist/src/cli-args.js +292 -0
- package/dist/src/cli-model.js +384 -0
- package/dist/src/codeflow-replay.js +81 -0
- package/dist/src/compact-structural.js +96 -0
- package/dist/src/compare.js +39 -0
- package/dist/src/competitive-cold-mcp.js +40 -0
- package/dist/src/competitive-constraints.js +21 -0
- package/dist/src/competitive-manifest.js +411 -0
- package/dist/src/competitive-measurement.js +183 -0
- package/dist/src/competitive-runner.js +487 -0
- package/dist/src/competitive-sandbox.js +108 -0
- package/dist/src/context-export.js +423 -0
- package/dist/src/context.js +102 -0
- package/dist/src/deterministic-random.js +34 -0
- package/dist/src/diagnostics-write-helper.js +473 -0
- package/dist/src/diagnostics.js +1476 -0
- package/dist/src/docs-sections.js +141 -0
- package/dist/src/doctor.js +382 -0
- package/dist/src/engine/ann-hnsw.js +261 -0
- package/dist/src/engine/embeddings.js +193 -0
- package/dist/src/engine/file-walker.js +49 -0
- package/dist/src/engine/git-history.js +289 -0
- package/dist/src/engine/index.js +14238 -0
- package/dist/src/engine/perf.js +115 -0
- package/dist/src/engine/prune.js +112 -0
- package/dist/src/engine/sarif-import.js +341 -0
- package/dist/src/engine/scip-import.js +423 -0
- package/dist/src/engine/source-policy.js +85 -0
- package/dist/src/engine/sqlite.js +71 -0
- package/dist/src/engine/state-paths.js +175 -0
- package/dist/src/engine/symbol-delete.js +58 -0
- package/dist/src/execution-profile.js +208 -0
- package/dist/src/failure-diagnosis.js +655 -0
- package/dist/src/fleet.js +7 -0
- package/dist/src/git-executable.js +31 -0
- package/dist/src/graph-layout.js +173 -0
- package/dist/src/graph-query-health.js +115 -0
- package/dist/src/hook-manager-integration.js +156 -0
- package/dist/src/index-activity.js +126 -0
- package/dist/src/init-progress-worker.js +106 -2
- package/dist/src/init-progress.js +155 -0
- package/dist/src/init.js +1295 -0
- package/dist/src/lifecycle-health.js +282 -0
- package/dist/src/lsp-readonly.js +217 -0
- package/dist/src/mcp-graph-worker.js +69 -0
- package/dist/src/mcp-reliability.js +154 -0
- package/dist/src/mcp-worker-supervisor.js +350 -0
- package/dist/src/mirror.js +290 -0
- package/dist/src/node-runtime.js +157 -0
- package/dist/src/output-compression.js +630 -0
- package/dist/src/output-telemetry.js +368 -0
- package/dist/src/pr-triage.js +638 -0
- package/dist/src/progressive-evidence.js +477 -0
- package/dist/src/pure-compression-cli.js +102 -0
- package/dist/src/relationship-adapters.js +377 -0
- package/dist/src/release-attestation.js +533 -0
- package/dist/src/release-preflight.js +513 -0
- package/dist/src/repair-lease.js +85 -0
- package/dist/src/repair-progress-worker.js +120 -2
- package/dist/src/repair-progress.js +262 -0
- package/dist/src/repository-init-process.js +177 -0
- package/dist/src/repository-management.js +1261 -0
- package/dist/src/response-budget.js +196 -0
- package/dist/src/server.js +217 -0
- package/dist/src/structural-fast-path.js +344 -0
- package/dist/src/structural-snapshot.js +37 -0
- package/dist/src/system-config.js +638 -0
- package/dist/src/terminal-help.js +83 -0
- package/dist/src/tools/knodin-tools.js +1640 -0
- package/dist/src/update-ceremony.js +162 -0
- package/dist/src/update-policy.js +944 -0
- package/dist/src/update-trust.js +504 -0
- package/dist/src/version.js +13 -0
- package/dist/src/visualization.js +515 -0
- package/dist/src/wait-for-fresh.js +98 -0
- package/dist/src/worktree-lifecycle.js +234 -0
- package/docs/BEHAVIORAL-CONTRACT.md +72 -0
- package/docs/CLI.md +20 -1
- package/docs/COMPARISON.md +403 -0
- package/docs/COMPETITIVE-LANDSCAPE-2026-08.md +267 -0
- package/docs/CONTAINED-EXECUTION.md +77 -0
- package/docs/DIAGNOSTICS.md +80 -0
- package/docs/GIT-HISTORY-REVIEW.md +39 -0
- package/docs/HANDOFF.md +180 -0
- package/docs/INSTALLATION.md +21 -18
- package/docs/MCP.md +59 -8
- package/docs/PROGRESSIVE-EVIDENCE.md +37 -0
- package/docs/PT-ACCESS-RECOMMENDATION.md +89 -0
- package/docs/RELEASE-0.3-EVIDENCE.md +73 -0
- package/docs/REPOSITORIES-AND-WORKTREES.md +18 -6
- package/docs/SCIP-IMPORT.md +62 -0
- package/docs/SIGNED-UPDATES.md +151 -0
- package/docs/TELEMETRY.md +46 -0
- package/docs/TOKEN-OPTIMIZER-SCORECARD.md +79 -0
- package/docs/assets/knodin-favicon.svg +4 -0
- package/docs/releases/0.3.0.md +46 -0
- package/docs/releases/0.4.0.md +68 -0
- package/docs/releases/0.4.1.md +28 -0
- package/docs/releases/0.4.2.md +27 -0
- package/docs/releases/0.4.3.md +23 -0
- package/docs/releases/0.5.0.md +29 -0
- package/docs/releases/0.5.1.md +17 -0
- package/docs/releases/0.6.0.md +18 -0
- package/docs/releases/0.7.0.md +24 -0
- package/docs/releases/0.7.1.md +21 -0
- package/docs/releases/0.7.2.md +21 -0
- package/docs/releases/0.7.3.md +23 -0
- package/docs/releases/0.7.4.md +17 -0
- package/docs/releases/0.7.5.md +20 -0
- package/docs/releases/0.8.0.md +74 -0
- package/docs/releases/0.8.2.md +34 -0
- package/package.json +127 -4
- package/roadmap/competitive-roadmap.md +3801 -0
- package/schemas/release-attestation-v1.schema.json +210 -0
- package/schemas/support-bundle-v2.schema.json +212 -0
- package/dist/chunks/chunk-DMQAGX77.js +0 -654
- package/dist/chunks/chunk-F4Z3Z766.js +0 -4
- package/dist/chunks/chunk-SIJAQVSX.js +0 -3
- package/dist/chunks/chunk-X6M4HUUE.js +0 -2
- package/dist/chunks/chunk-YPRMY2LP.js +0 -8
- package/dist/chunks/pure-compression-cli-4TA2TQD5.js +0 -5
- package/dist/chunks/server-7EDF4CBY.js +0 -14
- package/dist/chunks/structural-fast-path-KD5KQSPX.js +0 -4
- package/docs/releases/0.7.6.md +0 -25
|
@@ -0,0 +1,3801 @@
|
|
|
1
|
+
# knodin — Competitive Roadmap
|
|
2
|
+
|
|
3
|
+
This is the authoritative active work queue after the original parity program
|
|
4
|
+
closed on 2026-07-21. The completed R1–R48 history, including measured no-go and
|
|
5
|
+
deferred performance decisions, is preserved in
|
|
6
|
+
[`archive/parity-roadmap-2026-07-21.md`](archive/parity-roadmap-2026-07-21.md).
|
|
7
|
+
|
|
8
|
+
## Product constraint
|
|
9
|
+
|
|
10
|
+
knodin wins by being local, low-latency, and compact: one MCP tool, no hosted
|
|
11
|
+
service, no required credentials, and no source-code egress. Competitor features
|
|
12
|
+
belong here only when live evidence shows that they improve correctness or a
|
|
13
|
+
core developer loop without undermining that constraint.
|
|
14
|
+
|
|
15
|
+
## Evidence policy
|
|
16
|
+
|
|
17
|
+
Every competitive item must link to a reproducible result under
|
|
18
|
+
`benchmarks/competitors/`. A competitor having a feature is not sufficient by
|
|
19
|
+
itself. Items use four dispositions:
|
|
20
|
+
|
|
21
|
+
- **Must close** — correctness, safety, or a prerequisite for the stated product
|
|
22
|
+
position.
|
|
23
|
+
- **Should close** — meaningful capability or efficiency improvement with a
|
|
24
|
+
bounded implementation path.
|
|
25
|
+
- **Evaluate** — evidence is incomplete or implementation cost may outweigh the
|
|
26
|
+
benefit; do not implement before its gate passes.
|
|
27
|
+
- **Won't pursue** — conflicts with the product constraint or duplicates a
|
|
28
|
+
better knodin mechanism.
|
|
29
|
+
|
|
30
|
+
The first evidence set is
|
|
31
|
+
[`GitNexus vs knodin`](../benchmarks/competitors/gitnexus-vs-knodin.md), backed
|
|
32
|
+
by 49 local cases and their raw MCP results. Later competitor bake-offs append
|
|
33
|
+
new items or strengthen existing ones; they do not reopen the archived roadmap.
|
|
34
|
+
|
|
35
|
+
## Execution order
|
|
36
|
+
|
|
37
|
+
| ID | Priority | Disposition | Item | Depends on |
|
|
38
|
+
|---|---:|---|---|---|
|
|
39
|
+
| C1 | P0 | Must close | Declare a permissive license | — |
|
|
40
|
+
| C2 | P0 | Must close | Stable symbol identity and ambiguity-safe resolution | — |
|
|
41
|
+
| C3 | P0 | Must close | Enforce truthful response budgets and minimal modes | — |
|
|
42
|
+
| C4 | P1 | Must close | Symbol-level directional impact analysis | C2, C3 |
|
|
43
|
+
| C5 | P1 | Should close | Explicit git diff scopes | C3 |
|
|
44
|
+
| C6 | P1 | Should close | Deterministic import-cycle analysis | C2 |
|
|
45
|
+
| C7 | P2 | Should close | MCP tool and handler mapping | C2 |
|
|
46
|
+
| C8 | P2 | Evaluate | Local statement-level PDG | C2, C3 |
|
|
47
|
+
| C9 | P2 | Evaluate | API contract and response-shape analysis | C2, C3 |
|
|
48
|
+
| C10 | P3 | Evaluate | Bounded expert graph-query DSL | C2, C3 |
|
|
49
|
+
| C11 | P0 | Must close | Call/reference precision and structural identity | C2 |
|
|
50
|
+
| C12 | P1 | Should close | Uniform filters, pagination, and compact drill-downs | C2, C3 |
|
|
51
|
+
| C13 | P1 | Should close | Typed traversal and architecture facets | C2, C3, C11 |
|
|
52
|
+
| C14 | P2 | Should close | Portable budgeted context export | C3, C5 |
|
|
53
|
+
| C15 | P2 | Evaluate | LSP diagnostics and guarded symbol editing | C2, C11 |
|
|
54
|
+
| C16 | P2 | Should close | Index health, repair, and output telemetry | C3 |
|
|
55
|
+
| C17 | P3 | Evaluate | DFS and feature-path navigation | C2, C3, C11 |
|
|
56
|
+
| C18 | P1 | Should close | Unified competitive regression harness | — |
|
|
57
|
+
| C19 | P2 | Evaluate | Salesforce LWC bundles and Apex/data bridges | C2, C3, C11 |
|
|
58
|
+
| C20 | P2 | Evaluate | Salesforce metadata and Flow/automation model | C2, C3, C11 |
|
|
59
|
+
| C21 | P3 | Evaluate | Aura bundles and Visualforce/Aura/LWC interop | C19 |
|
|
60
|
+
| C22 | P3 | Evaluate | LWR and Experience Cloud topology | C19 |
|
|
61
|
+
| C23 | P2 | Evaluate | Apex platform-entry and asynchronous semantics | C2, C3, C11 |
|
|
62
|
+
| C24 | P0 | Must close | Competitive correctness contract and replay baseline | C18 |
|
|
63
|
+
| C25 | P0 | Must close | Prove symbol identity and reference precision leadership | C24, C2, C11 |
|
|
64
|
+
| C26 | P1 | Must close | Prove diff-review and directional traversal leadership | C24, C4, C5, C13 |
|
|
65
|
+
| C27 | P1 | Must close | Prove architecture and community-analysis leadership | C24, C3, C12, C13 |
|
|
66
|
+
| C28 | P1 | Must close | Prove dead-code and semantic-search quality | C24, C3, C11, C12 |
|
|
67
|
+
| C29 | P1 | Must close | Prove context-packing and export leadership | C24, C3, C5, C14 |
|
|
68
|
+
| C30 | P2 | Evaluate | Prove guarded editing and diagnostics leadership | C24, C2, C11, C15 |
|
|
69
|
+
| C31 | P1 | Must close | Prove index lifecycle and output-telemetry leadership | C24, C3, C16, C18 |
|
|
70
|
+
| C32 | P2 | Evaluate | Prove local API and statement-flow analysis leadership | C24, C2, C3, C8, C9 |
|
|
71
|
+
| C33 | P0 | Must close | Preserve local privacy and one-tool schema economy | C24 |
|
|
72
|
+
| C34 | P1 | Must close | Keep local graph indexes current across repository lifecycle events | C16, C24, C33 |
|
|
73
|
+
| C35 | P2 | implemented | Produce local interactive architecture and call-flow visualizations | C3, C13, C24, C33 |
|
|
74
|
+
| C36 | P0 | implemented | Instrument the current warm-operation performance paths | C18, C24 |
|
|
75
|
+
| C37 | P0 | implemented | Reuse generation-scoped graph analytics and traversal snapshots | C36 |
|
|
76
|
+
| C38 | P1 | implemented | Compose orientation context from one shared snapshot | C37 |
|
|
77
|
+
| C39 | P1 | implemented | Separate fast truthful status from explicit deep audit | C36 |
|
|
78
|
+
| C40 | P1 | implemented | Make context packing and artifact access linear and cache-safe | C36 |
|
|
79
|
+
| C41 | P0 | evaluated — deferred | Replay like-for-like performance after the implementation batch | C37, C38, C39, C40 |
|
|
80
|
+
| C42 | P0 | implemented | Enforce comparable competitive cases and correctness oracles | C41 |
|
|
81
|
+
| C43 | P0 | implemented | Measure cold/warm p50, p95, and independent process RSS | C42 |
|
|
82
|
+
| C44 | P0 | implemented | Restore low-latency diff review without weakening evidence | C43 |
|
|
83
|
+
| C45 | P1 | implemented | Add facet-selective architecture fast paths | C43 |
|
|
84
|
+
| C46 | P1 | implemented | Cache search state and hydrate only the requested page | C43 |
|
|
85
|
+
| C47 | P1 | implemented | Reduce status, repair, and measured lifecycle resource cost | C43 |
|
|
86
|
+
| C48 | P2 | evaluated — deferred | Decide an optional local project-memory boundary | C42 |
|
|
87
|
+
| C49 | P1 | implemented | Keep immutable runs lint-safe and make the full suite terminate | — |
|
|
88
|
+
| C50 | P0 | implemented | Make the single-worker full-suite gate start reliably | C49 |
|
|
89
|
+
| C51 | P2 | evaluated — rejected for production | Native training-free embedding quantization | C46 |
|
|
90
|
+
| C52 | P0 | implemented | Package repository lifecycle initialization and refresh | C16, C34, C50 |
|
|
91
|
+
| C53 | P0 | implemented | Make repair observable, cancellable, and resumable | C47, C52 |
|
|
92
|
+
| C54 | P0 | implemented | Certify the local 0.1.0 release artifact | C50, C52, C53 |
|
|
93
|
+
| C55 | P0 | implemented | Pin and audit the authorized Token Optimizer source | C24 |
|
|
94
|
+
| C56 | P0 | implemented | Replay all seven Token Optimizer workflows on shared fixtures | C55 |
|
|
95
|
+
| C57 | P0 | implemented | Add bounded recoverable diagnostic-output compression | C3, C55 |
|
|
96
|
+
| C58 | P0 | evaluated — rejected | Gate command execution on reviewed portable containment | C57 |
|
|
97
|
+
| C59 | P1 | implemented | Connect diagnostic failures to source-evidenced graph context | C57 |
|
|
98
|
+
| C60 | P1 | parked-external-evidence — 2026-08-02 (previous: active — legacy GHES 0.3.0 live; external mirror governance, workload identity, and native certification pending) | Complete two-minute cross-client and GHES onboarding | C52, C55 |
|
|
99
|
+
| C61 | P1 | implemented | Deliver private real-token ROI telemetry and dashboard | C16, C55 |
|
|
100
|
+
| C62 | P0 | parked-external-evidence — 2026-08-02 (previous: active — verifier implemented; production ceremony pending) | Define and implement signed update metadata trust | C55 |
|
|
101
|
+
| C63 | P0 | parked-external-evidence — 2026-08-02 (previous: active — signed-only client implemented; production activation gated) | Add safe update status, check, explain, apply, and rollback policy | C62 |
|
|
102
|
+
| C64 | P0 | parked-external-evidence — 2026-08-02 (previous: active — implementation and adversarial fixtures complete; live release evidence pending) | Produce provenance, SBOM, and cross-channel release attestation | C62 |
|
|
103
|
+
| C65 | P0 | parked-external-evidence — 2026-08-02 (previous: active — bounded multi-step root recovery implemented; production drills pending) | Prove compromised-channel, rollback, freeze, and recovery defenses | C63, C64 |
|
|
104
|
+
| C66 | P0 | parked-external-evidence — 2026-08-02 (previous: active — scorecard published; terminal refresh awaits active dependencies) | Publish an evidence-linked Token Optimizer scorecard | C56, C57, C59, C60, C61, C65, C68, C75 |
|
|
105
|
+
| C67 | P0 | parked-external-evidence — 2026-08-02 (previous: Must close) | Certify and distribute the completed competitive release | C60, C62, C63, C64, C65, C66, C75 |
|
|
106
|
+
| C68 | P0 | implemented | Close Token Optimizer structural output and latency gaps | C56 |
|
|
107
|
+
| C69 | P0 | implemented | Stream large repository inventories with per-repository degradation | C52 |
|
|
108
|
+
| C70 | P0 | implemented | Select portfolio repositories before inventory and reject unknown selectors | C69 |
|
|
109
|
+
| C71 | P0 | implemented | Make portfolio init bounded, resumable, and truthful in dry-run mode | C69, C70 |
|
|
110
|
+
| C72 | P0 | implemented | Make configure/init lifecycle claims atomic and drain compatible queued events | C52 |
|
|
111
|
+
| C73 | P1 | implemented | Distinguish portfolio doctor and installed-versus-active MCP states | C60, C69 |
|
|
112
|
+
| C74 | P1 | implemented | Bound and classify Salesforce metadata architecture candidates | C69 |
|
|
113
|
+
| C75 | P0 | certified with recorded degradations | Certify the 61-repository portfolio under the one-GB memory ceiling | C69, C70, C71, C72, C73, C74 |
|
|
114
|
+
| C76 | P1 | implemented | Replace duplicated CLI parsing/help with one responsive declarative command model | C73 |
|
|
115
|
+
| C77 | P2 | evaluated — retained knodin unchanged | Replay CodeFlow edge-provenance and architecture-export claims | C24, C27, C42 |
|
|
116
|
+
| C78 | P1 | evaluated — retained C8 unchanged | Add bounded source-to-sink resource reachability | C2, C3, C8, C24 |
|
|
117
|
+
| C79 | P1 | implemented | Evaluate or add bounded repository applicability signals | C52, C69 |
|
|
118
|
+
| C80 | P0 | implemented | Route retrieval through structural evidence before embeddings | C3, C24, C42, C43 |
|
|
119
|
+
| C81 | P0 | evaluated — retained knodin unchanged | Run a powered end-to-end benchmark with a diagnose arm | C42, C43, C59, C80 |
|
|
120
|
+
| C82 | P0 | implemented | Productize the retained compress diagnose workflow | C59, C81 |
|
|
121
|
+
| C83 | P0 | implemented | Deliver progressive evidence with a verified hash handshake | C80, C82 |
|
|
122
|
+
| C84 | P1 | implemented | Import optional SCIP evidence without a live-LSP dependency | C24, C80, C83 |
|
|
123
|
+
| C85 | P1 | implemented | Add bounded, itemized git-history review signals | C24, C44, C81 |
|
|
124
|
+
| C86 | P2 | evaluated — retained MiniLM unchanged | Evaluate static embeddings against retained MiniLM | C46, C51, C80, C81 |
|
|
125
|
+
| C87 | P0 | implemented | Add a truthful structural cold-start fast path | C56, C68 |
|
|
126
|
+
| C88 | P1 | implemented — macOS certified; Linux/Windows unavailable | Add profile-based contained execution | C57, C58, C59 |
|
|
127
|
+
| C89 | P0 | implemented | Harden MCP request reliability and recovery | C3, C16, C53 |
|
|
128
|
+
| C90 | P0 | implemented | Produce previewable privacy-safe support bundles | C16, C59, C89 |
|
|
129
|
+
| C91 | P0 | parked-external-evidence — 2026-08-05 (previous: owner-approved adoption priority) | Certify boring macOS installation and Node runtime handoff | C52, C73, C89, C90 |
|
|
130
|
+
| C92 | P0 | parked-external-evidence — 2026-08-06 (previous: proposed) | Unify fail-closed five-channel release orchestration | C64, C91 |
|
|
131
|
+
| C93 | P1 | implemented | Measure engineering task outcomes with and without knodin | C24, C81 |
|
|
132
|
+
| C94 | P1 | parked-external-evidence — 2026-08-06 (previous: blocked) | Make first use self-explanatory and demonstrable | C89, C90, C91 |
|
|
133
|
+
| C95 | P1 | implemented | Codify and replay the defensible behavioral contract | C3, C24, C83 |
|
|
134
|
+
| C96 | P2 | evaluated — deferred | Gate every new capability on a narrow outcome replay | C24, C93, C95 |
|
|
135
|
+
|
|
136
|
+
## Roadmap completion contract
|
|
137
|
+
|
|
138
|
+
This section is the control plane for completing this roadmap with a persistent
|
|
139
|
+
goal prompt. The execution-order table remains the full historical ledger. The
|
|
140
|
+
owner split remaining work on 2026-08-02 into Track A, active feature work whose
|
|
141
|
+
acceptance evidence is local and checked in, and Track B, production-release and
|
|
142
|
+
trusted-distribution work parked until its external evidence arrives. Parking
|
|
143
|
+
does not close, weaken, or simulate any gate.
|
|
144
|
+
|
|
145
|
+
Run the roadmap loop with:
|
|
146
|
+
|
|
147
|
+
```text
|
|
148
|
+
Complete Track A in roadmap/competitive-roadmap.md in dependency order. Preserve its
|
|
149
|
+
product constraint, evidence policy, limitations, user-owned changes, and
|
|
150
|
+
checked-in evidence. Work autonomously on every safe local step. For each item,
|
|
151
|
+
run its acceptance gates and the repository verification bar, update its Status
|
|
152
|
+
and evidence truthfully, and remove it from Track A only when its
|
|
153
|
+
terminal condition is met. Never turn an evaluation into an implementation or
|
|
154
|
+
a superiority claim unless its gate authorizes that outcome. When an external
|
|
155
|
+
credential, production channel, independent custodian, or platform is required,
|
|
156
|
+
route the item to Track B only by an explicit owner decision; do not simulate
|
|
157
|
+
production evidence. Continue until `npm run check:roadmap -- --complete`
|
|
158
|
+
passes. That command proves Track A completion only and reports Track B parked;
|
|
159
|
+
it is not a production-release-readiness claim.
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
### Status and closure rules
|
|
163
|
+
|
|
164
|
+
- `implemented`, `done`, `closed`, `certified`, `evaluated — deferred`,
|
|
165
|
+
`evaluated — rejected`, and `won't pursue` are terminal only when the item
|
|
166
|
+
links checked-in evidence satisfying its acceptance or evaluation gate.
|
|
167
|
+
- `proposed`, `active`, `permission-gated`, and `blocked` are non-terminal.
|
|
168
|
+
- An `Evaluate` item closes by recording one allowed disposition: retain the
|
|
169
|
+
current product unchanged, defer/reject the idea with evidence, or create a
|
|
170
|
+
separately numbered implementation item. Evaluation never silently expands
|
|
171
|
+
production scope.
|
|
172
|
+
- P0/P1/P2/P3 affect ordering, not whether an active item must be resolved.
|
|
173
|
+
Track A completion does not close Track B or the entire roadmap.
|
|
174
|
+
- External work is fail-closed. A missing credential, custodian, production
|
|
175
|
+
channel, Windows host, or approval may remain parked with exact evidence; it
|
|
176
|
+
may not produce a synthetic pass. Parked rows remain in Track B.
|
|
177
|
+
- Proposals in ADRs, landscape documents, task systems, or competitor reports
|
|
178
|
+
are outside this roadmap's completion boundary until assigned a C-number in
|
|
179
|
+
the execution-order table. Repository applicability signals are C79 and ADR
|
|
180
|
+
006 items 1–7 are promoted intact as C80–C86;
|
|
181
|
+
later ADR 006 proposals remain outside the boundary until separately approved.
|
|
182
|
+
- Closing an item requires updating its item-level `Status`, evidence links,
|
|
183
|
+
limitations, the execution-order disposition if needed, and its ledger in
|
|
184
|
+
the same change. Do not rewrite immutable raw benchmark artifacts.
|
|
185
|
+
- `npm run check:roadmap -- --publish-ready` is the ordinary 0.x publication
|
|
186
|
+
guard. It requires Track A to be empty while retaining and reporting Track B.
|
|
187
|
+
The owner clarified on 2026-08-03 that early Docusign usage is a rollout
|
|
188
|
+
cohort, not a separate artifact channel; ordinary releases may publish while
|
|
189
|
+
Track B remains parked. `npm run check:roadmap -- --release-ready` remains the
|
|
190
|
+
stricter GA/trusted-distribution guard and fails unless both tracks are empty.
|
|
191
|
+
- While Track B remains open, documentation and positioning must not claim
|
|
192
|
+
cross-platform certification, trusted distribution, or production-release
|
|
193
|
+
readiness. All existing qualifications and recorded limitations remain in
|
|
194
|
+
force.
|
|
195
|
+
|
|
196
|
+
### Track A — Active local-work ledger
|
|
197
|
+
|
|
198
|
+
Track A contains only feature or local-evidence work whose acceptance can be
|
|
199
|
+
proved by checked-in tests and replays. Rows are ordered by dependency and then
|
|
200
|
+
priority. The owner approved repository applicability signals as the top item
|
|
201
|
+
on 2026-08-03, followed by ADR 006 items 1–7 intact as C80–C86;
|
|
202
|
+
later ADR proposals remain outside the completion boundary until separately
|
|
203
|
+
approved and numbered.
|
|
204
|
+
|
|
205
|
+
| ID | State | Previous state | Depends on | Exit evidence |
|
|
206
|
+
|---|---|---|---|---|
|
|
207
|
+
|
|
208
|
+
### Track B — Parked external-evidence ledger
|
|
209
|
+
|
|
210
|
+
Track B parked (10 items, external). The owner parked C60–C67 on 2026-08-02,
|
|
211
|
+
parked C91 after its code-readiness increment on 2026-08-05, and parked C92
|
|
212
|
+
and C94 on 2026-08-06 because their acceptance depends on C91 evidence.
|
|
213
|
+
Every former state, dependency, exit condition, acceptance criterion, and
|
|
214
|
+
limitation remains binding. Re-enter this track when either the documented C60
|
|
215
|
+
packet or C62 packet arrives; completing one packet does not waive the other.
|
|
216
|
+
The durable handoff index is
|
|
217
|
+
`docs/evidence/handoff/README.md`, the C60 packet is
|
|
218
|
+
`docs/evidence/c60-autonomous-preparation-2026-08-02.md`, and the C62 packet is
|
|
219
|
+
`docs/ROOT-CEREMONY-RUNBOOK.md`.
|
|
220
|
+
|
|
221
|
+
| ID | State | Previous state | Depends on | Exit evidence |
|
|
222
|
+
|---|---|---|---|---|
|
|
223
|
+
| C60 | parked-external-evidence — 2026-08-02 | active | C52, C55 | Mirror automation, timed onboarding, and Linux/Windows certification evidence |
|
|
224
|
+
| C62 | parked-external-evidence — 2026-08-02 | active | C55 | Offline root ceremony, independent pin review, and recovery receipt |
|
|
225
|
+
| C63 | parked-external-evidence — 2026-08-02 | active | C62 | Production activation/rollback evidence or explicit safe deferred policy |
|
|
226
|
+
| C64 | parked-external-evidence — 2026-08-02 | active | C62 | Verified release-candidate provenance/SBOM bundles and attestation draft |
|
|
227
|
+
| C65 | parked-external-evidence — 2026-08-02 | active | C63, C64 | Production ceremony and channel/runner/manager defense drills |
|
|
228
|
+
| C66 | parked-external-evidence — 2026-08-02 | active | C56, C57, C59, C60, C61, C65, C68, C75 | Terminal scorecard refresh linked to closed dependencies |
|
|
229
|
+
| C67 | parked-external-evidence — 2026-08-02 | proposed | C60, C62, C63, C64, C65, C66, C75 | Certified release record and distribution receipts |
|
|
230
|
+
| C91 | parked-external-evidence — 2026-08-05 | owner-approved adoption priority | C52, C73, C89, C90 | Native macOS manager/client/Homebrew/Artifactory receipts plus transactional doctor repairs |
|
|
231
|
+
| C92 | parked-external-evidence — 2026-08-06 | owner-approved release priority | C64, C91 | Dry-run preflight, retry replay, immutable artifact proof, and receipt schema |
|
|
232
|
+
| C94 | parked-external-evidence — 2026-08-06 | owner-approved clarity priority | C89, C90, C91 | First-use copy tests and a checked-in five-minute real-repository demonstration |
|
|
233
|
+
|
|
234
|
+
### Goal-loop ordering
|
|
235
|
+
|
|
236
|
+
1. Keep Track A active work independent of release authority and prove every
|
|
237
|
+
item with local, checked-in acceptance evidence.
|
|
238
|
+
2. Re-enter Track B when either C60 or C62's documented evidence packet arrives.
|
|
239
|
+
Execute C60 and C62 in parallel where people and platforms permit.
|
|
240
|
+
3. After C62, produce C64's release-candidate artifacts. Close C63 with its
|
|
241
|
+
fail-closed deferred-activation policy, then use that terminal client as the
|
|
242
|
+
subject of C65's adversarial drills. C65, not C63, authorizes any later
|
|
243
|
+
production activation.
|
|
244
|
+
4. Produce C66's terminal prerelease snapshot only from terminal dependency
|
|
245
|
+
evidence. C67 may append release receipts without reopening C66.
|
|
246
|
+
5. Execute C67 last. GA and trusted-distribution claims remain prohibited until
|
|
247
|
+
Track B is empty and `npm run check:roadmap -- --release-ready` passes.
|
|
248
|
+
Ordinary 0.x publication requires `npm run check:roadmap -- --publish-ready`.
|
|
249
|
+
6. Keep C92 and C94 parked behind C91. Re-enter C91 only when the named native macOS
|
|
250
|
+
environments and credentialed channels are available and the transactional
|
|
251
|
+
doctor repair contract is implemented; deterministic fixtures alone do not
|
|
252
|
+
satisfy certification.
|
|
253
|
+
|
|
254
|
+
---
|
|
255
|
+
|
|
256
|
+
## C1 — Declare a permissive license
|
|
257
|
+
|
|
258
|
+
- Status: implemented
|
|
259
|
+
- Priority: P0
|
|
260
|
+
- Disposition: Must close
|
|
261
|
+
- Evidence: the GitNexus bake-off found GitNexus explicitly licensed under
|
|
262
|
+
PolyForm Noncommercial 1.0.0 while knodin has neither a `LICENSE` file nor a
|
|
263
|
+
`package.json` license field. “Local and free” is not a legally supportable
|
|
264
|
+
marketing claim until this is resolved.
|
|
265
|
+
- Touches: `LICENSE`, `package.json`, `README.md`, `docs/COMPARISON.md`
|
|
266
|
+
|
|
267
|
+
### Acceptance criteria
|
|
268
|
+
|
|
269
|
+
1. The repository owner selects and adds an OSI-approved permissive license;
|
|
270
|
+
implementation must not guess the owner's legal choice.
|
|
271
|
+
2. `package.json` declares the matching SPDX identifier.
|
|
272
|
+
3. README and comparison copy use only claims permitted by that license.
|
|
273
|
+
4. Package contents include the license after `bun pm pack --dry-run` or the
|
|
274
|
+
repository's equivalent packaging check.
|
|
275
|
+
|
|
276
|
+
---
|
|
277
|
+
|
|
278
|
+
## C2 — Stable symbol identity and ambiguity-safe resolution
|
|
279
|
+
|
|
280
|
+
- Status: implemented
|
|
281
|
+
- Priority: P0
|
|
282
|
+
- Disposition: Must close
|
|
283
|
+
- Evidence: GitNexus's `trace(main, createEngine)` returned three ranked `main`
|
|
284
|
+
candidates and required disambiguation. knodin silently returned
|
|
285
|
+
`main → createEngine`, with no indication of which definition it selected.
|
|
286
|
+
- Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
|
|
287
|
+
symbol/query/explain/rename/path tests, persisted schema if IDs must survive
|
|
288
|
+
rebuilds
|
|
289
|
+
|
|
290
|
+
### Required behavior
|
|
291
|
+
|
|
292
|
+
1. Every symbol-facing result exposes a stable repository-scoped identity.
|
|
293
|
+
2. `explain`, structured queries, rename preview/apply, and shortest path accept
|
|
294
|
+
optional identity, file, and kind selectors.
|
|
295
|
+
3. A bare-name request with multiple viable definitions returns an explicit
|
|
296
|
+
ambiguity result with ranked candidates; it never silently merges or chooses
|
|
297
|
+
definitions when that can change the answer.
|
|
298
|
+
4. Existing unambiguous bare-name calls remain source-compatible.
|
|
299
|
+
|
|
300
|
+
### Acceptance criteria
|
|
301
|
+
|
|
302
|
+
1. A fixture with at least three same-named functions reproduces the GitNexus
|
|
303
|
+
`main` case and proves no conflated path is returned.
|
|
304
|
+
2. Selecting each candidate by identity returns only that candidate's real
|
|
305
|
+
callers, callees, impact, and paths.
|
|
306
|
+
3. Rename by identity edits only the selected definition and its resolved
|
|
307
|
+
references.
|
|
308
|
+
4. Migration/rebuild behavior for persisted IDs is deterministic and tested.
|
|
309
|
+
|
|
310
|
+
---
|
|
311
|
+
|
|
312
|
+
## C3 — Enforce truthful response budgets and minimal modes
|
|
313
|
+
|
|
314
|
+
- Status: implemented
|
|
315
|
+
- Priority: P0
|
|
316
|
+
- Disposition: Must close
|
|
317
|
+
- Evidence: knodin's minimal symbol context still returned 22,843 bytes because
|
|
318
|
+
source was always present, and minimal `map` returned 254,102 bytes. GitNexus
|
|
319
|
+
was usually more compact when source was not requested. Conversely,
|
|
320
|
+
GitNexus's source-inclusive query returned 606,235 bytes, demonstrating why a
|
|
321
|
+
hard cap—not convention—is required.
|
|
322
|
+
- Reinforced by Graphify's 417-byte top-10 hubs versus knodin's 258,628-byte
|
|
323
|
+
mapped response, Aider's 425-byte mean at a 128-token budget versus knodin's
|
|
324
|
+
roughly 190 KB mapped mean, and Repomix/code2prompt's explicit token metrics.
|
|
325
|
+
- Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, response-size and
|
|
326
|
+
operation tests, docs
|
|
327
|
+
|
|
328
|
+
### Required behavior
|
|
329
|
+
|
|
330
|
+
1. Every operation has a tested serialized-byte ceiling and reports truncation
|
|
331
|
+
plus total counts when the ceiling binds.
|
|
332
|
+
2. Minimal `explain` omits source unless explicitly requested; standard explain
|
|
333
|
+
retains edit-ready source within its cap.
|
|
334
|
+
3. Minimal `map` returns counts, aggregates, and bounded top-N summaries only;
|
|
335
|
+
it never leaks full community member arrays or full edge lists.
|
|
336
|
+
4. Caps are applied after serialization-aware accounting so nested content
|
|
337
|
+
cannot bypass them.
|
|
338
|
+
5. Callers may choose token, byte, and item budgets; all three produce honest
|
|
339
|
+
totals, truncation state, and a continuation/drill-down route.
|
|
340
|
+
6. Hub, bridge, community, traversal, search, source, and export operations honor
|
|
341
|
+
their requested top-N or budget without constructing a full map response.
|
|
342
|
+
|
|
343
|
+
### Acceptance criteria
|
|
344
|
+
|
|
345
|
+
1. The GitNexus benchmark's `context-name-min` and `cypher-communities`
|
|
346
|
+
counterparts remain below documented budgets.
|
|
347
|
+
2. Large-symbol and large-map fixtures assert both byte ceilings and honest
|
|
348
|
+
`truncated`/total-count metadata.
|
|
349
|
+
3. No standard response silently loses edit-critical source without a clear
|
|
350
|
+
continuation or drill-down route.
|
|
351
|
+
|
|
352
|
+
---
|
|
353
|
+
|
|
354
|
+
## C4 — Symbol-level directional impact analysis
|
|
355
|
+
|
|
356
|
+
- Status: implemented
|
|
357
|
+
- Priority: P1
|
|
358
|
+
- Disposition: Must close
|
|
359
|
+
- DependsOn: C2, C3
|
|
360
|
+
- Evidence: GitNexus distinguished upstream/downstream/both, depths 1/3/6,
|
|
361
|
+
tests, relation confidence, processes, and modules for `createEngine`.
|
|
362
|
+
knodin's closest public impact query was file-level and returned the same
|
|
363
|
+
broad 31-row set for every directional comparison.
|
|
364
|
+
- Reinforced by codebase-memory-mcp's directional typed trace and argument-level
|
|
365
|
+
data-flow evidence, and Graphify's relation-filtered direct neighbors.
|
|
366
|
+
- Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
|
|
367
|
+
impact/query tests and docs
|
|
368
|
+
|
|
369
|
+
### Required behavior
|
|
370
|
+
|
|
371
|
+
1. Impact accepts a stable symbol selector or an explicit file-level mode.
|
|
372
|
+
2. Symbol mode supports direction, maximum depth, relation kinds, confidence
|
|
373
|
+
threshold, test inclusion, result limit, argument/data-flow evidence, and
|
|
374
|
+
compact summary.
|
|
375
|
+
3. Results preserve depth and edge evidence and identify affected flows and
|
|
376
|
+
communities without presenting heuristic reach as exact.
|
|
377
|
+
4. File-level impact remains available and is labeled distinctly.
|
|
378
|
+
|
|
379
|
+
### Acceptance criteria
|
|
380
|
+
|
|
381
|
+
1. The `createEngine` fixture produces meaningfully different upstream and
|
|
382
|
+
downstream results.
|
|
383
|
+
2. Depth limits are monotonic and enforced; confidence/relation filters have
|
|
384
|
+
targeted fixtures.
|
|
385
|
+
3. Ambiguous symbols use C2's candidate response rather than conflated reach.
|
|
386
|
+
4. Compact output satisfies C3's byte budget.
|
|
387
|
+
|
|
388
|
+
---
|
|
389
|
+
|
|
390
|
+
## C5 — Explicit git diff scopes
|
|
391
|
+
|
|
392
|
+
- Status: implemented
|
|
393
|
+
- Priority: P1
|
|
394
|
+
- Disposition: Should close
|
|
395
|
+
- DependsOn: C3
|
|
396
|
+
- Evidence: GitNexus supports `unstaged`, `staged`, `all`, and `compare`.
|
|
397
|
+
knodin cannot request staged-only review, and its result can omit changed
|
|
398
|
+
files when no indexed symbol maps to the diff.
|
|
399
|
+
- Reinforced by code-review-graph's caller-supplied changed-file review and
|
|
400
|
+
code2prompt's working-tree diff, branch diff, and branch log exports.
|
|
401
|
+
- Touches: git diff parsing, engine review contract, MCP/CLI schemas, review
|
|
402
|
+
tests and docs
|
|
403
|
+
|
|
404
|
+
### Acceptance criteria
|
|
405
|
+
|
|
406
|
+
1. Review accepts `unstaged|staged|all|compare`; compare requires or defaults a
|
|
407
|
+
documented base ref.
|
|
408
|
+
2. Each scope has a fixture containing different staged and unstaged edits.
|
|
409
|
+
3. Results always report changed files, including non-code/unindexed files,
|
|
410
|
+
separately from mapped changed symbols.
|
|
411
|
+
4. Existing `base` behavior has a documented compatibility mapping.
|
|
412
|
+
5. Review accepts an explicit file list or revision pair without requiring the
|
|
413
|
+
checkout to contain an equivalent live diff.
|
|
414
|
+
|
|
415
|
+
---
|
|
416
|
+
|
|
417
|
+
## C6 — Deterministic import-cycle analysis
|
|
418
|
+
|
|
419
|
+
- Status: implemented
|
|
420
|
+
- Priority: P1
|
|
421
|
+
- Disposition: Should close
|
|
422
|
+
- DependsOn: C2
|
|
423
|
+
- Evidence: GitNexus exposes a bounded read-only cycle check; knodin has no
|
|
424
|
+
direct equivalent despite already storing import dependencies.
|
|
425
|
+
- Touches: `src/engine/index.ts`, query schema/docs, cycle fixtures
|
|
426
|
+
|
|
427
|
+
### Acceptance criteria
|
|
428
|
+
|
|
429
|
+
1. A structured query returns canonical deterministic file-import cycles.
|
|
430
|
+
2. Equivalent rotations of one cycle are deduplicated.
|
|
431
|
+
3. Limits and truncation metadata prevent exponential output.
|
|
432
|
+
4. Fixtures cover no cycle, one cycle, overlapping cycles, and type-only/import
|
|
433
|
+
edge policy.
|
|
434
|
+
|
|
435
|
+
---
|
|
436
|
+
|
|
437
|
+
## C7 — MCP tool and handler mapping
|
|
438
|
+
|
|
439
|
+
- Status: implemented
|
|
440
|
+
- Priority: P2
|
|
441
|
+
- Disposition: Should close
|
|
442
|
+
- DependsOn: C2
|
|
443
|
+
- Evidence: GitNexus correctly detected knodin's `knodin` MCP tool and its
|
|
444
|
+
source file. knodin has endpoint patterns but no MCP registration-to-handler
|
|
445
|
+
map.
|
|
446
|
+
- Touches: language extraction, schema, structured queries, MCP fixtures/docs
|
|
447
|
+
|
|
448
|
+
### Acceptance criteria
|
|
449
|
+
|
|
450
|
+
1. Index common TypeScript MCP SDK registration forms and associate tool name,
|
|
451
|
+
description, schema declaration, handler symbol, and file.
|
|
452
|
+
2. A query supports all tools and lookup by tool name.
|
|
453
|
+
3. Dynamic/unresolved registrations are labeled heuristic rather than exact.
|
|
454
|
+
4. The repository's own `knodin` registration is found without matching prose
|
|
455
|
+
examples or benchmark fixture strings.
|
|
456
|
+
|
|
457
|
+
---
|
|
458
|
+
|
|
459
|
+
## C8 — On-demand local statement-level flow analysis
|
|
460
|
+
|
|
461
|
+
- Status: implemented (bounded on-demand analysis; no persisted PDG)
|
|
462
|
+
- Priority: P2
|
|
463
|
+
- Disposition: Evaluate
|
|
464
|
+
- DependsOn: C2, C3
|
|
465
|
+
- Evidence: GitNexus's opt-in local PDG returned concrete control-dependence and
|
|
466
|
+
reaching-definition rows, including variable-filtered `db` flows. knodin has
|
|
467
|
+
no statement-level answer. GitNexus also produced large 14–42 KB PDG
|
|
468
|
+
responses, so copying its output contract would conflict with knodin's
|
|
469
|
+
compactness goal.
|
|
470
|
+
- Touches if approved: parser/extractor architecture, persisted schema,
|
|
471
|
+
performance harness, query surface, multi-language fixtures
|
|
472
|
+
- Evaluation evidence: [`C8 summary and methodology`](../benchmarks/evaluations/c8-pdg/summary-methodology.md)
|
|
473
|
+
- Decision: the disposable prototype passed the mechanical gate at 100% labeled
|
|
474
|
+
fixture precision, with three real repositories measured and responses below
|
|
475
|
+
the C3 ceiling. The supported implementation is the narrower on-demand
|
|
476
|
+
`flow_analysis` query for one selected TS/JS symbol, with optional variable
|
|
477
|
+
filtering, compact source evidence, C3 budgets, and explicit unresolved cases;
|
|
478
|
+
it does not persist a broad PDG.
|
|
479
|
+
|
|
480
|
+
### Evaluation gate
|
|
481
|
+
|
|
482
|
+
1. Prototype TS/JS control-flow and reaching-definitions on disposable code,
|
|
483
|
+
not the production schema.
|
|
484
|
+
2. Measure clean-index time, incremental-index time, database growth, peak RSS,
|
|
485
|
+
query latency, precision, and response size against at least three real
|
|
486
|
+
repositories.
|
|
487
|
+
3. Require ≥90% precision on a labeled control/data-flow fixture and a bounded
|
|
488
|
+
response below C3's ceiling.
|
|
489
|
+
4. Approve implementation only if the value cannot be delivered more cheaply
|
|
490
|
+
through compiler/LSP facts or a narrower on-demand analysis.
|
|
491
|
+
|
|
492
|
+
---
|
|
493
|
+
|
|
494
|
+
## C9 — Evaluate API contract and response-shape analysis
|
|
495
|
+
|
|
496
|
+
- Status: implemented
|
|
497
|
+
- Priority: P2
|
|
498
|
+
- Disposition: Evaluate
|
|
499
|
+
- DependsOn: C2, C3
|
|
500
|
+
- Evidence: GitNexus exposes route maps, response-shape checking, and API impact,
|
|
501
|
+
but this repository had zero detected HTTP routes. The bake-off proves the
|
|
502
|
+
surface exists, not that it is correct or useful. knodin now persists
|
|
503
|
+
source-evidenced literal Express/Fastify route and fetch/Axios client facts,
|
|
504
|
+
with a bounded `api_contract_mismatches` query.
|
|
505
|
+
- Touches: route extraction, schema/shape model, cross-file/client usage
|
|
506
|
+
resolution, fixtures and query surface
|
|
507
|
+
- Evaluation evidence: [`C9 summary`](../benchmarks/evaluations/c9-api/summary.md)
|
|
508
|
+
- Decision: the evaluation gate passed. GitNexus found the actionable `profile`
|
|
509
|
+
response mismatch that current handler, endpoint, and impact queries missed.
|
|
510
|
+
The production implementation shipped after the checked-in fixture produced
|
|
511
|
+
3 true positives, 0 false positives, and 0 false negatives (precision 1.00,
|
|
512
|
+
recall 1.00); dynamic routes remain unresolved and absolute-origin clients do
|
|
513
|
+
not link to local endpoints. See [`C9 production acceptance`](../benchmarks/evaluations/c9-api/production-acceptance.md).
|
|
514
|
+
|
|
515
|
+
### Evaluation gate
|
|
516
|
+
|
|
517
|
+
1. Build a local fixture with multiple verbs on one route, middleware, typed and
|
|
518
|
+
inferred response bodies, internal clients, mismatches, and dynamic routes.
|
|
519
|
+
2. Run GitNexus and knodin endpoint primitives against the same fixture and
|
|
520
|
+
manually label precision/recall.
|
|
521
|
+
3. Approve implementation only if shape analysis finds actionable breakage that
|
|
522
|
+
current handler/endpoint/impact queries miss.
|
|
523
|
+
|
|
524
|
+
---
|
|
525
|
+
|
|
526
|
+
## C10 — Evaluate a bounded expert graph-query DSL
|
|
527
|
+
|
|
528
|
+
- Status: closed; general DSL will not be pursued
|
|
529
|
+
- Priority: P3
|
|
530
|
+
- Disposition: Evaluate
|
|
531
|
+
- DependsOn: C2, C3
|
|
532
|
+
- Evidence: GitNexus's parameterized Cypher expressed caller, community
|
|
533
|
+
aggregate, and custom dead-code queries outside a fixed enum. knodin's fixed
|
|
534
|
+
patterns are easier to route and safer but cannot answer arbitrary structural
|
|
535
|
+
questions. GitNexus's naive custom dead-code query also returned many false
|
|
536
|
+
positives, illustrating the cost of unconstrained power.
|
|
537
|
+
- Touches if reopened: parser/validator, query planner, response budget,
|
|
538
|
+
read-only safety tests and docs
|
|
539
|
+
- Evaluation evidence: [`C10 summary`](../benchmarks/evaluations/c10-dsl/summary.md)
|
|
540
|
+
- Decision: no DSL is justified in this roadmap slice. Three bounded fixed-pattern
|
|
541
|
+
candidates covered 12 of 13 provenance-backed questions (92.3%), so the DSL
|
|
542
|
+
prototype gate is a no-go.
|
|
543
|
+
|
|
544
|
+
### Evaluation gate
|
|
545
|
+
|
|
546
|
+
1. Collect at least ten real questions that existing patterns cannot answer.
|
|
547
|
+
2. Determine whether two or three new fixed patterns cover most demand.
|
|
548
|
+
3. If not, design a read-only bounded DSL with explicit node/edge allowlists,
|
|
549
|
+
mandatory limits, traversal-depth caps, and no raw SQL/Cypher execution.
|
|
550
|
+
4. Approve only with deterministic cost bounds and C3-compliant output.
|
|
551
|
+
|
|
552
|
+
---
|
|
553
|
+
|
|
554
|
+
## C11 — Call/reference precision and structural identity
|
|
555
|
+
|
|
556
|
+
- Status: implemented
|
|
557
|
+
- Priority: P0
|
|
558
|
+
- Disposition: Must close
|
|
559
|
+
- DependsOn: C2
|
|
560
|
+
- Evidence: Graphify found four credible call neighbors for `createEngine` while
|
|
561
|
+
knodin included a known name-collision false positive; Serena returned precise
|
|
562
|
+
LSP references and structural `KnodinEngine` implementations knodin's nominal
|
|
563
|
+
inheritance query missed. grepai and codebase-memory-mcp also demonstrated
|
|
564
|
+
that imports/files must not be mislabeled as callers.
|
|
565
|
+
|
|
566
|
+
### Acceptance criteria
|
|
567
|
+
|
|
568
|
+
1. Labeled fixtures distinguish call, import, usage, containment, inheritance,
|
|
569
|
+
and TypeScript structural implementation edges.
|
|
570
|
+
2. Callers/callees reach at least 95% precision and recall on those fixtures;
|
|
571
|
+
built-ins and same-name symbols never create cross-definition edges.
|
|
572
|
+
3. Structural implementations are queryable separately from nominal inheritors.
|
|
573
|
+
4. Tests cover top-level calls, aliases, barrels, object-literal interface
|
|
574
|
+
implementations, scripts, tests, and ambiguous names.
|
|
575
|
+
|
|
576
|
+
---
|
|
577
|
+
|
|
578
|
+
## C12 — Uniform filters, pagination, and compact drill-downs
|
|
579
|
+
|
|
580
|
+
- Status: implemented
|
|
581
|
+
- Priority: P1
|
|
582
|
+
- Disposition: Should close
|
|
583
|
+
- DependsOn: C2, C3
|
|
584
|
+
- Evidence: code-review-graph filters semantic kind, large-code threshold/kind,
|
|
585
|
+
flow/community sort and detail; codebase-memory-mcp adds label, qualified name,
|
|
586
|
+
file, relationship, degree, total/offset pagination; Claude Context filters
|
|
587
|
+
extensions. Graphify honors top-N hubs and relation filters.
|
|
588
|
+
|
|
589
|
+
### Acceptance criteria
|
|
590
|
+
|
|
591
|
+
1. Search accepts language/extension, kind, path, test/production, and source
|
|
592
|
+
inclusion filters plus offset/cursor pagination with total and `hasMore`.
|
|
593
|
+
2. Large-code queries accept explicit line/complexity thresholds, symbol kind,
|
|
594
|
+
and path; no requested parameter is silently ignored.
|
|
595
|
+
3. Hub, bridge, community, flow, and traversal drill-downs honor top-N, sort,
|
|
596
|
+
relation, and detail controls under C3.
|
|
597
|
+
4. Parameter-combination tests prove filters compose rather than overwrite one
|
|
598
|
+
another.
|
|
599
|
+
|
|
600
|
+
---
|
|
601
|
+
|
|
602
|
+
## C13 — Typed traversal and architecture facets
|
|
603
|
+
|
|
604
|
+
- Status: implemented
|
|
605
|
+
- Priority: P1
|
|
606
|
+
- Disposition: Should close
|
|
607
|
+
- DependsOn: C2, C3, C11
|
|
608
|
+
- Evidence: codebase-memory-mcp beat knodin on inbound/outbound/both traces,
|
|
609
|
+
edge types, argument expressions, tests, risk labels, packages, layers,
|
|
610
|
+
boundaries, hotspots, and path scoping. Graphify adds direct relation-filtered
|
|
611
|
+
neighbors.
|
|
612
|
+
|
|
613
|
+
### Acceptance criteria
|
|
614
|
+
|
|
615
|
+
1. Traversal supports direction and explicit edge types with per-hop evidence.
|
|
616
|
+
2. Call/data-flow results preserve argument expressions where extraction is
|
|
617
|
+
compiler-grounded and label heuristic facts honestly.
|
|
618
|
+
3. Architecture exposes independently selectable packages, layers, boundaries,
|
|
619
|
+
hotspots, entry points, languages, and path scope.
|
|
620
|
+
4. Results satisfy C11 precision fixtures and C3 budgets.
|
|
621
|
+
|
|
622
|
+
---
|
|
623
|
+
|
|
624
|
+
## C14 — Portable budgeted context export
|
|
625
|
+
|
|
626
|
+
- Status: implemented
|
|
627
|
+
- Priority: P2
|
|
628
|
+
- Disposition: Should close
|
|
629
|
+
- DependsOn: C3, C5
|
|
630
|
+
- Evidence: Repomix, code2prompt, and Aider provide portable source-shaped
|
|
631
|
+
context, deterministic formatting, token accounting, file policies, ranged
|
|
632
|
+
retrieval, mention/chat awareness, and Git context. These are complementary
|
|
633
|
+
to knodin's graph rather than replacements for it.
|
|
634
|
+
|
|
635
|
+
### Acceptance criteria
|
|
636
|
+
|
|
637
|
+
1. Export supports Markdown, JSON, and XML with relative paths, optional line
|
|
638
|
+
numbers/tree, deterministic ordering, and explicit tokenizer cost estimate.
|
|
639
|
+
2. Include/exclude rules and per-file `full|summary|structure-only` policies are
|
|
640
|
+
composable with a hard token/byte budget.
|
|
641
|
+
3. Already-present/chat files can be excluded from duplicated source context.
|
|
642
|
+
4. Optional diff/log sections use C5 scopes; packed artifacts support bounded
|
|
643
|
+
ranged reads and exact regex grep.
|
|
644
|
+
|
|
645
|
+
---
|
|
646
|
+
|
|
647
|
+
## C15 — Evaluate LSP diagnostics and guarded symbol editing
|
|
648
|
+
|
|
649
|
+
- Status: implemented (read-only TypeScript language-service queries)
|
|
650
|
+
- Priority: P2
|
|
651
|
+
- Disposition: Evaluate
|
|
652
|
+
- DependsOn: C2, C11
|
|
653
|
+
- Evidence: Serena successfully exposed diagnostics, declarations,
|
|
654
|
+
implementations, symbol-body replacement, before/after insertion, safe delete,
|
|
655
|
+
and guarded multi-file replacement. knodin's verified rename is safer but much
|
|
656
|
+
narrower. The 2026-07-22 disposable evaluation exercised TypeScript Language
|
|
657
|
+
Server, Pyright, JDTLS, Java/Javac, .NET compilation, and the official Roslyn
|
|
658
|
+
C# Language Server with guarded TypeScript previews and rollback. The final
|
|
659
|
+
raw result records a C# `textDocument/diagnostic` response and
|
|
660
|
+
`textDocument/implementation` navigation to the fixture's line-10
|
|
661
|
+
implementation; the passing compiler check remains corroboration rather than a
|
|
662
|
+
substitute for LSP evidence. knodin now ships optional local TypeScript
|
|
663
|
+
diagnostics, definition, declaration, and implementation queries with no
|
|
664
|
+
daemon, subprocess, workspace mutation, or index initialization. Missing or
|
|
665
|
+
unsupported adapters return explicit unavailable results; guarded editing and
|
|
666
|
+
cross-language mutation remain unapproved.
|
|
667
|
+
|
|
668
|
+
### Evaluation gate
|
|
669
|
+
|
|
670
|
+
1. Prototype a local TypeScript/LSP adapter for diagnostics, implementation, and
|
|
671
|
+
edit previews without adding a required daemon.
|
|
672
|
+
2. Require all edits to produce a preview and reuse rename's compiler check and
|
|
673
|
+
rollback guarantees; never expose unguarded filesystem mutation.
|
|
674
|
+
3. Measure startup/RSS/latency and verify behavior across TS, Python, Java, and
|
|
675
|
+
C# fixtures before approving a general editing surface.
|
|
676
|
+
|
|
677
|
+
---
|
|
678
|
+
|
|
679
|
+
## C16 — Index health, repair, and output telemetry
|
|
680
|
+
|
|
681
|
+
- Status: implemented
|
|
682
|
+
- Priority: P2
|
|
683
|
+
- Disposition: Should close
|
|
684
|
+
- DependsOn: C3
|
|
685
|
+
- Evidence: codebase-memory-mcp exposes portable lifecycle modes and schema;
|
|
686
|
+
Claude Context exposes clear/status lifecycle; mcp-codebase-index statically
|
|
687
|
+
defines health/repair; grepai records output/token savings. knodin has strong
|
|
688
|
+
freshness internals but limited public health and budget telemetry.
|
|
689
|
+
|
|
690
|
+
### Acceptance criteria
|
|
691
|
+
|
|
692
|
+
1. Status reports schema/model/version, file and symbol coverage, orphaned or
|
|
693
|
+
missing records, last successful reconciliation, and actionable repair steps.
|
|
694
|
+
2. A local repair command can rebuild only damaged/missing state and verify the
|
|
695
|
+
result without deleting healthy data.
|
|
696
|
+
3. Benchmark telemetry records operation, latency, serialized bytes, estimated
|
|
697
|
+
tokens, truncation, and detail mode without recording source content.
|
|
698
|
+
4. All telemetry is local, opt-in for persistence, and documented as such.
|
|
699
|
+
|
|
700
|
+
---
|
|
701
|
+
|
|
702
|
+
## C17 — Evaluate DFS and feature-path navigation
|
|
703
|
+
|
|
704
|
+
- Status: implemented
|
|
705
|
+
- Priority: P3
|
|
706
|
+
- Disposition: Evaluate
|
|
707
|
+
- DependsOn: C2, C3, C11
|
|
708
|
+
- Evidence: Graphify and code-review-graph expose DFS; grepai adds deterministic
|
|
709
|
+
functional feature paths. `query feature_path <symbol>` now supplies a compact,
|
|
710
|
+
deterministic downstream DFS over C2/C11-resolved references. It includes
|
|
711
|
+
file/line edge evidence, cycle guards, depth/node caps, and explicit
|
|
712
|
+
truncation without adding persistent identity state. The focused fixture-led
|
|
713
|
+
regression in `feature-path.spec.ts` proves a branching result observably
|
|
714
|
+
differs from BFS while remaining stable after harmless line shifts.
|
|
715
|
+
|
|
716
|
+
### Delivered behavior
|
|
717
|
+
|
|
718
|
+
1. `feature_path` follows only resolved downstream references in stable DFS
|
|
719
|
+
order; `traverse` remains the existing selectable-direction BFS neighborhood.
|
|
720
|
+
2. Standard-detail non-root nodes include C11-resolved identity; every hop
|
|
721
|
+
includes source file/line, kind, provenance, and confidence evidence.
|
|
722
|
+
3. Cycles are visited once, depth is clamped to 1–6, the item cap is enforced,
|
|
723
|
+
and a larger reachable graph reports `truncated: true`.
|
|
724
|
+
4. The feature is source-only: dynamic dispatch, reflection, unresolved calls,
|
|
725
|
+
and runtime execution order remain explicitly outside the contract.
|
|
726
|
+
|
|
727
|
+
---
|
|
728
|
+
|
|
729
|
+
## C18 — Unified competitive regression harness
|
|
730
|
+
|
|
731
|
+
- Status: implemented
|
|
732
|
+
- Priority: P1
|
|
733
|
+
- Disposition: Should close
|
|
734
|
+
- Evidence: the first program produced 12 independent protocols and 713 paired
|
|
735
|
+
cases. Without one version-pinned entry point, improvements to C1-C17 could
|
|
736
|
+
silently lose coverage or be compared against a different tool installation.
|
|
737
|
+
- Touches: `scripts/competitive-bakeoff.ts`, `src/competitive-manifest.ts`,
|
|
738
|
+
`src/competitive-runner.ts`, `benchmarks/competitors/README.md`
|
|
739
|
+
|
|
740
|
+
### Acceptance criteria
|
|
741
|
+
|
|
742
|
+
1. One manifest selects competitors directly or by C1-C17 roadmap coverage and
|
|
743
|
+
records exact versions, prerequisites, preparation, commands, and blockers.
|
|
744
|
+
2. Runs refuse dirty checkouts by default, validate installed versions, isolate
|
|
745
|
+
fixture mutations, and preserve immutable timestamp-and-commit artifacts.
|
|
746
|
+
3. Heterogeneous raw results normalize into one summary and compare against the
|
|
747
|
+
checked-in baseline with a configurable latency tolerance and nonzero
|
|
748
|
+
regression exit status.
|
|
749
|
+
4. List, prerequisite-check, and dry-run modes are non-mutating; commands use
|
|
750
|
+
argument arrays rather than shell interpolation.
|
|
751
|
+
5. Unit tests cover selection, immutable history, normalization, blockers, and
|
|
752
|
+
material-regression detection.
|
|
753
|
+
|
|
754
|
+
---
|
|
755
|
+
|
|
756
|
+
## Salesforce scope boundary and remaining limitations
|
|
757
|
+
|
|
758
|
+
The baseline parses Apex and Visualforce and now includes the bounded C19–C23
|
|
759
|
+
implementations for LWC, selected metadata/Flow automation, Aura interop, LWR
|
|
760
|
+
topology, and Apex platform entry/asynchronous semantics. Their individual
|
|
761
|
+
sections and fixtures define the supported static evidence. This remains
|
|
762
|
+
partial Salesforce coverage, not proof that every deployed-org dependency or
|
|
763
|
+
runtime dispatch is resolved. Static dead-code candidates still require
|
|
764
|
+
corroboration against deployed-org and platform dependency data before deletion.
|
|
765
|
+
|
|
766
|
+
OmniStudio and industry-cloud packages are intentionally out of scope: promote
|
|
767
|
+
them only after a concrete, legally usable Salesforce repository demonstrates a
|
|
768
|
+
workflow that C19–C23 do not cover.
|
|
769
|
+
|
|
770
|
+
Their proposed shared `src/engine/index.ts` footprint is a collision, so they
|
|
771
|
+
must be scheduled sequentially even where their logical dependencies are
|
|
772
|
+
independent.
|
|
773
|
+
|
|
774
|
+
---
|
|
775
|
+
|
|
776
|
+
## C19 — Evaluate Salesforce LWC bundles and Apex/data bridges
|
|
777
|
+
|
|
778
|
+
- Status: implemented
|
|
779
|
+
- Priority: P2
|
|
780
|
+
- Disposition: Evaluate
|
|
781
|
+
- DependsOn: C2, C3, C11
|
|
782
|
+
- Evidence: the version-pinned, source-only Salesforce DX fixture now records
|
|
783
|
+
13/13 exact bundle, component, Apex, schema, wire-adapter, and Lightning Data
|
|
784
|
+
Service labels (100% precision, 100% recall) in
|
|
785
|
+
`benchmarks/evaluations/c19-salesforce-lwc/raw-results-20260722T132842614Z.json`.
|
|
786
|
+
The replay copies the fixture to a temporary directory, needs no org login,
|
|
787
|
+
credentials, network access, or source egress, and records one-file LWC/Apex
|
|
788
|
+
reindex plus C3-bounded response measurements. Dynamic and namespaced forms
|
|
789
|
+
remain omitted rather than exact claims. The raw result hashes the evaluated
|
|
790
|
+
engine, runner, and label revision in addition to the fixture. This authorizes no broader Salesforce
|
|
791
|
+
product commitment beyond the evaluated static source edges.
|
|
792
|
+
- Touches if approved: `src/engine/index.ts`, `src/__tests__/unit/salesforce.spec.ts`,
|
|
793
|
+
`src/__tests__/unit/lwc.spec.ts` (new),
|
|
794
|
+
`benchmarks/evaluations/c19-salesforce-lwc/` (new),
|
|
795
|
+
`src/competitive-manifest.ts`
|
|
796
|
+
|
|
797
|
+
### Evaluation gate
|
|
798
|
+
|
|
799
|
+
1. Build a self-contained Salesforce DX fixture with LWC JavaScript, HTML, CSS,
|
|
800
|
+
and `*.js-meta.xml` bundle members; include component-to-component imports,
|
|
801
|
+
`@salesforce/apex`, `@salesforce/schema`, wire adapters, and Lightning Data
|
|
802
|
+
Service usage alongside resolvable Apex methods.
|
|
803
|
+
2. Label bundle containment plus LWC-to-Apex, LWC-to-object/field, and
|
|
804
|
+
component-to-component edges. Require at least 95% precision and 90% recall;
|
|
805
|
+
unresolved, dynamic, or namespaced references must be omitted or labeled
|
|
806
|
+
heuristic rather than reported as exact.
|
|
807
|
+
3. Re-index a one-file LWC and one-file Apex change, measure index/response
|
|
808
|
+
cost against C3, and prove the analysis needs neither an org login nor source
|
|
809
|
+
egress.
|
|
810
|
+
|
|
811
|
+
---
|
|
812
|
+
|
|
813
|
+
## C20 — Evaluate Salesforce metadata and Flow/automation model
|
|
814
|
+
|
|
815
|
+
- Status: implemented
|
|
816
|
+
- Priority: P2
|
|
817
|
+
- Disposition: Evaluate
|
|
818
|
+
- DependsOn: C2, C3, C11
|
|
819
|
+
- Evidence: the version-pinned source-only Salesforce DX fixture records 15/15
|
|
820
|
+
exact custom-object/field/record-type, permission-set, Flow, workflow, and
|
|
821
|
+
approval-process labels (100% precision and recall) in
|
|
822
|
+
`benchmarks/evaluations/c20-salesforce-metadata/raw-results-20260722T134732728Z.json`.
|
|
823
|
+
The replay copies the fixture to a temporary directory, runs a metadata-only
|
|
824
|
+
one-file update, stays within C3's response budget, and needs no org login,
|
|
825
|
+
deployment, credentials, network access, or source egress. Formula and
|
|
826
|
+
runtime-selected targets remain explicit unknowns rather than exact edges.
|
|
827
|
+
This authorizes no broader Salesforce product commitment beyond the evaluated
|
|
828
|
+
static source relationships.
|
|
829
|
+
- Touches if approved: `src/engine/index.ts`,
|
|
830
|
+
`src/__tests__/unit/salesforce-metadata.spec.ts` (new),
|
|
831
|
+
`benchmarks/evaluations/c20-salesforce-metadata/` (new),
|
|
832
|
+
`src/competitive-manifest.ts`
|
|
833
|
+
|
|
834
|
+
### Evaluation gate
|
|
835
|
+
|
|
836
|
+
1. Build a version-pinned Salesforce DX fixture containing custom objects and
|
|
837
|
+
fields, record types, permission sets, a record-triggered Flow, a screen or
|
|
838
|
+
autolaunched Flow, workflow/approval metadata, and a Flow-invoked Apex
|
|
839
|
+
action.
|
|
840
|
+
2. Label only statically declared metadata edges: Flow/automation to Apex,
|
|
841
|
+
object, field, permission, and referenced UI component. Require at least
|
|
842
|
+
95% precision and 90% recall; dynamic formulas and runtime-selected targets
|
|
843
|
+
must remain explicit unknowns.
|
|
844
|
+
3. Demonstrate that a metadata-only change updates the affected topology without
|
|
845
|
+
requiring deployment to an org, and keep output within C3's budget.
|
|
846
|
+
|
|
847
|
+
---
|
|
848
|
+
|
|
849
|
+
## C21 — Evaluate Aura bundles and Visualforce/Aura/LWC interop
|
|
850
|
+
|
|
851
|
+
- Status: implemented
|
|
852
|
+
- Priority: P3
|
|
853
|
+
- Disposition: Evaluate
|
|
854
|
+
- DependsOn: C19
|
|
855
|
+
- Evidence: the version-pinned, source-only Salesforce DX fixture records 10/10
|
|
856
|
+
exact Aura containment, Aura-to-Apex, Aura-to-LWC, and Visualforce controller
|
|
857
|
+
labels (100% precision and recall) in
|
|
858
|
+
`benchmarks/evaluations/c21-salesforce-interop/raw-results-20260722T161801709Z.json`.
|
|
859
|
+
The replay copies its fixture to a temporary directory, needs no org login,
|
|
860
|
+
credentials, network access, or source egress, and verifies source evidence
|
|
861
|
+
for every claimed cross-stack edge. knodin supports these static source
|
|
862
|
+
relationships; `$A` runtime lookup strings remain omitted rather than exact
|
|
863
|
+
claims.
|
|
864
|
+
- Touches: `src/engine/index.ts`,
|
|
865
|
+
`src/__tests__/unit/visualforce.spec.ts`,
|
|
866
|
+
`src/__tests__/unit/salesforce-interop.spec.ts` (new),
|
|
867
|
+
`benchmarks/evaluations/c21-salesforce-interop/` (new),
|
|
868
|
+
`src/competitive-manifest.ts`
|
|
869
|
+
|
|
870
|
+
### Evaluation gate
|
|
871
|
+
|
|
872
|
+
1. Build a fixture with Aura component, application, interface, event, and
|
|
873
|
+
design members; Visualforce page/component/controller bindings; and a
|
|
874
|
+
supported Aura-to-LWC or Visualforce-to-Lightning bridge.
|
|
875
|
+
2. Label containment and cross-stack edges, including Aura-to-Apex and the
|
|
876
|
+
declared bridge target. Require at least 95% precision and 90% recall, while
|
|
877
|
+
`$A` runtime lookup strings and unresolvable expressions remain heuristic.
|
|
878
|
+
3. Confirm that unchanged Apex/Visualforce behavior retains its current test
|
|
879
|
+
results and that cross-stack results expose the source evidence for every
|
|
880
|
+
claimed edge.
|
|
881
|
+
|
|
882
|
+
---
|
|
883
|
+
|
|
884
|
+
## C22 — Evaluate LWR and Experience Cloud topology
|
|
885
|
+
|
|
886
|
+
- Status: implemented
|
|
887
|
+
- Priority: P3
|
|
888
|
+
- Disposition: Evaluate
|
|
889
|
+
- DependsOn: C19
|
|
890
|
+
- Evidence: the version-pinned, source-only LWR and Experience Cloud fixture
|
|
891
|
+
records 9/9 exact site, route, view, LWC/Aura, theme, and SVG asset labels
|
|
892
|
+
(100% precision and recall) in
|
|
893
|
+
`benchmarks/evaluations/c22-salesforce-experience/raw-results-20260722T162757428Z.json`.
|
|
894
|
+
The replay copies its fixture to a temporary directory, measures a route-only
|
|
895
|
+
reindex, stays within C3's 65,536-byte map budget, and needs no org login,
|
|
896
|
+
credentials, network access, source egress, server, or deployment. knodin
|
|
897
|
+
supports this static source topology; runtime route expressions remain
|
|
898
|
+
unresolved rather than fabricated.
|
|
899
|
+
- Touches: `src/engine/index.ts`,
|
|
900
|
+
`src/__tests__/unit/salesforce-experience.spec.ts` (new),
|
|
901
|
+
`benchmarks/evaluations/c22-salesforce-experience/` (new),
|
|
902
|
+
`src/competitive-manifest.ts`
|
|
903
|
+
|
|
904
|
+
### Evaluation gate
|
|
905
|
+
|
|
906
|
+
1. Build a source-only fixture containing LWR configuration plus the supported
|
|
907
|
+
Experience Cloud metadata for one site, its routes, views, theme/assets, and
|
|
908
|
+
declared LWC or Aura targets.
|
|
909
|
+
2. Label site-to-route-to-view-to-component and static-asset relationships.
|
|
910
|
+
Require at least 95% precision and 90% recall; route values assembled at
|
|
911
|
+
runtime must be labeled unresolved rather than fabricated.
|
|
912
|
+
3. Measure a route/view-only incremental update and prove that topology output
|
|
913
|
+
is compact, locally computed, and does not require serving or deploying the
|
|
914
|
+
site.
|
|
915
|
+
|
|
916
|
+
---
|
|
917
|
+
|
|
918
|
+
## C23 — Evaluate Apex platform-entry and asynchronous semantics
|
|
919
|
+
|
|
920
|
+
- Status: implemented
|
|
921
|
+
- Priority: P2
|
|
922
|
+
- Disposition: Evaluate
|
|
923
|
+
- DependsOn: C2, C3, C11
|
|
924
|
+
- Evidence: the version-pinned, source-only Salesforce DX fixture records 20/20
|
|
925
|
+
exact REST/SOAP/invocable/trigger, async callback, enqueue, subscription, and
|
|
926
|
+
Salesforce Function labels (100% precision and recall) in
|
|
927
|
+
`benchmarks/evaluations/c23-apex-platform-semantics/raw-results-20260722T140720946Z.json`.
|
|
928
|
+
The replay copies the fixture to a temporary directory, reindexes one Apex
|
|
929
|
+
file, keeps map responses under C3's 65,536-byte budget, and requires no org
|
|
930
|
+
login, credentials, network access, or source egress. Transaction ordering,
|
|
931
|
+
bulk behavior, retries, and runtime payload shape remain unknown.
|
|
932
|
+
- Touches if approved: `src/engine/index.ts`,
|
|
933
|
+
`src/__tests__/unit/salesforce.spec.ts`,
|
|
934
|
+
`src/__tests__/unit/apex-platform-semantics.spec.ts` (new),
|
|
935
|
+
`benchmarks/evaluations/c23-apex-platform-semantics/` (new),
|
|
936
|
+
`src/competitive-manifest.ts`
|
|
937
|
+
|
|
938
|
+
### Evaluation gate
|
|
939
|
+
|
|
940
|
+
1. Build a version-pinned fixture covering `@RestResource` and verb handlers,
|
|
941
|
+
`webService`/SOAP, `@InvocableMethod`, trigger object/event declarations,
|
|
942
|
+
`@future`, Queueable, Batchable, Schedulable, Platform Event or change-event
|
|
943
|
+
subscribers, and statically resolvable Salesforce Function invocations.
|
|
944
|
+
2. Label entry-point, enqueue, callback, subscription, and invocation edges;
|
|
945
|
+
require at least 95% precision and 90% recall. Never infer transaction order,
|
|
946
|
+
bulk behavior, retries, or runtime payload shape from source alone.
|
|
947
|
+
3. Validate that each result carries an exact-versus-heuristic confidence label,
|
|
948
|
+
stays within C3's response budget, and is produced with no org credentials or
|
|
949
|
+
network access.
|
|
950
|
+
|
|
951
|
+
---
|
|
952
|
+
|
|
953
|
+
## Competitive leadership program
|
|
954
|
+
|
|
955
|
+
The competitive audit is evidence, not a second roadmap. C24-C33 turn every
|
|
956
|
+
row of [`COMPETITIVE-AUDIT.md`](../benchmarks/competitors/COMPETITIVE-AUDIT.md)
|
|
957
|
+
into a measurable delivery contract. A row is complete only when knodin meets
|
|
958
|
+
or exceeds the stated criterion against the pinned competitor and shared
|
|
959
|
+
fixture. A missing local prerequisite, incompatible license, hosted-service
|
|
960
|
+
requirement, or resource failure is a reproducible blocker, never a pass.
|
|
961
|
+
|
|
962
|
+
All replays must record the knodin commit, competitor version, command,
|
|
963
|
+
fixture revision, warm/cold mode, requested scope, response budget, and oracle
|
|
964
|
+
result. Correctness wins require a labeled oracle; latency, response size, or
|
|
965
|
+
feature presence alone cannot establish a win. A replay that proves a loss
|
|
966
|
+
must add a follow-on implementation item before the parent may close.
|
|
967
|
+
|
|
968
|
+
### Product-claim gate
|
|
969
|
+
|
|
970
|
+
Every roadmap item that changes a user-facing claim must state and verify all
|
|
971
|
+
four parts of the adoption case:
|
|
972
|
+
|
|
973
|
+
1. **Value:** the named developer or agent workflow and the concrete failure or
|
|
974
|
+
delay it removes.
|
|
975
|
+
2. **ROI:** a labeled correctness, latency, size, safety, or effort measure
|
|
976
|
+
with its baseline and measurement window; do not substitute a vendor claim
|
|
977
|
+
or feature count.
|
|
978
|
+
3. **Tomorrow:** a local, credential-free adoption path that is compatible with
|
|
979
|
+
existing Git, editor, and test workflows, including explicit unavailable or
|
|
980
|
+
degraded states.
|
|
981
|
+
4. **Defensibility:** the evidence, integration, or outcome that remains hard
|
|
982
|
+
to substitute after the underlying model capability becomes commonplace.
|
|
983
|
+
|
|
984
|
+
“Uses AI” and an unmeasured capability are never roadmap completion signals.
|
|
985
|
+
An item without a measurable adoption case remains evaluation-only; a replay
|
|
986
|
+
that finds a quality gap updates the roadmap and prioritizes the smallest
|
|
987
|
+
workflow correction before any new surface area is added.
|
|
988
|
+
|
|
989
|
+
| Audit row | Leadership item | Win target |
|
|
990
|
+
|---|---|---|
|
|
991
|
+
| Exact symbol identity and ambiguity | C25 | Correct selected identity or explicit ambiguity for every labeled duplicate-name case |
|
|
992
|
+
| Caller/callee/reference accuracy | C25 | ≥95% precision and recall; zero cross-definition false positives |
|
|
993
|
+
| Diff review and change scope | C26 | Exact changed-file/symbol oracle across each supported scope |
|
|
994
|
+
| Blast radius and traversal | C26 | Oracle-checked typed edges, depth, filters, and truncation |
|
|
995
|
+
| Architecture and communities | C27 | Oracle-checked topology plus deterministic bounded drill-downs |
|
|
996
|
+
| Dead code | C28 | Reported precision/recall and no unlabelled heuristic claim |
|
|
997
|
+
| Semantic orientation/search | C28 | Frozen relevance judgments with reported nDCG and recall@k |
|
|
998
|
+
| Code/context packing | C29 | Exact inclusion/range/token oracle under equivalent budgets |
|
|
999
|
+
| Source editing and diagnostics | C30 | Preview, compile/test, and rollback oracle on disposable fixtures |
|
|
1000
|
+
| Resource/index lifecycle | C31 | Comparable cold/warm/repair/RSS/disk evidence with successful recovery oracle |
|
|
1001
|
+
| API/PDG/specialized analysis | C32 | Precision/recall and bounded source-evidence oracle |
|
|
1002
|
+
| Local privacy and schema economy | C33 | Zero required credentials/egress/hosted service and one top-level MCP tool |
|
|
1003
|
+
|
|
1004
|
+
### C24 — Competitive correctness contract and replay baseline
|
|
1005
|
+
|
|
1006
|
+
- Status: implemented
|
|
1007
|
+
- Priority: P0
|
|
1008
|
+
- Disposition: Must close
|
|
1009
|
+
- DependsOn: C18
|
|
1010
|
+
- Evidence: the audit records 713 paired cases but explicitly says that most
|
|
1011
|
+
current comparisons lack a labeled correctness oracle. Existing label sets
|
|
1012
|
+
under `benchmarks/evaluations/` are fragmented and are not represented in
|
|
1013
|
+
the normalized competitive-run summary.
|
|
1014
|
+
- Competitors: all locally runnable entries in `COMPETITIVE_MANIFEST`; keep
|
|
1015
|
+
`mcp-codebase-index` blocked while it requires hosted Gemini embeddings.
|
|
1016
|
+
- Touches: `src/competitive-runner.ts`, `src/competitive-manifest.ts`,
|
|
1017
|
+
`scripts/competitive-bakeoff.ts`,
|
|
1018
|
+
`src/__tests__/unit/competitive-runner.spec.ts`,
|
|
1019
|
+
`benchmarks/evaluations/`, `benchmarks/competitors/COMPETITIVE-AUDIT.md`,
|
|
1020
|
+
`benchmarks/competitors/SYNTHESIS.md`, this roadmap
|
|
1021
|
+
- What: normalize competitor cases into a version-pinned, oracle-aware replay
|
|
1022
|
+
contract and reconcile every historical audit claim with the current code.
|
|
1023
|
+
|
|
1024
|
+
### Acceptance criteria
|
|
1025
|
+
|
|
1026
|
+
1. The normalized summary separately records correctness-oracle status,
|
|
1027
|
+
exact-match/precision/recall/false-positive results where applicable, and
|
|
1028
|
+
performance/resource measurements; unavailable fields remain explicit nulls.
|
|
1029
|
+
2. Every audit row has a fixture inventory that either reuses a checked-in
|
|
1030
|
+
labeled fixture or identifies the missing fixture as a blocked prerequisite.
|
|
1031
|
+
3. Every competitor run records pinned version, command, fixture revision,
|
|
1032
|
+
mode, budgets, and reproducible blockers; a partial run cannot be reported
|
|
1033
|
+
as a competitive win.
|
|
1034
|
+
4. The audit labels all existing rows `implemented—needs replay`, `unverified`,
|
|
1035
|
+
`demonstrated gap`, or `non-goal` instead of treating historical results as
|
|
1036
|
+
current product state.
|
|
1037
|
+
|
|
1038
|
+
### C25 — Prove symbol identity and reference precision leadership
|
|
1039
|
+
|
|
1040
|
+
- Status: implemented
|
|
1041
|
+
- Priority: P0
|
|
1042
|
+
- Disposition: Must close
|
|
1043
|
+
- DependsOn: C24, C2, C11
|
|
1044
|
+
- Competitors: GitNexus, Serena, codebase-memory-mcp, CodeGraph
|
|
1045
|
+
- Touches: `benchmarks/evaluations/competitive-symbols/`,
|
|
1046
|
+
`src/__tests__/unit/symbol-identity.spec.ts`,
|
|
1047
|
+
`src/__tests__/unit/reference-precision.spec.ts`,
|
|
1048
|
+
`src/competitive-manifest.ts`, relevant competitor bake-off scripts,
|
|
1049
|
+
`src/engine/index.ts` and `src/tools/knodin-tools.ts` only if replay fails
|
|
1050
|
+
- What: make exact target selection, ambiguity safety, and callers/callees/
|
|
1051
|
+
references objectively comparable on shared duplicate-name and structural
|
|
1052
|
+
implementation fixtures.
|
|
1053
|
+
- Replay evidence:
|
|
1054
|
+
`benchmarks/evaluations/competitive-symbols/raw-results-20260723.json`.
|
|
1055
|
+
knodin achieved 100% precision and recall with zero cross-definition false
|
|
1056
|
+
positives under the shared response budget; every pinned competitor completed
|
|
1057
|
+
against the same fixture and its oracle score is preserved.
|
|
1058
|
+
|
|
1059
|
+
### Acceptance criteria
|
|
1060
|
+
|
|
1061
|
+
1. A labeled fixture covers duplicate names, aliases, barrels, imports versus
|
|
1062
|
+
calls, structural implementations, tests, and at least one unresolved case.
|
|
1063
|
+
2. knodin returns the selected stable identity or an explicit ambiguity result;
|
|
1064
|
+
it never silently conflates viable candidates.
|
|
1065
|
+
3. knodin meets at least 95% precision and recall for labeled resolved
|
|
1066
|
+
call/reference edges, with zero cross-definition false positives.
|
|
1067
|
+
4. The pinned competitors and knodin run against the same fixture and budget;
|
|
1068
|
+
a win requires oracle parity or superiority, not a smaller response.
|
|
1069
|
+
|
|
1070
|
+
### C26 — Prove diff-review and directional traversal leadership
|
|
1071
|
+
|
|
1072
|
+
- Status: implemented
|
|
1073
|
+
- Priority: P1
|
|
1074
|
+
- Disposition: Must close
|
|
1075
|
+
- DependsOn: C24, C4, C5, C13
|
|
1076
|
+
- Competitors: GitNexus, code-review-graph, Graphify, codebase-memory-mcp
|
|
1077
|
+
- Touches: `benchmarks/evaluations/competitive-review/`,
|
|
1078
|
+
`src/__tests__/unit/review-scopes.spec.ts`,
|
|
1079
|
+
`src/__tests__/unit/impact.spec.ts`,
|
|
1080
|
+
`src/__tests__/unit/typed-traversal.spec.ts`, relevant bake-off scripts,
|
|
1081
|
+
engine/tool/CLI code only if replay fails
|
|
1082
|
+
- What: compare staged, unstaged, untracked, and revision-pair review results
|
|
1083
|
+
plus typed directional traversal using source-evidenced expected edges.
|
|
1084
|
+
- Replay evidence:
|
|
1085
|
+
`benchmarks/evaluations/competitive-review/raw-results-20260723.json`.
|
|
1086
|
+
knodin matched every diff-scope, unindexed-file, directional traversal, and
|
|
1087
|
+
edge-evidence oracle under the frozen budget. GitNexus reached traversal
|
|
1088
|
+
parity; Graphify and code-review-graph did not, and the unavailable
|
|
1089
|
+
codebase-memory query adapter remains an explicit blocker rather than a win.
|
|
1090
|
+
|
|
1091
|
+
### Acceptance criteria
|
|
1092
|
+
|
|
1093
|
+
1. The fixture labels changed files, mapped symbols, unindexed files, direct
|
|
1094
|
+
and transitive typed edges, depth, and expected truncation.
|
|
1095
|
+
2. knodin supports every documented diff scope and reports all changed files
|
|
1096
|
+
independently of index coverage.
|
|
1097
|
+
3. Upstream/downstream/both traversal obeys selected relation, confidence,
|
|
1098
|
+
depth, test, and budget controls with oracle-checked edge evidence.
|
|
1099
|
+
4. A replay declares success only when knodin meets the oracle and normalized
|
|
1100
|
+
budget; any competitor prerequisite failure is preserved as a blocker.
|
|
1101
|
+
|
|
1102
|
+
### C27 — Prove architecture and community-analysis leadership
|
|
1103
|
+
|
|
1104
|
+
- Status: implemented
|
|
1105
|
+
- Priority: P1
|
|
1106
|
+
- Disposition: Must close
|
|
1107
|
+
- DependsOn: C24, C3, C12, C13
|
|
1108
|
+
- Competitors: GitNexus, Graphify, codebase-memory-mcp, code-review-graph
|
|
1109
|
+
- Touches: `benchmarks/evaluations/competitive-architecture/`,
|
|
1110
|
+
`src/__tests__/unit/map.spec.ts`, `src/__tests__/unit/uniform-filters.spec.ts`,
|
|
1111
|
+
`src/__tests__/unit/typed-traversal.spec.ts`, relevant bake-off scripts,
|
|
1112
|
+
engine/tool code only if replay fails
|
|
1113
|
+
- What: test boundaries, layers, packages, hubs, bridges, communities,
|
|
1114
|
+
relation filters, pagination, and compact drill-downs against a labeled
|
|
1115
|
+
module-topology fixture.
|
|
1116
|
+
- Replay evidence:
|
|
1117
|
+
`benchmarks/evaluations/competitive-architecture/raw-results-20260723.json`.
|
|
1118
|
+
knodin met the labeled topology oracle, deterministic pagination, and C3
|
|
1119
|
+
budget. Pinned GitNexus, Graphify, and code-review-graph replays completed
|
|
1120
|
+
against the same fixture; codebase-memory remains an explicit policy blocker.
|
|
1121
|
+
|
|
1122
|
+
### Acceptance criteria
|
|
1123
|
+
|
|
1124
|
+
1. The fixture labels module boundaries, allowed cross-boundary edges,
|
|
1125
|
+
packages/layers, hub/bridge expectations, and known absent relationships.
|
|
1126
|
+
2. knodin produces deterministic, filterable, paginated results with explicit
|
|
1127
|
+
totals and C3-compliant byte/token/item budgets.
|
|
1128
|
+
3. Architecture claims retain exact-versus-heuristic confidence and never turn
|
|
1129
|
+
an inferred grouping into a precise boundary claim.
|
|
1130
|
+
4. A pinned replay proves oracle quality and budget parity or better.
|
|
1131
|
+
|
|
1132
|
+
### C28 — Prove dead-code and semantic-search quality
|
|
1133
|
+
|
|
1134
|
+
- Status: implemented
|
|
1135
|
+
- Priority: P1
|
|
1136
|
+
- Disposition: Must close
|
|
1137
|
+
- DependsOn: C24, C3, C11, C12
|
|
1138
|
+
- Competitors: GitNexus, CodeGraph, grepai, Claude Context
|
|
1139
|
+
- Touches: `benchmarks/evaluations/competitive-search/`,
|
|
1140
|
+
`src/__tests__/unit/query.spec.ts`, `src/__tests__/unit/search.spec.ts`,
|
|
1141
|
+
`src/__tests__/unit/response-budget.spec.ts`, relevant bake-off scripts,
|
|
1142
|
+
engine/tool code only if replay fails
|
|
1143
|
+
- What: measure dead-code precision/recall and semantic-orientation relevance
|
|
1144
|
+
using known-live/known-dead and judged-query fixtures.
|
|
1145
|
+
- Replay evidence:
|
|
1146
|
+
`benchmarks/evaluations/competitive-search/raw-results-20260723T214300000Z.json`
|
|
1147
|
+
preserves the frozen local oracle and exact version checks. The pinned
|
|
1148
|
+
GitNexus 1.6.9 adapter normalizes ranked function definitions and its
|
|
1149
|
+
read-only no-incoming-`CALLS` Cypher heuristic to the same `file::symbol`
|
|
1150
|
+
labels. knodin records dead-code precision/recall of 1.0/1.0 versus
|
|
1151
|
+
GitNexus's 0.667/1.0, and mean semantic nDCG/recall@3 of 0.855/1.0 versus
|
|
1152
|
+
0.677/1.0. Both stay within the 32 KiB response budget with no source egress.
|
|
1153
|
+
The remaining CodeGraph, grepai, and Claude Context capability blockers are
|
|
1154
|
+
retained honestly but no longer prevent a comparable pinned competitor replay.
|
|
1155
|
+
|
|
1156
|
+
### Acceptance criteria
|
|
1157
|
+
|
|
1158
|
+
1. Dead-code labels cover exports, imports, reflection-like unresolved uses,
|
|
1159
|
+
tests, barrels, and known live/dead symbols; results report precision,
|
|
1160
|
+
recall, and false positives.
|
|
1161
|
+
2. Search has frozen relevance judgments and reports nDCG and recall@k under a
|
|
1162
|
+
fixed source and response budget.
|
|
1163
|
+
3. knodin labels heuristic dead-code or semantic evidence honestly and offers a
|
|
1164
|
+
bounded drill-down for excluded evidence.
|
|
1165
|
+
4. The replay establishes a win only with oracle parity or better and no
|
|
1166
|
+
regression to C3 budgets.
|
|
1167
|
+
|
|
1168
|
+
### C29 — Prove context-packing and export leadership
|
|
1169
|
+
|
|
1170
|
+
- Status: implemented
|
|
1171
|
+
- Priority: P1
|
|
1172
|
+
- Disposition: Must close
|
|
1173
|
+
- DependsOn: C24, C3, C5, C14
|
|
1174
|
+
- Competitors: Repomix, Aider repo-map, code2prompt
|
|
1175
|
+
- Touches: `benchmarks/evaluations/competitive-context/`,
|
|
1176
|
+
`src/__tests__/unit/context-export.spec.ts`,
|
|
1177
|
+
`src/__tests__/unit/response-budget.spec.ts`, relevant bake-off scripts,
|
|
1178
|
+
`src/context-export.ts` and tool/CLI code only if replay fails
|
|
1179
|
+
- What: compare deterministic packing, exact token accounting, file policy,
|
|
1180
|
+
incremental/ranged retrieval, and optional Git context on a shared corpus.
|
|
1181
|
+
- Replay evidence:
|
|
1182
|
+
`benchmarks/evaluations/competitive-context/raw-results-20260723.json`.
|
|
1183
|
+
knodin met the shared inclusion, exclusion, range, Git, token, and byte
|
|
1184
|
+
oracle for deterministic Markdown, JSON, and XML exports. Repomix and
|
|
1185
|
+
code2prompt completed the bounded corpus replay with their unavailable
|
|
1186
|
+
controls recorded; Aider is explicitly blocked because it has no comparable
|
|
1187
|
+
artifact surface.
|
|
1188
|
+
|
|
1189
|
+
### Acceptance criteria
|
|
1190
|
+
|
|
1191
|
+
1. The fixture labels required/excluded source, tree, line ranges, Git changes,
|
|
1192
|
+
and expected token/byte bounds for each request.
|
|
1193
|
+
2. knodin produces deterministic Markdown, JSON, and XML exports with relative
|
|
1194
|
+
paths and a declared tokenizer estimate.
|
|
1195
|
+
3. Include/exclude and per-file policies compose with hard budgets; ranged
|
|
1196
|
+
reads and exact regex grep return verifiable source positions.
|
|
1197
|
+
4. The competitor replay compares equivalent scope and budget and preserves
|
|
1198
|
+
any unavailable command or format as an explicit blocker.
|
|
1199
|
+
|
|
1200
|
+
### C30 — Prove guarded editing and diagnostics leadership
|
|
1201
|
+
|
|
1202
|
+
- Status: implemented (evaluation found no production-surface gap)
|
|
1203
|
+
- Priority: P2
|
|
1204
|
+
- Disposition: Evaluate
|
|
1205
|
+
- DependsOn: C24, C2, C11, C15
|
|
1206
|
+
- Competitors: Serena, code-review-graph
|
|
1207
|
+
- Touches: `benchmarks/evaluations/competitive-editing/`,
|
|
1208
|
+
`src/__tests__/unit/lsp-readonly.spec.ts`,
|
|
1209
|
+
`src/__tests__/unit/rename-apply.spec.ts`, relevant bake-off scripts,
|
|
1210
|
+
`src/lsp-readonly.ts` and editing surfaces only if evaluation approves them
|
|
1211
|
+
- What: establish whether additional local, guarded edit previews deliver
|
|
1212
|
+
enough value beyond compiler-verified rename and read-only TypeScript
|
|
1213
|
+
language-service queries.
|
|
1214
|
+
- Replay evidence:
|
|
1215
|
+
`benchmarks/evaluations/competitive-editing/raw-results-20260723.json`.
|
|
1216
|
+
The disposable fixture verified diagnostic and navigation goldens,
|
|
1217
|
+
preview-first compiler/test validation, rejected-edit rollback, and explicit
|
|
1218
|
+
unavailable competitor adapters. It found no actionable gap requiring a new
|
|
1219
|
+
production editing surface.
|
|
1220
|
+
|
|
1221
|
+
### Acceptance criteria
|
|
1222
|
+
|
|
1223
|
+
1. Fixtures include diagnostic goldens, definition/implementation navigation,
|
|
1224
|
+
valid edits, rejected edits, rollback, and compile/test verification.
|
|
1225
|
+
2. Any proposed mutation is preview-first, bounded to a disposable copy, and
|
|
1226
|
+
produces a compiler/test result before it can be judged successful.
|
|
1227
|
+
3. No daemon, credential, network access, or unguarded workspace mutation is
|
|
1228
|
+
required; unsupported language adapters return explicit unavailable results.
|
|
1229
|
+
4. Promote an implementation only if the competitor comparison finds an
|
|
1230
|
+
actionable, locally deliverable gap that rename and read-only queries cannot
|
|
1231
|
+
cover.
|
|
1232
|
+
|
|
1233
|
+
### C31 — Prove index lifecycle and output-telemetry leadership
|
|
1234
|
+
|
|
1235
|
+
- Status: done
|
|
1236
|
+
- Priority: P1
|
|
1237
|
+
- Disposition: Must close
|
|
1238
|
+
- DependsOn: C24, C3, C16, C18
|
|
1239
|
+
- Competitors: codebase-memory-mcp, grepai, Claude Context
|
|
1240
|
+
- Touches: `benchmarks/evaluations/competitive-lifecycle/`,
|
|
1241
|
+
`src/__tests__/unit/index-health.spec.ts`,
|
|
1242
|
+
`src/__tests__/unit/perf.spec.ts`, `src/competitive-runner.ts`, relevant
|
|
1243
|
+
bake-off scripts, engine/tool/CLI code only if replay fails
|
|
1244
|
+
- What: compare cold index, incremental index, corruption repair, stale-index
|
|
1245
|
+
recovery, output telemetry, peak RSS, disk, and failure behavior locally.
|
|
1246
|
+
- Replay evidence:
|
|
1247
|
+
`benchmarks/evaluations/competitive-lifecycle/raw-results-20260723T220325Z.json`
|
|
1248
|
+
records disposable four-phase replays for pinned codebase-memory-mcp and
|
|
1249
|
+
grepai with verified versions and behavioral lifecycle oracles. Both expose
|
|
1250
|
+
stale indexed state after a unique symbol replacement, recover through their
|
|
1251
|
+
documented refresh/reindex mechanism, make the replacement queryable, remove
|
|
1252
|
+
the old symbol, and preserve the healthy symbol. Native telemetry inspection
|
|
1253
|
+
deterministically establishes the remaining capability gaps:
|
|
1254
|
+
codebase-memory reports graph/index counts and repository state but lacks
|
|
1255
|
+
detailed per-query output telemetry; grepai reports aggregate query/token
|
|
1256
|
+
savings but lacks per-query latency, truncation, and detail mode. knodin
|
|
1257
|
+
supplies those opt-in local fields. Claude Context remains a separately
|
|
1258
|
+
classified Ollama/Milvus setup-resource blocker rather than an inferred win.
|
|
1259
|
+
|
|
1260
|
+
### Acceptance criteria
|
|
1261
|
+
|
|
1262
|
+
1. The fixture records cold and warm timing, peak RSS, database/disk growth,
|
|
1263
|
+
failure injection, repair outcome, and telemetry fields without source
|
|
1264
|
+
content.
|
|
1265
|
+
2. knodin detects stale/damaged state, provides an actionable repair, and
|
|
1266
|
+
verifies the repaired result without deleting healthy state.
|
|
1267
|
+
3. Persisted telemetry is opt-in, local, and includes operation, latency,
|
|
1268
|
+
serialized bytes, estimated tokens, truncation, and detail mode.
|
|
1269
|
+
4. A replay compares equivalent lifecycle phases; setup or resource failures
|
|
1270
|
+
remain distinct from correctness and performance outcomes.
|
|
1271
|
+
|
|
1272
|
+
### C32 — Prove local API and statement-flow analysis leadership
|
|
1273
|
+
|
|
1274
|
+
- Status: implemented (evaluation found no production-surface gap)
|
|
1275
|
+
- Priority: P2
|
|
1276
|
+
- Disposition: Evaluate
|
|
1277
|
+
- DependsOn: C24, C2, C3, C8, C9
|
|
1278
|
+
- Competitors: GitNexus
|
|
1279
|
+
- Touches: `benchmarks/evaluations/competitive-api-flow/`,
|
|
1280
|
+
`src/__tests__/unit/flow-analysis.spec.ts`,
|
|
1281
|
+
`src/__tests__/unit/api-contracts.spec.ts`, relevant bake-off scripts,
|
|
1282
|
+
engine/tool code only if evaluation proves a gap
|
|
1283
|
+
- What: replay bounded API contract mismatch and statement-level flow results
|
|
1284
|
+
against a labeled fixture, retaining knodin's compact local-only contract.
|
|
1285
|
+
- Replay evidence:
|
|
1286
|
+
`benchmarks/evaluations/competitive-api-flow/raw-results-20260723.json`.
|
|
1287
|
+
The labeled local replay produced 1.00 precision and recall with zero false
|
|
1288
|
+
positives for both bounded API mismatches and variable-filtered statement
|
|
1289
|
+
flow. It verifies source evidence, an explicit C3 continuation at the flow
|
|
1290
|
+
budget, and the one-gateway/local-only constraints. No broader persisted PDG
|
|
1291
|
+
or API surface is justified by this evaluation.
|
|
1292
|
+
|
|
1293
|
+
### Acceptance criteria
|
|
1294
|
+
|
|
1295
|
+
1. The fixture labels true API mismatches, true control/data-flow rows,
|
|
1296
|
+
variable filtering, unresolved cases, and expected source evidence.
|
|
1297
|
+
2. knodin reports precision, recall, false positives, unresolved cases,
|
|
1298
|
+
truncation, and continuation routes under C3 budgets.
|
|
1299
|
+
3. The comparison never treats broad persisted PDG output, raw Cypher, or
|
|
1300
|
+
hosted analysis as required parity; it evaluates only proven local workflows.
|
|
1301
|
+
4. Promote broader implementation only when it finds actionable breakage that
|
|
1302
|
+
existing bounded API/flow queries miss.
|
|
1303
|
+
|
|
1304
|
+
### C33 — Preserve local privacy and one-tool schema economy
|
|
1305
|
+
|
|
1306
|
+
- Status: implemented
|
|
1307
|
+
- Priority: P0
|
|
1308
|
+
- Disposition: Must close
|
|
1309
|
+
- DependsOn: C24
|
|
1310
|
+
- Competitors: all; `mcp-codebase-index` and Claude Context supply negative
|
|
1311
|
+
evidence where hosted services or local service dependencies are required.
|
|
1312
|
+
- Touches: `src/tools/knodin-tools.ts`, `src/server.ts`,
|
|
1313
|
+
`src/competitive-manifest.ts`, `src/__tests__/unit/mcp-tools.spec.ts`,
|
|
1314
|
+
`benchmarks/evaluations/competitive-constraints/`, this roadmap
|
|
1315
|
+
- What: make the local/no-auth/no-egress and one-gateway constraints measurable
|
|
1316
|
+
release gates for every competitive claim.
|
|
1317
|
+
|
|
1318
|
+
### Acceptance criteria
|
|
1319
|
+
|
|
1320
|
+
1. A clean local replay requires no credentials, hosted service, network
|
|
1321
|
+
embeddings, source-code egress, or mandatory external daemon.
|
|
1322
|
+
2. The MCP surface remains one `knodin` gateway; a capability addition must
|
|
1323
|
+
document schema-token cost and cannot create a top-level tool.
|
|
1324
|
+
3. Any optional local service is reported as optional and unavailable states
|
|
1325
|
+
are explicit rather than silently degraded success.
|
|
1326
|
+
4. No item C24-C32 can close unless its replay records this constraint check.
|
|
1327
|
+
|
|
1328
|
+
### C34 — Keep local graph indexes current across repository lifecycle events
|
|
1329
|
+
|
|
1330
|
+
- Status: implemented
|
|
1331
|
+
- Priority: P1
|
|
1332
|
+
- Disposition: Must close
|
|
1333
|
+
- DependsOn: C16, C24, C33
|
|
1334
|
+
- Competitors: GitNexus, Graphify, codebase-memory-mcp
|
|
1335
|
+
- Touches: `bin/cli.ts`, `src/engine/index.ts`, `lefthook.yml`,
|
|
1336
|
+
`templates/hooks/`, `src/__tests__/unit/freshness.spec.ts`,
|
|
1337
|
+
`src/__tests__/unit/watcher-lifecycle.spec.ts`,
|
|
1338
|
+
`benchmarks/evaluations/competitive-lifecycle/`, docs
|
|
1339
|
+
- What: make fresh, reconciled, stale, and unavailable graph states explicit
|
|
1340
|
+
across file edits, branch switches, pulls, rebases, merges, and a restarted
|
|
1341
|
+
local process; provide safe, local refresh commands for every supported graph
|
|
1342
|
+
artifact without pretending an external index is current.
|
|
1343
|
+
- Evidence: `knodin refresh-artifacts [checkout|merge|code-change]` is an
|
|
1344
|
+
explicit opt-in, 30-second-bounded refresh for locally installed GitNexus and
|
|
1345
|
+
Graphify. It prefers the repository's GitNexus runner when present, records
|
|
1346
|
+
success, failure, or skipped state without source content, and is never wired
|
|
1347
|
+
into a commit hook. The lifecycle fixtures prove restart, replacement,
|
|
1348
|
+
branch-switch, merge, watcher-loss, and external-tool states; the replay
|
|
1349
|
+
records freshness latency, RSS, disk, and verified state.
|
|
1350
|
+
|
|
1351
|
+
### Acceptance criteria
|
|
1352
|
+
|
|
1353
|
+
1. knodin continues to reconcile its persisted index at cold start and before
|
|
1354
|
+
answers after offline drift, reporting `fresh`, `reconciled`, or `unknown`.
|
|
1355
|
+
2. The repository exposes a documented local command or opt-in hook path to
|
|
1356
|
+
refresh GitNexus after checkout/merge and Graphify after code changes; it
|
|
1357
|
+
records success, failure, or skipped state and never blocks a commit on an
|
|
1358
|
+
unbounded rebuild.
|
|
1359
|
+
3. A fixture proves correct behavior for edit, deletion, branch switch, merge,
|
|
1360
|
+
pull/rebase-equivalent replacement, watcher loss, and restart; no answer may
|
|
1361
|
+
silently claim an unverified index is fresh.
|
|
1362
|
+
4. The lifecycle replay reports freshness latency and resource cost separately
|
|
1363
|
+
from query quality and preserves any unavailable external tool as a blocker.
|
|
1364
|
+
|
|
1365
|
+
### C35 — Produce local interactive architecture and call-flow visualizations
|
|
1366
|
+
|
|
1367
|
+
- Status: implemented
|
|
1368
|
+
- Priority: P2
|
|
1369
|
+
- Disposition: Should close
|
|
1370
|
+
- DependsOn: C3, C13, C24, C33
|
|
1371
|
+
- Competitors: Graphify, GitNexus
|
|
1372
|
+
- Touches: `bin/cli.ts`, `src/visualization.ts`,
|
|
1373
|
+
`src/__tests__/unit/visualization.spec.ts`,
|
|
1374
|
+
`benchmarks/evaluations/competitive-visualization/`, docs
|
|
1375
|
+
- What: evaluate and, if the evidence justifies it, generate deterministic
|
|
1376
|
+
local HTML architecture and call-flow artifacts from knodin's bounded graph
|
|
1377
|
+
facts without adding a new MCP tool, hosted renderer, source egress, or a
|
|
1378
|
+
mandatory long-lived service.
|
|
1379
|
+
- Evaluation evidence:
|
|
1380
|
+
`benchmarks/evaluations/competitive-visualization/raw-results-20260723.json`
|
|
1381
|
+
records Graphify's pinned community-graph, tree, and call-flow artifacts and
|
|
1382
|
+
GitNexus's explicit lack of a deterministic static-artifact command. knodin's
|
|
1383
|
+
checked-in evaluation artifact is self-contained, deterministic, bounded,
|
|
1384
|
+
source-commit/index-freshness tagged, and source-evidenced. `knodin visualize
|
|
1385
|
+
<entry> --output <path.html>` now exposes that bounded artifact as a local
|
|
1386
|
+
CLI-only export with 1–6 depth, a 4–64 KiB generation ceiling, safe in-repo
|
|
1387
|
+
output paths, and stable identity/file/kind entry disambiguation. It adds no
|
|
1388
|
+
MCP operation, hosted renderer, credential, telemetry, or source egress.
|
|
1389
|
+
|
|
1390
|
+
### Acceptance criteria
|
|
1391
|
+
|
|
1392
|
+
1. The evaluation compares Graphify's community graph, tree, and call-flow
|
|
1393
|
+
artifacts with a knodin artifact generated from the same pinned fixture.
|
|
1394
|
+
2. Generated HTML is self-contained or uses only local static assets, records
|
|
1395
|
+
its source commit/index freshness, and links every displayed edge to bounded
|
|
1396
|
+
source evidence or an explicit heuristic label.
|
|
1397
|
+
3. The artifact supports at least subsystem drill-down, hub/bridge inspection,
|
|
1398
|
+
and bounded call-flow navigation while respecting C3 output and generation
|
|
1399
|
+
budgets.
|
|
1400
|
+
4. It remains an optional CLI/export capability behind the existing `knodin`
|
|
1401
|
+
gateway: no new top-level MCP tool, network service, credential, or source
|
|
1402
|
+
egress is introduced.
|
|
1403
|
+
|
|
1404
|
+
### C36 — Instrument the current warm-operation performance paths
|
|
1405
|
+
|
|
1406
|
+
- Status: implemented
|
|
1407
|
+
- Priority: P0
|
|
1408
|
+
- Disposition: Must close
|
|
1409
|
+
- DependsOn: C18, C24
|
|
1410
|
+
- Evidence: opt-in, source-free phase accounting now attributes freshness,
|
|
1411
|
+
graph analytics, traversal snapshots, architecture facets, context
|
|
1412
|
+
composition, status audit, pack walking/serialization, and artifact
|
|
1413
|
+
read/grep. `perf.spec.ts` verifies the complete phase record without latency
|
|
1414
|
+
thresholds.
|
|
1415
|
+
- Touches: `src/engine/perf.ts`, `src/engine/index.ts`, `src/context.ts`,
|
|
1416
|
+
`src/context-export.ts`, targeted unit tests
|
|
1417
|
+
- What: add opt-in phase accounting around the current production paths without
|
|
1418
|
+
changing response contracts or enabling a competitor replay.
|
|
1419
|
+
|
|
1420
|
+
### Acceptance criteria
|
|
1421
|
+
|
|
1422
|
+
1. Phase names separately cover freshness, graph analytics, traversal snapshot,
|
|
1423
|
+
architecture facets, context composition, status audit, pack walk,
|
|
1424
|
+
serialization, and artifact read/grep.
|
|
1425
|
+
2. Instrumentation remains inert unless a performance session is active and
|
|
1426
|
+
records no source content.
|
|
1427
|
+
3. Targeted tests verify phase accounting without wall-clock thresholds.
|
|
1428
|
+
4. No competitor or before/after comparison is run until C37-C40 are complete
|
|
1429
|
+
and the repository owner explicitly approves the replay.
|
|
1430
|
+
|
|
1431
|
+
### C37 — Reuse generation-scoped graph analytics and traversal snapshots
|
|
1432
|
+
|
|
1433
|
+
- Status: implemented
|
|
1434
|
+
- Priority: P0
|
|
1435
|
+
- Disposition: Must close
|
|
1436
|
+
- DependsOn: C36
|
|
1437
|
+
- Evidence: minimal and standard map now share one generation-keyed analytics
|
|
1438
|
+
snapshot; traversal uses a compact signed adjacency plus cached definition
|
|
1439
|
+
metadata; architecture membership/cohesion/coupling aggregation is one-pass.
|
|
1440
|
+
Option-keyed analytics/minimal-map caches and per-repository traversal/status
|
|
1441
|
+
caches have explicit entry ceilings rather than unbounded process growth.
|
|
1442
|
+
`typed-traversal.spec.ts` proves both map and traversal snapshots invalidate
|
|
1443
|
+
after an incremental index generation.
|
|
1444
|
+
- Touches: `src/engine/index.ts`, `src/__tests__/unit/map.spec.ts`,
|
|
1445
|
+
`src/__tests__/unit/typed-traversal.spec.ts`,
|
|
1446
|
+
`src/__tests__/unit/uniform-filters.spec.ts`
|
|
1447
|
+
- What: cache immutable graph analytics, bidirectional typed adjacency, and
|
|
1448
|
+
symbol metadata by repository/federation configuration plus index generation.
|
|
1449
|
+
|
|
1450
|
+
### Acceptance criteria
|
|
1451
|
+
|
|
1452
|
+
1. Minimal and standard map modes share one generation-scoped analytics source
|
|
1453
|
+
while retaining their existing response shapes and truthful totals.
|
|
1454
|
+
2. Traversal performs bounded BFS over cached adjacency with no per-result
|
|
1455
|
+
symbol lookup and preserves identity, direction, filters, evidence, and
|
|
1456
|
+
deterministic ordering.
|
|
1457
|
+
3. Architecture facets and coupling use one-pass identity-to-community indexes
|
|
1458
|
+
rather than repeated community-by-edge searches.
|
|
1459
|
+
4. Index, reconcile, watcher flush, repair, and close invalidate every affected
|
|
1460
|
+
cache; tests prove no stale result survives a generation change.
|
|
1461
|
+
|
|
1462
|
+
### C38 — Compose orientation context from one shared snapshot
|
|
1463
|
+
|
|
1464
|
+
- Status: implemented
|
|
1465
|
+
- Priority: P1
|
|
1466
|
+
- Disposition: Should close
|
|
1467
|
+
- DependsOn: C37
|
|
1468
|
+
- Evidence: `buildKnodinContext` establishes the standard map snapshot once,
|
|
1469
|
+
then composes stats, flows, and optional review over that initialized
|
|
1470
|
+
generation. `context-composition.spec.ts` verifies at most one probe and the
|
|
1471
|
+
existing compact caps/contracts without timing assertions.
|
|
1472
|
+
- Touches: `src/context.ts`, `src/engine/index.ts`,
|
|
1473
|
+
`src/__tests__/unit/knodin-tools.spec.ts`,
|
|
1474
|
+
`src/__tests__/unit/freshness.spec.ts`
|
|
1475
|
+
- What: acquire one freshness-qualified engine snapshot, reuse C37 analytics,
|
|
1476
|
+
and compose independent context fields without repeating graph construction.
|
|
1477
|
+
|
|
1478
|
+
### Acceptance criteria
|
|
1479
|
+
|
|
1480
|
+
1. One context request performs at most one freshness probe and one graph
|
|
1481
|
+
analytics construction for its repository generation.
|
|
1482
|
+
2. Stats, communities, hubs, flows, and optional risk retain their current
|
|
1483
|
+
contract, sorting, caps, and evidence.
|
|
1484
|
+
3. Independent work may run concurrently only after initialization is complete;
|
|
1485
|
+
no SQLite writer concurrency or duplicate engine singleton is introduced.
|
|
1486
|
+
4. Tests use counters and cache state, not timing thresholds.
|
|
1487
|
+
|
|
1488
|
+
### C39 — Separate fast truthful status from explicit deep audit
|
|
1489
|
+
|
|
1490
|
+
- Status: implemented
|
|
1491
|
+
- Priority: P1
|
|
1492
|
+
- Disposition: Should close
|
|
1493
|
+
- DependsOn: C36
|
|
1494
|
+
- Evidence: warm status reuses a generation/database-fingerprint keyed deep
|
|
1495
|
+
audit only after the existing bounded drift probe verifies the source state.
|
|
1496
|
+
Responses identify `deep-audit` versus
|
|
1497
|
+
`cached-after-freshness-probe` and include `verifiedAt`; `--deep` and MCP
|
|
1498
|
+
`statusAudit: deep` force the complete audit. `index-health.spec.ts` proves
|
|
1499
|
+
warm reuse, explicit deep mode, and file-change invalidation.
|
|
1500
|
+
- Touches: `src/engine/index.ts`, `src/tools/knodin-tools.ts`, `bin/cli.ts`,
|
|
1501
|
+
`src/__tests__/unit/index-health.spec.ts`
|
|
1502
|
+
- What: return an invalidation-safe cached health snapshot for warm status and
|
|
1503
|
+
expose the existing filesystem/database verification as an explicit deep
|
|
1504
|
+
audit, with verification age and mode reported truthfully.
|
|
1505
|
+
|
|
1506
|
+
### Acceptance criteria
|
|
1507
|
+
|
|
1508
|
+
1. Default warm status never claims a fresh audit when it is returning cached
|
|
1509
|
+
evidence; it reports verification mode and timestamp.
|
|
1510
|
+
2. Deep audit preserves current missing, damaged, orphaned, coverage, and repair
|
|
1511
|
+
behavior.
|
|
1512
|
+
3. Index generation, watcher changes, repair, schema change, and relevant file
|
|
1513
|
+
events invalidate the cached snapshot.
|
|
1514
|
+
4. Existing callers remain source-compatible and a caller can explicitly
|
|
1515
|
+
request the deep audit.
|
|
1516
|
+
|
|
1517
|
+
### C40 — Make context packing and artifact access linear and cache-safe
|
|
1518
|
+
|
|
1519
|
+
- Status: implemented
|
|
1520
|
+
- Priority: P1
|
|
1521
|
+
- Disposition: Should close
|
|
1522
|
+
- DependsOn: C36
|
|
1523
|
+
- Evidence: glob/policy matchers are bounded and compiled once, walking uses one
|
|
1524
|
+
accumulator, exact per-file serialization deltas replace accumulated
|
|
1525
|
+
reserialization, and artifact lines use a two-entry/four-MiB-per-file
|
|
1526
|
+
path/mtime/size cache. Context-export and response-budget tests preserve
|
|
1527
|
+
formats, policies, ordering, hard budgets, safe paths, reads, grep, and cache
|
|
1528
|
+
invalidation.
|
|
1529
|
+
- Touches: `src/context-export.ts`,
|
|
1530
|
+
`src/__tests__/unit/context-export.spec.ts`,
|
|
1531
|
+
`src/__tests__/unit/response-budget.spec.ts`
|
|
1532
|
+
- What: precompile selection policy, prune excluded directories, account bytes
|
|
1533
|
+
incrementally, serialize once, and reuse a bounded artifact line index keyed
|
|
1534
|
+
by path/mtime/size.
|
|
1535
|
+
|
|
1536
|
+
### Acceptance criteria
|
|
1537
|
+
|
|
1538
|
+
1. Candidate walking and pattern matching are linear in visited paths plus
|
|
1539
|
+
patterns; patterns are compiled once per request.
|
|
1540
|
+
2. Packing does not rebuild the entire accumulated artifact per candidate and
|
|
1541
|
+
still enforces exact response byte/token ceilings.
|
|
1542
|
+
3. Artifact read/grep reuse a bounded cache only when path, size, and mtime
|
|
1543
|
+
match; mutation and replacement invalidate it.
|
|
1544
|
+
4. Markdown, JSON, XML, policies, tree/Git sections, deterministic ordering,
|
|
1545
|
+
ranged reads, regex grep, and telemetry remain byte-for-byte compatible in
|
|
1546
|
+
golden fixtures.
|
|
1547
|
+
|
|
1548
|
+
### C41 — Replay like-for-like performance after the implementation batch
|
|
1549
|
+
|
|
1550
|
+
- Status: evaluated — deferred
|
|
1551
|
+
- Priority: P0
|
|
1552
|
+
- Disposition: Evaluate
|
|
1553
|
+
- DependsOn: C37, C38, C39, C40
|
|
1554
|
+
- Evidence: the terminal bounded replay and strengthened-oracle verdict are
|
|
1555
|
+
preserved at
|
|
1556
|
+
`benchmarks/competitors/runs/20260802T203822521Z-3e27547d095e/`. Its
|
|
1557
|
+
`c41-verdict.json` verifies all five competitors, explicit warm/cold p50/p95,
|
|
1558
|
+
independently attributed two-sided RSS, and grepai cached/deep lifecycle
|
|
1559
|
+
rows. Under the strengthened labels, 19 mapped rows passed, 10 failed, four
|
|
1560
|
+
grepai trace rows were unavailable, and six intentionally unmatched Repomix
|
|
1561
|
+
rows remained incomparable. Graphify passed its direct-call rows but not hub
|
|
1562
|
+
rows; codebase-memory passed six of ten search/traversal/package rows;
|
|
1563
|
+
Repomix passed seven pack/read/grep rows; code2prompt passed both pack rows;
|
|
1564
|
+
and grepai produced no oracle-qualified latency row because its local index
|
|
1565
|
+
remained empty. The bounded local-Ollama indexing attempts at
|
|
1566
|
+
`20260802T205026646Z-18c1107af915/`,
|
|
1567
|
+
`20260802T205625288Z-e97db4c32eb6/` preserve that setup limitation. Earlier
|
|
1568
|
+
runs retain the sandbox-launch, timeout, and uninitialized-index failures
|
|
1569
|
+
rather than overwriting them.
|
|
1570
|
+
- Evaluation disposition: retain the current product unchanged and defer any
|
|
1571
|
+
aggregate performance claim or optimization response. Only oracle-passing
|
|
1572
|
+
exact rows may support operation-specific observations; failures,
|
|
1573
|
+
unavailable rows, and setup limits remain separate dimensions and cannot
|
|
1574
|
+
become wins.
|
|
1575
|
+
- Touches if approved: competitor harness mappings, immutable run artifacts,
|
|
1576
|
+
`benchmarks/competitors/COMPETITIVE-AUDIT.md`, this roadmap
|
|
1577
|
+
- What: after explicit repository-owner permission, run warm/cold p50/p95/RSS
|
|
1578
|
+
measurements using equivalent operations, budgets, fixtures, and correctness
|
|
1579
|
+
oracles. A user instruction to complete this roadmap constitutes permission
|
|
1580
|
+
for the checked-in, zero-spend, local replay only; network access, installation
|
|
1581
|
+
of unpinned software, publication, or external mutation still requires its own
|
|
1582
|
+
authority.
|
|
1583
|
+
|
|
1584
|
+
### Evaluation gate
|
|
1585
|
+
|
|
1586
|
+
1. Graphify traversal/map rows use the same direction, relation, depth, item,
|
|
1587
|
+
token, and evidence scope.
|
|
1588
|
+
2. codebase-memory architecture/search rows use the same facets, pagination,
|
|
1589
|
+
source inclusion, and correctness labels.
|
|
1590
|
+
3. code2prompt and Repomix compare with knodin `pack`, packed read, and packed
|
|
1591
|
+
grep—not semantic `context`, `search`, or `file_summary`.
|
|
1592
|
+
4. grepai status compares cached summary with cached summary and deep audit with
|
|
1593
|
+
an equivalent verified lifecycle phase.
|
|
1594
|
+
5. Results record correctness before latency and treat RSS/setup failures as
|
|
1595
|
+
separate dimensions.
|
|
1596
|
+
|
|
1597
|
+
### C42 — Enforce comparable competitive cases and correctness oracles
|
|
1598
|
+
|
|
1599
|
+
- Status: implemented
|
|
1600
|
+
- Priority: P0
|
|
1601
|
+
- DependsOn: C41
|
|
1602
|
+
- Evidence: `src/competitive-contract.ts` and the competitor mappings enforce
|
|
1603
|
+
typed equivalence/oracle dispositions; the Repomix replay qualifies 18
|
|
1604
|
+
comparable rows and excludes wrong-fast, partial, unverified, and
|
|
1605
|
+
incomparable rows from wins.
|
|
1606
|
+
- Touches: `src/competitive-manifest.ts`, `src/competitive-runner.ts`,
|
|
1607
|
+
`scripts/competitive-bakeoff.ts`, competitor adapters, labeled competitive
|
|
1608
|
+
fixtures and runner tests
|
|
1609
|
+
- What: define one typed case contract for equivalent operation scope, budgets,
|
|
1610
|
+
fixtures, and correctness. Incomparable or wrong-fast rows must never count as
|
|
1611
|
+
latency wins.
|
|
1612
|
+
|
|
1613
|
+
### Acceptance criteria
|
|
1614
|
+
|
|
1615
|
+
1. Every timed pair records use case, fixture revision, mode, scope, budgets,
|
|
1616
|
+
and a comparable/incomparable disposition with a reason.
|
|
1617
|
+
2. Graphify, codebase-memory, Repomix, code2prompt, and grepai mappings satisfy
|
|
1618
|
+
C41's exact equivalence rules; unmatched rows are excluded explicitly.
|
|
1619
|
+
3. Comparable rows carry checked-in correctness oracles, and summaries
|
|
1620
|
+
distinguish oracle pass, oracle failure, unavailable, and incomparable.
|
|
1621
|
+
4. Partial runs and rows without passing oracles cannot produce a competitive
|
|
1622
|
+
win or aggregate performance claim.
|
|
1623
|
+
|
|
1624
|
+
### C43 — Measure cold/warm p50, p95, and independent process RSS
|
|
1625
|
+
|
|
1626
|
+
- Status: implemented
|
|
1627
|
+
- Priority: P0
|
|
1628
|
+
- DependsOn: C42
|
|
1629
|
+
- Evidence: `benchmarks/competitors/MEASUREMENT-CONTRACT.md` and shared sampling
|
|
1630
|
+
record three warmups, at least 20 warm samples, raw values, p50/p95/range,
|
|
1631
|
+
independently attributed RSS, descendant cleanup, and atomic partial
|
|
1632
|
+
summaries across all 11 adapters.
|
|
1633
|
+
- Touches: shared competitive measurement helper, competitor adapters,
|
|
1634
|
+
`scripts/competitive-bakeoff.ts`, lifecycle runner, normalized summaries and
|
|
1635
|
+
tests
|
|
1636
|
+
- What: replace copied three-sample loops and aggregate process-tree ceilings
|
|
1637
|
+
with a shared statistical and resource-measurement contract.
|
|
1638
|
+
|
|
1639
|
+
### Acceptance criteria
|
|
1640
|
+
|
|
1641
|
+
1. Warm and cold samples are separate and retain raw samples, warm-up count,
|
|
1642
|
+
p50, p95, minimum, and maximum; the documented warm sample minimum supports a
|
|
1643
|
+
meaningful p95.
|
|
1644
|
+
2. knodin and competitor RSS are attributed independently by process and phase;
|
|
1645
|
+
combined harness RSS is never assigned to either product.
|
|
1646
|
+
3. Per-call and per-competitor timeouts reap descendants, preserve partial
|
|
1647
|
+
evidence, identify the failed phase/case, and still write a terminal summary.
|
|
1648
|
+
4. Regression checks compare matching modes and enforce both p50 and p95
|
|
1649
|
+
tolerances after correctness passes.
|
|
1650
|
+
|
|
1651
|
+
### C44 — Restore low-latency diff review without weakening evidence
|
|
1652
|
+
|
|
1653
|
+
- Status: implemented
|
|
1654
|
+
- Priority: P0
|
|
1655
|
+
- DependsOn: C43
|
|
1656
|
+
- Evidence: `benchmarks/evaluations/competitive-review/` records an
|
|
1657
|
+
oracle-qualified bounded-gateway replay at 53.018 ms p50, 89.449 ms p95, and
|
|
1658
|
+
733 bytes; review tests preserve path-scoped churn/risk, diff scopes, flows,
|
|
1659
|
+
cache invalidation, and immutable evidence provenance.
|
|
1660
|
+
- Touches: `src/engine/index.ts`, `src/engine/perf.ts`, review/flow tests and
|
|
1661
|
+
competitive-review evidence
|
|
1662
|
+
- What: remove per-file Git history processes and whole-repository repeated work
|
|
1663
|
+
from warm review while preserving changed-file, risk, flow, and source
|
|
1664
|
+
evidence.
|
|
1665
|
+
|
|
1666
|
+
### Acceptance criteria
|
|
1667
|
+
|
|
1668
|
+
1. Review phase telemetry separates diff discovery, churn, parsing, flow lookup,
|
|
1669
|
+
and serialization.
|
|
1670
|
+
2. Churn/history is collected in one bounded pass or a repository/HEAD cache,
|
|
1671
|
+
and affected flows use a generation-scoped file-to-flow lookup.
|
|
1672
|
+
3. All diff scopes and existing correctness fixtures remain unchanged.
|
|
1673
|
+
4. The oracle-qualified warm minimal review meets the recorded C43 regression
|
|
1674
|
+
threshold on both p50 and p95.
|
|
1675
|
+
|
|
1676
|
+
### C45 — Add facet-selective architecture fast paths
|
|
1677
|
+
|
|
1678
|
+
- Status: implemented
|
|
1679
|
+
- Priority: P1
|
|
1680
|
+
- DependsOn: C43
|
|
1681
|
+
- Evidence: `benchmarks/evaluations/competitive-architecture/` records an
|
|
1682
|
+
oracle-qualified bounded-gateway replay at 0.082 ms p50 and 0.097 ms p95;
|
|
1683
|
+
architecture tests prove facet-selective construction, zero minimal edge
|
|
1684
|
+
materialization, mutation isolation, budgets, and generation/close
|
|
1685
|
+
invalidation.
|
|
1686
|
+
- Touches: `src/engine/index.ts`, architecture/map tests and
|
|
1687
|
+
competitive-architecture evidence
|
|
1688
|
+
- What: compute only requested architecture facets from reusable
|
|
1689
|
+
generation-scoped metadata instead of building a complete standard map.
|
|
1690
|
+
|
|
1691
|
+
### Acceptance criteria
|
|
1692
|
+
|
|
1693
|
+
1. Language, package, layer, entry-point, boundary, coupling, and hotspot
|
|
1694
|
+
requests construct only their required inputs.
|
|
1695
|
+
2. Minimal facet requests do not build full edge lists or hydrate unrelated
|
|
1696
|
+
source/symbol detail.
|
|
1697
|
+
3. Identity, filters, pagination, totals, budgets, and cache invalidation remain
|
|
1698
|
+
correct.
|
|
1699
|
+
4. Equivalent architecture rows pass their oracle and C43 thresholds.
|
|
1700
|
+
|
|
1701
|
+
### C46 — Cache search state and hydrate only the requested page
|
|
1702
|
+
|
|
1703
|
+
- Status: implemented
|
|
1704
|
+
- Priority: P1
|
|
1705
|
+
- DependsOn: C43
|
|
1706
|
+
- Evidence: `benchmarks/evaluations/competitive-search/` records a
|
|
1707
|
+
same-oracle 500-symbol offset-page improvement from 5.309/5.865 ms to
|
|
1708
|
+
0.310/0.481 ms p50/p95; tests preserve exact-default search, global ordering,
|
|
1709
|
+
totals, pagination, filters, ANN gating, federation, corruption, and budgets.
|
|
1710
|
+
- Touches: `src/engine/index.ts`, search/tool contracts, search and ANN tests,
|
|
1711
|
+
competitive-search evidence
|
|
1712
|
+
- What: reuse decoded generation-scoped search metadata, rank before hydration,
|
|
1713
|
+
and perform source/caller work only for the requested page.
|
|
1714
|
+
|
|
1715
|
+
### Acceptance criteria
|
|
1716
|
+
|
|
1717
|
+
1. Cheap filters run before vector scoring and source hydration; community
|
|
1718
|
+
lookup does not build a full map.
|
|
1719
|
+
2. Exact search remains the default unless the existing ANN recall gate passes.
|
|
1720
|
+
3. Global ordering, offset pagination, totals, corruption handling, federation,
|
|
1721
|
+
and response budgets remain deterministic.
|
|
1722
|
+
4. A separately specified opaque cursor may be added only if it binds query,
|
|
1723
|
+
filters, order, and generation and truthfully rejects stale cursors.
|
|
1724
|
+
|
|
1725
|
+
### C47 — Reduce status, repair, and measured lifecycle resource cost
|
|
1726
|
+
|
|
1727
|
+
- Status: implemented
|
|
1728
|
+
- Priority: P1
|
|
1729
|
+
- DependsOn: C43
|
|
1730
|
+
- Evidence: `benchmarks/evaluations/competitive-lifecycle/` records isolated,
|
|
1731
|
+
ownership-labeled lifecycle stages and immutable provenance; health tests
|
|
1732
|
+
prove lease expiry guards, single-flight deduplicated batch repair, one
|
|
1733
|
+
invalidation, and final forced deep verification.
|
|
1734
|
+
- Touches: `src/engine/index.ts`, `src/engine/perf.ts`, lifecycle runner,
|
|
1735
|
+
freshness/index-health/watcher tests
|
|
1736
|
+
- What: eliminate duplicate freshness/audit work, batch repair indexing, and
|
|
1737
|
+
reduce only resource retention demonstrated by isolated C43 measurements.
|
|
1738
|
+
|
|
1739
|
+
### Acceptance criteria
|
|
1740
|
+
|
|
1741
|
+
1. Cached status reuses a valid freshness lease and deep-audit snapshot without
|
|
1742
|
+
claiming unverified freshness; lease expiry still performs the missed-watcher
|
|
1743
|
+
guard.
|
|
1744
|
+
2. Repair reuses valid audit evidence, deduplicates paths, batches indexing and
|
|
1745
|
+
orphan cleanup, invalidates once, and completes one final deep verification.
|
|
1746
|
+
3. The lifecycle runner records isolated baseline, post-index, post-query,
|
|
1747
|
+
post-close, peak RSS, and disk measurements for each product.
|
|
1748
|
+
4. Cache/parser changes are accepted only when attribution proves a retained
|
|
1749
|
+
owner and lifecycle correctness remains intact.
|
|
1750
|
+
|
|
1751
|
+
### C48 — Decide an optional local project-memory boundary
|
|
1752
|
+
|
|
1753
|
+
- Status: evaluated — deferred
|
|
1754
|
+
- Priority: P2
|
|
1755
|
+
- DependsOn: C42
|
|
1756
|
+
- Decision: defer production work. The evaluated workflows do not yet
|
|
1757
|
+
demonstrate value beyond repository-owned documentation, while a memory
|
|
1758
|
+
surface would add a second source of truth, retention/privacy obligations,
|
|
1759
|
+
and an output-budget surface. See
|
|
1760
|
+
[`docs/adr/004-optional-project-memory.md`](../docs/adr/004-optional-project-memory.md)
|
|
1761
|
+
and the
|
|
1762
|
+
[C48 evaluation](../benchmarks/evaluations/c48-project-memory/evaluation.md).
|
|
1763
|
+
- Touches: roadmap/ADR and Serena memory evaluation evidence; no production
|
|
1764
|
+
code because the evaluation gate did not pass
|
|
1765
|
+
- What: evaluate concrete durable-note workflows without turning generalized
|
|
1766
|
+
agent memory into a core graph/index requirement.
|
|
1767
|
+
|
|
1768
|
+
### Evaluation gate
|
|
1769
|
+
|
|
1770
|
+
1. Record user workflows, path containment, retention/backup, privacy, output
|
|
1771
|
+
budgets, and RSS/disk cost against Serena's memory operations.
|
|
1772
|
+
2. Prefer an optional local `.reckon/notes/` adapter that does not affect graph
|
|
1773
|
+
identity, indexing, freshness, or the one-tool gateway count.
|
|
1774
|
+
3. Approve production work only when a labeled workflow demonstrates value
|
|
1775
|
+
beyond ordinary repository documentation.
|
|
1776
|
+
4. If approved, require atomic writes, traversal protection, deterministic
|
|
1777
|
+
list/read/edit/rename/delete behavior, and explicit size budgets.
|
|
1778
|
+
|
|
1779
|
+
### C49 — Keep immutable runs lint-safe and make the full suite terminate
|
|
1780
|
+
|
|
1781
|
+
- Status: implemented
|
|
1782
|
+
- Priority: P1
|
|
1783
|
+
- DependsOn: —
|
|
1784
|
+
- Evidence: Biome excludes only immutable `benchmarks/competitors/runs/**`;
|
|
1785
|
+
handle diagnostics and teardown coverage close watchers, databases, queues,
|
|
1786
|
+
and child processes. The full coverage suite terminates naturally within the
|
|
1787
|
+
documented six-minute ceiling (474/474 in the final C47 gate).
|
|
1788
|
+
- Touches: `biome.json`, `package.json`, Vitest configuration/setup,
|
|
1789
|
+
watcher/freshness teardown tests, competitive-run documentation
|
|
1790
|
+
- What: exclude immutable generated evidence from formatting gates while
|
|
1791
|
+
diagnosing and fixing the open handle that prevents the full suite from
|
|
1792
|
+
exiting.
|
|
1793
|
+
|
|
1794
|
+
### Acceptance criteria
|
|
1795
|
+
|
|
1796
|
+
1. Biome ignores `benchmarks/competitors/runs/**` but continues checking
|
|
1797
|
+
generators, baselines, labels, schemas, and source.
|
|
1798
|
+
2. A hanging-process diagnostic command identifies active handles without
|
|
1799
|
+
masking them through force-exit.
|
|
1800
|
+
3. Watchers, databases, queues, and spawned test children are asserted closed in
|
|
1801
|
+
teardown, including injected-failure paths.
|
|
1802
|
+
4. `bun run lint` and the full test suite terminate naturally within the
|
|
1803
|
+
repository's documented ceiling.
|
|
1804
|
+
|
|
1805
|
+
### C50 — Make the single-worker full-suite gate start reliably
|
|
1806
|
+
|
|
1807
|
+
- Status: implemented
|
|
1808
|
+
- Priority: P0
|
|
1809
|
+
- DependsOn: C49
|
|
1810
|
+
- Touches: Vitest configuration/setup, embedding test doubles and teardown,
|
|
1811
|
+
gate documentation, and a worker-start regression fixture
|
|
1812
|
+
- What: eliminate repeated native ONNX initialization or worker churn that can
|
|
1813
|
+
exhaust Vitest's 600-second pool ceiling before all test files start, without
|
|
1814
|
+
skipping files, forcing exit, weakening isolation blindly, or changing
|
|
1815
|
+
production embedding behavior.
|
|
1816
|
+
|
|
1817
|
+
### Acceptance criteria
|
|
1818
|
+
|
|
1819
|
+
1. A diagnostic reproduces and attributes the failed fork-worker startup for
|
|
1820
|
+
`knodin.test.ts`, `map.spec.ts`, and the lifecycle runner instead of
|
|
1821
|
+
increasing the timeout.
|
|
1822
|
+
2. Tests use an explicit deterministic embedder where semantic model quality is
|
|
1823
|
+
not under test; dedicated embedding tests retain production-path coverage.
|
|
1824
|
+
3. All test files start, the full suite passes, and each fresh-worker shard
|
|
1825
|
+
exits naturally under its bounded five-minute test ceiling with immutable
|
|
1826
|
+
run artifacts present in the checkout.
|
|
1827
|
+
4. Lint, typecheck, build, coverage thresholds, native-resource teardown, and
|
|
1828
|
+
production local-embedding behavior remain unchanged.
|
|
1829
|
+
|
|
1830
|
+
### Evidence and corrected diagnosis
|
|
1831
|
+
|
|
1832
|
+
- Root cause (corrected): the handoff attributed the startup failures to repeated
|
|
1833
|
+
ONNX init, but a diagnostic reproduced them with the three ONNX-free canary
|
|
1834
|
+
files alone. The real cause is `isolate: true` (Vitest default): the pool
|
|
1835
|
+
re-forks a **fresh process per test file** — the config's stated "keep one
|
|
1836
|
+
long-lived worker" intent was never achieved — and under host load a per-file
|
|
1837
|
+
fork spawn intermittently exceeds Vitest's 60 s worker-start handshake
|
|
1838
|
+
(`START_TIMEOUT`), failing a neighbouring file. Actual test work is ~200 s
|
|
1839
|
+
(well under the ceiling); the overage was pure per-file fork churn.
|
|
1840
|
+
- Fix: `isolate: false` + `singleFork` (one reused process — the documented
|
|
1841
|
+
intent, not a blind isolation weakening), so native modules init once and only
|
|
1842
|
+
one fork is ever spawned. Cross-file safety is preserved by explicit guards:
|
|
1843
|
+
the shared deterministic lexical embedder is installed by setup and restored by
|
|
1844
|
+
every teardown; a per-test env/cwd snapshot-restore plus `restoreMocks`/
|
|
1845
|
+
`unstubEnvs`; the real ONNX model is isolated to `embeddings.spec` (loaded once,
|
|
1846
|
+
disposed after) and guarded by a `getRealEmbedderLoadCount()` regression; the
|
|
1847
|
+
six competitive benchmark `runner.spec.ts` files spawn their sub-benchmarks
|
|
1848
|
+
asynchronously (no synchronous `execFileSync` blocking the worker); and removed
|
|
1849
|
+
per-test repo state (SQLite handles + FSWatcher/FSEvent descriptors) is evicted
|
|
1850
|
+
each test.
|
|
1851
|
+
- Verified: the gate uses 16 strictly sequential, fresh-worker shards so
|
|
1852
|
+
native SQLite/watcher state cannot accumulate across the whole suite. The
|
|
1853
|
+
two semantic-ranking specs pass on the deterministic lexical embedder; lint,
|
|
1854
|
+
typecheck, and build pass. Final integration certification passed all 87
|
|
1855
|
+
files/503 tests in 1,471.79 seconds with 597,917,696 bytes peak RSS, and its
|
|
1856
|
+
merged 503-test coverage run passed at 89.5% statements, 79.25% branches,
|
|
1857
|
+
90.62% functions, and 91.8% lines (901,840,896 bytes peak RSS). No tests or
|
|
1858
|
+
files are quarantined or skipped to obtain these results.
|
|
1859
|
+
- Release certification still reruns the fully accumulated integration checkout
|
|
1860
|
+
under C54; this distinction prevents cumulative-branch evidence from being
|
|
1861
|
+
mislabeled as a final integration-branch run.
|
|
1862
|
+
|
|
1863
|
+
### C51 — Native training-free embedding quantization (memory-efficient vector search)
|
|
1864
|
+
|
|
1865
|
+
- Status: evaluated — rejected for production
|
|
1866
|
+
- Priority: P2
|
|
1867
|
+
- DependsOn: C46
|
|
1868
|
+
- Touches: `src/engine/embeddings.ts`, `src/engine/index.ts` (exact scan + ANN
|
|
1869
|
+
path), the `symbol_embeddings` storage, and a recall-gate fixture
|
|
1870
|
+
- What: add an opt-in, in-process, data-oblivious quantization of the 384-dim
|
|
1871
|
+
symbol embeddings — normalize → fixed deterministic rotation → precomputed
|
|
1872
|
+
bucketing (the TurboQuant concept, implemented natively in TypeScript) —
|
|
1873
|
+
shrinking stored vectors ~8× (4-bit) and speeding the cosine scan, with float32
|
|
1874
|
+
remaining authoritative and the safe default. This is the memory-efficient
|
|
1875
|
+
vector-search lever that keeps the zero-auth / local / no-egress / one-process
|
|
1876
|
+
guarantees intact, versus competitors that reach for a hosted vector DB. It is
|
|
1877
|
+
explicitly **not** a sidecar, external index, Rust dependency, or hosted store.
|
|
1878
|
+
See [ADR 005](../docs/adr/005-native-embedding-quantization.md).
|
|
1879
|
+
- Evidence: `benchmarks/evaluations/c51-quantization/raw-results.json` binds a
|
|
1880
|
+
2,048-vector real MiniLM fixture and preserves per-query top-5/top-10 overlap,
|
|
1881
|
+
rank correlation, estimated database/resident payload bytes, conversion time,
|
|
1882
|
+
3+20 warm and five cold samples, peak RSS, constraints, gates, and disposition. The fixed-seed
|
|
1883
|
+
block-Hadamard 4-bit prototype used 20 unique queries per tier and achieved
|
|
1884
|
+
mean top-10 recall 96%, 93.5%, and 92% at 128, 512, and 2,048 vectors. It
|
|
1885
|
+
therefore failed the >=95% gate at two engagement sizes.
|
|
1886
|
+
Payloads were 8× smaller, but warm p50 was not materially better and
|
|
1887
|
+
conversion-inclusive cold p50 was worse at every size; 99,876,864-byte peak
|
|
1888
|
+
process RSS did not establish an attributable end-to-end RSS win.
|
|
1889
|
+
- Disposition: reject production quantization from this design. Float32 remains
|
|
1890
|
+
authoritative and unchanged. Representation shrink without the recall gate
|
|
1891
|
+
and a material end-to-end latency or memory win does not authorize production
|
|
1892
|
+
work, so no separately numbered implementation item is proposed.
|
|
1893
|
+
|
|
1894
|
+
### Acceptance criteria
|
|
1895
|
+
|
|
1896
|
+
Evaluation outcome: criteria 1 and 4 hold for the prototype; criterion 3 fails
|
|
1897
|
+
at 512 and 2,048 vectors. Criterion 2 is deliberately not implemented because
|
|
1898
|
+
the failed evaluation does not authorize a production path or environment gate.
|
|
1899
|
+
|
|
1900
|
+
1. Encode/decode is deterministic (hard-coded rotation seed) and reproducible
|
|
1901
|
+
across machines, with no trained or per-repo codebook and no calibration pass.
|
|
1902
|
+
2. Quantized codes are opt-in behind an env gate (default off); float32 stays the
|
|
1903
|
+
source of truth and is rebuildable/invalidated by the existing generation
|
|
1904
|
+
counter.
|
|
1905
|
+
3. A committed real-vector fixture proves at least 95% mean top-10 recall against
|
|
1906
|
+
the float32 exact scan at every size where the path could engage, mirroring
|
|
1907
|
+
the R18/R48 gate. Report top-5/top-10 overlap, rank correlation, database and
|
|
1908
|
+
resident bytes, conversion time, cold/warm p50/p95, and peak RSS. Eligibility
|
|
1909
|
+
additionally requires a material end-to-end latency or memory win; an inner-
|
|
1910
|
+
loop-only improvement is insufficient.
|
|
1911
|
+
4. No external index, sidecar, Rust/native dependency, hosted DB, or source/vector
|
|
1912
|
+
egress; the embedding model, its dimensionality, and its normalization are
|
|
1913
|
+
unchanged.
|
|
1914
|
+
|
|
1915
|
+
### C52 — Package repository lifecycle initialization and refresh
|
|
1916
|
+
|
|
1917
|
+
- Status: implemented
|
|
1918
|
+
- Priority: P0
|
|
1919
|
+
- DependsOn: C16, C34, C50
|
|
1920
|
+
- Touches: `src/init.ts`, `src/repository-management.ts`, the deprecated
|
|
1921
|
+
`src/fleet.ts` compatibility adapter, `src/indexable-paths.ts`,
|
|
1922
|
+
`bin/cli.ts`, `src/engine/index.ts`, `scripts/pack-install-smoke.ts`,
|
|
1923
|
+
repository-init/CLI acceptance tests, package build output
|
|
1924
|
+
- What: make the compiled package usable across existing repositories without a
|
|
1925
|
+
knodin source checkout. Resolve the installed executable in hooks, share one
|
|
1926
|
+
canonical indexable-path policy, refresh after commit/checkout/merge/rewrite,
|
|
1927
|
+
and initialize repository checkouts idempotently.
|
|
1928
|
+
|
|
1929
|
+
### Acceptance and evidence
|
|
1930
|
+
|
|
1931
|
+
1. `knodin init` preserves existing hook content, supports worktrees and
|
|
1932
|
+
`core.hooksPath`, and repeated initialization produces no duplicate knodin
|
|
1933
|
+
blocks.
|
|
1934
|
+
2. Hook refresh selects paths through the engine's canonical policy, including
|
|
1935
|
+
TypeScript and supported Salesforce Apex, Aura, LWC, and metadata inputs.
|
|
1936
|
+
3. `knodin repos init` discovers Git roots deterministically, deduplicates real
|
|
1937
|
+
paths/worktrees, supports dry-run and JSON output, and records individual
|
|
1938
|
+
failures without abandoning the remaining repositories. The historical
|
|
1939
|
+
`knodin fleet init` spelling remains a time-bounded compatibility alias.
|
|
1940
|
+
4. `scripts/pack-install-smoke.ts` installs the exact tarball in a clean external
|
|
1941
|
+
repository and proves package-resolved hooks across post-commit,
|
|
1942
|
+
post-checkout, post-merge, post-rewrite, rename, and deletion lifecycle
|
|
1943
|
+
events.
|
|
1944
|
+
|
|
1945
|
+
Evidence: packaged hooks landed in `c3a10a5`, fleet initialization in `72211ab`,
|
|
1946
|
+
and the clean-consumer gate in `1a6dbca`. The 0.1.0 candidate gate completed 49
|
|
1947
|
+
sequential commands with 465 MiB peak aggregate RSS.
|
|
1948
|
+
|
|
1949
|
+
### C53 — Make repair observable, cancellable, and resumable
|
|
1950
|
+
|
|
1951
|
+
- Status: implemented
|
|
1952
|
+
- Priority: P0
|
|
1953
|
+
- DependsOn: C47, C52
|
|
1954
|
+
- Touches: `src/engine/index.ts`, `bin/repair-progress.ts`, `bin/cli.ts`,
|
|
1955
|
+
`docs/specs/repair-progress-telemetry.md`, index-health and CLI progress tests,
|
|
1956
|
+
packed-install repair acceptance
|
|
1957
|
+
- What: expose truthful phase/file progress for long repairs, preserve completed
|
|
1958
|
+
transactions on cancellation, and provide composable human and machine CLI
|
|
1959
|
+
output.
|
|
1960
|
+
|
|
1961
|
+
### Acceptance and evidence
|
|
1962
|
+
|
|
1963
|
+
1. Engine events carry monotonic sequences and counters, repository-relative
|
|
1964
|
+
source-safe paths, Salesforce metadata families, and terminal completion,
|
|
1965
|
+
cancellation, or failure without allowing callback errors to fail repair.
|
|
1966
|
+
2. `AbortSignal` boundaries leave committed files recoverable, avoid recording
|
|
1967
|
+
a successful reconciliation on cancellation, and let the next audit resume
|
|
1968
|
+
from `index_state`; concurrent callers receive their own event stream.
|
|
1969
|
+
3. CLI progress supports automatic TTY, plain, JSON, JSONL, and silent modes;
|
|
1970
|
+
progress uses stderr except for JSONL envelopes, and final results retain
|
|
1971
|
+
deterministic stdout.
|
|
1972
|
+
4. Unit/CLI tests and the installed-tarball smoke gate cover cancellation,
|
|
1973
|
+
resumption, throttling, heartbeat, stream separation, and argument
|
|
1974
|
+
validation.
|
|
1975
|
+
|
|
1976
|
+
Evidence: the engine event foundation landed in `e0dcb5f`, cancellation and
|
|
1977
|
+
resumption in `1dc74cb`, and CLI rendering/validation in `a05edcd` and
|
|
1978
|
+
`c0d6745`; the installed-tarball gate exercises all five CLI output modes.
|
|
1979
|
+
|
|
1980
|
+
### C54 — Certify the local 0.1.0 release artifact
|
|
1981
|
+
|
|
1982
|
+
- Status: implemented
|
|
1983
|
+
- Priority: P0
|
|
1984
|
+
- DependsOn: C50, C52, C53
|
|
1985
|
+
- Touches: `package.json`, `bun.lock`, `README.md`,
|
|
1986
|
+
`docs/releases/0.1.0.md`, this roadmap,
|
|
1987
|
+
`docs/specs/repair-progress-telemetry.md`,
|
|
1988
|
+
`scripts/pack-install-smoke.ts`, compiled package/tarball evidence
|
|
1989
|
+
- What: version the first fleet-consumable package, prove it from a clean
|
|
1990
|
+
consumer under the 3 GiB development cap, reconcile roadmap/spec truth, and
|
|
1991
|
+
produce a local release commit and tag without claiming a remote or registry
|
|
1992
|
+
publication.
|
|
1993
|
+
|
|
1994
|
+
### Acceptance and evidence
|
|
1995
|
+
|
|
1996
|
+
1. Package metadata identifies `0.1.0`, and `bun.lock` is regenerated from that
|
|
1997
|
+
manifest (Bun's lock format does not duplicate the workspace version); the
|
|
1998
|
+
exact `knodin-0.1.0.tgz` passes the clean-consumer acceptance gate with
|
|
1999
|
+
an aggregate process-tree limit of 2,750 MiB.
|
|
2000
|
+
2. Lint, typecheck, build, documentation checks, and the full suite pass
|
|
2001
|
+
sequentially under the memory cap; C50 is not closed through quarantine,
|
|
2002
|
+
skipped files, or forced exit. Fresh-worker shards use a bounded five-minute
|
|
2003
|
+
default test ceiling for loaded-host graph builds while replay subprocesses
|
|
2004
|
+
retain tighter explicit limits.
|
|
2005
|
+
3. README and release notes distinguish source-checkout use, local packed
|
|
2006
|
+
installation, and actual npm/remote publication state.
|
|
2007
|
+
4. Tasks 1–3 in the repair telemetry specification are recorded complete while
|
|
2008
|
+
the MCP progress bridge, `--plan`, embedding detail, large Salesforce replay,
|
|
2009
|
+
and incremental embedding work remain explicitly pending.
|
|
2010
|
+
|
|
2011
|
+
Evidence to date: the final exact `knodin-0.1.0.tgz` replay passed 49
|
|
2012
|
+
sequential clean-consumer commands in 43,917 ms with 478 MiB peak aggregate RSS
|
|
2013
|
+
against the 2,750 MiB fail-closed limit. Final integration lint, typecheck,
|
|
2014
|
+
build, 503-test suite, and merged coverage gates pass under 3 GiB (the highest
|
|
2015
|
+
observed gate RSS is 901,840,896 bytes). The local release commit is tagged
|
|
2016
|
+
`v0.1.0`; no remote or registry publication is implied.
|
|
2017
|
+
|
|
2018
|
+
### C55 — Pin and audit the authorized Token Optimizer source
|
|
2019
|
+
|
|
2020
|
+
- Status: implemented
|
|
2021
|
+
- Priority: P0
|
|
2022
|
+
- DependsOn: C24
|
|
2023
|
+
- Touches:
|
|
2024
|
+
`benchmarks/evaluations/token-optimizer-20260730/comparison-manifest.json`,
|
|
2025
|
+
`scripts/token-optimizer-source-audit.ts`, immutable raw audit evidence, and
|
|
2026
|
+
manifest-consistency tests
|
|
2027
|
+
- What: replace the demo-only source assumption with an authorized,
|
|
2028
|
+
commit-bound GHES audit while preserving the distinction between source
|
|
2029
|
+
evidence and equivalent behavioral measurement.
|
|
2030
|
+
|
|
2031
|
+
### Acceptance and evidence
|
|
2032
|
+
|
|
2033
|
+
1. The comparison pins repository, revision, Git tree, file blobs, README hash,
|
|
2034
|
+
and demo hash; branch movement cannot alter the evaluated input.
|
|
2035
|
+
2. An executable audit fetches blobs by immutable revision and fails closed on
|
|
2036
|
+
a different commit, truncated tree, missing file, contradicted finding, or
|
|
2037
|
+
failed behavioral probe.
|
|
2038
|
+
3. The audit verifies parser strategy, checked-in test inventory, shell and
|
|
2039
|
+
containment posture, path scope, telemetry fields/defaults, token/baseline
|
|
2040
|
+
math, compression budget behavior, dependency range, update/provenance
|
|
2041
|
+
inventory, and installation steps.
|
|
2042
|
+
4. The manifest remains `behavior-partial`: source verification must not be
|
|
2043
|
+
presented as production routing, savings, or end-to-end parity evidence.
|
|
2044
|
+
|
|
2045
|
+
Evidence:
|
|
2046
|
+
`benchmarks/evaluations/token-optimizer-20260730/raw-source-audit-20260730.json`
|
|
2047
|
+
pins commit `a2b9cef5efdfcf0393e9e7a5ded229ffd9f613d7`, 17 files, and 16 verified
|
|
2048
|
+
findings. Its behavioral probe returns 104 lines for 100 matching input lines
|
|
2049
|
+
despite `max_lines=50`.
|
|
2050
|
+
|
|
2051
|
+
### C56 — Replay all seven Token Optimizer workflows
|
|
2052
|
+
|
|
2053
|
+
- Status: implemented
|
|
2054
|
+
- Priority: P0
|
|
2055
|
+
- DependsOn: C55
|
|
2056
|
+
- Evidence:
|
|
2057
|
+
`benchmarks/evaluations/token-optimizer-20260730/raw-behavior-replay-20260730.json`
|
|
2058
|
+
|
|
2059
|
+
The immutable replay executes 20 warm and five cold-process samples per
|
|
2060
|
+
operation, uses the real `o200k_base` tokenizer, and applies shared correctness
|
|
2061
|
+
oracles to all five structural workflows. Both products pass those oracles.
|
|
2062
|
+
The final C68 replay records knodin as smaller in real tokens and faster at warm
|
|
2063
|
+
p50/p95 for all five structural cases on the small fixture; its cold-process
|
|
2064
|
+
startup remains materially slower, and command execution remains gated. The competitor compressor
|
|
2065
|
+
preserves the labeled diagnostic strings but returns 29 lines for a requested
|
|
2066
|
+
maximum of 20. knodin's C57 compressor preserves 7/7 detected signals in 14
|
|
2067
|
+
lines with hard byte/line budgets, exact omission accounting, and retained
|
|
2068
|
+
local drill-down. Its selected content is smaller; its complete recovery
|
|
2069
|
+
envelope remains 18 tokens larger on this fixture.
|
|
2070
|
+
|
|
2071
|
+
### C57 — Bounded recoverable diagnostic-output compression
|
|
2072
|
+
|
|
2073
|
+
- Status: implemented
|
|
2074
|
+
- Priority: P0
|
|
2075
|
+
- DependsOn: C3, C55
|
|
2076
|
+
- Evidence:
|
|
2077
|
+
`src/output-compression.ts`,
|
|
2078
|
+
`src/__tests__/unit/output-compression.spec.ts`,
|
|
2079
|
+
`src/__tests__/unit/output-compression-cli.spec.ts`,
|
|
2080
|
+
`benchmarks/evaluations/token-optimizer-20260730/raw-behavior-replay-20260730.json`,
|
|
2081
|
+
and `docs/COMMAND-OUTPUT-COMPRESSION.md`
|
|
2082
|
+
|
|
2083
|
+
knodin accepts already-produced combined text, ordered stdout/stderr events, or
|
|
2084
|
+
a repository-contained log artifact. It deterministically applies `smart`,
|
|
2085
|
+
`head-tail`, or `errors-only` selection with generic and ecosystem-specific
|
|
2086
|
+
adapters; enforces hard rendered-line and UTF-8 content-byte budgets; preserves
|
|
2087
|
+
exit/signal metadata; reports unpreserved signals rather than overstating
|
|
2088
|
+
fidelity; records exact omitted line/source-byte ranges; and retains bounded
|
|
2089
|
+
private drill-down by SHA-256 identity without rerunning a command.
|
|
2090
|
+
|
|
2091
|
+
The pinned adversarial replay measures 604 raw tokens. knodin returns 109
|
|
2092
|
+
content tokens and a 205-token recovery envelope while preserving 7/7 detected
|
|
2093
|
+
signals in 14/20 lines. The competitor returns 187 tokens and preserves every
|
|
2094
|
+
golden string, but emits 29 lines for a requested maximum of 20. This supports
|
|
2095
|
+
the narrow `verified-stronger-budget-and-recovery` classification, not a broad
|
|
2096
|
+
claim that every compression dimension is stronger.
|
|
2097
|
+
|
|
2098
|
+
Command execution is not part of C57. C58's portable-containment evaluation was
|
|
2099
|
+
rejected, and direct MCP `text` input does not prove context savings when raw
|
|
2100
|
+
text has already crossed the model boundary.
|
|
2101
|
+
|
|
2102
|
+
### C58 — Gate command execution on reviewed portable containment
|
|
2103
|
+
|
|
2104
|
+
- Status: evaluated — rejected
|
|
2105
|
+
- Priority: P0
|
|
2106
|
+
- Disposition: evaluated — rejected
|
|
2107
|
+
- DependsOn: C57
|
|
2108
|
+
- Owner roles: security reviewer, release owner, and one platform certifier for
|
|
2109
|
+
each supported operating system
|
|
2110
|
+
- Touches only if approved: a disposable command-runner prototype, adversarial
|
|
2111
|
+
containment fixtures, platform evidence, threat model, and a separately
|
|
2112
|
+
reviewed production implementation item
|
|
2113
|
+
- Evidence: `benchmarks/evaluations/c58-containment/` contains the disposable
|
|
2114
|
+
prototype, adversarial fixture, native macOS raw result, methodology, threat
|
|
2115
|
+
model, tests, and independent `c58-security-review-r1` dated 2026-08-02.
|
|
2116
|
+
- Terminal disposition: `evaluated — rejected`. Native macOS evidence passed
|
|
2117
|
+
the bounded defense-in-depth cases, but repository cwd/path checks are not a
|
|
2118
|
+
filesystem or network sandbox, process polling retains escape/PID races, and
|
|
2119
|
+
Linux/Windows are explicitly unavailable. Production execution remains
|
|
2120
|
+
absent and no implementation item is created.
|
|
2121
|
+
- Post-evaluation gate hardening: the three-level fixture now starts its 100 ms
|
|
2122
|
+
containment window only after an exact bounded readiness handshake, with a
|
|
2123
|
+
separate startup deadline and silent-tree cleanup test. This removes a
|
|
2124
|
+
coverage-timing race without changing raw evidence, platform omissions,
|
|
2125
|
+
security limitations, or the rejected disposition.
|
|
2126
|
+
|
|
2127
|
+
The evaluation decides whether knodin can safely execute a caller-supplied
|
|
2128
|
+
command before compressing its output. C57's already-produced text and artifact
|
|
2129
|
+
inputs remain the production default. Merely matching a competitor workflow is
|
|
2130
|
+
not sufficient to authorize execution.
|
|
2131
|
+
|
|
2132
|
+
Evaluation gate:
|
|
2133
|
+
|
|
2134
|
+
1. Build no production command surface. First produce a disposable prototype
|
|
2135
|
+
using no-shell argv execution, repository-root binding, a minimal sanitized
|
|
2136
|
+
environment, stdin closure, output/time/RSS/process-count limits, and
|
|
2137
|
+
process-tree termination on every proposed supported platform.
|
|
2138
|
+
2. Check in adversarial fixtures for shell metacharacters, traversal, symlink
|
|
2139
|
+
escape, inherited secrets, descriptor inheritance, child/grandchild escape,
|
|
2140
|
+
detached processes, output floods, timeout races, signals, and cancellation.
|
|
2141
|
+
3. Record native macOS, Linux, and Windows results independently. Unsupported
|
|
2142
|
+
containment primitives or an unavailable platform are explicit omissions,
|
|
2143
|
+
not inferred portability.
|
|
2144
|
+
4. Require an independent security review of the threat model and residual
|
|
2145
|
+
risks. Retain its revision, date, findings, and disposition without including
|
|
2146
|
+
credentials or private source.
|
|
2147
|
+
5. Close with exactly one disposition: `evaluated — rejected`, `evaluated —
|
|
2148
|
+
deferred`, or a newly numbered implementation item with its own acceptance
|
|
2149
|
+
criteria. C58 itself never authorizes shipping from prototype evidence.
|
|
2150
|
+
6. If rejected or deferred, keep command execution absent and update the
|
|
2151
|
+
scorecard limitation. This is a valid terminal outcome and does not block
|
|
2152
|
+
C67 when the exclusion remains explicit.
|
|
2153
|
+
|
|
2154
|
+
### C57–C67 — Token Optimizer parity and trusted-distribution program
|
|
2155
|
+
|
|
2156
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2157
|
+
- Previous status: active
|
|
2158
|
+
- Source objective: the checked-in comparison manifest and the DTS competitive
|
|
2159
|
+
goal accepted on 2026-07-30
|
|
2160
|
+
- Ordering: behavioral replay (C56) precedes claims; safe compression (C57)
|
|
2161
|
+
precedes failure-to-code intelligence (C59); command execution (C58) remains
|
|
2162
|
+
excluded after its rejected evaluation; signed trust (C62) precedes update
|
|
2163
|
+
mutation (C63), release attestation
|
|
2164
|
+
(C64), and adversarial supply-chain proof (C65); the scorecard and release are
|
|
2165
|
+
terminal gates.
|
|
2166
|
+
|
|
2167
|
+
The program must preserve these invariants:
|
|
2168
|
+
|
|
2169
|
+
1. Every competitor capability is either verified equivalent/stronger, recorded
|
|
2170
|
+
weaker, or explicitly excluded by a reviewed safety gate.
|
|
2171
|
+
2. Compression has a hard truthful budget, exact omission accounting, retained
|
|
2172
|
+
local drill-down data, and adversarial diagnostic-fidelity fixtures.
|
|
2173
|
+
3. No command runner ships without no-shell argv execution, repository binding,
|
|
2174
|
+
sanitized environment, process-tree/resource/output containment, audit, and
|
|
2175
|
+
supported-platform evidence.
|
|
2176
|
+
4. ROI telemetry is local and opt-in, uses a real tokenizer, separates measured
|
|
2177
|
+
and modeled baselines, and excludes source, raw output, command text,
|
|
2178
|
+
usernames, and absolute paths by default.
|
|
2179
|
+
5. Update availability is non-blocking and privacy-preserving. Mutation requires
|
|
2180
|
+
expiring threshold-signed metadata, rollback/freeze/mix-and-match defenses,
|
|
2181
|
+
digest/length/provenance checks, quarantine, health validation, atomic
|
|
2182
|
+
activation, and rollback.
|
|
2183
|
+
6. A compromised npm, Homebrew, Artifactory, GitHub SaaS, or GHES account alone
|
|
2184
|
+
cannot authorize installation.
|
|
2185
|
+
7. The terminal scorecard links every cell to source, executable tests,
|
|
2186
|
+
benchmark output, and a limitation. Broad superiority language remains
|
|
2187
|
+
prohibited until C66 and C67 close.
|
|
2188
|
+
|
|
2189
|
+
### C59 — Source-evidenced failure-to-code diagnosis
|
|
2190
|
+
|
|
2191
|
+
- Status: implemented
|
|
2192
|
+
- Priority: P1
|
|
2193
|
+
- DependsOn: C57
|
|
2194
|
+
- Evidence:
|
|
2195
|
+
`src/failure-diagnosis.ts`,
|
|
2196
|
+
`src/__tests__/unit/failure-diagnosis.spec.ts`,
|
|
2197
|
+
`src/__tests__/unit/output-compression-cli.spec.ts`,
|
|
2198
|
+
`src/__tests__/unit/knodin-tools.spec.ts`, and
|
|
2199
|
+
`docs/COMMAND-OUTPUT-COMPRESSION.md`
|
|
2200
|
+
|
|
2201
|
+
The retained compression artifact can be diagnosed through
|
|
2202
|
+
`knodin compress diagnose <artifact-id>` or
|
|
2203
|
+
`operation: "compress", compressionAction: "diagnose"` without rerunning the
|
|
2204
|
+
failed command or adding a second MCP tool. Direct already-produced text is
|
|
2205
|
+
also accepted behind the same input, redaction, and byte limits.
|
|
2206
|
+
|
|
2207
|
+
The resolver extracts common JavaScript/TypeScript, Python, Java, C#, Go, and
|
|
2208
|
+
Rust source locations; binds exact or unique-suffix paths to tracked files;
|
|
2209
|
+
refuses traversal, external symlinks, untracked paths, and ambiguous basenames;
|
|
2210
|
+
and returns source-evidenced owning symbols, stable identities, nearest package
|
|
2211
|
+
manifests, tests, upstream callers, downstream dependencies, recent changes,
|
|
2212
|
+
bounded snippets, and indexed/current commit freshness.
|
|
2213
|
+
|
|
2214
|
+
The aggregate source-context byte budget is hard. Incomplete input, exhausted
|
|
2215
|
+
context, unresolved locations, or non-current graph state makes the result
|
|
2216
|
+
`partial` rather than complete. Static relationships remain diagnostic
|
|
2217
|
+
candidates rather than runtime-causality proof; source maps, generated paths,
|
|
2218
|
+
dynamic dispatch, and framework wiring are recorded limitations.
|
|
2219
|
+
|
|
2220
|
+
### C60 — Cross-client and GHES onboarding
|
|
2221
|
+
|
|
2222
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2223
|
+
- Previous status: active — the legacy GHES source and versioned 0.3.0 release path are
|
|
2224
|
+
live; local fail-closed mirror preparation is complete, while enterprise
|
|
2225
|
+
governance/workload identity, timed independent cross-platform onboarding,
|
|
2226
|
+
and Linux/Windows named-client certification remain external gates
|
|
2227
|
+
- Priority: P1
|
|
2228
|
+
- DependsOn: C52, C55
|
|
2229
|
+
- Evidence: `docs/PT-ACCESS-RECOMMENDATION.md`, `docs/INSTALLATION.md`,
|
|
2230
|
+
`docs/evidence/pt-ghes-onboarding-2026-07-31.md`,
|
|
2231
|
+
`docs/evidence/c60-macos-codex-native-2026-08-02.json`,
|
|
2232
|
+
`docs/evidence/c60-macos-native-rss-fix-2026-08-02.json`,
|
|
2233
|
+
`schemas/c60-named-client-evidence-v2.schema.json`,
|
|
2234
|
+
`scripts/c60-native-certify.ts`, `scripts/c60-native-platform.ts`,
|
|
2235
|
+
`scripts/c60-macos-codex-certify.ts`,
|
|
2236
|
+
`src/__tests__/unit/c60-native-platform.spec.ts`,
|
|
2237
|
+
`src/__tests__/unit/agent-integration.spec.ts`,
|
|
2238
|
+
`src/__tests__/unit/doctor.spec.ts`, `scripts/pack-install-smoke.ts`,
|
|
2239
|
+
`config/enterprise-mirror-policy.json`, and
|
|
2240
|
+
`docs/evidence/c60-autonomous-preparation-2026-08-02.md`
|
|
2241
|
+
|
|
2242
|
+
The `Enterprise-Apps/knodin` GHES repository is the DTS / Application Engineering
|
|
2243
|
+
read-only downstream mirror of the authoritative GitHub SaaS repository
|
|
2244
|
+
`DTS-Productivity-Engineering/knodin`. P&T engineers without GitHub SaaS
|
|
2245
|
+
access can download an exact versioned tarball from its
|
|
2246
|
+
release page, verify the recorded SHA-256, install with lifecycle scripts
|
|
2247
|
+
disabled, initialize an existing checkout, and diagnose the single MCP gateway.
|
|
2248
|
+
The observed 0.3.0 GHES asset is byte-identical to the retained CI artifact and
|
|
2249
|
+
public npm artifact. Docusign Artifactory currently stores the same bytes as a raw
|
|
2250
|
+
artifact, not npm registry metadata; the docs do not misrepresent that endpoint
|
|
2251
|
+
as an npm registry.
|
|
2252
|
+
|
|
2253
|
+
Named Claude, Codex, Gemini, and Antigravity configuration plus a real generic
|
|
2254
|
+
MCP handshake are executable tests. The immutable schema-v2 macOS arm64
|
|
2255
|
+
Codex-configuration plus generic-MCP record certifies the exact 0.6.0 tarball in
|
|
2256
|
+
13.739 seconds using pre-populated offline npm/model caches: lifecycle-disabled
|
|
2257
|
+
install, existing-checkout init, project-local configuration without user-home
|
|
2258
|
+
mutation, bounded MCP initialize/list/status/query/doctor, and aggregate
|
|
2259
|
+
RSS/output evidence all pass.
|
|
2260
|
+
An isolated offline build and npm pack from the recorded source commit is
|
|
2261
|
+
byte-identical to the consumed tarball, with its prerequisite commands
|
|
2262
|
+
separately bounded and recorded. It does not prove cold-cache installation, an
|
|
2263
|
+
actual Codex process or active UI
|
|
2264
|
+
session, or the eventual distributed release artifact. Linux packed consumption
|
|
2265
|
+
remains release-gated. A first real Linux attempt falsified the original
|
|
2266
|
+
portability claim because zero-RSS kernel rows made RSS parsing fail before the
|
|
2267
|
+
first command completed; Windows PID 0 had the same defect. The fixed parser
|
|
2268
|
+
retains structural and integer validation, filters non-process/zero-resident
|
|
2269
|
+
rows, and rejects an all-filtered snapshot. A new provenance-bound macOS replay
|
|
2270
|
+
of the exact fix commit passed in 11.212 seconds with byte-identical package
|
|
2271
|
+
evidence, proving no macOS regression but not Linux/Windows readiness. The
|
|
2272
|
+
native Linux and Windows runs must restart from the fixed commit without source
|
|
2273
|
+
changes. GHES `main` is not protected, GHES releases are not
|
|
2274
|
+
immutable, and native Linux/Windows named-client runs, independent cross-platform review,
|
|
2275
|
+
and non-personal automated mirror synchronization are missing, so C60 remains
|
|
2276
|
+
active and no cross-platform-complete or universal two-minute claim is allowed.
|
|
2277
|
+
|
|
2278
|
+
### C61 — Private real-token ROI telemetry and dashboard
|
|
2279
|
+
|
|
2280
|
+
- Status: implemented
|
|
2281
|
+
- Priority: P1
|
|
2282
|
+
- DependsOn: C16, C55
|
|
2283
|
+
- Evidence: `src/output-telemetry.ts`, `src/cli-model.ts`, `bin/cli.ts`,
|
|
2284
|
+
`src/tools/knodin-tools.ts`, `src/__tests__/unit/output-telemetry.spec.ts`,
|
|
2285
|
+
`src/__tests__/unit/output-telemetry-cli.spec.ts`,
|
|
2286
|
+
`src/__tests__/unit/knodin-tools.spec.ts`, and `docs/TELEMETRY.md`
|
|
2287
|
+
|
|
2288
|
+
Persistence remains disabled by default and local. Opt-in JSONL records use the
|
|
2289
|
+
pinned real tokenizer, executed full-file baselines only when bounded source is
|
|
2290
|
+
available, explicit unknowns otherwise, cold/warm and outcome dimensions,
|
|
2291
|
+
agent-round-trip, latency, RSS, freshness, compression-fidelity, indexing-time,
|
|
2292
|
+
and index-storage measurements. Repository grouping uses a one-way identifier;
|
|
2293
|
+
source, raw output, command text, usernames, and absolute paths are excluded.
|
|
2294
|
+
|
|
2295
|
+
The CLI and single MCP gateway expose status, static HTML report, sanitized JSON
|
|
2296
|
+
export, retention, and explicit clear operations. Inputs and outputs are
|
|
2297
|
+
repository-contained and symlink-refusing, persistence is atomic/private, and
|
|
2298
|
+
the self-contained dashboard reports summary, operation, time-series, private
|
|
2299
|
+
repository, latency p50/p95, indexing cost, freshness, and fidelity dimensions
|
|
2300
|
+
with light/dark rendering. Metrics with no defensible observation remain
|
|
2301
|
+
unavailable rather than modeled.
|
|
2302
|
+
|
|
2303
|
+
### C62 — Signed update metadata trust
|
|
2304
|
+
|
|
2305
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2306
|
+
- Previous status: active — verification core and fail-closed local ceremony preparation
|
|
2307
|
+
implemented; production offline custodians, independent pin review, recovery
|
|
2308
|
+
evidence, and trust-pin activation remain external gates
|
|
2309
|
+
- Priority: P0
|
|
2310
|
+
- DependsOn: C55
|
|
2311
|
+
- Evidence: `src/update-trust.ts`,
|
|
2312
|
+
`src/__tests__/unit/update-trust.spec.ts`, `src/update-ceremony.ts`,
|
|
2313
|
+
`src/__tests__/unit/update-ceremony.spec.ts`,
|
|
2314
|
+
`scripts/root-ceremony-preflight.ts`,
|
|
2315
|
+
`schemas/root-ceremony-manifest-v1.schema.json`,
|
|
2316
|
+
`schemas/root-pin-review-receipt-v1.schema.json`,
|
|
2317
|
+
`docs/ROOT-CEREMONY-RUNBOOK.md`, and `docs/SIGNED-UPDATES.md`
|
|
2318
|
+
|
|
2319
|
+
The pure verifier implements deterministic signed payloads, content-addressed
|
|
2320
|
+
Ed25519 keys, disjoint root/targets/snapshot/timestamp roles, production
|
|
2321
|
+
thresholds for root and targets, bounded expiry, independent root pinning,
|
|
2322
|
+
dual-threshold one-step root rotation, version-and-digest rollback/equivocation
|
|
2323
|
+
state, timestamp→snapshot→targets length/hash/version binding, safe target
|
|
2324
|
+
paths, and signed artifact SHA-256/length verification. It performs no network
|
|
2325
|
+
or installation mutation.
|
|
2326
|
+
|
|
2327
|
+
Adversarial fixtures reject missing thresholds, field and artifact
|
|
2328
|
+
substitution, expired timestamp freezes, rollback, same-version equivocation,
|
|
2329
|
+
mix-and-match metadata, role-key reuse, root-pin substitution, future metadata,
|
|
2330
|
+
and rotations not signed by the old root threshold.
|
|
2331
|
+
|
|
2332
|
+
C62 is not closed merely because the verification code exists. No production
|
|
2333
|
+
private key is checked in or generated by tests. Release owners must complete
|
|
2334
|
+
the offline multi-custodian root ceremony, independently review and embed the
|
|
2335
|
+
initial root digest, and record rotation/recovery evidence. C63 now makes all
|
|
2336
|
+
doctor, CLI, and MCP update state consume only verified metadata; the former
|
|
2337
|
+
npm lookup has been removed. Production authorization still waits for the
|
|
2338
|
+
offline ceremony.
|
|
2339
|
+
|
|
2340
|
+
### C63 — Signed-only update client and safe policy
|
|
2341
|
+
|
|
2342
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2343
|
+
- Previous status: active — implementation and adversarial unit/CLI evidence complete;
|
|
2344
|
+
closure requires a recorded fail-closed deferred-activation policy after C62
|
|
2345
|
+
- Priority: P0
|
|
2346
|
+
- DependsOn: C62
|
|
2347
|
+
- Evidence: `src/update-policy.ts`,
|
|
2348
|
+
`src/__tests__/unit/update-policy.spec.ts`,
|
|
2349
|
+
`src/__tests__/unit/update-policy-cli.spec.ts`,
|
|
2350
|
+
`docs/DOCTOR-AND-UPDATES.md`, and `docs/SIGNED-UPDATES.md`
|
|
2351
|
+
|
|
2352
|
+
The five `knodin update` actions work outside a repository. Periodic checks use
|
|
2353
|
+
an atomic background lease and never put network latency on ordinary commands.
|
|
2354
|
+
Fixed signed metadata paths are HTTPS-origin/path bound, redirect-free, capped,
|
|
2355
|
+
timed, privacy-preserving, and optionally byte-identical across mirrors.
|
|
2356
|
+
Verified state is atomic/private and retains monotonic evidence across failed
|
|
2357
|
+
checks.
|
|
2358
|
+
|
|
2359
|
+
Apply enforces notify/download/patch/minor policy, minimum age, channels,
|
|
2360
|
+
enterprise mirrors, maintenance windows, exact artifact length/digest, signed
|
|
2361
|
+
provenance digest, candidate/rollback quarantine, no-shell exact local npm
|
|
2362
|
+
artifact invocation, symlink-safe exclusive quarantine, signed-version path
|
|
2363
|
+
validation, allowlisted enterprise cross-check origins, health validation,
|
|
2364
|
+
health-verified automatic rollback, and offline re-verification for explicit
|
|
2365
|
+
rollback. Mutable registry tags are never used.
|
|
2366
|
+
|
|
2367
|
+
The production CLI remains fail-closed as `trust-unconfigured`: a root and its
|
|
2368
|
+
pin from one mutable user file are not accepted. No certified sandbox smoke
|
|
2369
|
+
adapter or Homebrew/Node-manager rollback adapter is enabled yet. C62's offline
|
|
2370
|
+
ceremony must embed the independent anchor. C63 closes by recording that
|
|
2371
|
+
automatic production activation remains disabled; it does not wait for C65.
|
|
2372
|
+
C65 then uses the terminal signed-only client to certify the smoke sandbox,
|
|
2373
|
+
multi-step rotation/recovery, compromised channels, and manager-specific
|
|
2374
|
+
activation before any later auto-apply enablement. This ordering removes a
|
|
2375
|
+
C63↔C65 closure cycle without weakening the fail-closed policy.
|
|
2376
|
+
|
|
2377
|
+
### C64 — Provenance, SBOM, and cross-channel release attestation
|
|
2378
|
+
|
|
2379
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2380
|
+
- Previous status: active — implementation and adversarial fixtures complete; a live
|
|
2381
|
+
five-channel release attestation remains a C67 publication gate
|
|
2382
|
+
- Priority: P0
|
|
2383
|
+
- DependsOn: C62
|
|
2384
|
+
- Evidence: `src/release-attestation.ts`,
|
|
2385
|
+
`src/__tests__/unit/release-attestation.spec.ts`,
|
|
2386
|
+
`src/__tests__/unit/release-attestation-cli.spec.ts`,
|
|
2387
|
+
`src/__tests__/unit/release-workflow.spec.ts`,
|
|
2388
|
+
`src/__tests__/unit/release-candidate-workflow.spec.ts`,
|
|
2389
|
+
`scripts/release-attestation.ts`, `.github/workflows/publish.yml`,
|
|
2390
|
+
`.github/workflows/release-candidate.yml`, and `docs/RELEASING.md`
|
|
2391
|
+
|
|
2392
|
+
The release workflow installs dependencies with lifecycle scripts disabled,
|
|
2393
|
+
runs tests, and packs the artifact in an unprivileged job with read-only
|
|
2394
|
+
repository access. A separate job with OIDC and attestation authority downloads
|
|
2395
|
+
only the retained artifact, generates a production-dependency CycloneDX SBOM,
|
|
2396
|
+
and uses full-commit-pinned official GitHub actions to sign SLSA build provenance
|
|
2397
|
+
and the SBOM. It independently verifies both retained bundles against the exact
|
|
2398
|
+
artifact, source digest, tag ref, repository, workflow certificate identity,
|
|
2399
|
+
and GitHub OIDC issuer before npm publication. A manual retry must run at the
|
|
2400
|
+
tag ref; checked-out commit, workflow SHA, tag, and package version disagreement
|
|
2401
|
+
is a hard failure. The exact tarball, signed evidence, and verification receipts
|
|
2402
|
+
are retained for 90 days before that tarball is passed to npm and Artifactory.
|
|
2403
|
+
|
|
2404
|
+
The separate manual-only release-candidate workflow now structurally stops
|
|
2405
|
+
after the exact artifact and rollback bytes, provenance and SBOM bundles,
|
|
2406
|
+
verification and identity receipts, bounded input manifest, and incomplete
|
|
2407
|
+
five-channel attestation draft are retained for 90 days. Its build job is
|
|
2408
|
+
unprivileged and lifecycle-disabled; its OIDC job downloads the retained input
|
|
2409
|
+
set without a source checkout or dependency install and has no npm/JFrog/channel
|
|
2410
|
+
credential, publishing command, reusable publisher, or distribution mutation.
|
|
2411
|
+
This is local implementation evidence only: C64 remains active until a live OIDC
|
|
2412
|
+
run at an immutable tag produces the retained candidate bundle and independent
|
|
2413
|
+
verification. C67 alone authorizes distribution after that candidate is
|
|
2414
|
+
independently verified.
|
|
2415
|
+
|
|
2416
|
+
The canonical schema derives, rather than accepts, `verified`, `incomplete`, or
|
|
2417
|
+
`quarantined` status. It records the exact artifact SHA-256/SHA-512/npm
|
|
2418
|
+
integrity, source commit, build identity, provenance and SBOM digests,
|
|
2419
|
+
publication timestamps and locations, all five approved destinations, and an
|
|
2420
|
+
exact rollback target whose digest the CLI derives from downloaded bytes rather
|
|
2421
|
+
than accepting from the manifest. Missing evidence is incomplete; any artifact, version,
|
|
2422
|
+
or source disagreement is quarantined. The CLI binds inputs beneath one
|
|
2423
|
+
directory, refuses direct or nested symlinks/traversal/unbounded files/output replacement,
|
|
2424
|
+
writes visible failure evidence, and exits nonzero for an incomplete final or
|
|
2425
|
+
any mismatch. It ignores a manifest's claimed verification state and invokes
|
|
2426
|
+
an explicitly approved absolute `gh attestation verify` executable without a
|
|
2427
|
+
shell or inherited credentials/environment against both local bundles and the
|
|
2428
|
+
exact artifact/source/ref/workflow/certificate policy before deriving verified
|
|
2429
|
+
signature evidence.
|
|
2430
|
+
|
|
2431
|
+
Signed targets may bind the provenance, SBOM, SBOM-attestation, and release-
|
|
2432
|
+
attestation digests. Update quarantine and explicit rollback re-fetch or re-read
|
|
2433
|
+
all four objects, validate their exact bytes and attested channel state, and
|
|
2434
|
+
refuse missing or inconsistent evidence before a package manager is invoked.
|
|
2435
|
+
Apply cross-checks the candidate's attested rollback version, path, and digest
|
|
2436
|
+
against the separately quarantined rollback artifact.
|
|
2437
|
+
|
|
2438
|
+
C64 is not a claim that the current 0.3.0 channels have been retroactively
|
|
2439
|
+
attested. GitHub/Sigstore build signatures are independently useful evidence,
|
|
2440
|
+
but do not replace C62's offline-controlled threshold root. C64 closes when the
|
|
2441
|
+
exact release-candidate artifact has verified provenance and SBOM bundles plus
|
|
2442
|
+
a complete attestation draft containing the five intended destinations. The
|
|
2443
|
+
draft remains `incomplete` until distribution. C67 publishes those exact bytes,
|
|
2444
|
+
fills the destination receipts, derives the final `verified` record, and
|
|
2445
|
+
TUF-authorizes it. This split prevents C64 and C67 from depending on each
|
|
2446
|
+
other's terminal evidence.
|
|
2447
|
+
|
|
2448
|
+
The signature proves workflow identity and exact bytes, not that a compromised
|
|
2449
|
+
self-hosted runner built those bytes faithfully. Independent build reproduction
|
|
2450
|
+
and runner-compromise drills remain C65/C67 gates. The JFrog reusable workflow's
|
|
2451
|
+
`@master` reference also remains an explicit external constraint because its
|
|
2452
|
+
Vault OIDC policy rejects commit-pinned `job_workflow_ref` claims; this is not
|
|
2453
|
+
misreported as full workflow immutability.
|
|
2454
|
+
|
|
2455
|
+
### C65 — Compromised-channel, rollback, freeze, and recovery defenses
|
|
2456
|
+
|
|
2457
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2458
|
+
- Previous status: active — bounded multi-step root recovery is implemented; production
|
|
2459
|
+
key ceremony, smoke sandbox, manager-specific activation, and live drills
|
|
2460
|
+
remain required
|
|
2461
|
+
- Priority: P0
|
|
2462
|
+
- DependsOn: C63, C64
|
|
2463
|
+
- Evidence: `src/update-trust.ts`, `src/update-policy.ts`,
|
|
2464
|
+
`src/__tests__/unit/update-trust.spec.ts`,
|
|
2465
|
+
`src/__tests__/unit/update-policy.spec.ts`, `docs/SIGNED-UPDATES.md`, and
|
|
2466
|
+
`docs/evidence/0.4.3-update-defense-drill.json`
|
|
2467
|
+
|
|
2468
|
+
The update client can recover from up to 32 missed root rotations per check by
|
|
2469
|
+
fetching every numbered intermediate root, cross-checking its exact bytes across
|
|
2470
|
+
configured mirrors, and verifying each version under both the previous and new
|
|
2471
|
+
root thresholds. It never accepts a direct version jump or promotes a root from
|
|
2472
|
+
mutable state. Missing, malformed, wrong-version, skipped, or mirror-divergent
|
|
2473
|
+
intermediates fail before new trust is persisted; the rotation bound prevents a
|
|
2474
|
+
malicious terminal version from causing unbounded requests.
|
|
2475
|
+
|
|
2476
|
+
This closes the implemented multi-step retrieval gap only. C65 still requires
|
|
2477
|
+
the offline multi-custodian production root/recovery rehearsal, certified smoke
|
|
2478
|
+
sandbox, npm/Homebrew/manager activation and rollback drills, compromised
|
|
2479
|
+
transport and runner exercises, and checked-in evidence from those executions.
|
|
2480
|
+
No automatic production activation or complete supply-chain claim is allowed
|
|
2481
|
+
until those gates and C67's live release are complete.
|
|
2482
|
+
|
|
2483
|
+
The checked-in 0.4.3 isolated drill records one combined execution of 52 trust,
|
|
2484
|
+
policy, CLI, and attestation scenarios, including compromised-fixture refusal,
|
|
2485
|
+
multi-step recovery, automatic rollback, and unhealthy-rollback refusal. It
|
|
2486
|
+
strengthens repeatable local evidence but deliberately leaves the production
|
|
2487
|
+
ceremony, certified adapters/sandbox, real manager mutation, and live
|
|
2488
|
+
transport/runner exercises open.
|
|
2489
|
+
|
|
2490
|
+
### C66 — Evidence-linked Token Optimizer scorecard
|
|
2491
|
+
|
|
2492
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2493
|
+
- Previous status: active — scorecard published; terminal prerelease snapshot awaits
|
|
2494
|
+
active dependencies
|
|
2495
|
+
- Priority: P0
|
|
2496
|
+
- DependsOn: C56, C57, C59, C60, C61, C65, C68, C75
|
|
2497
|
+
- Evidence: `docs/TOKEN-OPTIMIZER-SCORECARD.md`,
|
|
2498
|
+
`benchmarks/evaluations/token-optimizer-20260730/scorecard.json`,
|
|
2499
|
+
`src/__tests__/unit/docs-integrity.spec.ts`, and the pinned audit/replay
|
|
2500
|
+
artifacts under `benchmarks/evaluations/token-optimizer-20260730/`
|
|
2501
|
+
|
|
2502
|
+
The human scorecard and machine manifest classify every advertised competitor
|
|
2503
|
+
workflow plus the cross-cutting product dimensions. Every row links checked-in
|
|
2504
|
+
implementation, verification source/test paths, benchmark or live evidence,
|
|
2505
|
+
scope, readiness, and a material limitation. The manifest pins the competitor,
|
|
2506
|
+
the raw replay digest, its declared base revision, every evaluated source hash,
|
|
2507
|
+
and the first merged knodin commit containing that exact composite. Its integrity test
|
|
2508
|
+
rejects dimension omissions/duplicates, revision or replay-digest drift,
|
|
2509
|
+
missing evidence paths, or empty scope/limitation fields.
|
|
2510
|
+
|
|
2511
|
+
The scorecard supports a qualified “more complete local code-intelligence
|
|
2512
|
+
system” statement. It explicitly prohibits “better across all dimensions”:
|
|
2513
|
+
command execution remains intentionally excluded, Windows/timed onboarding is
|
|
2514
|
+
not certified, cold startup is slower, and production trust ceremony, drills,
|
|
2515
|
+
and five-channel release evidence remain open. C66 publication does not close
|
|
2516
|
+
its active dependencies or C67. C66 becomes terminal when every listed
|
|
2517
|
+
dependency is terminal and the scorecard records the prerelease truth. C67
|
|
2518
|
+
appends the final release record and receipts as a non-gating amendment; C66
|
|
2519
|
+
does not depend on the release it gates.
|
|
2520
|
+
|
|
2521
|
+
### C67 — Certify and distribute the completed competitive release
|
|
2522
|
+
|
|
2523
|
+
- Status: parked-external-evidence — 2026-08-02
|
|
2524
|
+
- Previous status: proposed
|
|
2525
|
+
- Priority: P0
|
|
2526
|
+
- Disposition: Must close
|
|
2527
|
+
- DependsOn: C60, C62, C63, C64, C65, C66, C75
|
|
2528
|
+
- Owner roles: release owner, offline-root custodians meeting the recorded
|
|
2529
|
+
threshold, security reviewer, five channel operators, and Windows/macOS/Linux
|
|
2530
|
+
certifiers
|
|
2531
|
+
- Evidence directory: `docs/evidence/releases/<version>/` with immutable raw
|
|
2532
|
+
receipts under `benchmarks/evaluations/releases/<version>/`
|
|
2533
|
+
- Preflight implementation: `src/release-preflight.ts`,
|
|
2534
|
+
`scripts/release-preflight.ts`, `schemas/release-plan-v1.schema.json`,
|
|
2535
|
+
`src/__tests__/unit/release-preflight.spec.ts`, and `docs/RELEASING.md`
|
|
2536
|
+
|
|
2537
|
+
C67 is the terminal release and roadmap gate. It certifies one exact version
|
|
2538
|
+
from one source commit and distributes one byte-identical package through all
|
|
2539
|
+
approved channels without relaxing any unresolved limitation.
|
|
2540
|
+
|
|
2541
|
+
The pure C67 preflight is implemented without release authority. It validates
|
|
2542
|
+
a bounded plan, exact clean annotated-tag identity, the terminal roadmap and
|
|
2543
|
+
ledger state of every dependency, tracked repository-contained nonsymlink
|
|
2544
|
+
evidence including an immutable terminal artifact per dependency, certified
|
|
2545
|
+
native targets, exact verification commands, an independently reviewed
|
|
2546
|
+
production root/ceremony/digest tuple, and a new version-bound evidence path. It performs no publishing,
|
|
2547
|
+
tagging, evidence-directory creation, credential access, or other external
|
|
2548
|
+
mutation and does not change C64's nonpublishing candidate workflow. C67 remains
|
|
2549
|
+
proposed and open: against the current authoritative ledger the preflight must
|
|
2550
|
+
fail until C60, C62, C63, C64, C65, and C66 close; a passing local unit suite is
|
|
2551
|
+
not a release authorization or distribution receipt.
|
|
2552
|
+
|
|
2553
|
+
Acceptance and execution contract:
|
|
2554
|
+
|
|
2555
|
+
1. Preflight fails unless C60, C62, C63, C64, C65, C66, and C75 have terminal
|
|
2556
|
+
statuses and checked-in evidence. Record package version, tag, source commit,
|
|
2557
|
+
workflow commit, runtimes, target platforms, trust-root digest, and exact
|
|
2558
|
+
verification commands before mutation.
|
|
2559
|
+
2. From a clean tag checkout, run `npm ci --ignore-scripts`, `npm run lint`,
|
|
2560
|
+
`npm run typecheck`, `npm run build`, `npm run test`,
|
|
2561
|
+
`npm run test:pack-install`, `npm run check:roadmap`, and the release-specific
|
|
2562
|
+
competitive, trust, defense, and attestation replays. Record exit status,
|
|
2563
|
+
elapsed time, peak RSS, and immutable artifact digests without overwriting an
|
|
2564
|
+
earlier run.
|
|
2565
|
+
3. Build once in the unprivileged release job. Verify the retained package,
|
|
2566
|
+
provenance, SBOM, SBOM attestation, and release-attestation draft before any
|
|
2567
|
+
channel receives the artifact. A retry must use the same tag and bytes.
|
|
2568
|
+
4. Publish SHA-256/length-identical bytes to npm, Homebrew, Artifactory, GitHub
|
|
2569
|
+
SaaS, and GHES. Record immutable channel identifiers, timestamps, downloaded
|
|
2570
|
+
digests, and independently fetched receipts. Any missing or divergent channel
|
|
2571
|
+
fails the release and triggers quarantine/recovery rather than success.
|
|
2572
|
+
5. Produce threshold-signed targets/snapshot/timestamp metadata bound to the
|
|
2573
|
+
artifact and C64 evidence digests. Exercise update status, check, explain,
|
|
2574
|
+
apply, health validation, and rollback from clean supported consumers.
|
|
2575
|
+
6. Re-run named-client onboarding on certified macOS, Linux, and Windows and
|
|
2576
|
+
record C60's two-minute measurement. Verify `knodin init`, lifecycle refresh,
|
|
2577
|
+
CLI query, and the single MCP gateway against an existing checkout.
|
|
2578
|
+
7. Publish a release record listing every gate, receipt, degradation,
|
|
2579
|
+
unsupported case, rollback target, and authorized product claim. Append those
|
|
2580
|
+
receipts to C66's already-terminal prerelease scorecard without reopening its
|
|
2581
|
+
gate; preserve every limitation whose independent gate did not pass.
|
|
2582
|
+
8. Mark C67 terminal and remove its ledger row in the same change, then require
|
|
2583
|
+
a fresh `npm run check:roadmap -- --complete` to succeed. If authority or
|
|
2584
|
+
infrastructure is unavailable, retain `blocked` with the exact missing
|
|
2585
|
+
authority and safe next action; never fabricate a receipt.
|
|
2586
|
+
|
|
2587
|
+
### C68 — Close structural output and latency gaps
|
|
2588
|
+
|
|
2589
|
+
- Status: implemented
|
|
2590
|
+
- Priority: P0
|
|
2591
|
+
- DependsOn: C56
|
|
2592
|
+
- Evidence: the initial C56 replay identified compactness and warm-latency gaps;
|
|
2593
|
+
the final pinned replay proves the C68 compact mode passes its narrow shared
|
|
2594
|
+
oracles and beats competitor response tokens and warm p50/p95 for all five
|
|
2595
|
+
structural operations on the small fixture. Cold startup remains slower.
|
|
2596
|
+
|
|
2597
|
+
Acceptance:
|
|
2598
|
+
|
|
2599
|
+
1. Add a purpose-built compact structural response mode that preserves stable
|
|
2600
|
+
identity, source location, ambiguity, freshness, and truthful budgets without
|
|
2601
|
+
returning unrelated graph fields or telemetry envelopes.
|
|
2602
|
+
2. Measure gateway schema overhead separately from per-call output.
|
|
2603
|
+
3. Avoid repeated freshness/index startup work inside one session and document
|
|
2604
|
+
the irreducible cold cost of a persistent graph separately from operation
|
|
2605
|
+
latency.
|
|
2606
|
+
4. Re-run the exact C56 fixture. knodin must match or beat competitor response
|
|
2607
|
+
tokens for all five structural workflows and must not be slower in warm
|
|
2608
|
+
operation p50/p95 without an explicitly approved, evidence-backed tradeoff.
|
|
2609
|
+
5. Correctness, ambiguity, and stale-index oracles must remain passing; size or
|
|
2610
|
+
latency cannot be improved by weakening evidence.
|
|
2611
|
+
|
|
2612
|
+
Implementation evidence: `detailLevel: "compact"` now selects a purpose-built
|
|
2613
|
+
compact structural contract for file summary, exact source, structural search,
|
|
2614
|
+
batch outline, and project overview. Stable `~` identity prefixes are accepted
|
|
2615
|
+
as ambiguity-safe selectors, returned/total counts remain explicit, and no
|
|
2616
|
+
generic telemetry or response-budget envelope is added. Generation-scoped
|
|
2617
|
+
structural caching plus a watcher-invalidated freshness lease removes repeated
|
|
2618
|
+
session work; a no-subprocess HEAD check prevents Git movement from reusing the
|
|
2619
|
+
lease. The pinned C56 replay records gateway schema cost separately and passes
|
|
2620
|
+
all five correctness, token, warm-p50, and warm-p95 gates. Cold-process cost
|
|
2621
|
+
remains explicitly reported as a persistent-graph tradeoff.
|
|
2622
|
+
|
|
2623
|
+
### C69–C75 — Large-portfolio lifecycle hardening
|
|
2624
|
+
|
|
2625
|
+
- Status: certified with recorded degradations
|
|
2626
|
+
- Priority: P0/P1 as listed above
|
|
2627
|
+
- Source objective:
|
|
2628
|
+
`knodin-Portfolio-Integration-Issues-2026-07-30.md` and the tested candidate
|
|
2629
|
+
commit `c981480db89ffbbb646206a73660e9ab842686ba`
|
|
2630
|
+
- Scope: portfolio discovery, selection, initialization, lifecycle truth,
|
|
2631
|
+
diagnostics, and Salesforce metadata candidate quality.
|
|
2632
|
+
- C75 Evidence: `docs/evidence/portfolio-initialization-2026-07-31.md`
|
|
2633
|
+
|
|
2634
|
+
The candidate commit proves that a 256 MiB subprocess buffer, early selection,
|
|
2635
|
+
and sequential inventory can discover the current 61-repository portfolio, but
|
|
2636
|
+
that commit is evidence rather than the terminal design. It still materializes
|
|
2637
|
+
SalesforceCI's roughly 105 MB `git ls-files` response and portfolio dry-run
|
|
2638
|
+
peaked near the one-GB incremental-memory ceiling.
|
|
2639
|
+
|
|
2640
|
+
Acceptance:
|
|
2641
|
+
|
|
2642
|
+
1. Inventory tracked files with a streaming or equivalently bounded design.
|
|
2643
|
+
SalesforceCI's 904,072 tracked files cannot abort the portfolio. A large,
|
|
2644
|
+
malformed, timed-out, or otherwise failed repository produces a
|
|
2645
|
+
machine-readable degraded record while other repositories continue.
|
|
2646
|
+
2. Apply include/exclude selection before tracked-file inventory. An include
|
|
2647
|
+
selector that matches no discovered repository returns an explicit
|
|
2648
|
+
diagnostic rather than a successful empty result.
|
|
2649
|
+
3. `repos discover ~/code --json` returns valid JSON for all 61 current
|
|
2650
|
+
repositories while remaining below one GB of incremental RSS. Selected
|
|
2651
|
+
searches do not inventory excluded repositories.
|
|
2652
|
+
4. Make portfolio init and dry-run sequential, resource-releasing, resumable,
|
|
2653
|
+
and bounded. Dry-run performs no model or database work, estimates scale and
|
|
2654
|
+
memory risk, and reports init/update/repair actions accurately. A
|
|
2655
|
+
per-process memory ceiling degrades and continues rather than killing the
|
|
2656
|
+
portfolio run.
|
|
2657
|
+
5. `knodin configure` is either clearly configuration-only or atomically
|
|
2658
|
+
produces a usable graph and hooks. `knodin init` drains compatible queued
|
|
2659
|
+
lifecycle events after graph health is established; incompatible events
|
|
2660
|
+
receive an exact remediation. A `knodin wait --fresh` deadline must not
|
|
2661
|
+
terminate a healthy lifecycle processor; dead or failed attempts are
|
|
2662
|
+
retried at a bounded interval. Immediate status and search must agree with
|
|
2663
|
+
the success message.
|
|
2664
|
+
6. Running `knodin doctor` on a non-Git portfolio parent warns and directs the
|
|
2665
|
+
user to portfolio diagnostics. MCP diagnostics separately report on-disk
|
|
2666
|
+
configuration, subprocess handshake, active-client exposure being unknown,
|
|
2667
|
+
reload requirements, and the exact configuration path.
|
|
2668
|
+
7. Architecture and roadmap candidates use basename/directory-aware evidence,
|
|
2669
|
+
bounded lists, reasons, and confidence. Salesforce object or metadata names
|
|
2670
|
+
containing words such as `Design`, `WizardDesign`, or `Backlog` are not
|
|
2671
|
+
promoted without qualifying path/content evidence.
|
|
2672
|
+
8. Certification records the 61-repository result, peak RSS, elapsed time,
|
|
2673
|
+
degraded repositories, selector behavior, immediate queryability, and all
|
|
2674
|
+
regression/typecheck/build/package-install gates. C66 and C67 cannot make a
|
|
2675
|
+
portfolio-readiness claim until C75 closes.
|
|
2676
|
+
|
|
2677
|
+
C72 lifecycle evidence is recorded in
|
|
2678
|
+
`docs/evidence/lifecycle-event-drain-2026-07-31.md`. Generated events now bind
|
|
2679
|
+
to the actual worktree root and status separately exposes a pending event and
|
|
2680
|
+
`idle`, `running`, or `stale-lock` processor state. The checked-in wait path
|
|
2681
|
+
retries a stranded compatible processor. The evidence explicitly does not
|
|
2682
|
+
claim that the installed 0.3.0 Homebrew release contains those source changes;
|
|
2683
|
+
distribution remains a C67 gate.
|
|
2684
|
+
|
|
2685
|
+
C71 is implemented by `src/repository-init-process.ts` and the sequential
|
|
2686
|
+
orchestration in `src/repository-management.ts`. Every live repository runs in
|
|
2687
|
+
a disposable process with a 768 MiB RSS ceiling, independent parent-side RSS
|
|
2688
|
+
polling plus worker heartbeats, process-group termination on macOS/Linux,
|
|
2689
|
+
timeout escalation, repository-scoped degradation, aggregate resource evidence,
|
|
2690
|
+
and resumable manifests. `src/__tests__/unit/repository-init-process.spec.ts`
|
|
2691
|
+
and `src/__tests__/unit/repository-management.spec.ts` prove termination,
|
|
2692
|
+
continuation, and single-worker sequencing. That implementation evidence is
|
|
2693
|
+
separate from the following live C75 certification.
|
|
2694
|
+
|
|
2695
|
+
C75 live evidence is recorded in
|
|
2696
|
+
`docs/evidence/portfolio-initialization-2026-07-31.md`. The 69-record run stayed
|
|
2697
|
+
below one GiB, completed 48 selected repositories, degraded 16 workers at the
|
|
2698
|
+
memory ceiling, preserved tracked team integration in four repositories, and
|
|
2699
|
+
continued after every failure. C75 therefore certifies bounded portfolio
|
|
2700
|
+
behavior with named limitations; it does not claim that all repositories are
|
|
2701
|
+
healthy or initialized.
|
|
2702
|
+
|
|
2703
|
+
### C76 — Responsive declarative CLI model
|
|
2704
|
+
|
|
2705
|
+
- Status: implemented
|
|
2706
|
+
- Priority: P1
|
|
2707
|
+
- DependsOn: C73
|
|
2708
|
+
- Evidence: PR #29 replaced the hand-authored parser/help template with the
|
|
2709
|
+
declarative Commander model and pseudo-terminal width acceptance coverage.
|
|
2710
|
+
|
|
2711
|
+
Acceptance:
|
|
2712
|
+
|
|
2713
|
+
1. Immediately reflow help to the detected TTY width with deterministic
|
|
2714
|
+
non-TTY output and narrow/wide pseudo-terminal tests.
|
|
2715
|
+
2. Migrate parsing, validation, global options, nested subcommands, and help to
|
|
2716
|
+
one declarative command model so documentation cannot drift from accepted
|
|
2717
|
+
arguments.
|
|
2718
|
+
3. Preserve every checked-in legacy invocation and JSON stdout contract.
|
|
2719
|
+
4. Prefer Commander unless an executable bakeoff disproves the choice:
|
|
2720
|
+
current Commander supplies width-aware help, nested commands, TypeScript
|
|
2721
|
+
support, a documented security policy, and zero runtime dependencies.
|
|
2722
|
+
5. Pin the exact dependency and integrity through the normal lockfiles and
|
|
2723
|
+
release provenance. Do not adopt a broader dependency tree solely for
|
|
2724
|
+
cosmetic terminal rendering.
|
|
2725
|
+
|
|
2726
|
+
Implementation evidence: `src/cli-model.ts` is the single Commander-backed
|
|
2727
|
+
grammar for global options, root and nested commands, positional contracts,
|
|
2728
|
+
numeric parsing, unknown-option rejection, and root or command help.
|
|
2729
|
+
`bin/cli.ts` validates every non-help invocation through that model and reads
|
|
2730
|
+
shared option values from its parsed result before compatibility dispatch.
|
|
2731
|
+
Commander 15.0.0 is exact-pinned in both lockfiles and has no transitive runtime
|
|
2732
|
+
dependencies. `src/__tests__/unit/cli-model.spec.ts`,
|
|
2733
|
+
`src/__tests__/unit/cli-pty-help.spec.ts`, the existing CLI subprocess suite,
|
|
2734
|
+
and `docs/CLI.md` cover narrow/wide deterministic output, a real Unix
|
|
2735
|
+
pseudo-terminal when available, nested help, invalid syntax, legacy invocation
|
|
2736
|
+
compatibility, packaging, and the Windows-safe non-PTY path.
|
|
2737
|
+
|
|
2738
|
+
### C77 — Replay CodeFlow edge-provenance and architecture-export claims
|
|
2739
|
+
|
|
2740
|
+
- Status: evaluated — retained knodin unchanged
|
|
2741
|
+
- Priority: P2
|
|
2742
|
+
- Disposition: Evaluate
|
|
2743
|
+
- DependsOn: C24, C27, C42
|
|
2744
|
+
- Motivation: CodeFlow is a browser-local, MIT architecture mapper with broad
|
|
2745
|
+
language coverage, blast-radius views, and raw JSON export. A secondary
|
|
2746
|
+
LinkedIn post says each dependency identifies its extraction mechanism, but
|
|
2747
|
+
that claim was not found in the current repository or README. Knodin already
|
|
2748
|
+
exposes provenance, confidence, exact-versus-heuristic labels, source evidence,
|
|
2749
|
+
and source lines; only a shared replay can establish whether CodeFlow presents
|
|
2750
|
+
uncertain edges or architecture exports more usefully.
|
|
2751
|
+
- In scope: add a pinned CodeFlow adapter or documented browser-local protocol;
|
|
2752
|
+
compare dependency precision/recall, unsupported edges, uncertainty labels,
|
|
2753
|
+
source evidence, blast-radius completeness, export fidelity, output size,
|
|
2754
|
+
latency, and peak RSS on the existing architecture fixtures.
|
|
2755
|
+
- Out of scope: adopting CodeFlow, copying its browser UI, adding a new service,
|
|
2756
|
+
or changing Knodin's extraction architecture solely from vendor or social-post
|
|
2757
|
+
claims.
|
|
2758
|
+
- Touches: `benchmarks/competitors/TRACKER.md`, a CodeFlow comparison report and
|
|
2759
|
+
immutable raw result, the shared competitive harness only if its existing
|
|
2760
|
+
adapter contract is insufficient, and this roadmap.
|
|
2761
|
+
- Collision risk: competitor tracker and shared harness mappings; schedule apart
|
|
2762
|
+
from other competitive-replay edits.
|
|
2763
|
+
- Source: https://github.com/braedonsaunders/codeflow
|
|
2764
|
+
- Source task: `6h8V3JGJMhV3m63h`
|
|
2765
|
+
- Evidence: `benchmarks/competitors/codeflow-vs-knodin.raw.json` records the
|
|
2766
|
+
exact commit and MIT license, local-only protocol, original/normalized edges,
|
|
2767
|
+
false-positive/false-negative sets, provenance coverage, export fidelity,
|
|
2768
|
+
response bytes, 3+20 warm and five-process cold samples, and independently
|
|
2769
|
+
measured RSS. Both products achieved precision/recall 1.0, unsupported-edge
|
|
2770
|
+
rate 0, and blast completeness 1.0. CodeFlow used 105,955,328 bytes peak RSS
|
|
2771
|
+
and had 122.788/186.788 ms cold p50/p95; knodin used 591,446,016 bytes and had
|
|
2772
|
+
1,056.246/6,332.857 ms. Knodin warm p50/p95 was 0.532/0.736 ms versus
|
|
2773
|
+
CodeFlow's 3.101/8.202 ms. CodeFlow supplied none of the six evaluated
|
|
2774
|
+
file-relationship provenance categories; knodin supplied all six for all
|
|
2775
|
+
three edges. Architecture block exports remain explicitly incomparable.
|
|
2776
|
+
- Disposition: retain knodin unchanged. CodeFlow demonstrates a real cold-start
|
|
2777
|
+
and memory advantage, while knodin already supplies the evaluated provenance
|
|
2778
|
+
and warm-query behavior. Neither the incomparable architecture exports nor
|
|
2779
|
+
feature presence authorizes an architecture or presentation change, and this
|
|
2780
|
+
evaluation makes no aggregate superiority claim.
|
|
2781
|
+
|
|
2782
|
+
Acceptance:
|
|
2783
|
+
|
|
2784
|
+
1. Pin the tested CodeFlow commit and record its MIT license, local execution
|
|
2785
|
+
path, setup, and limitations; do not require an account, hosted service,
|
|
2786
|
+
paid API, source egress, or new Knodin runtime dependency.
|
|
2787
|
+
2. Run both products against identical checked-in fixtures and correctness
|
|
2788
|
+
oracles. Preserve false positives, false negatives, unresolved edges, and
|
|
2789
|
+
unavailable cases in immutable raw results.
|
|
2790
|
+
3. For every exported relationship, record whether the tool supplies relation
|
|
2791
|
+
kind, extractor identity, exact/heuristic confidence, source file and line,
|
|
2792
|
+
and bounded source evidence. Treat the LinkedIn extractor-identity claim as
|
|
2793
|
+
unverified unless the pinned implementation emits it.
|
|
2794
|
+
4. Report precision/recall, unsupported-edge rate, blast-radius completeness,
|
|
2795
|
+
export fidelity, response bytes, cold/warm p50 and p95, and independently
|
|
2796
|
+
measured peak RSS. Do not publish a win from incomparable or oracle-free rows.
|
|
2797
|
+
5. Keep incremental spend at zero and incremental memory below 1 GB. Keep all
|
|
2798
|
+
fixture source local; if CodeFlow cannot run without GitHub egress, record a
|
|
2799
|
+
blocker rather than uploading private source.
|
|
2800
|
+
6. Close with one evidence-based disposition: retain Knodin unchanged, improve
|
|
2801
|
+
provenance presentation/export, or propose a separately reviewed roadmap
|
|
2802
|
+
item. Do not infer an architecture change from feature presence alone.
|
|
2803
|
+
|
|
2804
|
+
### C78 — Add bounded source-to-sink resource reachability
|
|
2805
|
+
|
|
2806
|
+
- Status: evaluated — retained C8 unchanged
|
|
2807
|
+
- Priority: P1
|
|
2808
|
+
- Disposition: Evaluate
|
|
2809
|
+
- DependsOn: C2, C3, C8, C24
|
|
2810
|
+
- Motivation: Code-Graph-RAG's opt-in `READS_FROM`, `WRITES_TO`, and `FLOWS_TO`
|
|
2811
|
+
edges answer a useful provenance question Knodin cannot currently answer:
|
|
2812
|
+
whether a value from an environment variable, file, database, socket, or
|
|
2813
|
+
network source can reach a log, file, database, or outbound-network sink.
|
|
2814
|
+
Knodin's C8 `flow_analysis` is deliberately limited to bounded, on-demand
|
|
2815
|
+
TS/JS facts inside one selected symbol; call-argument evidence is explicitly
|
|
2816
|
+
heuristic and does not compose source-to-sink resource reachability.
|
|
2817
|
+
- In scope: prototype synthetic resource identities plus source-evidenced read,
|
|
2818
|
+
write, assignment, argument, return, and kill facts; begin with environment or
|
|
2819
|
+
local configuration flowing to logging, network, and database sinks; expose a
|
|
2820
|
+
bounded query with exact coverage and omission metadata.
|
|
2821
|
+
- Out of scope: adopting Code-Graph-RAG, Memgraph, Qdrant, Docker, a hosted
|
|
2822
|
+
model, unrestricted graph queries, autonomous vulnerability verdicts, a full
|
|
2823
|
+
persisted PDG, SSA, alias-complete or path-sensitive analysis, or claims of
|
|
2824
|
+
runtime reachability.
|
|
2825
|
+
- Touches if approved: language extraction, optional persisted resource/flow
|
|
2826
|
+
schema or a cheaper on-demand representation, one compact query pattern,
|
|
2827
|
+
source/sink registry, documentation, and labeled multi-language fixtures.
|
|
2828
|
+
- Collision risk: `src/engine/index.ts`, graph schema/migrations, MCP/CLI query
|
|
2829
|
+
enums, and shared language extractors; schedule apart from other graph-schema
|
|
2830
|
+
or query-surface work.
|
|
2831
|
+
- Source: https://github.com/vitali87/code-graph-rag/blob/d0b257b402bb25b54cf8f1dee0af3a71623c8d9e/docs/architecture/data-flow-edges.md
|
|
2832
|
+
- Source task: `6hC5Cv934v8xj2Gh`
|
|
2833
|
+
- Evidence: `benchmarks/evaluations/c78-resource-reachability/` contains the
|
|
2834
|
+
labeled TS/JS oracle, disposable prototype, immutable raw result, tests, and
|
|
2835
|
+
methodology. It achieved 100% precision but 62.5% recall, missing argument,
|
|
2836
|
+
return, and recursion/cycle handoff. Three local repositories measured
|
|
2837
|
+
bounded on-demand analysis against temporary-file persisted JSON; peak RSS was
|
|
2838
|
+
129,859,584 bytes with zero spend and no egress.
|
|
2839
|
+
- Disposition: retain C8's bounded statement flow and call-argument evidence
|
|
2840
|
+
unchanged. The precision gate passed, but recall failed the 80% minimum useful
|
|
2841
|
+
threshold, so neither persistence nor a production query is authorized and no
|
|
2842
|
+
separately numbered implementation item is proposed.
|
|
2843
|
+
|
|
2844
|
+
Acceptance:
|
|
2845
|
+
|
|
2846
|
+
1. First build a disposable prototype and labeled oracle containing true flows,
|
|
2847
|
+
clean overwrites, shadowing, unrelated source/sink co-occurrence, dynamic
|
|
2848
|
+
resource names, argument handoff, return handoff, recursion/cycles, and
|
|
2849
|
+
deliberately unsupported constructs. Do not change the production schema
|
|
2850
|
+
until the prototype passes the gate.
|
|
2851
|
+
2. Require at least 95% precision and report recall per supported language and
|
|
2852
|
+
source/sink class. False positives are a hard failure because the result may
|
|
2853
|
+
inform security review; unsupported or ambiguous paths must remain explicit.
|
|
2854
|
+
3. Every result identifies the source and sink, relation kind, bounded path,
|
|
2855
|
+
source file and line evidence, extraction provenance, confidence, supported
|
|
2856
|
+
language/registry coverage, omissions, truncation, and freshness. Label every
|
|
2857
|
+
path static and heuristic; never call it proof of exploitability, secrecy, or
|
|
2858
|
+
runtime reachability.
|
|
2859
|
+
4. Compare a bounded on-demand implementation with persisted resource/flow
|
|
2860
|
+
edges. Choose the smallest design meeting correctness and warm-latency goals;
|
|
2861
|
+
preserve stable identity, ambiguity safety, monotonic bounds, deterministic
|
|
2862
|
+
ordering, hard item/byte/token budgets, and fail-closed stale-index behavior.
|
|
2863
|
+
5. Measure clean and incremental index time, database growth, cold/warm p50 and
|
|
2864
|
+
p95 query latency, response bytes, and peak RSS on at least three real local
|
|
2865
|
+
repositories. Incremental spend remains zero and incremental memory remains
|
|
2866
|
+
below 1 GB, with no hosted service, credentials, source egress, or required
|
|
2867
|
+
production model.
|
|
2868
|
+
6. Start with a deliberately narrow language and source/sink matrix justified by
|
|
2869
|
+
fixtures. Adding a registry entry requires positive and negative acceptance
|
|
2870
|
+
cases; a language without a verified registry emits an explicit omission,
|
|
2871
|
+
not inferred coverage.
|
|
2872
|
+
7. If the gate fails, retain C8's current bounded statement flow and call-
|
|
2873
|
+
argument evidence unchanged. Do not broaden or persist the feature merely to
|
|
2874
|
+
match a competitor's schema.
|
|
2875
|
+
|
|
2876
|
+
### C79 — Bounded repository applicability signals
|
|
2877
|
+
|
|
2878
|
+
- Status: implemented
|
|
2879
|
+
- Priority: P1
|
|
2880
|
+
- Disposition: Evaluate
|
|
2881
|
+
- DependsOn: C52, C69
|
|
2882
|
+
- Touches: `repos discover` CLI option/schema, bounded repository inspection,
|
|
2883
|
+
Git-remote parsing, worktree-aware discovery fixtures, CLI help/reference, and
|
|
2884
|
+
compatibility evidence. It does not change `doctor` or the MCP surface.
|
|
2885
|
+
- What: decide whether repository marker and configuration detection belongs in
|
|
2886
|
+
knodin's repository-discovery charter. If it does, add an additive, opt-in
|
|
2887
|
+
`knodin repos discover <roots...> --json --signals` facet so consumers can use
|
|
2888
|
+
knodin's existing repository/worktree classification instead of duplicating a
|
|
2889
|
+
second traversal and a less reliable definition of a repository.
|
|
2890
|
+
- Scope: without `--signals`, output must remain byte-identical. With the flag,
|
|
2891
|
+
every repository gains `signals`, using `{}` for honest absence. Detection is
|
|
2892
|
+
read-only, offline, credential-free, bounded by existing bytes/tokens/items
|
|
2893
|
+
budgets, and limited to a documented allowlist of marker paths rather than a
|
|
2894
|
+
tree glob. Never read potentially secret file contents; the only content-read
|
|
2895
|
+
exception is allowlisted hook-manager configuration needed to report whether
|
|
2896
|
+
it references `aidev-track`. Linked worktrees inspect their own checkout and
|
|
2897
|
+
preserve existing skip/include behavior.
|
|
2898
|
+
- Evidence: either a checked-in charter decision closing the item unchanged, or
|
|
2899
|
+
a checked-in implementation with compatibility bytes, help/docs, unit
|
|
2900
|
+
fixtures, bounded-resource results, and verifier output.
|
|
2901
|
+
|
|
2902
|
+
Charter decision (2026-08-03): accepted. These path-derived, repository-scoped
|
|
2903
|
+
signals belong in discovery because they reuse knodin's existing bounded
|
|
2904
|
+
repository and linked-worktree classification. The facet remains opt-in and
|
|
2905
|
+
descriptive: it does not decide whether an agent skill applies or whether a
|
|
2906
|
+
detected service is configured correctly.
|
|
2907
|
+
|
|
2908
|
+
Acceptance:
|
|
2909
|
+
|
|
2910
|
+
1. First record the charter decision. If marker/config policy is outside
|
|
2911
|
+
knodin's charter, close C79 unchanged with checked-in rationale; no feature
|
|
2912
|
+
implementation is required. Otherwise implement only the opt-in facet.
|
|
2913
|
+
2. A golden compatibility test proves `repos discover <roots...> --json`
|
|
2914
|
+
remains byte-identical without the flag. With `--signals`, every repository
|
|
2915
|
+
entry has a deterministic `signals` object, including `{}` when the
|
|
2916
|
+
allowlisted inspection finds nothing; schema version changes only if the
|
|
2917
|
+
existing compatibility policy requires it.
|
|
2918
|
+
3. When detected, the deterministic contract must return `hookManager`,
|
|
2919
|
+
allowlisted `markerFiles` (including Sonar, CI/PR-workflow, agent, and MCP
|
|
2920
|
+
client markers), sanitized `remotes` with name/host/owner, `ciProviders`,
|
|
2921
|
+
`agentConfigs`, and `aidevTrackReferenced: true|false` when an allowlisted
|
|
2922
|
+
hook-manager config is present. Detection performs
|
|
2923
|
+
no writes, network access, credential access, source egress, or secret-file
|
|
2924
|
+
content reads, except the narrow allowlisted hook-manager parse for the
|
|
2925
|
+
`aidev-track` reference. Incremental spend is zero.
|
|
2926
|
+
4. Unit fixtures cover lefthook, husky, pre-commit, no hook manager,
|
|
2927
|
+
`aidev-track` reference detection, multiple remotes, multiple remote hosts,
|
|
2928
|
+
missing `.git`, every marker absent, and linked worktrees. Tests prove fixed
|
|
2929
|
+
allowlists and existing bytes/tokens/items budgets bound inspection and
|
|
2930
|
+
serialization.
|
|
2931
|
+
5. `knodin doctor` behavior is unchanged. The repos CLI reference and
|
|
2932
|
+
`repos discover --help` document the flag, returned fields, allowlist,
|
|
2933
|
+
omissions, privacy constraints, and worktree behavior.
|
|
2934
|
+
|
|
2935
|
+
Implementation evidence: the deterministic verifier and bounded-resource
|
|
2936
|
+
result are checked in at
|
|
2937
|
+
[`docs/evidence/c79-repository-signals-2026-08-03.md`](../docs/evidence/c79-repository-signals-2026-08-03.md)
|
|
2938
|
+
and
|
|
2939
|
+
[`docs/evidence/c79-repository-signals-2026-08-03.json`](../docs/evidence/c79-repository-signals-2026-08-03.json).
|
|
2940
|
+
The unit acceptance fixtures are in
|
|
2941
|
+
`src/__tests__/unit/repository-management.spec.ts`; CLI help coverage is in
|
|
2942
|
+
`src/__tests__/unit/cli-model.spec.ts`, and the unchanged doctor regression is
|
|
2943
|
+
`src/__tests__/unit/doctor.spec.ts`.
|
|
2944
|
+
|
|
2945
|
+
Known limitations and decision gate: marker presence describes repository
|
|
2946
|
+
configuration; it does not prove that a skill applies or that a service is
|
|
2947
|
+
configured correctly. If the charter review rejects this application-policy
|
|
2948
|
+
facet, closing C79 unchanged is a valid terminal outcome and downstream tools
|
|
2949
|
+
may inspect only the repository paths knodin returns.
|
|
2950
|
+
|
|
2951
|
+
### C80 — Structural-first retrieval routing
|
|
2952
|
+
|
|
2953
|
+
- Status: implemented
|
|
2954
|
+
- Priority: P0
|
|
2955
|
+
- Disposition: Must close
|
|
2956
|
+
- DependsOn: C3, C24, C42, C43
|
|
2957
|
+
- Touches: retrieval dispatch and telemetry, search/context tests, a checked-in
|
|
2958
|
+
cross-file oracle, and deterministic ablation replay artifacts.
|
|
2959
|
+
- What: route stable identity and exact name/signature/path evidence through
|
|
2960
|
+
lexical/FTS and bounded graph expansion before loading embeddings. Serialize
|
|
2961
|
+
the selected route and `embeddingsUsed` so behavior is explainable.
|
|
2962
|
+
- Scope: preserve the one-tool gateway, stable identities, ambiguity safety,
|
|
2963
|
+
freshness, deterministic ordering, and current hard response budgets.
|
|
2964
|
+
- Evidence: a local checked-in implementation, fixtures, raw three-arm replay,
|
|
2965
|
+
verifier, and report covering factual precision/recall, unsupported claims,
|
|
2966
|
+
latency, and independently measured peak RSS.
|
|
2967
|
+
|
|
2968
|
+
Acceptance:
|
|
2969
|
+
|
|
2970
|
+
1. Exact symbol, path, and source-evidenced cross-file questions do not load the
|
|
2971
|
+
embedding model when structural evidence is sufficient; tests assert the
|
|
2972
|
+
route and `embeddingsUsed: false` through CLI and MCP-compatible responses.
|
|
2973
|
+
2. A checked-in multi-hop oracle replays lexical/vector-only, bounded graph
|
|
2974
|
+
expansion, and the complete route deterministically. Graph expansion becomes
|
|
2975
|
+
default only if correctness improves without increasing unsupported,
|
|
2976
|
+
ambiguous, or stale claims; otherwise retain the narrower passing route.
|
|
2977
|
+
3. Preserve raw results and report precision, recall, omissions/refusals,
|
|
2978
|
+
latency, response size, and peak RSS. Incremental spend is zero, source stays
|
|
2979
|
+
local, no network or account is required, and incremental memory stays below
|
|
2980
|
+
1 GB.
|
|
2981
|
+
|
|
2982
|
+
Delivered evidence: structural-first dispatch and request-local retrieval
|
|
2983
|
+
telemetry are implemented in `src/engine/index.ts`; CLI/MCP parity, stable
|
|
2984
|
+
identity, exact path/name, lexical sufficiency, embedding fallback, bounded
|
|
2985
|
+
work, and duplicate-name safety are covered by
|
|
2986
|
+
`src/__tests__/unit/structural-routing.spec.ts`. The checked-in multi-hop and
|
|
2987
|
+
duplicate-name oracle, fixture, deterministic four-route raw replay, and
|
|
2988
|
+
verifier outputs are under
|
|
2989
|
+
`benchmarks/evaluations/c80-structural-routing/`; reproduce them with
|
|
2990
|
+
`npm run bench:c80 && npm run verify:c80`. The limitations-aware precision,
|
|
2991
|
+
recall, unsupported/ambiguous/stale, omission, latency, response-size, spend,
|
|
2992
|
+
locality, and independently isolated RSS report is
|
|
2993
|
+
`docs/evidence/c80-structural-routing-2026-08-03.md`.
|
|
2994
|
+
|
|
2995
|
+
Decision: bounded graph expansion is the default structural route because the
|
|
2996
|
+
checked-in gate improves mean oracle recall from 0.389 to 1.000 while preserving
|
|
2997
|
+
1.000 precision, zero unsupported and stale claims, and the same one surfaced
|
|
2998
|
+
duplicate-name ambiguity. The complete route preserves those results and loads
|
|
2999
|
+
embeddings only when exact or all-token structural evidence is insufficient.
|
|
3000
|
+
The largest isolated incremental RSS sample was 160,284,672 bytes; incremental
|
|
3001
|
+
spend was zero and no network, account, credential, or hosted service was used.
|
|
3002
|
+
|
|
3003
|
+
Known limitations and decision gate: structural routing is not proof that
|
|
3004
|
+
semantic retrieval is unnecessary. Embeddings remain a bounded fallback, and
|
|
3005
|
+
token reduction alone cannot authorize a route that weakens correctness. Graph
|
|
3006
|
+
expansion follows resolved source relationships only, is capped at two hops and
|
|
3007
|
+
200 selected nodes, and cannot recover runtime-only or unresolved relationships.
|
|
3008
|
+
The three-case deterministic oracle proves this decision only on its checked-in
|
|
3009
|
+
cross-file chains; sampled RSS can miss short-lived peaks between samples.
|
|
3010
|
+
|
|
3011
|
+
### C81 — Powered end-to-end benchmark with a diagnose arm
|
|
3012
|
+
|
|
3013
|
+
- Status: evaluated — retained knodin unchanged
|
|
3014
|
+
- Priority: P0
|
|
3015
|
+
- Disposition: Must close
|
|
3016
|
+
- DependsOn: C42, C43, C59, C80
|
|
3017
|
+
- Touches: competitive benchmark corpus/runner, task labels, deterministic
|
|
3018
|
+
patch/test oracles, raw trajectories, statistics, and evidence report.
|
|
3019
|
+
- What: compare ordinary filesystem tools, knodin, and one pinned local
|
|
3020
|
+
competitor on retrieval and failing-build/test diagnosis using a powered,
|
|
3021
|
+
paired design.
|
|
3022
|
+
- Scope: at least 40 tasks, at least three repetitions, pinned repositories,
|
|
3023
|
+
commits, prompts, models, versions, seeds, and budgets; confidence intervals
|
|
3024
|
+
and pre-registered analysis are mandatory.
|
|
3025
|
+
- Evidence: checked-in corpus and labels, immutable raw trajectories, verifier
|
|
3026
|
+
output, confidence intervals, and a limitations-aware report.
|
|
3027
|
+
|
|
3028
|
+
Acceptance:
|
|
3029
|
+
|
|
3030
|
+
1. Each arm runs the same tasks and deterministic correctness oracles. The
|
|
3031
|
+
diagnose subset scores turns to a correct patch and passing test, not an LLM
|
|
3032
|
+
judge alone; unavailable and incomparable rows remain explicit.
|
|
3033
|
+
2. Cache behavior is controlled and recorded. Do not report billed-cost claims
|
|
3034
|
+
unless cache economics are comparable; this item must incur zero model/API
|
|
3035
|
+
spend and require no hosted service, credentials, or source egress.
|
|
3036
|
+
Incremental spend is zero.
|
|
3037
|
+
3. Record patch-application rate, turns to correct edit, factual correctness,
|
|
3038
|
+
tokens where measured, cold/warm latency, and peak RSS under 1 GB. Preserve
|
|
3039
|
+
losing cases and raw data without overwriting prior evidence.
|
|
3040
|
+
|
|
3041
|
+
Known limitations and decision gate: the benchmark may justify retaining the
|
|
3042
|
+
product unchanged. It authorizes no superiority claim beyond oracle-qualified,
|
|
3043
|
+
comparable rows and no production feature outside a separately numbered item.
|
|
3044
|
+
|
|
3045
|
+
Delivered evidence: the pinned 40-task corpus, three-repetition paired runner,
|
|
3046
|
+
deterministic retrieval and patch/test oracles, immutable accepted and excluded
|
|
3047
|
+
pilot trajectories, integrity hashes, and verifier are under
|
|
3048
|
+
`benchmarks/evaluations/c81-powered-benchmark/`; reproduce validation and
|
|
3049
|
+
confidence intervals with `npm run verify:c81`. The limitations-aware report is
|
|
3050
|
+
[`C81 powered benchmark evidence`](../docs/evidence/c81-powered-benchmark-2026-08-03.md).
|
|
3051
|
+
|
|
3052
|
+
Decision: retain knodin unchanged. All three available arms achieved 1.000
|
|
3053
|
+
factual correctness and patch-plus-passing-test rate on this narrow synthetic
|
|
3054
|
+
corpus, with paired correctness differences of 0.000. The result supports no
|
|
3055
|
+
superiority, token, billed-cost, or production claim. The conservative
|
|
3056
|
+
worker-plus-descendant peak-RSS upper bound was 351,469,568 bytes; spend,
|
|
3057
|
+
hosted-service use, credentials, and source egress were zero.
|
|
3058
|
+
|
|
3059
|
+
### C82 — Productize `compress diagnose`
|
|
3060
|
+
|
|
3061
|
+
- Status: implemented
|
|
3062
|
+
- Priority: P0
|
|
3063
|
+
- Disposition: Must close
|
|
3064
|
+
- DependsOn: C59, C81
|
|
3065
|
+
- Touches: existing compress diagnose CLI/MCP dispatch, discoverability and
|
|
3066
|
+
documentation, stable response schema, fixtures, telemetry, and replay.
|
|
3067
|
+
- What: make the retained-failure-to-graph workflow discoverable and robust so
|
|
3068
|
+
an artifact resolves to owning symbols, tests, callers, and bounded source
|
|
3069
|
+
without rerunning the failed command.
|
|
3070
|
+
- Scope: reuse the existing single MCP tool and retained C59 primitive; preserve
|
|
3071
|
+
stable artifact identity, freshness, ambiguity, traversal, and budget bounds.
|
|
3072
|
+
- Evidence: checked-in adversarial diagnostic fixtures, CLI/MCP parity tests,
|
|
3073
|
+
golden responses, raw replay, verifier, and local evidence report.
|
|
3074
|
+
|
|
3075
|
+
Acceptance:
|
|
3076
|
+
|
|
3077
|
+
1. Representative compiler, test, stack-trace, multiline, truncated, malformed,
|
|
3078
|
+
stale, ambiguous, and missing-artifact cases return stable schema fields and
|
|
3079
|
+
exact owning source evidence or an explicit bounded refusal.
|
|
3080
|
+
2. CLI and MCP produce equivalent source identities, callers/tests, omissions,
|
|
3081
|
+
freshness, confidence, and telemetry without a new MCP tool, command rerun,
|
|
3082
|
+
hosted service, credential, source egress, or model/API spend.
|
|
3083
|
+
Incremental spend is zero.
|
|
3084
|
+
3. The C81 diagnose corpus demonstrates the workflow end to end with checked-in
|
|
3085
|
+
deterministic oracles, hard item/byte/token caps, recoverable continuation,
|
|
3086
|
+
and peak incremental memory below 1 GB.
|
|
3087
|
+
|
|
3088
|
+
Known limitations and decision gate: static graph evidence cannot prove runtime
|
|
3089
|
+
causality. Documentation must say that diagnoses are bounded source-evidenced
|
|
3090
|
+
candidates, not guaranteed root causes or automatic fixes.
|
|
3091
|
+
|
|
3092
|
+
Delivered evidence: the existing CLI and single-tool MCP dispatch now share an
|
|
3093
|
+
additive stable diagnosis envelope with confidence, exact diagnostic/relation
|
|
3094
|
+
omissions, retained-artifact continuation, and command-rerun/cap telemetry.
|
|
3095
|
+
Focused parity and refusal coverage is in `src/__tests__/unit/`; adversarial
|
|
3096
|
+
fixtures, golden responses, immutable raw replay, hard-cap verification, and
|
|
3097
|
+
resource evidence are under
|
|
3098
|
+
`benchmarks/evaluations/c82-compress-diagnose/`. Reproduce all ten adversarial
|
|
3099
|
+
cases and all ten C81 diagnose tasks with `npm run verify:c82`. The evidence
|
|
3100
|
+
report is [`C82 compress diagnose evidence`](../docs/evidence/c82-compress-diagnose-2026-08-03.md).
|
|
3101
|
+
|
|
3102
|
+
The conservative measured maximum RSS was 309,116,928 bytes, below 1 GB.
|
|
3103
|
+
Incremental spend, command reruns, network/hosted service use, credentials, and
|
|
3104
|
+
source egress were zero. Static candidates remain neither runtime root-cause
|
|
3105
|
+
proof nor automatic fixes; the synthetic replay does not cover every toolchain.
|
|
3106
|
+
|
|
3107
|
+
### C83 — Progressive evidence delivery and hash handshake
|
|
3108
|
+
|
|
3109
|
+
- Status: implemented
|
|
3110
|
+
- Priority: P0
|
|
3111
|
+
- Disposition: Must close
|
|
3112
|
+
- DependsOn: C80, C82
|
|
3113
|
+
- Touches: gateway response contracts, locate/outline/evidence/expand dispatch,
|
|
3114
|
+
continuation/evidence handles, content hashing, CLI/MCP tests, and docs.
|
|
3115
|
+
- What: deliver recoverable evidence levels and allow delta/omission only when
|
|
3116
|
+
the client supplies an exact content hash for the baseline it holds.
|
|
3117
|
+
- Scope: every level reports returned and omitted material, reason, `more`, a
|
|
3118
|
+
stable continuation, freshness, confidence, and truthful hard budgets.
|
|
3119
|
+
- Evidence: checked-in golden protocol fixtures, tamper/staleness/rename tests,
|
|
3120
|
+
raw bounded replay, verifier, and compatibility report.
|
|
3121
|
+
|
|
3122
|
+
Delivered evidence: the additive `evidence` operation and CLI command share the
|
|
3123
|
+
source-preserving implementation in `src/progressive-evidence.ts`. Golden and
|
|
3124
|
+
raw 14-case protocol replays, fixture, resource record, compatibility report,
|
|
3125
|
+
and limitations are under `benchmarks/evaluations/c83-progressive-evidence/`;
|
|
3126
|
+
reproduce them with `npm run verify:c83`. Unit acceptance covers complete
|
|
3127
|
+
omission recovery, adversarial hashes/handles/paths, stale and renamed handles,
|
|
3128
|
+
and byte-identical CLI/MCP edit anchors. The evidence report is
|
|
3129
|
+
[`C83 progressive evidence evidence`](../docs/evidence/c83-progressive-evidence-2026-08-03.md).
|
|
3130
|
+
|
|
3131
|
+
Acceptance:
|
|
3132
|
+
|
|
3133
|
+
1. Exact hash match permits the documented delta or omission path; missing,
|
|
3134
|
+
stale, truncated, renamed, or tampered baselines return full required source
|
|
3135
|
+
and state which path was taken. Client assertion alone is never trusted.
|
|
3136
|
+
2. Locate, outline, evidence, and expand remain deterministic and recoverable
|
|
3137
|
+
under item/byte/token caps; tests prove omitted evidence can be requested and
|
|
3138
|
+
that edit anchors remain intact across CLI and MCP.
|
|
3139
|
+
3. The protocol remains local, zero-spend, offline, under 1 GB incremental RSS,
|
|
3140
|
+
backward-compatible where promised, and fail-closed on stale or ambiguous
|
|
3141
|
+
evidence. Checked-in fixtures include adversarial hash and handle inputs.
|
|
3142
|
+
|
|
3143
|
+
Known limitations and decision gate: this is not lossy summarization and does
|
|
3144
|
+
not authorize token-savings or task-success claims until C81 measures them.
|
|
3145
|
+
The V1 surface addresses one exact regular repository file at a time, refuses
|
|
3146
|
+
symlinks and ambiguous/path-traversal input, and uses a lightweight top-level
|
|
3147
|
+
outline rather than claiming the engine's complete relationship model. The
|
|
3148
|
+
synthetic replay authorizes no production-readiness or cross-platform claim.
|
|
3149
|
+
|
|
3150
|
+
### C84 — Optional SCIP import
|
|
3151
|
+
|
|
3152
|
+
- Status: implemented
|
|
3153
|
+
- Priority: P1
|
|
3154
|
+
- Disposition: Must close
|
|
3155
|
+
- DependsOn: C24, C80, C83
|
|
3156
|
+
- Touches: optional SCIP ingestion, stable identity/provenance mapping, conflict
|
|
3157
|
+
handling, LSIF/native fallbacks, fixtures, lifecycle bounds, and docs.
|
|
3158
|
+
- What: import compiler-produced SCIP facts as an opt-in precision tier while
|
|
3159
|
+
retaining LSIF compatibility and the native AST/structural/heuristic ladder.
|
|
3160
|
+
- Scope: SCIP files are local inputs; import is deterministic, bounded, and
|
|
3161
|
+
optional. Live LSP integration is explicitly not part of C84.
|
|
3162
|
+
- Evidence: checked-in SCIP fixtures and expected graph facts, malformed/path/
|
|
3163
|
+
oversize/conflict tests, fallback replay, verifier, and resource report.
|
|
3164
|
+
|
|
3165
|
+
Delivered evidence: `knodin index --scip <file>` explicitly imports a bounded
|
|
3166
|
+
local SCIP protobuf snapshot without discovering or running a producer. The
|
|
3167
|
+
npm-integrity-pinned `@sourcegraph/scip-typescript@0.4.0` fixture, binary hash,
|
|
3168
|
+
source, configuration, expected facts, duplicate document-local symbols, and
|
|
3169
|
+
adversarial cases are under `fixtures/c84-scip/` and `src/__tests__/unit/`.
|
|
3170
|
+
The raw offline replay, distinct relationship kinds, native/LSIF fallback,
|
|
3171
|
+
stable identity refresh, 64 MiB/10,000-file/250,000-fact/30-second bounds, and
|
|
3172
|
+
502,431,744-byte incremental RSS record are under
|
|
3173
|
+
`benchmarks/evaluations/c84-scip-import/`; reproduce them with
|
|
3174
|
+
`npm run verify:c84`. The evidence report is
|
|
3175
|
+
[`C84 optional SCIP import evidence`](../docs/evidence/c84-optional-scip-import-2026-08-03.md).
|
|
3176
|
+
|
|
3177
|
+
Acceptance:
|
|
3178
|
+
|
|
3179
|
+
1. Pinned fixtures produce deterministic stable identities, relationships,
|
|
3180
|
+
source ranges, and `scip` provenance. Conflicting tiers surface ambiguity and
|
|
3181
|
+
never silently replace higher-confidence or fresher evidence.
|
|
3182
|
+
2. Missing, malformed, hostile-path, stale, and oversized inputs fail safely;
|
|
3183
|
+
repositories without SCIP retain current native/LSIF behavior with no daemon,
|
|
3184
|
+
account, network, source egress, or model/API spend.
|
|
3185
|
+
Incremental spend is zero.
|
|
3186
|
+
3. Import and query replays remain within explicit file/fact/time limits and
|
|
3187
|
+
below 1 GB incremental RSS. Raw local results, verifier output, and known
|
|
3188
|
+
language/indexer coverage are checked in.
|
|
3189
|
+
|
|
3190
|
+
Known limitations and decision gate: SCIP precision is bounded by the producer
|
|
3191
|
+
and languages represented in the imported index. Any live LSP tier requires a
|
|
3192
|
+
separately approved roadmap item and must never become a baseline dependency.
|
|
3193
|
+
|
|
3194
|
+
### C85 — Bounded git-history review signals
|
|
3195
|
+
|
|
3196
|
+
- Status: implemented
|
|
3197
|
+
- Priority: P1
|
|
3198
|
+
- Disposition: Must close
|
|
3199
|
+
- DependsOn: C24, C44, C81
|
|
3200
|
+
- Touches: review risk serialization, bounded Git history queries, unavailable/
|
|
3201
|
+
shallow states, synthetic repository fixtures, telemetry, and docs.
|
|
3202
|
+
- What: add churn, co-change, and coupling as separately itemized review signals
|
|
3203
|
+
beside graph impact, test gaps, and structural centrality; forbid opaque scores.
|
|
3204
|
+
- Scope: all Git operations have explicit commit/file/time bounds and preserve
|
|
3205
|
+
deterministic ordering, source evidence, and graceful unavailable states.
|
|
3206
|
+
- Evidence: checked-in synthetic Git DAG, exact signal oracles, bounded-command
|
|
3207
|
+
tests, raw review replay, verifier, and resource report.
|
|
3208
|
+
- Evidence delivered: `src/engine/git-history.ts` and
|
|
3209
|
+
`src/__tests__/unit/git-history-signals.spec.ts` cover rename-aware itemized
|
|
3210
|
+
facts, merge/unrelated separation, shallow deepen, unborn, missing Git,
|
|
3211
|
+
non-repository, arbitrary HEAD failure, timeout, path containment, hard bounds,
|
|
3212
|
+
deterministic truncation, cache validity, and CLI/MCP serialization.
|
|
3213
|
+
`benchmarks/evaluations/c85-git-history/` checks in the exact synthetic DAG,
|
|
3214
|
+
raw replay, latency/command/resource evidence, limitations, and verifier;
|
|
3215
|
+
`npm run verify:c85` reproduces churn 2, co-change 2, coupling 1, three bounded
|
|
3216
|
+
history commands, zero locality costs, and peak RSS below 1 GB.
|
|
3217
|
+
|
|
3218
|
+
Acceptance:
|
|
3219
|
+
|
|
3220
|
+
1. Fixtures cover renames, merges, unrelated co-occurrence, shallow history,
|
|
3221
|
+
unborn repositories, missing Git, and bounded truncation. Each output exposes
|
|
3222
|
+
the contributing facts and omissions rather than only a composite number.
|
|
3223
|
+
2. Review remains deterministic, fail-closed, and useful when history is absent;
|
|
3224
|
+
commands cannot escape the repository or exceed configured commit/file/time
|
|
3225
|
+
limits. CLI/MCP-compatible evidence retains freshness and confidence.
|
|
3226
|
+
3. Checked-in replay reports correctness, latency, command counts, and peak RSS
|
|
3227
|
+
below 1 GB with zero spend, no network, no credentials, and no source egress.
|
|
3228
|
+
|
|
3229
|
+
Known limitations and decision gate: co-change is correlation, not causation.
|
|
3230
|
+
History signals may influence itemized review context but cannot independently
|
|
3231
|
+
assert defect likelihood, ownership, or mandatory remediation.
|
|
3232
|
+
|
|
3233
|
+
### C86 — Static-embedding bake-off
|
|
3234
|
+
|
|
3235
|
+
- Status: evaluated — retained MiniLM unchanged
|
|
3236
|
+
- Priority: P2
|
|
3237
|
+
- Disposition: Evaluate
|
|
3238
|
+
- DependsOn: C46, C51, C80, C81
|
|
3239
|
+
- Touches: isolated Model2Vec/MiniLM evaluation runner, fixed corpus and labels,
|
|
3240
|
+
model provenance/digests, raw measurements, verifier, and decision record.
|
|
3241
|
+
- What: evaluate a pinned static-embedding candidate against retained MiniLM
|
|
3242
|
+
before authorizing any runtime, index-format, or default-search change.
|
|
3243
|
+
- Scope: measure cold/warm indexing and query performance, RSS, storage, recall,
|
|
3244
|
+
NDCG, and downstream task correctness with at least three repetitions.
|
|
3245
|
+
- Evidence: checked-in corpus/query labels, model identifiers and digests, raw
|
|
3246
|
+
repeated results, verifier output, confidence/variance report, and disposition.
|
|
3247
|
+
- Evidence delivered: `benchmarks/evaluations/c86-static-embeddings/` contains
|
|
3248
|
+
the pinned 20-document corpus, ten relevance/task labels, model revisions and
|
|
3249
|
+
ONNX digests, six fresh-process raw repetitions, cold/warm timing, RSS and
|
|
3250
|
+
storage measurements, descriptive confidence/variance, unsupported cases,
|
|
3251
|
+
limitations, and the retain-MiniLM decision. `npm run verify:c86` checks the
|
|
3252
|
+
identical-route/budget oracle, three repetitions per arm, local/offline and
|
|
3253
|
+
zero-spend contract, sub-1-GB incremental RSS, artifact pins, and evaluation
|
|
3254
|
+
isolation without requiring either model artifact.
|
|
3255
|
+
|
|
3256
|
+
Acceptance:
|
|
3257
|
+
|
|
3258
|
+
1. Both candidates run on identical local fixtures, retrieval routes, budgets,
|
|
3259
|
+
and deterministic relevance/task oracles. Record setup, unsupported cases,
|
|
3260
|
+
cold/warm timing, index size, recall/NDCG, correctness, and peak RSS.
|
|
3261
|
+
2. The replay is offline after documented local preparation, adds no production
|
|
3262
|
+
dependency, uses no hosted API/account/source egress, incurs zero spend, and
|
|
3263
|
+
stays below 1 GB incremental memory. Pin model identity and artifact digest.
|
|
3264
|
+
3. Retain MiniLM unless the pre-registered decision gate shows a reproducible
|
|
3265
|
+
correctness-preserving win. Close by retaining MiniLM, deferring/rejecting
|
|
3266
|
+
Model2Vec, or proposing a separately reviewed implementation item; evaluation
|
|
3267
|
+
alone must not change production behavior or schema.
|
|
3268
|
+
|
|
3269
|
+
Known limitations and decision gate: corpus results may not generalize to every
|
|
3270
|
+
language or repository. Publish only measured local outcomes, never a universal
|
|
3271
|
+
embedding-quality or product-superiority claim.
|
|
3272
|
+
|
|
3273
|
+
Decision: retain MiniLM unchanged. On the checked-in fixture both models reached
|
|
3274
|
+
1.000 recall@5, 1.000 NDCG@10, and 10/10 top-1 task correctness in every
|
|
3275
|
+
repetition. Model2Vec was repeatably faster and smaller locally, but the corpus
|
|
3276
|
+
does not resolve repository-scale or multilingual-query quality, so the
|
|
3277
|
+
pre-registered gate's no-material-unsupported-category condition did not pass.
|
|
3278
|
+
No production runtime, dependency, index format, schema, default, or MCP surface
|
|
3279
|
+
changed, and no separately numbered implementation item is authorized.
|
|
3280
|
+
|
|
3281
|
+
### C87 — Structural cold-start fast path
|
|
3282
|
+
|
|
3283
|
+
- Status: implemented
|
|
3284
|
+
- Priority: P0
|
|
3285
|
+
- Disposition: Must close
|
|
3286
|
+
- DependsOn: C56, C68
|
|
3287
|
+
- Touches: packaged CLI launcher, structural snapshot publication, direct-file
|
|
3288
|
+
fallback, compression routing, benchmarks, and truth-bound response fields.
|
|
3289
|
+
- What: bypass graph/model initialization for structural-only cold requests
|
|
3290
|
+
while preserving the persistent graph route for relationship intelligence.
|
|
3291
|
+
- Scope: file outline, exact file-scoped extraction, batch outline, project
|
|
3292
|
+
overview, and pure output compression only.
|
|
3293
|
+
- Evidence: `src/structural-snapshot.ts`, `src/structural-fast-path.ts`,
|
|
3294
|
+
`src/pure-compression-cli.ts`, focused unit tests, and
|
|
3295
|
+
`benchmarks/evaluations/c87-structural-cold-start/raw-results.json`.
|
|
3296
|
+
|
|
3297
|
+
Indexing now publishes an atomic compact file/symbol snapshot. The packaged
|
|
3298
|
+
launcher routes file outlines, exact file-scoped extraction, batch outlines,
|
|
3299
|
+
project overview, and pure compression without importing the graph or embedding
|
|
3300
|
+
modules. A missing or fingerprint-stale requested file is parsed directly and
|
|
3301
|
+
labeled `lexical-structural-fallback`; graph-backed search, impact, review,
|
|
3302
|
+
relationships, and architecture retain their existing initialization and
|
|
3303
|
+
freshness contract.
|
|
3304
|
+
|
|
3305
|
+
Acceptance:
|
|
3306
|
+
|
|
3307
|
+
1. File-outline, exact-symbol, and project-overview fresh-process p50 are
|
|
3308
|
+
at most 50 ms and p95 at most 75 ms after separating benchmark-parent spawn/
|
|
3309
|
+
scheduling overhead from measured child-process lifetime. Preserve both raw
|
|
3310
|
+
observations; do not hide an application or runtime-floor miss.
|
|
3311
|
+
2. Warm p50 regresses by no more than 10%, response tokens by no more than 5%,
|
|
3312
|
+
and the existing shared correctness oracles remain 100% for both fast and
|
|
3313
|
+
graph routes.
|
|
3314
|
+
3. Every fast answer includes path/fingerprint evidence, exact locations,
|
|
3315
|
+
evidence quality, `graphEnriched: false`, file-scoped uniqueness limits, and
|
|
3316
|
+
an upgrade handle. Dependency evidence proves no graph/model import.
|
|
3317
|
+
4. Graph-backed operations retain stable identity, relationships, and freshness
|
|
3318
|
+
evidence. Snapshot absence/staleness never becomes a repository-wide
|
|
3319
|
+
freshness claim. Checked-in replay evidence requires zero hosted spend,
|
|
3320
|
+
network, account, credential, or source egress.
|
|
3321
|
+
|
|
3322
|
+
Known limitations and decision gate: the direct fallback is a bounded lexical
|
|
3323
|
+
structural parser, not the full Tree-sitter graph extractor. Its classification
|
|
3324
|
+
is explicit and a later graph operation is the supported enrichment path. Keep
|
|
3325
|
+
the fast route only while all checked-in acceptance gates pass.
|
|
3326
|
+
|
|
3327
|
+
### C88 — Profile-based contained execution
|
|
3328
|
+
|
|
3329
|
+
- Status: implemented — macOS certified; Linux/Windows unavailable
|
|
3330
|
+
- Priority: P1
|
|
3331
|
+
- Disposition: Must close
|
|
3332
|
+
- DependsOn: C57, C58, C59
|
|
3333
|
+
- Touches: one-tool schema/dispatcher, profile configuration, macOS Seatbelt
|
|
3334
|
+
adapter, output compression/diagnosis, audit records, adversarial evidence,
|
|
3335
|
+
and security documentation.
|
|
3336
|
+
- What: permit only preapproved single-process profiles behind enforceable
|
|
3337
|
+
macOS filesystem/network/process boundaries and fail closed elsewhere.
|
|
3338
|
+
- Scope: test, lint, typecheck, and build-style profiles with immutable argv;
|
|
3339
|
+
arbitrary commands, child processes, and portable containment claims remain
|
|
3340
|
+
excluded.
|
|
3341
|
+
- Evidence: `src/execution-profile.ts`, `docs/CONTAINED-EXECUTION.md`, gateway/
|
|
3342
|
+
adversarial unit tests, and `benchmarks/evaluations/c88-contained-execution/`.
|
|
3343
|
+
|
|
3344
|
+
C58's portable polling prototype remains rejected. C88 authorizes a narrower
|
|
3345
|
+
production contract: independent global and repository enablement, immutable
|
|
3346
|
+
repository profile identity, absolute configured executable and fixed argv,
|
|
3347
|
+
no shell/caller suffix/cwd/environment, sanitized environment, repository cwd,
|
|
3348
|
+
hard timeout/output/fork limits, private profile-only audit, recoverable
|
|
3349
|
+
compression, and failure-to-code diagnosis.
|
|
3350
|
+
|
|
3351
|
+
Acceptance:
|
|
3352
|
+
|
|
3353
|
+
1. macOS native evidence proves literal metacharacter argv, denied ambient
|
|
3354
|
+
credentials/user-data reads, denied network sockets, denied forks,
|
|
3355
|
+
repository-only authorization, closed stdin, output/timeout termination, and
|
|
3356
|
+
no raw command text in audit/telemetry.
|
|
3357
|
+
2. `fully-contained` is reported only by the certified macOS Seatbelt adapter.
|
|
3358
|
+
Linux and Windows report `containment-unavailable` and spawn nothing until
|
|
3359
|
+
independent native adapters and evidence exist.
|
|
3360
|
+
3. MCP accepts only `profile`; it cannot accept an executable, arguments, cwd,
|
|
3361
|
+
environment, or limits. Every result names the profile, actual containment,
|
|
3362
|
+
configured bounds, exit state, compressed artifact, and diagnosis state.
|
|
3363
|
+
4. Preserve the one-tool schema reduction gate and local-only operation. The
|
|
3364
|
+
checked-in evidence requires zero hosted spend and zero source egress while
|
|
3365
|
+
preserving the original C58 limitations. Do not describe deprecated
|
|
3366
|
+
`sandbox-exec` as a supported Apple public security API.
|
|
3367
|
+
|
|
3368
|
+
Known limitations and decision gate: profiles are single-process
|
|
3369
|
+
(`maxProcesses: 1`), system runtime paths remain readable, approved repository
|
|
3370
|
+
code can modify its checkout, and Linux/Windows execution is unavailable. Keep
|
|
3371
|
+
the surface only while native evidence proves the stated boundary; any broader
|
|
3372
|
+
profile or platform requires a separately certified adapter. Preserve no unrestricted command surface.
|
|
3373
|
+
|
|
3374
|
+
### C89 — MCP request reliability and recovery
|
|
3375
|
+
|
|
3376
|
+
- Status: implemented
|
|
3377
|
+
- Priority: P0
|
|
3378
|
+
- Disposition: Must close
|
|
3379
|
+
- DependsOn: C3, C16, C53
|
|
3380
|
+
- Touches: MCP server lifecycle, request dispatcher, cancellation/progress
|
|
3381
|
+
bridge, durable diagnostic log, error taxonomy, and adversarial transport tests.
|
|
3382
|
+
- What: prevent one expensive or stuck operation from ending the useful MCP
|
|
3383
|
+
session and make every failure attributable and recoverable.
|
|
3384
|
+
- Scope: per-operation deadlines, client cancellation, progress notifications,
|
|
3385
|
+
request and trace IDs, bounded exit-safe logs, stuck-request isolation, and
|
|
3386
|
+
distinct crash, lock, memory-pressure, timeout, and client-disconnect errors.
|
|
3387
|
+
- Evidence: `src/mcp-worker-supervisor.ts`, `src/mcp-graph-worker.ts`,
|
|
3388
|
+
`src/mcp-reliability.ts`, `src/server.ts`, `docs/MCP.md`, and the checked-in
|
|
3389
|
+
supervisor/lifecycle/CLI replays prove versioned IPC, real held-lock handling,
|
|
3390
|
+
abrupt worker exit, the observed `Transport closed` disconnect cancellation
|
|
3391
|
+
race, hard termination, bounded restart, predecessor-to-next-success
|
|
3392
|
+
correlation, warm reuse, source-safe
|
|
3393
|
+
bounded journals, and linked-worktree isolation without hosted dependencies.
|
|
3394
|
+
|
|
3395
|
+
Acceptance:
|
|
3396
|
+
|
|
3397
|
+
1. Every request receives stable request/trace correlation, an operation-aware
|
|
3398
|
+
deadline, and cancellation that releases its resources without corrupting the
|
|
3399
|
+
graph; expensive operations emit bounded truthful progress when supported.
|
|
3400
|
+
2. A killed worker, held graph lock, memory-pressure sentinel, deadline, and
|
|
3401
|
+
disconnected client each produce a distinct actionable error. The server
|
|
3402
|
+
remains usable or restarts automatically without retry loops or lost logs.
|
|
3403
|
+
3. Durable logs survive abrupt process exit, rotate within documented byte/time
|
|
3404
|
+
bounds, omit repository source and secrets, and correlate progress, failure,
|
|
3405
|
+
cleanup, restart, and the next successful request.
|
|
3406
|
+
4. Checked-in CLI/MCP parity, concurrency, cancellation-race, forced-exit, and
|
|
3407
|
+
recovery replays pass locally with zero hosted spend, credentials, network,
|
|
3408
|
+
or source egress.
|
|
3409
|
+
|
|
3410
|
+
Known limitations and decision gate: deadlines and memory-pressure attribution
|
|
3411
|
+
are bounded diagnoses, not proof of an operating-system root cause. Do not claim
|
|
3412
|
+
transparent recovery unless the next independent request succeeds in replay.
|
|
3413
|
+
|
|
3414
|
+
### C90 — Privacy-safe diagnostic support bundle
|
|
3415
|
+
|
|
3416
|
+
- Status: implemented
|
|
3417
|
+
- Priority: P0
|
|
3418
|
+
- Disposition: Must close
|
|
3419
|
+
- DependsOn: C16, C59, C89
|
|
3420
|
+
- Touches: diagnostics CLI, MCP trace store, lifecycle/graph health, redaction,
|
|
3421
|
+
archive manifest, preview/report rendering, and privacy adversarial tests.
|
|
3422
|
+
- What: turn local diagnostics into one support-ready bundle whose exact shared
|
|
3423
|
+
contents users can inspect before export.
|
|
3424
|
+
- Scope: recent MCP traces/failures, lifecycle and graph health, versions and
|
|
3425
|
+
runtime, repository scale without source, manifest, human report, and preview.
|
|
3426
|
+
- Evidence: `src/diagnostics.ts`, `src/diagnostics-write-helper.ts`,
|
|
3427
|
+
`schemas/support-bundle-v2.schema.json`, the C90 adversarial fixture and unit
|
|
3428
|
+
replays, and packed-install preview/archive/inspect smoke prove a fail-closed
|
|
3429
|
+
allowlist, exact content-addressed preview parity, bounded C89 trace recovery,
|
|
3430
|
+
honest unavailable fields, linked-worktree isolation, bounded retention, and
|
|
3431
|
+
repository-bound no-follow writes without hosted dependencies or source
|
|
3432
|
+
egress.
|
|
3433
|
+
|
|
3434
|
+
Acceptance:
|
|
3435
|
+
|
|
3436
|
+
1. Preview and final manifest enumerate every file and field, byte size,
|
|
3437
|
+
retention window, redaction applied, and unavailable section before sharing;
|
|
3438
|
+
the concise report includes request/trace IDs and actionable recovery steps.
|
|
3439
|
+
2. Adversarial fixtures containing source, diffs, paths outside the repository,
|
|
3440
|
+
credentials, environment values, usernames, and command output prove none can
|
|
3441
|
+
enter the default bundle. Repository scale is aggregate metadata only.
|
|
3442
|
+
3. The bundle includes exit-surviving C89 traces and recent classified failures,
|
|
3443
|
+
lifecycle/graph health, knodin/Node/platform versions, and honest missing or
|
|
3444
|
+
stale states without initiating network activity.
|
|
3445
|
+
4. Checked-in deterministic CLI/MCP tests prove preview/archive parity, bounded
|
|
3446
|
+
size and retention, zero hosted spend, no credentials, and no source egress.
|
|
3447
|
+
|
|
3448
|
+
Known limitations and decision gate: automated redaction cannot certify arbitrary
|
|
3449
|
+
future fields. New diagnostic fields fail closed until included in the explicit
|
|
3450
|
+
allowlist and privacy oracle; sending a bundle remains a user action.
|
|
3451
|
+
|
|
3452
|
+
### C91 — Boring macOS installation and runtime handoff
|
|
3453
|
+
|
|
3454
|
+
- Status: parked-external-evidence — 2026-08-05
|
|
3455
|
+
- Previous status: owner-approved adoption priority
|
|
3456
|
+
- Priority: P0
|
|
3457
|
+
- Disposition: Must close
|
|
3458
|
+
- DependsOn: C52, C73, C89, C90
|
|
3459
|
+
- Touches: package/install CI, runtime resolver, doctor repairs, client setup
|
|
3460
|
+
guides, macOS acceptance fixtures, and package compatibility checks.
|
|
3461
|
+
- What: make install through first useful MCP call repeatable on environments the
|
|
3462
|
+
project can actually validate.
|
|
3463
|
+
- Scope: macOS; Node 20 project handoff to an available Node 24 runtime; mise,
|
|
3464
|
+
nvm, fnm, Volta, asdf, Homebrew, npm, and Artifactory; Copilot, Claude, Gemini,
|
|
3465
|
+
and Codex; safe doctor repairs.
|
|
3466
|
+
- Evidence: checked-in automated install matrix, package-content/compatibility
|
|
3467
|
+
gates, isolated runtime-manager fixtures, and timed acceptance receipts.
|
|
3468
|
+
|
|
3469
|
+
Acceptance:
|
|
3470
|
+
|
|
3471
|
+
1. Each named installation path is exercised from a clean macOS fixture through
|
|
3472
|
+
install, `knodin init`, `knodin status`, one useful MCP call, and uninstall or
|
|
3473
|
+
rollback; package compatibility is checked automatically in CI.
|
|
3474
|
+
2. A Node 20 project reliably hands execution to an already available compatible
|
|
3475
|
+
Node 24 runtime without modifying the project's declared runtime. Missing or
|
|
3476
|
+
conflicting runtimes produce one actionable diagnostic.
|
|
3477
|
+
3. Copilot, Claude, Gemini, and Codex each complete the same acceptance flow in
|
|
3478
|
+
no more than two minutes under the documented warm/cold boundary, with exact
|
|
3479
|
+
client/version/configuration evidence and no universal client claim.
|
|
3480
|
+
4. Doctor applies only enumerated idempotent repairs with preview and rollback;
|
|
3481
|
+
checked-in tests require zero hosted spend for local fixtures and no source
|
|
3482
|
+
egress. Credentialed Artifactory evidence is recorded only when available.
|
|
3483
|
+
|
|
3484
|
+
Known limitations and decision gate: Linux and Windows certification are not a
|
|
3485
|
+
goal and unavailable platform evidence must stay unavailable. Runtime-manager
|
|
3486
|
+
fixtures prove supported configurations, not every shell customization.
|
|
3487
|
+
|
|
3488
|
+
Code-readiness on 2026-08-05 added a merge-preserving GitHub Copilot/VS Code
|
|
3489
|
+
MCP adapter, doctor visibility, idempotent CLI-only reversal, and fail-closed
|
|
3490
|
+
malformed/symlink handling. Certification remains parked until native macOS
|
|
3491
|
+
install-to-uninstall receipts exist for nvm, fnm, Volta, and asdf; an exact
|
|
3492
|
+
GitHub Copilot/VS Code timed receipt joins the available Claude, Gemini, and
|
|
3493
|
+
Codex versions; a published Homebrew formula and credentialed Artifactory path
|
|
3494
|
+
are exercised where authorized; and the same four clients complete the bounded
|
|
3495
|
+
install, `init`, `status`, useful MCP call, and rollback flow. Doctor also still
|
|
3496
|
+
needs an enumerated transactional preview/apply/rollback implementation and
|
|
3497
|
+
acceptance replay. No fixture result may be promoted to this missing native
|
|
3498
|
+
evidence, and C94 remains dependency-blocked.
|
|
3499
|
+
|
|
3500
|
+
### C92 — Unified five-channel release orchestration
|
|
3501
|
+
|
|
3502
|
+
- Status: parked-external-evidence — 2026-08-06
|
|
3503
|
+
- Previous status: owner-approved release priority
|
|
3504
|
+
- Priority: P0
|
|
3505
|
+
- Disposition: Must close
|
|
3506
|
+
- DependsOn: C64, C91
|
|
3507
|
+
- Touches: release workflows, preflight, immutable package staging, npm,
|
|
3508
|
+
Artifactory, Homebrew, GitHub SaaS/GHES synchronization, receipts, and retry tests.
|
|
3509
|
+
- What: replace fragmented publication paths with one fail-closed orchestrator
|
|
3510
|
+
that promotes one immutable tarball everywhere.
|
|
3511
|
+
- Scope: npm, Artifactory, Homebrew, GitHub SaaS, and the read-only GHES mirror;
|
|
3512
|
+
availability preflight, private-repository attestations, retries, and receipts.
|
|
3513
|
+
- Evidence: checked-in dry-run channel adapters, fault matrix, immutable digest
|
|
3514
|
+
oracle, retry replay, receipt schema, and private-repository attestation fixture.
|
|
3515
|
+
|
|
3516
|
+
Acceptance:
|
|
3517
|
+
|
|
3518
|
+
1. Before any publication, preflight proves every required channel, credential,
|
|
3519
|
+
permission, target version, and attestation mode is available; one failure
|
|
3520
|
+
publishes nothing and leaves a machine-readable failure receipt.
|
|
3521
|
+
2. One staged tarball digest and length are promoted to all five channels.
|
|
3522
|
+
Retries reconcile idempotently without moving or recreating source tags and
|
|
3523
|
+
refuse any channel whose existing bytes differ.
|
|
3524
|
+
3. Private-repository attestations follow the supported trust path rather than
|
|
3525
|
+
assuming public-repository identity behavior. The final receipt records each
|
|
3526
|
+
channel's artifact identity, digest, status, attempt, and authoritative URL.
|
|
3527
|
+
4. Checked-in fault-injection tests cover every preflight/publish boundary with
|
|
3528
|
+
zero real publication or spend. Live channel receipts remain Track B evidence
|
|
3529
|
+
and cannot be simulated by this local implementation item.
|
|
3530
|
+
|
|
3531
|
+
Known limitations and decision gate: dry-run adapters do not certify production
|
|
3532
|
+
availability. GitHub SaaS remains authoritative; GHES is a read-only mirror and
|
|
3533
|
+
must never be forced. Stop on divergence.
|
|
3534
|
+
|
|
3535
|
+
Parking decision: C92 remains open behind C64 and C91. Ordinary 0.x publishing
|
|
3536
|
+
continues through the existing fail-closed tag workflow; this does not claim the
|
|
3537
|
+
unimplemented unified five-channel orchestrator or manufacture channel receipts.
|
|
3538
|
+
|
|
3539
|
+
### C93 — Paired engineering-outcome measurement
|
|
3540
|
+
|
|
3541
|
+
- Status: implemented
|
|
3542
|
+
- Priority: P1
|
|
3543
|
+
- Disposition: Must close
|
|
3544
|
+
- DependsOn: C24, C81
|
|
3545
|
+
- Touches: competitive runner, task corpus, patch evaluator, dependency oracle,
|
|
3546
|
+
agent telemetry import, cost accounting, and outcome report.
|
|
3547
|
+
- What: measure whether knodin improves completed engineering work, not whether
|
|
3548
|
+
it merely emits fewer tokens.
|
|
3549
|
+
- Scope: identical with/without-knodin tasks measuring task success,
|
|
3550
|
+
missed-dependency rate, corrective round trips, time to first correct edit,
|
|
3551
|
+
patch-application success, total billed cost, and performance.
|
|
3552
|
+
- Evidence: preregistered checked-in corpus/oracles, paired raw runs, variance and
|
|
3553
|
+
failure accounting, verifier, and a limitations-bound report.
|
|
3554
|
+
- Evidence delivered: `benchmarks/evaluations/c93-engineering-outcomes/`
|
|
3555
|
+
contains the preregistered two-arm corpus, executable dependency/patch oracle,
|
|
3556
|
+
source-free raw event trajectories, deterministic local runner, paired
|
|
3557
|
+
variance/uncertainty report, and visible limitations. `npm run verify:c93`
|
|
3558
|
+
recreates all 12 runs from the identical fixture commit on isolated linked
|
|
3559
|
+
worktrees and non-default branches, verifies independent arm commits and a
|
|
3560
|
+
clean base checkout, and records zero hosted spend, observed network requests,
|
|
3561
|
+
or source egress plus explicit process-isolation bounds and unavailable—not
|
|
3562
|
+
inferred—token and billed-cost data.
|
|
3563
|
+
|
|
3564
|
+
Acceptance:
|
|
3565
|
+
|
|
3566
|
+
1. Both arms use identical repositories, commits, tasks, model/client versions,
|
|
3567
|
+
prompts, permissions, time limits, repetitions, and correctness oracles; the
|
|
3568
|
+
only intended treatment difference is knodin availability/instructions.
|
|
3569
|
+
2. Every named outcome is reported per task and in aggregate, including failures,
|
|
3570
|
+
censored time, unavailable token/cost data, and confidence/variance. Token
|
|
3571
|
+
counts may be diagnostic but are never the success metric.
|
|
3572
|
+
3. Dependency and patch oracles inspect the resulting edit and tests rather than
|
|
3573
|
+
grading prose. Raw events support auditable round-trip and first-correct-edit
|
|
3574
|
+
reconstruction without storing repository source in telemetry.
|
|
3575
|
+
4. The local harness and synthetic fixtures are checked in and run with zero hosted
|
|
3576
|
+
spend; any paid model replay records actual total billed cost and is
|
|
3577
|
+
never inferred from tokens alone.
|
|
3578
|
+
Checked-in evidence requires zero hosted spend.
|
|
3579
|
+
|
|
3580
|
+
Known limitations and decision gate: a bounded corpus cannot prove universal
|
|
3581
|
+
product superiority. Publish only measured effect sizes and uncertainty; a null
|
|
3582
|
+
or negative outcome must remain visible and may require product rollback.
|
|
3583
|
+
|
|
3584
|
+
### C94 — First-use clarity and five-minute demonstration
|
|
3585
|
+
|
|
3586
|
+
- Status: parked-external-evidence — 2026-08-06
|
|
3587
|
+
- Previous status: owner-approved clarity priority
|
|
3588
|
+
- Blocked by: C91 external certification
|
|
3589
|
+
- Priority: P1
|
|
3590
|
+
- Disposition: Must close
|
|
3591
|
+
- DependsOn: C89, C90, C91
|
|
3592
|
+
- Touches: MCP/CLI help, status responses, client guides, examples, demo fixture,
|
|
3593
|
+
freshness/error copy, and documentation integrity tests.
|
|
3594
|
+
- What: make the first session explain the tool identity, next action, evidence
|
|
3595
|
+
limits, freshness, and recovery without requiring prior product knowledge.
|
|
3596
|
+
- Scope: `knodin.knodin` naming, operation discovery/examples, next-useful-action
|
|
3597
|
+
recommendations, concise limitations/recovery, and a real-repository demo.
|
|
3598
|
+
- Evidence: checked-in golden help/status/error output, four-client guide checks,
|
|
3599
|
+
timed demo transcript, and source-evidence links.
|
|
3600
|
+
|
|
3601
|
+
Acceptance:
|
|
3602
|
+
|
|
3603
|
+
1. Setup docs explain that clients display `knodin.knodin` as server name plus
|
|
3604
|
+
tool name. Schema/help gives copyable examples and makes operation discovery
|
|
3605
|
+
possible without adding MCP tools.
|
|
3606
|
+
2. Status recommends a context-sensitive next useful operation and presents
|
|
3607
|
+
freshness, availability, known bounds, and repair/reindex actions concisely;
|
|
3608
|
+
no healthy, stale, ambiguous, or repair-needed state is conflated.
|
|
3609
|
+
3. A checked-in five-minute demonstration runs `knodin init`, status, orientation,
|
|
3610
|
+
source-evidenced impact/review, and bounded context on a real repository and
|
|
3611
|
+
explicitly demonstrates value, ROI, tomorrow's action, and secret sauce.
|
|
3612
|
+
4. Golden CLI/MCP and docs-integrity tests run locally with zero hosted spend,
|
|
3613
|
+
no credentials, and no source egress.
|
|
3614
|
+
|
|
3615
|
+
Known limitations and decision gate: a scripted demonstration is onboarding
|
|
3616
|
+
evidence, not an outcome or superiority benchmark. Client UI labels may vary by
|
|
3617
|
+
version and must name the version actually checked.
|
|
3618
|
+
|
|
3619
|
+
Parking decision: C94 remains open behind C91. The existing local help and
|
|
3620
|
+
documentation remain available, but no fixture is promoted into the missing
|
|
3621
|
+
native client demonstration or certification evidence.
|
|
3622
|
+
|
|
3623
|
+
### C95 — Defensible behavioral-contract replay
|
|
3624
|
+
|
|
3625
|
+
- Status: implemented
|
|
3626
|
+
- Priority: P1
|
|
3627
|
+
- Disposition: Must close
|
|
3628
|
+
- DependsOn: C3, C24, C83
|
|
3629
|
+
- Touches: contract manifest, cross-operation fixtures, freshness/identity/
|
|
3630
|
+
evidence/budget/compression/diagnosis/privacy tests, and product positioning.
|
|
3631
|
+
- What: make the integrated behavior—not individual graph feature count—the
|
|
3632
|
+
product's checked and release-gated contract.
|
|
3633
|
+
- Scope: truthful budgets, explicit freshness/availability, source evidence,
|
|
3634
|
+
ambiguity-safe stable identity, recoverable compression, failure-to-symbol
|
|
3635
|
+
diagnosis, and local operation without source egress.
|
|
3636
|
+
- Evidence: `contracts/behavior-contract-v1.json`,
|
|
3637
|
+
`benchmarks/evaluations/c95-behavior-contract/runner.ts`,
|
|
3638
|
+
`scripts/verify-c95.ts`, `docs/BEHAVIORAL-CONTRACT.md`, and the release
|
|
3639
|
+
preflight gate map and dispatch all 22 production operations plus every
|
|
3640
|
+
schema-declared query/action discriminator against linked-worktree fixtures.
|
|
3641
|
+
The replays cover hard-budget continuation, stable evidence,
|
|
3642
|
+
recoverable compression, freshness and ambiguity qualification, diagnosis,
|
|
3643
|
+
and observed local privacy across fetch, HTTP(S), TCP/TLS, DNS, native-addon,
|
|
3644
|
+
and subprocess boundaries. All 458 mapped fixture/clause rows have isolated
|
|
3645
|
+
pre-serialization production-response mutations that fail the same release
|
|
3646
|
+
oracle; omitted or mutated clauses fail closed.
|
|
3647
|
+
|
|
3648
|
+
Acceptance:
|
|
3649
|
+
|
|
3650
|
+
1. Every production operation is mapped to the applicable contract clauses and
|
|
3651
|
+
adversarial fixtures prove honest truncation, stale/unavailable states,
|
|
3652
|
+
evidence provenance, ambiguity, recovery handles, diagnosis confidence, and
|
|
3653
|
+
privacy behavior.
|
|
3654
|
+
2. The replay verifies composition: a bounded response can be expanded by a
|
|
3655
|
+
stable handle without changing identity/evidence, and stale or ambiguous
|
|
3656
|
+
evidence cannot become a confident diagnosis or impact claim.
|
|
3657
|
+
3. Release preflight fails on a contract regression. Product documentation
|
|
3658
|
+
answers value, ROI, tomorrow, and secret sauce with checked-in evidence and
|
|
3659
|
+
preserves every roadmap limitation.
|
|
3660
|
+
4. Tests and replays run locally with zero hosted spend, credentials, network,
|
|
3661
|
+
or source egress; no universal superiority claim is inferred.
|
|
3662
|
+
|
|
3663
|
+
Known limitations and decision gate: the contract proves the checked operation
|
|
3664
|
+
and action fixtures, not all repositories or clients. Network construction is
|
|
3665
|
+
denied before request-body streaming, and opaque transport inside an
|
|
3666
|
+
already-loaded native library remains unobservable. A new operation or action
|
|
3667
|
+
must declare its clauses, positive response semantics, and isolated production
|
|
3668
|
+
mutation before release.
|
|
3669
|
+
|
|
3670
|
+
### C96 — Replay-gated capability discipline
|
|
3671
|
+
|
|
3672
|
+
- Status: evaluated — deferred
|
|
3673
|
+
- Priority: P2
|
|
3674
|
+
- Disposition: Evaluate
|
|
3675
|
+
- DependsOn: C24, C93, C95
|
|
3676
|
+
- Touches: roadmap policy, proposal template, outcome replay, one narrow
|
|
3677
|
+
cross-substrate fixture, verifier, and decision record.
|
|
3678
|
+
- What: require evidence of correctness or workflow improvement before expanding
|
|
3679
|
+
production capability, beginning with one bounded cross-substrate impact case.
|
|
3680
|
+
- Scope: one relationship crossing two explicitly named substrates; no broad
|
|
3681
|
+
platform, language, repository, or superiority promise.
|
|
3682
|
+
- Evidence: checked-in baseline/treatment fixture, exact dependency and task
|
|
3683
|
+
oracles, cost/resource report, limitations, and retain/defer/implement decision.
|
|
3684
|
+
|
|
3685
|
+
Acceptance:
|
|
3686
|
+
|
|
3687
|
+
1. Preregister one concrete missed-dependency or workflow failure and the exact
|
|
3688
|
+
improvement threshold before implementation. Compare unchanged knodin with a
|
|
3689
|
+
bounded prototype on the same task using C93 outcomes and C95 contracts.
|
|
3690
|
+
2. The fixture names supported substrates, edge provenance, ambiguity behavior,
|
|
3691
|
+
freshness, budgets, and false-positive/negative oracles; unsupported crossings
|
|
3692
|
+
remain explicit rather than generalized.
|
|
3693
|
+
3. Close by retaining knodin unchanged, deferring/rejecting the capability, or
|
|
3694
|
+
creating a separately numbered implementation item only when the gate passes.
|
|
3695
|
+
4. Evaluation artifacts and verifier are checked in, require zero hosted spend,
|
|
3696
|
+
credentials, network, or source egress, and preserve losing results.
|
|
3697
|
+
|
|
3698
|
+
Known limitations and decision gate: one passing cross-substrate fixture does
|
|
3699
|
+
not authorize a broad platform claim. Implementation requires a separately
|
|
3700
|
+
reviewed item and must preserve the behavioral contract.
|
|
3701
|
+
|
|
3702
|
+
Decision: defer production expansion. C93 found no correctness advantage on its
|
|
3703
|
+
narrow paired fixture, and C95 already requires every new operation or action to
|
|
3704
|
+
declare and replay its behavioral clauses. No preregistered cross-substrate
|
|
3705
|
+
missed-dependency case currently justifies a prototype, so knodin remains
|
|
3706
|
+
unchanged. Reopen under a separately numbered item only when a concrete failure,
|
|
3707
|
+
supported substrates, exact oracle, and improvement threshold are preregistered.
|
|
3708
|
+
Evidence: `benchmarks/evaluations/c93-engineering-outcomes/verification.json`,
|
|
3709
|
+
`contracts/behavior-contract-v1.json`, and
|
|
3710
|
+
`benchmarks/evaluations/c95-behavior-contract/runner.ts`.
|
|
3711
|
+
|
|
3712
|
+
## Explicit non-goals from the GitNexus comparison
|
|
3713
|
+
|
|
3714
|
+
- Do not split knodin's gateway into one MCP schema per capability; the one-tool
|
|
3715
|
+
surface is a measured differentiator.
|
|
3716
|
+
- Do not copy unrestricted Cypher directly into the MCP surface.
|
|
3717
|
+
- Do not require Milvus, Qdrant, Ollama, credentials, or any hosted service;
|
|
3718
|
+
Claude Context and mcp-codebase-index validate this differentiator.
|
|
3719
|
+
- Do not add one MCP tool per capability; Serena's 29-tool and
|
|
3720
|
+
code-review-graph's 30-tool surfaces reinforce the one-gateway advantage.
|
|
3721
|
+
- Do not make vector visualization, agent memory, or skill generation core graph
|
|
3722
|
+
requirements. C35 is limited to deterministic local architecture/call-flow
|
|
3723
|
+
artifacts with a concrete developer workflow; it does not authorize a hosted
|
|
3724
|
+
visualization service or a generalized vector-visualization feature.
|
|
3725
|
+
- Do not enable network embeddings, hosted indexing, telemetry, or source-code
|
|
3726
|
+
egress to chase competitor features.
|
|
3727
|
+
- Do not claim API-shape or cross-repository superiority until dedicated local
|
|
3728
|
+
fixtures produce evidence.
|
|
3729
|
+
- Do not interpret every competitor option as a product requirement; parameters
|
|
3730
|
+
that merely select nonexistent groups, branches, or services need real
|
|
3731
|
+
fixtures before roadmap promotion.
|
|
3732
|
+
|
|
3733
|
+
## Verification bar
|
|
3734
|
+
|
|
3735
|
+
Every implementation item must run the repository's discovered gates and add a
|
|
3736
|
+
targeted acceptance fixture. At minimum:
|
|
3737
|
+
|
|
3738
|
+
```bash
|
|
3739
|
+
bun run typecheck
|
|
3740
|
+
bun run lint
|
|
3741
|
+
bun run test
|
|
3742
|
+
```
|
|
3743
|
+
|
|
3744
|
+
Competitive claims must also re-run the relevant harness under
|
|
3745
|
+
`benchmarks/competitors/` and update its result file without deleting losing
|
|
3746
|
+
cases.
|
|
3747
|
+
|
|
3748
|
+
## Competitive replay — 2026-07-24 (Phase 2)
|
|
3749
|
+
|
|
3750
|
+
Fresh, provenance-bound replays of every dimension were run against the real
|
|
3751
|
+
competitor CLIs and saved as **new** immutable artifacts (never overwriting
|
|
3752
|
+
prior results): search `raw-results-20260724T182246Z.json`, review + architecture
|
|
3753
|
+
`…T182448Z.json`, and symbols/editing/visualization/context/lifecycle
|
|
3754
|
+
`…T182707Z.json` under `benchmarks/evaluations/competitive-*/`.
|
|
3755
|
+
|
|
3756
|
+
Applying the measurement contract (a WIN needs equivalent operation/fixture/mode/
|
|
3757
|
+
scope/budget/oracle **and** an oracle-qualified row; unavailable/unverified/
|
|
3758
|
+
incomparable rows are excluded from win claims):
|
|
3759
|
+
|
|
3760
|
+
### Where knodin is ahead or level (oracle-qualified, comparable)
|
|
3761
|
+
|
|
3762
|
+
- **Diff review — WIN.** traversal recall 1.0 and 733 B vs gitnexus (recall 1.0,
|
|
3763
|
+
1458 B), code-review-graph (recall 0, 2269 B), graphify (recall 0.667).
|
|
3764
|
+
- **Architecture — WIN.** full package coverage + hubs + bridges; every completed
|
|
3765
|
+
competitor is at best equal on coverage, none better.
|
|
3766
|
+
- **Search — WIN vs the only comparable competitor.** nDCG 0.855 vs gitnexus
|
|
3767
|
+
0.677; dead-code precision 1.0 vs 0.667. codegraph (probe-only, unscored),
|
|
3768
|
+
grepai (MCP probe blocked), claude-context (needs Ollama+Milvus) are
|
|
3769
|
+
contract-excluded.
|
|
3770
|
+
- **Symbols — WIN/level.** precision 1 / recall 1: ties gitnexus (1/1), beats
|
|
3771
|
+
codegraph (0/0), serena (0/0), codebase-memory (1/0.5).
|
|
3772
|
+
- **Lifecycle capability — WIN.** indexed/queryable/closed/stale-detection/
|
|
3773
|
+
per-query telemetry all qualified; codebase-memory and grepai advertise
|
|
3774
|
+
neither stale detection nor per-query telemetry.
|
|
3775
|
+
|
|
3776
|
+
### Where the comparison is not oracle-qualified (excluded, not losses)
|
|
3777
|
+
|
|
3778
|
+
- **Editing / Visualization / Context.** knodin passes its golden/export oracles;
|
|
3779
|
+
the competitors either did not run (serena, code-review-graph, aider
|
|
3780
|
+
unavailable/blocked) or ran without a shared correctness oracle (graphify,
|
|
3781
|
+
repomix, code2prompt "complete" but unscored). No comparable head-to-head.
|
|
3782
|
+
- **Project memory.** knodin has no memory feature (deferred, ADR 004); Serena's
|
|
3783
|
+
own value-beyond-docs gate also failed. No contest.
|
|
3784
|
+
|
|
3785
|
+
### Where knodin still costs more — and why
|
|
3786
|
+
|
|
3787
|
+
- **Lifecycle peak RSS: 83.8 MB vs codebase-memory 31.3 MB, grepai 22.9 MB.**
|
|
3788
|
+
This is the one axis where knodin's raw number is worse. It is **not** an
|
|
3789
|
+
oracle-qualified head-to-head: knodin is measured as a dedicated, long-lived
|
|
3790
|
+
product process holding the real semantic model + graph resident (which is
|
|
3791
|
+
precisely what wins the search-quality contest above), while the competitors
|
|
3792
|
+
are measured at spawn-per-command peak. The RSS cost and the search-quality
|
|
3793
|
+
lead are two sides of the same resident-model design; knodin cannot undercut
|
|
3794
|
+
them on RSS without abandoning the model that beats them on search. C51
|
|
3795
|
+
(native training-free quantization) targets the *stored-embedding* portion of
|
|
3796
|
+
that footprint; the model-runtime portion is inherent to local semantic search.
|
|
3797
|
+
|
|
3798
|
+
Bottom line: on every dimension with an oracle-qualified head-to-head against an
|
|
3799
|
+
available competitor, knodin is **as good or better** (wins or ties). The only
|
|
3800
|
+
place a competitor's raw number beats knodin is lifecycle RSS, which is an
|
|
3801
|
+
incomparable isolation model and the direct cost of knodin's search-quality lead.
|