@jmtrin/opencode-kevin 0.6.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/README.md +582 -580
  2. package/dist/migrations/001_initial.sql +91 -91
  3. package/dist/migrations/003_v02_signal.sql +57 -57
  4. package/dist/migrations/004_v03_knowledge.sql +138 -138
  5. package/dist/migrations/005_v04_signal.sql +57 -57
  6. package/dist/migrations/006_v05_glassbox.sql +118 -118
  7. package/dist/migrations/007_v06_pull.sql +144 -144
  8. package/dist/migrations/008_v07_truth.sql +124 -0
  9. package/dist/plugin/ConflictDetector.d.ts +35 -0
  10. package/dist/plugin/ConflictDetector.js +283 -0
  11. package/dist/plugin/ConflictDetector.js.map +1 -0
  12. package/dist/plugin/ConventionMiner.d.ts +35 -0
  13. package/dist/plugin/ConventionMiner.js +243 -0
  14. package/dist/plugin/ConventionMiner.js.map +1 -0
  15. package/dist/plugin/Curator.js +45 -14
  16. package/dist/plugin/Curator.js.map +1 -1
  17. package/dist/plugin/MemoryService.d.ts +18 -0
  18. package/dist/plugin/MemoryService.js +92 -2
  19. package/dist/plugin/MemoryService.js.map +1 -1
  20. package/dist/plugin/Migrate.js +31 -2
  21. package/dist/plugin/Migrate.js.map +1 -1
  22. package/dist/plugin/Reflector.js +13 -0
  23. package/dist/plugin/Reflector.js.map +1 -1
  24. package/dist/plugin/RepoTruth.d.ts +80 -0
  25. package/dist/plugin/RepoTruth.js +600 -0
  26. package/dist/plugin/RepoTruth.js.map +1 -0
  27. package/dist/plugin/Retrospective.js +8 -0
  28. package/dist/plugin/Retrospective.js.map +1 -1
  29. package/dist/plugin/index.d.ts +7 -2
  30. package/dist/plugin/index.js +135 -2
  31. package/dist/plugin/index.js.map +1 -1
  32. package/dist/plugin/kevin_audit.d.ts +41 -1
  33. package/dist/plugin/kevin_audit.js +98 -3
  34. package/dist/plugin/kevin_audit.js.map +1 -1
  35. package/dist/plugin/kevin_conflicts.d.ts +9 -0
  36. package/dist/plugin/kevin_conflicts.js +51 -0
  37. package/dist/plugin/kevin_conflicts.js.map +1 -0
  38. package/dist/plugin/kevin_facts.d.ts +42 -0
  39. package/dist/plugin/kevin_facts.js +37 -0
  40. package/dist/plugin/kevin_facts.js.map +1 -0
  41. package/dist/plugin/metrics.d.ts +3 -3
  42. package/dist/plugin/metrics.js +12 -2
  43. package/dist/plugin/metrics.js.map +1 -1
  44. package/dist/plugin/replay-types.d.ts +2 -2
  45. package/migrations/001_initial.sql +91 -91
  46. package/migrations/003_v02_signal.sql +57 -57
  47. package/migrations/004_v03_knowledge.sql +138 -138
  48. package/migrations/005_v04_signal.sql +57 -57
  49. package/migrations/006_v05_glassbox.sql +118 -118
  50. package/migrations/007_v06_pull.sql +144 -144
  51. package/migrations/008_v07_truth.sql +124 -0
  52. package/package.json +3 -2
package/README.md CHANGED
@@ -1,580 +1,582 @@
1
- # Kevin
2
-
3
- > Observe and learn: the learning layer OpenCode was missing.
4
-
5
- Kevin is an [OpenCode](https://opencode.ai) plugin that **observes** every agent tool call, **learns** from failures by generating lessons, and **shares** what it learned proactively in future sessions. It does not plan, orchestrate, or compete with the plugin ecosystem. It only learns.
6
-
7
- - **Local-first**: SQLite + FTS5, no external services, no network calls.
8
- - **Global memory**: a single `~/.opencode-kevin/kevin.db` shared across all your projects (WAL mode → safe for concurrent sessions). No per-project folders.
9
- - **Knowledge + Causality**: causal failure→fix chains, `kevin_why` explanations, OKF export/import, a supersede model, and human-in-the-loop AGENTS.md suggestions.
10
- - **Signal over Noise**: a quality gate that stores weak lessons without injecting them, an injection ledger with honest `precision_rate`, and two-sided confidence.
11
- - **Glass Box**: honest measurement replaces estimates — three-way injection settlement (`effective` / `ineffective` / `inconclusive`), human feedback that actually moves confidence, a strict dry-run `kevin_trace`, a read-only `kevin_audit`, and a hermetic replay harness.
12
- - **Pull**: knowledge earns its way into files the model actually reads — `kevin_propose` generates a reviewable diff, a human approves, and **only then** does Kevin write, inside a frozen marker block, preserving your file's CRLF/BOM/formatting byte-for-byte outside it. Plus three distribution channels (AGENTS.md, skills, references) and a push budget gated by a confidence floor.
13
- - **Audited**: the v0.4.0 bug catalog (`docs/Kevin_v0.4.0_Bugs.md`) is fully closed — 16/16 bugs fixed and regression-tested.
14
- - **Standalone**: works without any other plugin. With the ecosystem, it learns more richly.
15
-
16
- ---
17
-
18
- ## Contents
19
-
20
- - [Installation](#installation)
21
- - [How Kevin works](#how-kevin-works)
22
- - [Tools](#tools)
23
- - [How Kevin measures itself](#how-kevin-measures-itself)
24
- - [Curation & Pull](#curation--pull)
25
- - [Replay harness](#replay-harness)
26
- - [Hooks](#hooks)
27
- - [Configuration](#configuration)
28
- - [Development](#development)
29
- - [License](#license)
30
-
31
- ---
32
-
33
- ## Installation
34
-
35
- ### 1. Declare the plugin
36
-
37
- Add Kevin to your OpenCode config. For **all projects** (global):
38
-
39
- ```jsonc
40
- // ~/.config/opencode/opencode.jsonc
41
- {
42
- "$schema": "https://opencode.ai/config.json",
43
- "plugin": [
44
- "@jmtrin/opencode-kevin@latest"
45
- ]
46
- }
47
- ```
48
-
49
- For a **single project**, put the same `plugin` array in `./opencode.json` or `.opencode/opencode.json` at the project root.
50
-
51
- ### 2. Restart OpenCode
52
-
53
- Config is loaded once at startup and is **not hot-reloaded** — quit and reopen OpenCode after editing. On start, OpenCode resolves the npm spec, caches the plugin in `~/.cache/opencode/packages/@jmtrin/opencode-kevin/`, and exposes sixteen tools: `kevin_save`, `kevin_query`, `kevin_get`, `kevin_recall`, `kevin_status`, `kevin_retrospective`, `kevin_why`, `kevin_export`, `kevin_import`, `kevin_config`, `kevin_feedback`, `kevin_trace`, `kevin_audit`, `kevin_propose`, `kevin_approve`, `kevin_publish`.
54
-
55
- ### 3. Where data lives
56
-
57
- Kevin stores everything in a single **global, shared** location under your home directory — no per-project `.kevin/` folders:
58
-
59
- | Path | Content |
60
- |---|---|
61
- | `~/.opencode-kevin/kevin.db` | SQLite database (memories, tool calls, retrospectives). WAL mode → safe for concurrent OpenCode sessions across projects. |
62
- | `~/.opencode-kevin/retrospectives/<session>.md` | Per-session retrospective markdown. |
63
-
64
- Migrations run automatically on startup.
65
-
66
- ### Requirements
67
-
68
- - **Node.js >= 22.5** (uses `node:sqlite`, the built-in SQLite module — no native binaries to compile).
69
- - OpenCode with plugin support (`@opencode-ai/plugin` >= 1.17).
70
-
71
- > **Runtimes**:
72
- > - **Bun**: uses `bun:sqlite` (built-in).
73
- > - **Node 24+**: uses `node:sqlite` directly, no flags needed (emits an experimental warning, harmless).
74
- > - **Node 22/23 without `--experimental-sqlite` flag** or **Node 20**: falls back to `better-sqlite3`, declared as `optionalDependencies`. If you need it, install it manually in your opencode config directory (`~/.config/opencode/`): `npm install better-sqlite3`.
75
-
76
- ### Verification
77
-
78
- ```bash
79
- npm run verify
80
- ```
81
-
82
- Checks Node version, SQLite, migration, MemoryService save/query, Reflector, ContextInjector, and TypeScript strict mode.
83
-
84
- ### Advanced (optional)
85
-
86
- Override defaults via the plugin tuple form:
87
-
88
- ```jsonc
89
- {
90
- "plugin": [
91
- ["@jmtrin/opencode-kevin", {
92
- "dbPath": "/custom/path/kevin.db",
93
- "retrospectivesDir": "/custom/path/retrospectives",
94
- "throttleMs": 120000
95
- }]
96
- ]
97
- }
98
- ```
99
-
100
- Use `:memory:` for `dbPath` in tests.
101
-
102
- ---
103
-
104
- ## How Kevin works
105
-
106
- Every tool call is observed; every failure becomes a lesson; every lesson is either pushed into the next prompt, written into an artifact a human approved, or retired when it stops earning its place.
107
-
108
- ```
109
- Tool call (success or failure)
110
-
111
-
112
- ┌─────────────────────────┐ OBSERVE
113
- │ ToolCallObserver │ records every call (tool, redacted args,
114
- └───────────┬─────────────┘ success, duration, error type, dedup)
115
-
116
- failure │ success
117
- ┌──────────▼───────────┐ ┌─────────────────────────┐
118
- Reflector │ │ CausalChain │
119
- │ heuristic lesson │ │ links the fix to the │
120
- per error code │ │ failure within 10 calls
121
- (throttled per └────────────┬────────────┘
122
- fingerprint)
123
- └──────────┬───────────┘ session.idle
124
-
125
- ▼ ┌─────────────────────────┐
126
- ┌───────────────────────┐ promotes recurring │
127
- │ ContextInjector │◄─┤ errors → causal │
128
- SHARE: injects │ patterns (cumulative
129
- <kevin-context> │ evidence) │
130
- ≤400 tokens/prompt └─────────────────────────┘
131
- └───────────┬───────────┘
132
- session.idle
133
-
134
- ┌─────────────────────────┐ RETROSPECTIVE: <session>.md with
135
- │ Retrospective │ lessons, metrics snapshot, causal
136
- └─────────────────────────┘ promotion, pattern mining (opt-in)
137
- ```
138
-
139
- At `session.idle` Kevin also settles injection outcomes, retires stale memories, and — when curation is enabled — drafts pull proposals for your review (see [Curation & Pull](#curation--pull)).
140
-
141
- ---
142
-
143
- ## Tools
144
-
145
- Kevin exposes 16 tools callable by the agent.
146
-
147
- ### `kevin_save`
148
-
149
- Saves an explicit memory.
150
-
151
- ```
152
- kevin_save({ type: "decision", content: "We use vitest for tests", scope: "project" })
153
- // → { "id": "0195a3b2-..." }
154
- ```
155
-
156
- `type`: `error` | `pattern` | `decision` | `context` | `rule` | `solution`. `scope`: `project` (persists) | `session` (TTL 24h).
157
-
158
- Saving a `decision` or `rule` with the same `fingerprint` as an existing active row supersedes the old one (`status='superseded'`, hidden from default queries).
159
-
160
- ### `kevin_query`
161
-
162
- Searches memories by text (FTS5 + bm25). Returns a **slim** payload by default; pass `full: true` for the complete content, or `evidence: true` to include `confidence`, `evidence_count` and `last_verified_at`.
163
-
164
- ```
165
- kevin_query({ query: "typecheck", type: "error", limit: 5 })
166
- // → [{ "id": "...", "type": "error", "scope": "project", "score": -0.87,
167
- // "snippet": "When bash fails with typecheck:..." }, ...]
168
- ```
169
-
170
- ### `kevin_get`
171
-
172
- Fetches a **single full memory** by id (progressive disclosure) — use it when `kevin_query` returned a slim snippet and you need the complete content.
173
-
174
- ```
175
- kevin_get({ id: "0195a3b2-..." })
176
- // → { "id": "...", "type": "error", "content": "...", "scope": "project",
177
- // "relevanceScore": 0.55, "origin": "reflector", "fingerprint": "cbf29ce484222325",
178
- // "projectId": null, "metadata": null,
179
- // "evidenceCount": 2, "recurrenceCount": 1, "lastVerifiedAt": "2026-08-01 10:00:00",
180
- // "status": "active", "confidence": 0.55, "fixArgs": "npm i -g rg" }
181
- ```
182
-
183
- ### `kevin_recall`
184
-
185
- Retrieves relevant memories (greedy fill by relevance). Without `query`, returns all memories in scope. Pass `includeSuperseded: true` to include superseded rows.
186
-
187
- ```
188
- kevin_recall({ query: "auth", limit: 3 })
189
- // → [{ "id": "...", "type": "decision", ... }, ...]
190
- ```
191
-
192
- ### `kevin_status`
193
-
194
- Global counts and metrics: memory census, the precision block, the six blocked-gate counters, feedback totals, and the v0.6 block (`schema_version`, `curation_enabled`, emission states, `proposals_pending` — omitted on pre-007 databases).
195
-
196
- ```
197
- kevin_status({})
198
- // → { "memories": 42, "memories_reflector": 12, "memories_agent": 30,
199
- // "memories_pattern": 0, "memories_causal": 1, "tool_calls": 318,
200
- // "retrospectives": 7, "tool_count": 16,
201
- // "metrics": { "tokens_injected_pre_prompt": 51, "tokens_injected_compacting": 0,
202
- // "reflections_throttled": 3, "duplicate_suppressions": 2,
203
- // "tool_calls_deduped": 0, "patterns_mined": 0,
204
- // "patterns_causal": 1, "causal_links": 2, "memories_superseded": 0,
205
- // "injections_inconclusive": 9, ... },
206
- // "injections_total": 14, "injections_effective": 2, "injections_ineffective": 3,
207
- // "injections_inconclusive": 9, "precision_rate": 0.40, "coverage_rate": 0.36,
208
- // "blocked": { "seen": 1, "weak": 0, "recurrence": 2, "stale": 0,
209
- // "ignored": 1, "confidence": 2 },
210
- // "memories_ignored": 1, "memories_archived": 4,
211
- // "feedback": { "positive": 2, "negative": 1 },
212
- // "patterns_promoted_new": 2, "recurrence_by_origin": { "reflector": 3, "causal": 1 },
213
- // "v06": { "schema_version": "007", "curation_enabled": "1",
214
- // "skill_emission": "off", "reference_emission": "off",
215
- // "proposals_pending": 2 } }
216
- ```
217
-
218
- ### `kevin_retrospective`
219
-
220
- Generates a retrospective for a session (uses the current session if `session_id` is omitted).
221
-
222
- ```
223
- kevin_retrospective({ session_id: "sess-abc" })
224
- // → { "file_path": "~/.opencode-kevin/retrospectives/sess-abc.md" }
225
- // or → { "message": "No failures in session sess-abc." }
226
- ```
227
-
228
- ### `kevin_why`
229
-
230
- Explains *why* a failure keeps happening: looks up causal patterns for the query and builds a failure → fix trace from memories + tool_calls, including related TypeScript error-code rules.
231
-
232
- ```
233
- kevin_why({ query: "TS2304 cannot find name" })
234
- // → { "summary": "TS2304 recurs because ... Confirmed by 2 fixes.",
235
- // "confidence": 0.7, "evidence_count": 2, "last_verified": "2026-08-01 10:00:00",
236
- // "trace": [ { "type": "error", "summary": "..." }, { "type": "fix", "tool": "bash" } ],
237
- // "related_rules": [ { "code": "TS2304", "suggestion": "import or typo" } ] }
238
- ```
239
-
240
- ### `kevin_export`
241
-
242
- Exports knowledge for sharing: `decision`/`rule`/`pattern` memories (active only, no raw errors) as YAML-frontmatter blocks (`format: "okf"`) or markdown (`format: "markdown"`). Includes `id`, `type`, `confidence`, `evidence_count`, `recurrence_count`, `last_verified_at`, `fingerprint`. Timestamps are treated as UTC — a re-import reproduces the exact source values.
243
-
244
- ### `kevin_import`
245
-
246
- Ingests an exported bundle. Each entry becomes a `context` memory with `origin='imported'`; a fingerprint collision with an existing `decision`/`rule` supersedes the old row. Returns `{ imported, superseded }`.
247
-
248
- ### `kevin_config`
249
-
250
- Reads/writes `kevin_settings` without SQL. `action: "list"` returns every setting; `action: "set"` upserts a value (default `"1"` when omitted) and rejects unknown keys unless `strict: false`.
251
-
252
- ```
253
- kevin_config({ action: "list" })
254
- // → { "quality_gate_enabled": "1", "lesson_snippet_injection": "1",
255
- // "llm_reflection_enabled": "0", "pre_prompt_budget_tokens": "400",
256
- // "injection_confidence_floor": "0.6", ... }
257
-
258
- kevin_config({ action: "set", key: "quality_gate_enabled", value: "0" })
259
- // → { "ok": true, "key": "quality_gate_enabled", "value": "0" }
260
- ```
261
-
262
- All settings and their defaults are listed in [Configuration](#configuration).
263
-
264
- ### `kevin_feedback`
265
-
266
- Rates an injected memory and makes the rating count. `verdict` is `useful` | `wrong` | `outdated` | `ignore`. The first three are stored in `memory_feedback` and move `kevin_why`'s confidence (`+0.05` / `-0.1` per count); **`ignore` is a hard action** — the memory is stamped `ignored = 1` and excluded from retrieval, queries and injection.
267
-
268
- ```
269
- kevin_feedback({ memory_id: "0195a3b2-...", verdict: "wrong", note: "the fix was wrong" })
270
- // → { "ok": true, "verdict": "wrong" }
271
- ```
272
-
273
- ### `kevin_trace`
274
-
275
- Strict dry-run: predicts exactly which memories `onSystemTransform` WOULD inject for a query (optionally `session_id`, `tag` and `cap`), with **zero side effects** — no counters, no ledger rows, no seen-set writes, no relevance bumps. Rejected items carry their `GateReason` (`seen_this_session` | `weak` | `recurrence` | `stale` | `ignored` | `confidence`).
276
-
277
- ```
278
- kevin_trace({ query: "tsc error" })
279
- // → { "query": "tsc error", "tag": "context", "cap": 400, "would_inject": true,
280
- // "total_tokens": 82,
281
- // "admitted": [ { "id": "...", "type": "error", "decision": "admitted", "tokens": 62 } ],
282
- // "blocked": [ { "id": "...", "type": "error", "decision": "blocked",
283
- // "reason": "confidence", "tokens": 20 } ] }
284
- ```
285
-
286
- ### `kevin_audit`
287
-
288
- Read-only report of the whole system state: memories by `status`/`origin`/`type`, injection outcomes with `precision_rate`/`coverage_rate`, the six `blocked` counters, feedback by verdict, tokens injected, the push-vs-pull `channels` comparison and the `curation` scoreboard. `verbose: true` adds the settings block. No writes, no LLM; on pre-007 databases it omits the v0.6 blocks and reports `"partial": true`.
289
-
290
- ```
291
- kevin_audit({})
292
- // → { "memories": { "total": 42, "by_status": { "active": 37, "stale": 1, "archived": 4 },
293
- // "by_origin": { "reflector": 12, "agent": 30 }, "by_type": { "error": 20, ... },
294
- // "ignored": 1, "with_feedback": 3 },
295
- // "injections": { "total": 14, "effective": 2, "ineffective": 3, "inconclusive": 9,
296
- // "unmeasured": 0, "precision_rate": 0.40, "coverage_rate": 0.36 },
297
- // "blocked": { "seen": 1, "weak": 0, "recurrence": 2, "stale": 0,
298
- // "ignored": 1, "confidence": 2 },
299
- // "feedback": { "positive": 2, "negative": 1, "by_verdict": { "useful": 2, "wrong": 1 } },
300
- // "tokens": { "pre_prompt": 51, "compacting": 0 }, "partial": false,
301
- // "channels": { "push": { "tokens_pre_prompt": 51, "injections_total": 14,
302
- // "precision_rate": 0.40, "coverage_rate": 0.36,
303
- // "budget_tokens": 400 },
304
- // "pull": { "proposals_created": 6, "proposals_approved": 1,
305
- // "proposals_rejected": 2, "artifact_writes_total": 2,
306
- // "artifact_writes_noop": 1, "references_registered": 0,
307
- // "skills_registered": 0,
308
- // "skill_emission": "off", "reference_emission": "off" } },
309
- // "curation": { "eligible": 5, "curated": 1, "inferable": 3, "non_inferable": 2,
310
- // "unknown": 1, "proposals_by_status": { "pending": 2, "applied": 1, ... } } }
311
- ```
312
-
313
- ### `kevin_propose`
314
-
315
- Creates curation proposals as `pending` rows with unified diffs — **a strict dry run**. Reads the eligible memories (`inferable != 1`), renders what would go into the artifact, and returns the minimal diff. No disk write, no `curated` marks, no side effects. Only `kevin_approve` may write.
316
-
317
- ```
318
- kevin_propose({ kind: "agents_md" }) // kind: "agents_md" | "skill" | "reference"
319
- // → { "proposals": [ { "id": "...", "kind": "agents_md", "targetPath": "AGENTS.md",
320
- // "memoryIds": ["mem-1"], "status": "pending",
321
- // "createdAt": "2026-08-14 10:00:00",
322
- // "diff": "--- a/AGENTS.md\n+++ b/AGENTS.md\n@@ ..." } ] }
323
- ```
324
-
325
- ### `kevin_approve`
326
-
327
- The **only** code path that writes a file. `approve` applies the proposal's diff atomically (temp file + rename, CRLF/BOM preserved), records an `artifact_writes` audit row, marks the proposal `applied` and its memories `curated`. `reject` records the human decision and touches nothing. Refusals and noops are audited, never silent.
328
-
329
- ```
330
- kevin_approve({ proposal_id: "...", decision: "approve" }) // or "reject"
331
- // → { "proposalId": "...", "status": "applied", "outcome": "written", "curated": 1 }
332
- // ("outcome": "noop" when the artifact already matches, "refused" when the
333
- // marker block is malformed; a rejected proposal returns
334
- // { "proposalId": "...", "status": "rejected" })
335
- ```
336
-
337
- ### `kevin_publish`
338
-
339
- Regenerates the pull-channel bundles under `~/.opencode-kevin/` — `skills/project-knowledge.md` and `refs/<topic>.md` — reporting per-bundle outcome and the emission state (`on` / `off` / `unavailable`). Registration with the host happens at plugin startup; this tool only materializes and reports.
340
-
341
- ---
342
-
343
- ## How Kevin measures itself
344
-
345
- ### Injection outcomes
346
-
347
- Every injection is settled at `session.idle` into one of **four outcomes**:
348
-
349
- | Outcome | Meaning | Counts toward precision? |
350
- |---|---|---|
351
- | `effective` | A linked fix was observed after the injection | yes (numerator) |
352
- | `ineffective` | The same error recurred after the injection | yes (denominator) |
353
- | `inconclusive` | Neither the error did not recur, but no fix was seen either | no |
354
- | `unmeasured` | Session went idle before settlement could run | no |
355
-
356
- - **`precision_rate`** = `effective / (effective + ineffective)`. Measuring *effect*, not absence of recurrence: a lesson that was injected and never contradicted counts as `inconclusive`, not success. **Your precision rate will look lower than before v0.5.0. That is the honest number.**
357
- - **`coverage_rate`** = `(effective + ineffective) / total` — the share of injections that were actually measured. Reported alongside precision so a low measurable fraction stays visible instead of hiding behind a large total.
358
- - **`blocked`** counts every gate rejection by reason `seen_this_session`, `weak`, `recurrence`, `stale`, `ignored`, `confidence` a rejection you did not count did not happen.
359
-
360
- ### The quality gate
361
-
362
- Weak lessons errors the reflector cannot dispatch to a deterministic rule — are **stored but never injected** while `quality_gate_enabled = '1'` (default). Recurrences demote lessons (`recurrence_count` → `stale`) and lower confidence. Debug mode: `kevin_config({ action: "set", key: "quality_gate_enabled", value: "0" })` re-injects weak lessons with a `(low confidence)` marker.
363
-
364
- ### Seeing the whole picture
365
-
366
- `kevin_trace` shows you the plan *before* it happens (dry run, zero side effects); `kevin_audit` reads the whole state after; `kevin_feedback` lets a human correct it — and the correction moves the confidence number `kevin_why` reports.
367
-
368
- ---
369
-
370
- ## Curation & Pull
371
-
372
- ### The marker contract
373
-
374
- Kevin never edits your files directly. Every artifact write happens inside a frozen marker block, delimited verbatim by:
375
-
376
- ```
377
- <!-- kevin:begin — curated by opencode-kevin, safe to edit -->
378
- <!-- kevin:end -->
379
- ```
380
-
381
- These exact strings are **frozen for the v0.x line** — README, tests and the v1.0.0 migration plan all depend on their byte sequences. What Kevin guarantees:
382
-
383
- - **Only the block between the markers may change.** Bytes outside them are byte-identical after every write — including line endings (a CRLF file stays CRLF everywhere, even inside the generated block), a leading UTF-8 BOM, and the file's final newline.
384
- - **Malformed markers are refused, never repaired.** If the file contains a `begin` without an `end` (or vice versa), Kevin refuses the write with an explicit reason and the file is untouched. Repairing would mean guessing at user intent; refusing means the state stays visible and auditable.
385
- - **Idempotent**: applying an unchanged plan is a counted `noop`no temp file, no write, no mtime churn.
386
-
387
- ### The propose review approve flow
388
-
389
- ```
390
- eligible memories (inferable != 1)
391
-
392
-
393
- kevin_propose({ kind }) ── creates pending rows + unified diffs.
394
- │ NO disk write, NO curated marks.
395
-
396
- HUMAN REVIEWS THE DIFF ── this is the entire safety model:
397
- │ a memory earns its way into a file
398
- ▼ only after a human said yes.
399
- kevin_approve({ proposal_id, decision })
400
-
401
- ├── "approve" ── ArtifactWriter.apply() (the ONLY write path)
402
- atomic temp+rename, audit row in artifact_writes,
403
- │ memory marked curated, proposal marked applied
404
- └── "reject" ── recorded, nothing touches disk
405
- ```
406
-
407
- Rejection history is never deleted: it is the evidence base for the roadmap's kill criterion "proposals rejected more often than approved".
408
-
409
- ### Three distribution channels
410
-
411
- | Channel | Artifact | Cost when unused |
412
- |---|---|---|
413
- | **Push** | per-prompt `<kevin-context>` injection | charges on every prompt — now capped at 400 tokens by default |
414
- | **Pull — AGENTS.md** | marker block in the project's `AGENTS.md` | zero |
415
- | **Pull — skills** | `~/.opencode-kevin/skills/project-knowledge.md` (`skill_emission_enabled`) | zero |
416
- | **Pull — references** | `~/.opencode-kevin/refs/<topic>.md` (`reference_emission_enabled`) | zero |
417
-
418
- `kevin_audit`'s `channels` block compares push vs pull on the same axes, and reports each emission channel as `"on"`, `"off"` (setting `'0'` on a capable host) or `"unavailable"` (host without the v2 domain).
419
-
420
- ### The confidence floor gate
421
-
422
- `injection_confidence_floor` (default `'0.6'`) rejects memories whose computed confidence is below the floor, counted as `injections_blocked_confidence` — the sixth gate rejection reason, measured exactly like the first five. Single-observation memories (base confidence 0.5, no confirmed evidence) stop being pushed by default; `kevin_config({ action: "set", key: "injection_confidence_floor", value: "0" })` restores v0.5 behaviour exactly.
423
-
424
- ---
425
-
426
- ## Replay harness
427
-
428
- `npm run replay` runs every transcript in `tests/replay/fixtures/` through the plugin against an in-memory database with a frozen clock and prints one table row per transcript (memories created, injection outcomes, `precision_rate`, `coverage_rate`, tokens). Record your own session as a JSON array of typed events (`session.created`, `chat.message`, `tool.before`, `tool.after`, `system.transform`, `compacting`, `session.idle`) with ISO-8601 `at` timestamps, drop it into `tests/replay/fixtures/`, and re-run. The `at` timestamps are the only source of time during replay.
429
-
430
- ---
431
-
432
- ## Hooks
433
-
434
- Kevin subscribes to 6 OpenCode hooks:
435
-
436
- | Hook | What Kevin does |
437
- |---|---|
438
- | `tool.execute.before` | Records tool call start (callID + redacted args) |
439
- | `tool.execute.after` | Records result (id = callID); on failure → Reflector.invoke async (throttled); on success → CausalChain links the fix |
440
- | `experimental.chat.system.transform` | Injects relevant lessons in `<kevin-context>` (400 tokens by default, configurable) + optional `<kevin-suggestion>` |
441
- | `experimental.session.compacting` | Re-injects lessons in `<kevin-memory>` after compacting (2000 tokens) + optional `<kevin-suggestion>` |
442
- | `event` (`session.created`) | Captures current `sessionID` (skill/reference emissions register at plugin startup, not per session) |
443
- | `event` (`session.idle`) | Settles injection outcomes; generates the retrospective; boosts positive lessons; penalizes recurring failures; promotes causal patterns and mines patterns (opt-in); drafts curation proposals (`curation_enabled`); flushes metrics |
444
-
445
- **Redaction**: absolute paths (`C:\Users\...`, `/home/...`) `<path>` and secrets (`API_KEY=`, `Bearer`, `token`) `<redacted>` before persisting anything. `<private>…</private>` blocks are swept from tool call args and output before persistence and replaced with `<private: redacted N chars>`.
446
-
447
- **Throttle**: Reflector generates at most 1 lesson per minute per unique fingerprint (per-fingerprint, not global). Configurable via `throttleMs`.
448
-
449
- **Truncation**: content > 4KB keeps the lesson searchable; only the additional context is truncated (`metadata.truncated = true`).
450
-
451
- ---
452
-
453
- ## Configuration
454
-
455
- ### Plugin options
456
-
457
- Kevin accepts options via the plugin's tuple form (see Installation → Advanced). Programmatic defaults:
458
-
459
- ```ts
460
- import { KevinPlugin } from "@jmtrin/opencode-kevin";
461
-
462
- // defaults
463
- KevinPlugin(input, {
464
- dbPath: "~/.opencode-kevin/kevin.db", // or ":memory:" for tests
465
- migrationsDir: "<package>/dist/migrations", // resolved automatically
466
- retrospectivesDir: "~/.opencode-kevin/retrospectives",
467
- throttleMs: 60_000,
468
- });
469
- ```
470
-
471
- ### Settings
472
-
473
- Read/write via `kevin_config({ action: "list" | "set", ... })`. All values are TEXT; booleans compare against `"1"`.
474
-
475
- | Setting | Default | Effect |
476
- |---|---|---|
477
- | `quality_gate_enabled` | `"1"` | Weak lessons are stored but never injected while enabled |
478
- | `lesson_snippet_injection` | `"1"` | Injects the rescued errorType snippet with each lesson |
479
- | `llm_reflection_enabled` | `"0"` | Opt-in LLM enrichment of reflector lessons |
480
- | `cross_project_enabled` | `"0"` | `kevin_query` includes imported cross-project memories |
481
- | `patternminer_enabled` | `"0"` | Opt-in deterministic 2-gram/3-gram pattern miner at `session.idle` |
482
- | `tool_calls_dedup_enabled` | `"0"` | Opt-in dedup of repeated tool calls |
483
- | `deterministic_retrieval` | `"0"` | Freezes Kevin's internal clock (recency factor 1.0, no relevance bumps) — for hermetic tests and the replay harness |
484
- | `pre_prompt_budget_tokens` | `"400"` | Pre-prompt injection cap, clamped to `[0, 4000]`; `0` turns push off |
485
- | `archive_after_days` | `"30"` | Age at which stale non-pattern memories are retired to `archived` on `session.idle` |
486
- | `curation_enabled` | `"1"` | Generates curation proposals at `session.idle` |
487
- | `agents_md_path` | `"AGENTS.md"` | Where the AGENTS.md channel writes (project-relative) |
488
- | `skill_emission_enabled` | `"0"` | Registers the curated skill with the host at startup (v2 hosts only) |
489
- | `reference_emission_enabled` | `"0"` | Registers `@kevin/<topic>` references at startup (v2 hosts only) |
490
- | `injection_confidence_floor` | `"0.6"` | Push gate: memories below this confidence are counted and rejected |
491
-
492
- ---
493
-
494
- ## Development
495
-
496
- ```bash
497
- git clone https://github.com/jmtrin/opencode-kevin.git
498
- cd opencode-kevin
499
- npm install
500
- npm run typecheck # tsc --noEmit (strict)
501
- npm run lint # biome check .
502
- npm test # vitest run (unit + integration + e2e + replay)
503
- npm run verify # post-install verification
504
- npm run replay # replay report over tests/replay/fixtures
505
- ```
506
-
507
- ### Publishing (maintainer)
508
-
509
- ```bash
510
- npm login # as the jmtrin account that owns the @jmtrin scope
511
- npm publish --access public
512
- ```
513
-
514
- `prepublishOnly` runs `npm run build` (tsc + copy migrations) automatically. The `files` field ships only `dist/plugin`, `dist/migrations`, and `migrations`. `dist/` is gitignored and rebuilt on publish.
515
-
516
- ### Structure
517
-
518
- ```
519
- plugin/
520
- index.ts # Entry point: KevinPlugin (wires hooks, tools, emissions)
521
- Store.ts # SQLite wrapper (node:sqlite / bun:sqlite / better-sqlite3 fallback)
522
- sqlite-adapter.ts # Runtime-agnostic SQLite adapter behind Store
523
- Migrate.ts # Idempotent migrations + post-apply hooks
524
- MemoryService.ts # save/query/getRelevant (FTS5 + bm25 + origin-aware rank + supersede)
525
- ToolCallObserver.ts # onBefore/onAfter + redact + inferErrorType + dedup (opt-in)
526
- Reflector.ts # Heuristic lessons + per-fingerprint throttle + LLM enrich (opt-in)
527
- ContextInjector.ts # deriveQuery + pre-prompt/compacting injection + <kevin-suggestion>
528
- Retrospective.ts # Generates retrospective.md + FP recap + metrics snapshot
529
- Feedback.ts # kevin_feedback: verdicts, confidence terms, ignored stamp
530
- Archiver.ts # Retires stale non-pattern memories past archive_after_days
531
- CausalChain.ts # Links fixes to failures + promotes causal patterns
532
- QualityGate.ts # Weak-lesson gate (stored, not injected by default)
533
- InjectionLedger.ts # Injection ledger + settle precision_rate
534
- LessonFixer.ts # Deterministic fix_args capture + promotion enrichment
535
- PatternMiner.ts # Opt-in deterministic 2-gram/3-gram miner
536
- Curator.ts # Curation candidates + propose/approve lifecycle
537
- ArtifactWriter.ts # The SINGLE write path (markers, atomic, noop, audit rows)
538
- Materializer.ts # Pull-channel topic bundles (skills, refs)
539
- inferability.ts # Deterministic inferable/non-inferable/unknown classifier
540
- capabilities.ts # v2 domain probe (skills / references)
541
- diff.ts # Minimal unified diff for proposal review
542
- replay.ts # Hermetic replay driver over recorded transcripts
543
- replay-types.ts # Transcript/result types for the replay harness
544
- kevin_propose.ts # kevin_propose tool (strict dry run)
545
- kevin_approve.ts # kevin_approve tool (only writer call site)
546
- kevin_publish.ts # kevin_publish tool (bundle regeneration)
547
- kevin_audit.ts # Read-only audit + channels/curation blocks
548
- kevin_why.ts # kevin_why tool: failure→fix traces + related rules
549
- okf-export.ts # kevin_export: OKF/markdown export
550
- okf-import.ts # kevin_import: bundle parser + import
551
- confidence.ts # Two-sided computeConfidence (evidence + recurrence + feedback)
552
- query-tokenizer.ts # FTS5 tokenizer for query sanitization
553
- memory-format.ts # escapeInjectedText, formatMemories, <protect> + id: line wrappers
554
- redact.ts # redactPaths + stripPrivate
555
- fingerprint.ts # FNV-1a 64-bit (in-house, no node:crypto)
556
- metrics.ts # In-memory counters + debounced flush to kevin_metrics
557
- uuid.ts # UUIDv7
558
- migrations/
559
- 001_initial.sql # schema: memories, tool_calls, retrospectives
560
- 002_indexes.sql # FTS5 + indexes
561
- 003_v02_signal.sql # fingerprint, origin, metrics, dedup indexes
562
- 004_v03_knowledge.sql # evidence/status/supersede, error_fingerprint
563
- 005_v04_signal.sql # recurrence_count, fix_args, last_injected_at
564
- 006_v05_glassbox.sql # ignored/archived/superseded_by, feedback, metrics
565
- 007_v06_pull.sql # curation_proposals, artifact_writes, curated/inferable
566
- tests/
567
- unit/ # component tests
568
- integration/ # tool-level tests through real components
569
- e2e/ # closed-loop tests through the host hooks
570
- replay/ # transcript fixtures + replay harness tests
571
- scripts/
572
- copy-migrations.mjs # build step: copies *.sql to dist/migrations
573
- verify-install.ts # npm run verify
574
- ```
575
-
576
- ---
577
-
578
- ## License
579
-
580
- MIT
1
+ # Kevin
2
+
3
+ > Observe and learn: the learning layer OpenCode was missing.
4
+
5
+ Kevin is an [OpenCode](https://opencode.ai) plugin that **observes** every agent tool call, **learns** from failures by generating lessons, and **shares** what it learned proactively in future sessions. It does not plan, orchestrate, or compete with the plugin ecosystem. It only learns.
6
+
7
+ - **Local-first**: SQLite + FTS5, no external services, no network calls.
8
+ - **Global memory**: a single `~/.opencode-kevin/kevin.db` shared across all your projects (WAL mode → safe for concurrent sessions). No per-project folders.
9
+ - **Knowledge + Causality**: causal failure→fix chains, `kevin_why` explanations, OKF export/import, a supersede model, and human-in-the-loop AGENTS.md suggestions.
10
+ - **Signal over Noise**: a quality gate that stores weak lessons without injecting them, an injection ledger with honest `precision_rate`, and two-sided confidence.
11
+ - **Glass Box**: honest measurement replaces estimates — three-way injection settlement (`effective` / `ineffective` / `inconclusive`), human feedback that actually moves confidence, a strict dry-run `kevin_trace`, a read-only `kevin_audit`, and a hermetic replay harness.
12
+ - **Pull**: knowledge earns its way into files the model actually reads — `kevin_propose` generates a reviewable diff, a human approves, and **only then** does Kevin write, inside a frozen marker block, preserving your file's CRLF/BOM/formatting byte-for-byte outside it. Plus three distribution channels (AGENTS.md, skills, references) and a push budget gated by a confidence floor.
13
+ - **Audited**: the v0.4.0 bug catalog (`docs/Kevin_v0.4.0_Bugs.md`) is fully closed — 16/16 bugs fixed and regression-tested.
14
+ - **Standalone**: works without any other plugin. With the ecosystem, it learns more richly.
15
+
16
+ ---
17
+
18
+ ## Contents
19
+
20
+ - [Installation](#installation)
21
+ - [How Kevin works](#how-kevin-works)
22
+ - [Tools](#tools)
23
+ - [How Kevin measures itself](#how-kevin-measures-itself)
24
+ - [Curation & Pull](#curation--pull)
25
+ - [Replay harness](#replay-harness)
26
+ - [Hooks](#hooks)
27
+ - [Configuration](#configuration)
28
+ - [Development](#development)
29
+ - [License](#license)
30
+
31
+ ---
32
+
33
+ ## Installation
34
+
35
+ ### 1. Declare the plugin
36
+
37
+ Add Kevin to your OpenCode config. For **all projects** (global):
38
+
39
+ ```jsonc
40
+ // ~/.config/opencode/opencode.jsonc
41
+ {
42
+ "$schema": "https://opencode.ai/config.json",
43
+ "plugin": [
44
+ "@jmtrin/opencode-kevin@latest"
45
+ ]
46
+ }
47
+ ```
48
+
49
+ For a **single project**, put the same `plugin` array in `./opencode.json` or `.opencode/opencode.json` at the project root.
50
+
51
+ ### 2. Restart OpenCode
52
+
53
+ Config is loaded once at startup and is **not hot-reloaded** — quit and reopen OpenCode after editing. On start, Kevin exposes 18 tools, including `kevin_facts` and `kevin_conflicts`.
54
+
55
+ Contradictions de-rank memories and surface conflicts. They never delete, stale, archive, or auto-resolve a memory.
56
+
57
+ ### 3. Where data lives
58
+
59
+ Kevin stores everything in a single **global, shared** location under your home directory — no per-project `.kevin/` folders:
60
+
61
+ | Path | Content |
62
+ |---|---|
63
+ | `~/.opencode-kevin/kevin.db` | SQLite database (memories, tool calls, retrospectives). WAL mode → safe for concurrent OpenCode sessions across projects. |
64
+ | `~/.opencode-kevin/retrospectives/<session>.md` | Per-session retrospective markdown. |
65
+
66
+ Migrations run automatically on startup.
67
+
68
+ ### Requirements
69
+
70
+ - **Node.js >= 22.5** (uses `node:sqlite`, the built-in SQLite module — no native binaries to compile).
71
+ - OpenCode with plugin support (`@opencode-ai/plugin` >= 1.17).
72
+
73
+ > **Runtimes**:
74
+ > - **Bun**: uses `bun:sqlite` (built-in).
75
+ > - **Node 24+**: uses `node:sqlite` directly, no flags needed (emits an experimental warning, harmless).
76
+ > - **Node 22/23 without `--experimental-sqlite` flag** or **Node 20**: falls back to `better-sqlite3`, declared as `optionalDependencies`. If you need it, install it manually in your opencode config directory (`~/.config/opencode/`): `npm install better-sqlite3`.
77
+
78
+ ### Verification
79
+
80
+ ```bash
81
+ npm run verify
82
+ ```
83
+
84
+ Checks Node version, SQLite, migration, MemoryService save/query, Reflector, ContextInjector, and TypeScript strict mode.
85
+
86
+ ### Advanced (optional)
87
+
88
+ Override defaults via the plugin tuple form:
89
+
90
+ ```jsonc
91
+ {
92
+ "plugin": [
93
+ ["@jmtrin/opencode-kevin", {
94
+ "dbPath": "/custom/path/kevin.db",
95
+ "retrospectivesDir": "/custom/path/retrospectives",
96
+ "throttleMs": 120000
97
+ }]
98
+ ]
99
+ }
100
+ ```
101
+
102
+ Use `:memory:` for `dbPath` in tests.
103
+
104
+ ---
105
+
106
+ ## How Kevin works
107
+
108
+ Every tool call is observed; every failure becomes a lesson; every lesson is either pushed into the next prompt, written into an artifact a human approved, or retired when it stops earning its place.
109
+
110
+ ```
111
+ Tool call (success or failure)
112
+
113
+
114
+ ┌─────────────────────────┐ OBSERVE
115
+ ToolCallObserver │ records every call (tool, redacted args,
116
+ └───────────┬─────────────┘ success, duration, error type, dedup)
117
+
118
+ failure success
119
+ ┌──────────▼───────────┐ ┌─────────────────────────┐
120
+ Reflector │ │ CausalChain
121
+ heuristic lesson │ links the fix to the │
122
+ per error code │ failure within 10 calls
123
+ (throttled per │ └────────────┬────────────┘
124
+ fingerprint) │ │
125
+ └──────────┬───────────┘ │ session.idle
126
+
127
+ ▼ ┌─────────────────────────┐
128
+ ┌───────────────────────┐promotes recurring
129
+ ContextInjector │◄─┤ errors → causal
130
+ SHARE: injects │ │ patterns (cumulative
131
+ │ <kevin-context> │ │ evidence) │
132
+ ≤400 tokens/prompt │ └─────────────────────────┘
133
+ └───────────┬───────────┘
134
+ session.idle
135
+
136
+ ┌─────────────────────────┐ RETROSPECTIVE: <session>.md with
137
+ │ Retrospective │ lessons, metrics snapshot, causal
138
+ └─────────────────────────┘ promotion, pattern mining (opt-in)
139
+ ```
140
+
141
+ At `session.idle` Kevin also settles injection outcomes, retires stale memories, and — when curation is enabled — drafts pull proposals for your review (see [Curation & Pull](#curation--pull)).
142
+
143
+ ---
144
+
145
+ ## Tools
146
+
147
+ Kevin exposes 16 tools callable by the agent.
148
+
149
+ ### `kevin_save`
150
+
151
+ Saves an explicit memory.
152
+
153
+ ```
154
+ kevin_save({ type: "decision", content: "We use vitest for tests", scope: "project" })
155
+ // → { "id": "0195a3b2-..." }
156
+ ```
157
+
158
+ `type`: `error` | `pattern` | `decision` | `context` | `rule` | `solution`. `scope`: `project` (persists) | `session` (TTL 24h).
159
+
160
+ Saving a `decision` or `rule` with the same `fingerprint` as an existing active row supersedes the old one (`status='superseded'`, hidden from default queries).
161
+
162
+ ### `kevin_query`
163
+
164
+ Searches memories by text (FTS5 + bm25). Returns a **slim** payload by default; pass `full: true` for the complete content, or `evidence: true` to include `confidence`, `evidence_count` and `last_verified_at`.
165
+
166
+ ```
167
+ kevin_query({ query: "typecheck", type: "error", limit: 5 })
168
+ // → [{ "id": "...", "type": "error", "scope": "project", "score": -0.87,
169
+ // "snippet": "When bash fails with typecheck:..." }, ...]
170
+ ```
171
+
172
+ ### `kevin_get`
173
+
174
+ Fetches a **single full memory** by id (progressive disclosure) — use it when `kevin_query` returned a slim snippet and you need the complete content.
175
+
176
+ ```
177
+ kevin_get({ id: "0195a3b2-..." })
178
+ // → { "id": "...", "type": "error", "content": "...", "scope": "project",
179
+ // "relevanceScore": 0.55, "origin": "reflector", "fingerprint": "cbf29ce484222325",
180
+ // "projectId": null, "metadata": null,
181
+ // "evidenceCount": 2, "recurrenceCount": 1, "lastVerifiedAt": "2026-08-01 10:00:00",
182
+ // "status": "active", "confidence": 0.55, "fixArgs": "npm i -g rg" }
183
+ ```
184
+
185
+ ### `kevin_recall`
186
+
187
+ Retrieves relevant memories (greedy fill by relevance). Without `query`, returns all memories in scope. Pass `includeSuperseded: true` to include superseded rows.
188
+
189
+ ```
190
+ kevin_recall({ query: "auth", limit: 3 })
191
+ // → [{ "id": "...", "type": "decision", ... }, ...]
192
+ ```
193
+
194
+ ### `kevin_status`
195
+
196
+ Global counts and metrics: memory census, the precision block, the six blocked-gate counters, feedback totals, and the v0.6 block (`schema_version`, `curation_enabled`, emission states, `proposals_pending` — omitted on pre-007 databases).
197
+
198
+ ```
199
+ kevin_status({})
200
+ // → { "memories": 42, "memories_reflector": 12, "memories_agent": 30,
201
+ // "memories_pattern": 0, "memories_causal": 1, "tool_calls": 318,
202
+ // "retrospectives": 7, "tool_count": 16,
203
+ // "metrics": { "tokens_injected_pre_prompt": 51, "tokens_injected_compacting": 0,
204
+ // "reflections_throttled": 3, "duplicate_suppressions": 2,
205
+ // "tool_calls_deduped": 0, "patterns_mined": 0,
206
+ // "patterns_causal": 1, "causal_links": 2, "memories_superseded": 0,
207
+ // "injections_inconclusive": 9, ... },
208
+ // "injections_total": 14, "injections_effective": 2, "injections_ineffective": 3,
209
+ // "injections_inconclusive": 9, "precision_rate": 0.40, "coverage_rate": 0.36,
210
+ // "blocked": { "seen": 1, "weak": 0, "recurrence": 2, "stale": 0,
211
+ // "ignored": 1, "confidence": 2 },
212
+ // "memories_ignored": 1, "memories_archived": 4,
213
+ // "feedback": { "positive": 2, "negative": 1 },
214
+ // "patterns_promoted_new": 2, "recurrence_by_origin": { "reflector": 3, "causal": 1 },
215
+ // "v06": { "schema_version": "007", "curation_enabled": "1",
216
+ // "skill_emission": "off", "reference_emission": "off",
217
+ // "proposals_pending": 2 } }
218
+ ```
219
+
220
+ ### `kevin_retrospective`
221
+
222
+ Generates a retrospective for a session (uses the current session if `session_id` is omitted).
223
+
224
+ ```
225
+ kevin_retrospective({ session_id: "sess-abc" })
226
+ // → { "file_path": "~/.opencode-kevin/retrospectives/sess-abc.md" }
227
+ // or → { "message": "No failures in session sess-abc." }
228
+ ```
229
+
230
+ ### `kevin_why`
231
+
232
+ Explains *why* a failure keeps happening: looks up causal patterns for the query and builds a failure → fix trace from memories + tool_calls, including related TypeScript error-code rules.
233
+
234
+ ```
235
+ kevin_why({ query: "TS2304 cannot find name" })
236
+ // { "summary": "TS2304 recurs because ... Confirmed by 2 fixes.",
237
+ // "confidence": 0.7, "evidence_count": 2, "last_verified": "2026-08-01 10:00:00",
238
+ // "trace": [ { "type": "error", "summary": "..." }, { "type": "fix", "tool": "bash" } ],
239
+ // "related_rules": [ { "code": "TS2304", "suggestion": "import or typo" } ] }
240
+ ```
241
+
242
+ ### `kevin_export`
243
+
244
+ Exports knowledge for sharing: `decision`/`rule`/`pattern` memories (active only, no raw errors) as YAML-frontmatter blocks (`format: "okf"`) or markdown (`format: "markdown"`). Includes `id`, `type`, `confidence`, `evidence_count`, `recurrence_count`, `last_verified_at`, `fingerprint`. Timestamps are treated as UTC — a re-import reproduces the exact source values.
245
+
246
+ ### `kevin_import`
247
+
248
+ Ingests an exported bundle. Each entry becomes a `context` memory with `origin='imported'`; a fingerprint collision with an existing `decision`/`rule` supersedes the old row. Returns `{ imported, superseded }`.
249
+
250
+ ### `kevin_config`
251
+
252
+ Reads/writes `kevin_settings` without SQL. `action: "list"` returns every setting; `action: "set"` upserts a value (default `"1"` when omitted) and rejects unknown keys unless `strict: false`.
253
+
254
+ ```
255
+ kevin_config({ action: "list" })
256
+ // → { "quality_gate_enabled": "1", "lesson_snippet_injection": "1",
257
+ // "llm_reflection_enabled": "0", "pre_prompt_budget_tokens": "400",
258
+ // "injection_confidence_floor": "0.6", ... }
259
+
260
+ kevin_config({ action: "set", key: "quality_gate_enabled", value: "0" })
261
+ // → { "ok": true, "key": "quality_gate_enabled", "value": "0" }
262
+ ```
263
+
264
+ All settings and their defaults are listed in [Configuration](#configuration).
265
+
266
+ ### `kevin_feedback`
267
+
268
+ Rates an injected memory and makes the rating count. `verdict` is `useful` | `wrong` | `outdated` | `ignore`. The first three are stored in `memory_feedback` and move `kevin_why`'s confidence (`+0.05` / `-0.1` per count); **`ignore` is a hard action** — the memory is stamped `ignored = 1` and excluded from retrieval, queries and injection.
269
+
270
+ ```
271
+ kevin_feedback({ memory_id: "0195a3b2-...", verdict: "wrong", note: "the fix was wrong" })
272
+ // → { "ok": true, "verdict": "wrong" }
273
+ ```
274
+
275
+ ### `kevin_trace`
276
+
277
+ Strict dry-run: predicts exactly which memories `onSystemTransform` WOULD inject for a query (optionally `session_id`, `tag` and `cap`), with **zero side effects** — no counters, no ledger rows, no seen-set writes, no relevance bumps. Rejected items carry their `GateReason` (`seen_this_session` | `weak` | `recurrence` | `stale` | `ignored` | `confidence`).
278
+
279
+ ```
280
+ kevin_trace({ query: "tsc error" })
281
+ // { "query": "tsc error", "tag": "context", "cap": 400, "would_inject": true,
282
+ // "total_tokens": 82,
283
+ // "admitted": [ { "id": "...", "type": "error", "decision": "admitted", "tokens": 62 } ],
284
+ // "blocked": [ { "id": "...", "type": "error", "decision": "blocked",
285
+ // "reason": "confidence", "tokens": 20 } ] }
286
+ ```
287
+
288
+ ### `kevin_audit`
289
+
290
+ Read-only report of the whole system state: memories by `status`/`origin`/`type`, injection outcomes with `precision_rate`/`coverage_rate`, the six `blocked` counters, feedback by verdict, tokens injected, the push-vs-pull `channels` comparison and the `curation` scoreboard. `verbose: true` adds the settings block. No writes, no LLM; on pre-007 databases it omits the v0.6 blocks and reports `"partial": true`.
291
+
292
+ ```
293
+ kevin_audit({})
294
+ // → { "memories": { "total": 42, "by_status": { "active": 37, "stale": 1, "archived": 4 },
295
+ // "by_origin": { "reflector": 12, "agent": 30 }, "by_type": { "error": 20, ... },
296
+ // "ignored": 1, "with_feedback": 3 },
297
+ // "injections": { "total": 14, "effective": 2, "ineffective": 3, "inconclusive": 9,
298
+ // "unmeasured": 0, "precision_rate": 0.40, "coverage_rate": 0.36 },
299
+ // "blocked": { "seen": 1, "weak": 0, "recurrence": 2, "stale": 0,
300
+ // "ignored": 1, "confidence": 2 },
301
+ // "feedback": { "positive": 2, "negative": 1, "by_verdict": { "useful": 2, "wrong": 1 } },
302
+ // "tokens": { "pre_prompt": 51, "compacting": 0 }, "partial": false,
303
+ // "channels": { "push": { "tokens_pre_prompt": 51, "injections_total": 14,
304
+ // "precision_rate": 0.40, "coverage_rate": 0.36,
305
+ // "budget_tokens": 400 },
306
+ // "pull": { "proposals_created": 6, "proposals_approved": 1,
307
+ // "proposals_rejected": 2, "artifact_writes_total": 2,
308
+ // "artifact_writes_noop": 1, "references_registered": 0,
309
+ // "skills_registered": 0,
310
+ // "skill_emission": "off", "reference_emission": "off" } },
311
+ // "curation": { "eligible": 5, "curated": 1, "inferable": 3, "non_inferable": 2,
312
+ // "unknown": 1, "proposals_by_status": { "pending": 2, "applied": 1, ... } } }
313
+ ```
314
+
315
+ ### `kevin_propose`
316
+
317
+ Creates curation proposals as `pending` rows with unified diffs — **a strict dry run**. Reads the eligible memories (`inferable != 1`), renders what would go into the artifact, and returns the minimal diff. No disk write, no `curated` marks, no side effects. Only `kevin_approve` may write.
318
+
319
+ ```
320
+ kevin_propose({ kind: "agents_md" }) // kind: "agents_md" | "skill" | "reference"
321
+ // → { "proposals": [ { "id": "...", "kind": "agents_md", "targetPath": "AGENTS.md",
322
+ // "memoryIds": ["mem-1"], "status": "pending",
323
+ // "createdAt": "2026-08-14 10:00:00",
324
+ // "diff": "--- a/AGENTS.md\n+++ b/AGENTS.md\n@@ ..." } ] }
325
+ ```
326
+
327
+ ### `kevin_approve`
328
+
329
+ The **only** code path that writes a file. `approve` applies the proposal's diff atomically (temp file + rename, CRLF/BOM preserved), records an `artifact_writes` audit row, marks the proposal `applied` and its memories `curated`. `reject` records the human decision and touches nothing. Refusals and noops are audited, never silent.
330
+
331
+ ```
332
+ kevin_approve({ proposal_id: "...", decision: "approve" }) // or "reject"
333
+ // { "proposalId": "...", "status": "applied", "outcome": "written", "curated": 1 }
334
+ // ("outcome": "noop" when the artifact already matches, "refused" when the
335
+ // marker block is malformed; a rejected proposal returns
336
+ // { "proposalId": "...", "status": "rejected" })
337
+ ```
338
+
339
+ ### `kevin_publish`
340
+
341
+ Regenerates the pull-channel bundles under `~/.opencode-kevin/` — `skills/project-knowledge.md` and `refs/<topic>.md` — reporting per-bundle outcome and the emission state (`on` / `off` / `unavailable`). Registration with the host happens at plugin startup; this tool only materializes and reports.
342
+
343
+ ---
344
+
345
+ ## How Kevin measures itself
346
+
347
+ ### Injection outcomes
348
+
349
+ Every injection is settled at `session.idle` into one of **four outcomes**:
350
+
351
+ | Outcome | Meaning | Counts toward precision? |
352
+ |---|---|---|
353
+ | `effective` | A linked fix was observed after the injection | yes (numerator) |
354
+ | `ineffective` | The same error recurred after the injection | yes (denominator) |
355
+ | `inconclusive` | Neither — the error did not recur, but no fix was seen either | no |
356
+ | `unmeasured` | Session went idle before settlement could run | no |
357
+
358
+ - **`precision_rate`** = `effective / (effective + ineffective)`. Measuring *effect*, not absence of recurrence: a lesson that was injected and never contradicted counts as `inconclusive`, not success. **Your precision rate will look lower than before v0.5.0. That is the honest number.**
359
+ - **`coverage_rate`** = `(effective + ineffective) / total` — the share of injections that were actually measured. Reported alongside precision so a low measurable fraction stays visible instead of hiding behind a large total.
360
+ - **`blocked`** counts every gate rejection by reason — `seen_this_session`, `weak`, `recurrence`, `stale`, `ignored`, `confidence` — a rejection you did not count did not happen.
361
+
362
+ ### The quality gate
363
+
364
+ Weak lessons — errors the reflector cannot dispatch to a deterministic rule — are **stored but never injected** while `quality_gate_enabled = '1'` (default). Recurrences demote lessons (`recurrence_count` → `stale`) and lower confidence. Debug mode: `kevin_config({ action: "set", key: "quality_gate_enabled", value: "0" })` re-injects weak lessons with a `(low confidence)` marker.
365
+
366
+ ### Seeing the whole picture
367
+
368
+ `kevin_trace` shows you the plan *before* it happens (dry run, zero side effects); `kevin_audit` reads the whole state after; `kevin_feedback` lets a human correct it — and the correction moves the confidence number `kevin_why` reports.
369
+
370
+ ---
371
+
372
+ ## Curation & Pull
373
+
374
+ ### The marker contract
375
+
376
+ Kevin never edits your files directly. Every artifact write happens inside a frozen marker block, delimited verbatim by:
377
+
378
+ ```
379
+ <!-- kevin:begin — curated by opencode-kevin, safe to edit -->
380
+ <!-- kevin:end -->
381
+ ```
382
+
383
+ These exact strings are **frozen for the v0.x line** README, tests and the v1.0.0 migration plan all depend on their byte sequences. What Kevin guarantees:
384
+
385
+ - **Only the block between the markers may change.** Bytes outside them are byte-identical after every write including line endings (a CRLF file stays CRLF everywhere, even inside the generated block), a leading UTF-8 BOM, and the file's final newline.
386
+ - **Malformed markers are refused, never repaired.** If the file contains a `begin` without an `end` (or vice versa), Kevin refuses the write with an explicit reason and the file is untouched. Repairing would mean guessing at user intent; refusing means the state stays visible and auditable.
387
+ - **Idempotent**: applying an unchanged plan is a counted `noop` — no temp file, no write, no mtime churn.
388
+
389
+ ### The propose → review → approve flow
390
+
391
+ ```
392
+ eligible memories (inferable != 1)
393
+
394
+
395
+ kevin_propose({ kind }) ── creates pending rows + unified diffs.
396
+ │ NO disk write, NO curated marks.
397
+
398
+ HUMAN REVIEWS THE DIFF ── this is the entire safety model:
399
+ │ a memory earns its way into a file
400
+ ▼ only after a human said yes.
401
+ kevin_approve({ proposal_id, decision })
402
+
403
+ ├── "approve" ── ArtifactWriter.apply() (the ONLY write path)
404
+ │ atomic temp+rename, audit row in artifact_writes,
405
+ │ memory marked curated, proposal marked applied
406
+ └── "reject" ── recorded, nothing touches disk
407
+ ```
408
+
409
+ Rejection history is never deleted: it is the evidence base for the roadmap's kill criterion "proposals rejected more often than approved".
410
+
411
+ ### Three distribution channels
412
+
413
+ | Channel | Artifact | Cost when unused |
414
+ |---|---|---|
415
+ | **Push** | per-prompt `<kevin-context>` injection | charges on every prompt — now capped at 400 tokens by default |
416
+ | **Pull — AGENTS.md** | marker block in the project's `AGENTS.md` | zero |
417
+ | **Pull — skills** | `~/.opencode-kevin/skills/project-knowledge.md` (`skill_emission_enabled`) | zero |
418
+ | **Pull references** | `~/.opencode-kevin/refs/<topic>.md` (`reference_emission_enabled`) | zero |
419
+
420
+ `kevin_audit`'s `channels` block compares push vs pull on the same axes, and reports each emission channel as `"on"`, `"off"` (setting `'0'` on a capable host) or `"unavailable"` (host without the v2 domain).
421
+
422
+ ### The confidence floor gate
423
+
424
+ `injection_confidence_floor` (default `'0.6'`) rejects memories whose computed confidence is below the floor, counted as `injections_blocked_confidence` — the sixth gate rejection reason, measured exactly like the first five. Single-observation memories (base confidence 0.5, no confirmed evidence) stop being pushed by default; `kevin_config({ action: "set", key: "injection_confidence_floor", value: "0" })` restores v0.5 behaviour exactly.
425
+
426
+ ---
427
+
428
+ ## Replay harness
429
+
430
+ `npm run replay` runs every transcript in `tests/replay/fixtures/` through the plugin against an in-memory database with a frozen clock and prints one table row per transcript (memories created, injection outcomes, `precision_rate`, `coverage_rate`, tokens). Record your own session as a JSON array of typed events (`session.created`, `chat.message`, `tool.before`, `tool.after`, `system.transform`, `compacting`, `session.idle`) with ISO-8601 `at` timestamps, drop it into `tests/replay/fixtures/`, and re-run. The `at` timestamps are the only source of time during replay.
431
+
432
+ ---
433
+
434
+ ## Hooks
435
+
436
+ Kevin subscribes to 6 OpenCode hooks:
437
+
438
+ | Hook | What Kevin does |
439
+ |---|---|
440
+ | `tool.execute.before` | Records tool call start (callID + redacted args) |
441
+ | `tool.execute.after` | Records result (id = callID); on failure → Reflector.invoke async (throttled); on success CausalChain links the fix |
442
+ | `experimental.chat.system.transform` | Injects relevant lessons in `<kevin-context>` (400 tokens by default, configurable) + optional `<kevin-suggestion>` |
443
+ | `experimental.session.compacting` | Re-injects lessons in `<kevin-memory>` after compacting (2000 tokens) + optional `<kevin-suggestion>` |
444
+ | `event` (`session.created`) | Captures current `sessionID` (skill/reference emissions register at plugin startup, not per session) |
445
+ | `event` (`session.idle`) | Settles injection outcomes; generates the retrospective; boosts positive lessons; penalizes recurring failures; promotes causal patterns and mines patterns (opt-in); drafts curation proposals (`curation_enabled`); flushes metrics |
446
+
447
+ **Redaction**: absolute paths (`C:\Users\...`, `/home/...`) `<path>` and secrets (`API_KEY=`, `Bearer`, `token`) `<redacted>` before persisting anything. `<private>…</private>` blocks are swept from tool call args and output before persistence and replaced with `<private: redacted N chars>`.
448
+
449
+ **Throttle**: Reflector generates at most 1 lesson per minute per unique fingerprint (per-fingerprint, not global). Configurable via `throttleMs`.
450
+
451
+ **Truncation**: content > 4KB keeps the lesson searchable; only the additional context is truncated (`metadata.truncated = true`).
452
+
453
+ ---
454
+
455
+ ## Configuration
456
+
457
+ ### Plugin options
458
+
459
+ Kevin accepts options via the plugin's tuple form (see Installation → Advanced). Programmatic defaults:
460
+
461
+ ```ts
462
+ import { KevinPlugin } from "@jmtrin/opencode-kevin";
463
+
464
+ // defaults
465
+ KevinPlugin(input, {
466
+ dbPath: "~/.opencode-kevin/kevin.db", // or ":memory:" for tests
467
+ migrationsDir: "<package>/dist/migrations", // resolved automatically
468
+ retrospectivesDir: "~/.opencode-kevin/retrospectives",
469
+ throttleMs: 60_000,
470
+ });
471
+ ```
472
+
473
+ ### Settings
474
+
475
+ Read/write via `kevin_config({ action: "list" | "set", ... })`. All values are TEXT; booleans compare against `"1"`.
476
+
477
+ | Setting | Default | Effect |
478
+ |---|---|---|
479
+ | `quality_gate_enabled` | `"1"` | Weak lessons are stored but never injected while enabled |
480
+ | `lesson_snippet_injection` | `"1"` | Injects the rescued errorType snippet with each lesson |
481
+ | `llm_reflection_enabled` | `"0"` | Opt-in LLM enrichment of reflector lessons |
482
+ | `cross_project_enabled` | `"0"` | `kevin_query` includes imported cross-project memories |
483
+ | `patternminer_enabled` | `"0"` | Opt-in deterministic 2-gram/3-gram pattern miner at `session.idle` |
484
+ | `tool_calls_dedup_enabled` | `"0"` | Opt-in dedup of repeated tool calls |
485
+ | `deterministic_retrieval` | `"0"` | Freezes Kevin's internal clock (recency factor 1.0, no relevance bumps) for hermetic tests and the replay harness |
486
+ | `pre_prompt_budget_tokens` | `"400"` | Pre-prompt injection cap, clamped to `[0, 4000]`; `0` turns push off |
487
+ | `archive_after_days` | `"30"` | Age at which stale non-pattern memories are retired to `archived` on `session.idle` |
488
+ | `curation_enabled` | `"1"` | Generates curation proposals at `session.idle` |
489
+ | `agents_md_path` | `"AGENTS.md"` | Where the AGENTS.md channel writes (project-relative) |
490
+ | `skill_emission_enabled` | `"0"` | Registers the curated skill with the host at startup (v2 hosts only) |
491
+ | `reference_emission_enabled` | `"0"` | Registers `@kevin/<topic>` references at startup (v2 hosts only) |
492
+ | `injection_confidence_floor` | `"0.6"` | Push gate: memories below this confidence are counted and rejected |
493
+
494
+ ---
495
+
496
+ ## Development
497
+
498
+ ```bash
499
+ git clone https://github.com/jmtrin/opencode-kevin.git
500
+ cd opencode-kevin
501
+ npm install
502
+ npm run typecheck # tsc --noEmit (strict)
503
+ npm run lint # biome check .
504
+ npm test # vitest run (unit + integration + e2e + replay)
505
+ npm run verify # post-install verification
506
+ npm run replay # replay report over tests/replay/fixtures
507
+ ```
508
+
509
+ ### Publishing (maintainer)
510
+
511
+ ```bash
512
+ npm login # as the jmtrin account that owns the @jmtrin scope
513
+ npm publish --access public
514
+ ```
515
+
516
+ `prepublishOnly` runs `npm run build` (tsc + copy migrations) automatically. The `files` field ships only `dist/plugin`, `dist/migrations`, and `migrations`. `dist/` is gitignored and rebuilt on publish.
517
+
518
+ ### Structure
519
+
520
+ ```
521
+ plugin/
522
+ index.ts # Entry point: KevinPlugin (wires hooks, tools, emissions)
523
+ Store.ts # SQLite wrapper (node:sqlite / bun:sqlite / better-sqlite3 fallback)
524
+ sqlite-adapter.ts # Runtime-agnostic SQLite adapter behind Store
525
+ Migrate.ts # Idempotent migrations + post-apply hooks
526
+ MemoryService.ts # save/query/getRelevant (FTS5 + bm25 + origin-aware rank + supersede)
527
+ ToolCallObserver.ts # onBefore/onAfter + redact + inferErrorType + dedup (opt-in)
528
+ Reflector.ts # Heuristic lessons + per-fingerprint throttle + LLM enrich (opt-in)
529
+ ContextInjector.ts # deriveQuery + pre-prompt/compacting injection + <kevin-suggestion>
530
+ Retrospective.ts # Generates retrospective.md + FP recap + metrics snapshot
531
+ Feedback.ts # kevin_feedback: verdicts, confidence terms, ignored stamp
532
+ Archiver.ts # Retires stale non-pattern memories past archive_after_days
533
+ CausalChain.ts # Links fixes to failures + promotes causal patterns
534
+ QualityGate.ts # Weak-lesson gate (stored, not injected by default)
535
+ InjectionLedger.ts # Injection ledger + settle → precision_rate
536
+ LessonFixer.ts # Deterministic fix_args capture + promotion enrichment
537
+ PatternMiner.ts # Opt-in deterministic 2-gram/3-gram miner
538
+ Curator.ts # Curation candidates + propose/approve lifecycle
539
+ ArtifactWriter.ts # The SINGLE write path (markers, atomic, noop, audit rows)
540
+ Materializer.ts # Pull-channel topic bundles (skills, refs)
541
+ inferability.ts # Deterministic inferable/non-inferable/unknown classifier
542
+ capabilities.ts # v2 domain probe (skills / references)
543
+ diff.ts # Minimal unified diff for proposal review
544
+ replay.ts # Hermetic replay driver over recorded transcripts
545
+ replay-types.ts # Transcript/result types for the replay harness
546
+ kevin_propose.ts # kevin_propose tool (strict dry run)
547
+ kevin_approve.ts # kevin_approve tool (only writer call site)
548
+ kevin_publish.ts # kevin_publish tool (bundle regeneration)
549
+ kevin_audit.ts # Read-only audit + channels/curation blocks
550
+ kevin_why.ts # kevin_why tool: failure→fix traces + related rules
551
+ okf-export.ts # kevin_export: OKF/markdown export
552
+ okf-import.ts # kevin_import: bundle parser + import
553
+ confidence.ts # Two-sided computeConfidence (evidence + recurrence + feedback)
554
+ query-tokenizer.ts # FTS5 tokenizer for query sanitization
555
+ memory-format.ts # escapeInjectedText, formatMemories, <protect> + id: line wrappers
556
+ redact.ts # redactPaths + stripPrivate
557
+ fingerprint.ts # FNV-1a 64-bit (in-house, no node:crypto)
558
+ metrics.ts # In-memory counters + debounced flush to kevin_metrics
559
+ uuid.ts # UUIDv7
560
+ migrations/
561
+ 001_initial.sql # schema: memories, tool_calls, retrospectives
562
+ 002_indexes.sql # FTS5 + indexes
563
+ 003_v02_signal.sql # fingerprint, origin, metrics, dedup indexes
564
+ 004_v03_knowledge.sql # evidence/status/supersede, error_fingerprint
565
+ 005_v04_signal.sql # recurrence_count, fix_args, last_injected_at
566
+ 006_v05_glassbox.sql # ignored/archived/superseded_by, feedback, metrics
567
+ 007_v06_pull.sql # curation_proposals, artifact_writes, curated/inferable
568
+ tests/
569
+ unit/ # component tests
570
+ integration/ # tool-level tests through real components
571
+ e2e/ # closed-loop tests through the host hooks
572
+ replay/ # transcript fixtures + replay harness tests
573
+ scripts/
574
+ copy-migrations.mjs # build step: copies *.sql to dist/migrations
575
+ verify-install.ts # npm run verify
576
+ ```
577
+
578
+ ---
579
+
580
+ ## License
581
+
582
+ MIT