@herjarsa/omo-meta-governor 0.18.0 → 0.19.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,565 +1,565 @@
1
- # @herjarsa/omo-meta-governor
2
-
3
- Self-judging agent orchestration layer for OpenCode. Observes tool executions,
4
- reads session state, scores progress, and dispatches decisions. Includes **15 custom tools**
5
- that the agent can invoke across CodeGraph, Graphify, AFT, AgentMemory, Magic Context, and SQLite.
6
-
7
- ## Install
8
-
9
- ```bash
10
- npm install @herjarsa/omo-meta-governor
11
- ```
12
-
13
- ## Usage
14
-
15
- Add as a plugin in your OpenCode config:
16
-
17
- ```jsonc
18
- {
19
- "plugins": ["@herjarsa/omo-meta-governor"]
20
- }
21
- ```
22
-
23
- The 15 custom tools register automatically (even without setting enabled:true).
24
- To also enable the governance pipeline (intervention, protocol enforcement):
25
-
26
- ```jsonc
27
- {
28
- "meta_governor": {
29
- "enabled": true,
30
- "intervention": {
31
- "mode": "message",
32
- "minActionForMessage": "warn"
33
- }
34
- }
35
- }
36
- ```
37
-
38
- ## 15 Custom Tools
39
-
40
- The plugin registers 15 tools the LLM can invoke. All available immediately on install.
41
-
42
- ### Code Search & Navigation
43
-
44
- | Tool | What it does | Use case |
45
- |------|-------------|----------|
46
- | `omo_search` | Semantic code search via codegraph/graphify with AFT fallback | Architecture questions, finding features — USE THIS FIRST |
47
- | `omo_find` | Exact symbol lookup (definition + direct callers) via codegraph node | "Find the function `validateToken`" |
48
- | `omo_impact` | Impact analysis: callers, transitive callers, test files, doc files | Run BEFORE modifying a function |
49
- | `omo_path` | Shortest conceptual path between two concepts via graphify | "How does auth connect to database?" |
50
- | `omo_explain` | Plain-language explanation of a concept via graphify | "What is the SwinTransformer?" |
51
- | `omo_outline` | Structural outline of files/directories via AFT | Understanding a new file's structure |
52
-
53
- ### Lesson & Memory
54
-
55
- | Tool | What it does | Use case |
56
- |------|-------------|----------|
57
- | `omo_recall` | Search past lessons via local SQLite FTS5 (fast, always available) | "How did we set up auth before?" |
58
- | `omo_recall_mcp` | Search cross-session memory via AgentMemory | "What did we learn about X in previous sessions?" |
59
- | `omo_remember` | Save a fact/observation to cross-session AgentMemory | "Remember this bug pattern for next time" |
60
-
61
- ### Rules & Notes
62
-
63
- | Tool | What it does | Use case |
64
- |------|-------------|----------|
65
- | `omo_rule` | Save a durable rule to Magic Context (ctx_memory) | "Always use bun:sqlite, not better-sqlite3" |
66
- | `omo_history` | Search git history + past messages via ctx_search | "When did we add this feature?" |
67
- | `omo_note` | Write ephemeral session note via ctx_note | "Currently debugging auth in module X" |
68
-
69
- ### Safety & Status
70
-
71
- | Tool | What it does | Use case |
72
- |------|-------------|----------|
73
- | `omo_checkpoint` | Create a named AFT snapshot before risky changes | Undo protection before refactoring |
74
- | `omo_undo` | Revert to most recent AFT checkpoint | "That broke things, revert it" |
75
- | `omo_health` | Show plugin runtime status: metrics, decisions, errors | "Is the plugin working?" |
76
-
77
- ## Health & Observability
78
-
79
- The plugin exposes a health JSON file at `~/.config/opencode/meta-governor-health.json`:
80
-
81
- ```bash
82
- cat ~/.config/opencode/meta-governor-health.json
83
- ```
84
-
85
- Or the agent can call `omo_health` directly to get a formatted report.
86
-
87
- Structured JSONL logs at `~/.config/opencode/meta-governor.log` with size-based rotation
88
- (10MB max, 5 rotated files).
89
-
90
- ## Persistence
91
-
92
- Lessons learned by the plugin persist in **SQLite** at `~/.omo-meta-governor/meta-governor.db`
93
- with full-text search (FTS5) for fast recall. Zero dependencies needed — uses Bun's built-in
94
- `bun:sqlite`.
95
-
96
- Optionally, the Opción A tools (`omo_remember`, `omo_recall_mcp`, `omo_rule`, `omo_history`,
97
- `omo_note`) can bridge to AgentMemory and Magic Context via `session.prompt()` — the LLM
98
- receives a structured instruction to call the appropriate MCP tool.
99
-
100
- ## Graph Sync (v0.11.0)
101
-
102
- MetaGovernor wires the plugin into the native git hooks of **codegraph** and
103
- **graphify** so each commit automatically reindexes both graphs.
104
-
105
- ### What it does on first load in a project
106
-
107
- 1. **Auto-install** codegraph via `npm i -D @colbymchenry/codegraph` and
108
- graphify via `pip install graphifyy` (falls back to `uv tool install
109
- graphifyy`) if they're not already on PATH.
110
- 2. **Run `codegraph init`** + **`graphify . --no-viz`** to build the initial
111
- indexes for the project.
112
- 3. **Run `graphify hook install`** to wire up the native `post-commit` and
113
- `post-checkout` git hooks.
114
-
115
- ### What it does on each `git commit`
116
-
117
- - **Primary path** (native git hook): `graphify update` runs in background.
118
- - **Backup path** (plugin's `tool.execute.after`): detects `git commit` in
119
- bash commands and runs `codegraph sync -q [path]`.
120
-
121
- ## Intervention
122
-
123
- MetaGovernor can inject governance decisions into the agent's context.
124
- Enabled when `meta_governor.enabled: true` in config.
125
-
126
- ### Modes
127
-
128
- | Mode | Mechanism | Effect |
129
- |------|-----------|--------|
130
- | `silent` | (none) | Decision is logged only |
131
- | `message` | `experimental.chat.messages.transform` | Injects a synthetic user message visible to the LLM |
132
- | `system` | `experimental.chat.system.transform` | Appends guidance to the system prompt |
133
-
134
- ### Configuration
135
-
136
- ```jsonc
137
- {
138
- "meta_governor": {
139
- "enabled": true,
140
- "intervention": {
141
- "mode": "message",
142
- "minActionForMessage": "warn",
143
- "maxInterventionsPerSession": 3,
144
- "respectDoneSignal": true,
145
- "phaseAwareDoneSignal": true // v0.15.0: multi-phase plan support
146
- }
147
- }
148
- }
149
- ```
150
-
151
- ### Fields
152
-
153
- | Field | Default | Description |
154
- |-------|---------|-------------|
155
- | `mode` | `"message"` | How to inject: `"silent"`, `"message"`, or `"system"` |
156
- | `minActionForMessage` | `"warn"` | Minimum action: `"warn"`, `"escalate"`, or `"stop"` |
157
- | `maxInterventionsPerSession` | `3` | Hard cap on injections per session |
158
- | `respectDoneSignal` | `true` | Stop injecting after terminal signal + Oracle verified |
159
- | `phaseAwareDoneSignal` | `false` | **v0.15.0**: when `true`, only `<promise>PLAN-COMPLETE</promise>` latches intervention. DONE/PHASE-N-COMPLETE are per-phase hints. Recommended for multi-phase plans. |
160
-
161
- ### Multi-phase plans (v0.15.0)
162
-
163
- For work plans with multiple phases (e.g. Sisyphus/Prometheus work plans),
164
- configure `phaseAwareDoneSignal: true` and emit `<promise>PLAN-COMPLETE</promise>`
165
- only when the **entire** plan is verified done by Oracle. The new markers:
166
-
167
- | Marker | Effect |
168
- |--------|--------|
169
- | `<promise>DONE</promise>` | Per-phase hint. Logged but does NOT latch intervention (when `phaseAwareDoneSignal: true`). |
170
- | `<promise>PHASE-N-COMPLETE</promise>` | Per-phase hint (e.g. `<promise>PHASE-1-COMPLETE</promise>`). Same as DONE — logged, does NOT latch. |
171
- | `<promise>PLAN-COMPLETE</promise>` | Terminal. Latches intervention when Oracle has verified. |
172
-
173
- **Migration**: existing v0.10.0–v0.14.x users keep working without changes (default
174
- `phaseAwareDoneSignal: false` preserves the legacy single-task behavior). Set the
175
- flag to `true` and switch your terminal marker to `PLAN-COMPLETE` to enable
176
- multi-phase governance.
177
-
178
- ## v0.16.0 — Audit remediation: memory hygiene, dead code, tool coverage, CI
179
-
180
- v0.16.0 closes the 50+ findings from the multi-front audit at `.omo/ulw-research/20260727-000530/plan-audit-v0.15.0.md`. The release is **additive in behavior, no breaking API changes** for users — only internal cleanup, dead code removal, and CI hardening.
181
-
182
- ### Highlights
183
-
184
- #### Memory hygiene (F1)
185
-
186
- - **`AuditStateCache`** (`src/audit-state-cache.ts`) — TTL+LRU bounded cache (100 entries, 1h TTL) replaces the bare `Map` that accumulated audit state without bounds. Stale sessions are evicted automatically.
187
- - **`TTLQueue`** (`src/ttl-queue.ts`) — TTL-based expiration for `pendingBotFeedback` and `pendingViolations` queues. Previously unbounded.
188
- - Removed dynamic `require("node:fs")` inside `shouldInjectPlanReminder` — replaced with static ESM imports (no more runtime module resolution failures).
189
-
190
- #### Dead code elimination (F2)
191
-
192
- - `takeAnyDecision()` — deprecated; removed from the active governance pipeline.
193
- - `systemInjection` — now awaited eagerly instead of fire-and-forget, eliminating a silent failure route.
194
- - `logToFile` in `graph-sync.ts` — wired to the real JSONL file logger (was a no-op stub).
195
- - Plugin version — derived from `package.json` at runtime instead of hardcoded "0.13.0" (closes the version-drift bug where `omo_health` reported stale versions).
196
-
197
- #### Tool bug fixes (F3)
198
-
199
- - **AFT checkpoint/undo**: args split on whitespace broke names with spaces. Rewrote arg construction with proper quoting.
200
- - **AFT subcommand**: now uses `options.projectDir` instead of `process.cwd()`.
201
- - **graphify binary override**: `omo_path` / `omo_explain` honored the `graphifyBin` option (was hardcoded).
202
- - **`as never` cast** on `setClient` → proper runtime guard that validates client shape.
203
- - **`session-bridge`**: replaced module-level `_client` with `AsyncLocalStorage` for per-request isolation. Concurrent sessions no longer race on the same client reference.
204
-
205
- #### Test coverage (F4)
206
-
207
- - 22 tests covering all 15 custom tools (`src/custom-tools.test.ts`). Previously the entire public tool surface had zero test coverage.
208
- - 12 tests for `decision-store` (previously untested).
209
-
210
- #### Type/token pipeline (F5)
211
-
212
- - `token-predictor` refactor: dead code (`delegate`/`switch-model`) removed; output is now informational-only as designed.
213
- - Type alignment across `types.ts`, `token-predictor.ts`, `orchestrator.ts`.
214
-
215
- #### CI matrix (F6)
216
-
217
- - `bun run typecheck` now runs on **macos-latest** and **windows-latest** (was Ubuntu-only).
218
- - Removed `package-lock.json` (bun project — canonical is `bun.lock`).
219
- - Secret redaction layer in `logToFile` (JWT, OpenAI keys, Bearer tokens, GitHub PATs, generic key:value patterns).
220
- - Implementation plan renamed `IMPLEMENTATION_PLAN.md` → `ARCHITECTURE.md`.
221
-
222
- #### Final refactors (F7)
223
-
224
- - Score formula documented (header doc with full formula spec).
225
- - **NaN guard** in `score()` — defaults to neutral continue when `iterationRatio` or `ambient.iteration/maxIterations` produce NaN.
226
- - `ACTION_SEVERITY` keyed by `DecisionHandlerOutput["action"]` union literal (was bare `Record<string, number>`).
227
- - `projectHasCodegraph` / `projectHasGraphify` IIFE booleans replaced with lookup-time calls to `graphRetrieval.hasCodegraphDir(cwd)`.
228
- - `extractConcepts` includes file basename for FTS lookup by tool/file name.
229
- - Backup graph-sync uses `triggerReindex` (was `triggerCodegraphSync`) — reindexes both codegraph AND graphify backends.
230
-
231
- ### Test & build status
232
-
233
- - **495/495 tests pass** (up from 487 in v0.15.0/0.15.1).
234
- - `bun run typecheck` clean.
235
- - `bun build.ts` clean (0.34 MB dist).
236
- - `npm pack --dry-run` validated (no forbidden artifacts).
237
-
238
- ### Migration
239
-
240
- No user action required. All changes are internal. The default `phaseAwareDoneSignal` is still `false` for backward compatibility; the v0.15.0 multi-phase behavior is preserved when explicitly enabled.
241
-
242
- ### Deferred to v0.17.0
243
-
244
- - F5.1 — wiring `escalate` action to a real dispatcher (Oracle is recommended but not yet wired).
245
- - F5.4 — `maxLessonsPerSession` enforcement (config field exists but is not enforced).
246
- - F3.6 — Bridge tools lying about delivery (5 tools still return "dispatched" without polling). Recommend the user explicitly request this if delivery verification is critical.
247
-
248
-
249
-
250
-
251
- ## v0.17.0 — Wire escalate to Oracle, enforce lesson cap, verify bridge delivery
252
-
253
- v0.17.0 closes the 3 deferred items from the v0.16.0 audit: **F5.1** (escalate → Oracle), **F5.4** (`maxLessonsPerSession` enforcement), and **F3.6** (bridge tool delivery verification).
254
-
255
- ### Highlights
256
-
257
- #### F5.1 — Escalate action now fires Oracle (v0.17.0)
258
-
259
- When the scoring engine produces an `escalate` action with target `oracle`, the plugin's `tool.execute.after` hook now fires a `session.prompt()` instructing the LLM to invoke `task(subagent_type=oracle)`. The prompt includes the decision reasoning, evidence count, and a verification pass directive. New `buildEscalationPrompt()` function in `session-bridge.ts` is the pure prompt builder (testable in isolation). User-targeted escalations get a separate prompt asking the LLM to summarize for human input.
260
-
261
- ```ts
262
- // Decision flow when score lands in escalate band:
263
- score ≤ -escalateThreshold (default -0.6)
264
- → decision.action = "escalate"
265
- → decision.shouldEscalateTo = "oracle" (or "user" for grave deviations)
266
- → plugin fires session.prompt with buildEscalationPrompt(...)
267
- → LLM invokes Oracle (or summarizes for user)
268
- → Oracle verifies → oracleInvoked=true → governance continues
269
- ```
270
-
271
- #### F5.4 — `maxLessonsPerSession` is now enforced
272
-
273
- The cap (default 20) was a config field that was never enforced. v0.17.0 adds:
274
- - `currentLessonCount` on `LearnFromOutcomeInput` and `MetaGovernorInput`
275
- - `lessonCount` tracked in per-session `AuditState`
276
- - `observeAndLearn()` short-circuits when `currentLessonCount >= maxLessonsPerSession`
277
- - The orchestrator increments `sessionState.lessonCount` after each successful save
278
- - **Cap semantics: inclusive** — when count equals cap, no more lessons are saved
279
-
280
- #### F3.6 — Bridge tool delivery verification
281
-
282
- The 5 bridge tools (`omo_remember`, `omo_recall_mcp`, `omo_rule`, `omo_history`, `omo_note`) previously returned "dispatched" after the `session.prompt()` was queued — without verifying the LLM actually called the MCP tool. v0.17.0 adds:
283
-
284
- - **New `PendingDeliveryRegistry` module** (`src/delivery-registry.ts`) — tracks pending dispatches per session with TTL-based cleanup.
285
- - **`tool.execute.after` hook** marks deliveries when a matching MCP tool call is observed.
286
- - **All 5 bridge tools** now report `deliveryStatus: "delivered" | "pending"` in their tool result and metadata, and briefly poll (1.5s) for fast deliveries.
287
- - When the LLM follows the prompt, the tool returns immediately with `"delivered"`. When it doesn't, the tool returns `"pending"` and the entry expires silently after 10s.
288
-
289
- ```ts
290
- // Bridge tool result metadata now includes:
291
- {
292
- tool: "omo_remember",
293
- ok: true,
294
- deliveryStatus: "delivered" | "pending",
295
- messageID: "...",
296
- durationMs: 1234,
297
- contentLength: 256
298
- }
299
- ```
300
-
301
- ### Test & build status
302
-
303
- - **514/514 tests pass** (up from 495 in v0.16.0 — 5 + 4 + 10 new tests across F5.4, F5.1, F3.6).
304
- - `bun run typecheck` clean.
305
- - `bun build.ts` clean (0.34 MB dist).
306
- - `npm pack --dry-run` validated.
307
-
308
- ### Migration
309
-
310
- No user action required. All changes are internal or additive:
311
- - `deliveryStatus` is an additive metadata field — existing consumers ignore it.
312
- - `maxLessonsPerSession` is now actually enforced — if you have sessions that previously saved more than 20 lessons (e.g. from before the cap was added), this may surprise you. Bump the cap in your config if needed.
313
- - `escalate` action now actively fires Oracle — this is the first version where Oracle is auto-invoked, not just manually invoked by the LLM.
314
-
315
- ### Audit roadmap (status as of v0.17.0)
316
-
317
- | Release | Status | Scope |
318
- |---------|--------|-------|
319
- | v0.15.1 (F0) | ✅ Shipped | Hotfix self-dep + npm pack gate |
320
- | v0.16.0 (F1-F7) | ✅ Shipped | Memory hygiene, dead code, tool coverage, CI |
321
- | v0.17.0 (deferred) | ✅ Shipped | F5.1 escalate, F5.4 cap, F3.6 delivery verify |
322
-
323
- All audit findings are now closed. Future work focuses on new features and user-driven feedback.
324
-
325
-
326
-
327
- ## v0.17.1 — Audit args fix (patch)
328
-
329
- v0.17.1 is a single-bug patch release. The fix addresses an issue discovered during v0.17.0 verification:
330
-
331
- ### The bug
332
-
333
- The `tool.execute.before` hook passed an empty `{}` object as the second argument to `auditToolCall()`. This meant the audit function never saw the tool's args (e.g. file content for write tools) and could never detect:
334
-
335
- - `@ts-ignore` / `@ts-expect-error` directives
336
- - `as any` type assertions
337
- - `catch(e) {}` empty catch blocks
338
-
339
- The hook signature was also incomplete — it didn't receive the `output` parameter that contains the mutable args, even though the SDK provides it.
340
-
341
- ### The fix
342
-
343
- Two changes in `src/plugin.ts`:
344
-
345
- 1. **Hook signature updated** to receive the `output` parameter:
346
- ```ts
347
- "tool.execute.before": async (
348
- toolInput: { tool: string; sessionID: string; callID: string },
349
- _output: { args: unknown },
350
- ): Promise<void> => {
351
- ```
352
-
353
- 2. **Audit call** now passes `_output.args` instead of `{}`:
354
- ```ts
355
- const violations = auditToolCall(toolInput.tool, _output.args, { ... })
356
- ```
357
-
358
- ### Tests
359
-
360
- Added 4 new tests in `src/plugin.test.ts`:
361
- - `@ts-ignore + as any` in args → `no-type-suppression` violation detected and injected
362
- - `catch(e) {}` in args → `no-empty-catch` violation detected and injected
363
- - Clean code → no violation injected (false-positive guard)
364
- - `auditToolCalls: false` → audit short-circuits (regression check)
365
-
366
- ### Test & build status
367
-
368
- - **518/518 tests pass** (up from 514 in v0.17.0 — 4 new audit tests).
369
- - `bun run typecheck` clean.
370
- - `bun build.ts` clean (0.34 MB dist).
371
-
372
- ### Migration
373
-
374
- No user action required. The audit detection now correctly fires when the agent writes forbidden patterns. This means:
375
-
376
- - **If your agent previously wrote `@ts-ignore` without being flagged**: it will now be flagged with `[GRAVE] no-type-suppression: ...` injected as a synthetic user message.
377
- - **If you want to disable the audit**: set `protocolEnforcement.auditToolCalls: false` (already supported).
378
-
379
-
380
-
381
- ## v0.17.2 — Fix escalation dead code + 4 audit gaps
382
-
383
- v0.17.2 closes 4 gaps discovered during live verification of v0.17.0/v0.17.1. The most important: F5.1 (escalate → Oracle) was effectively dead in production due to two compounding bugs.
384
-
385
- ### Highlights
386
-
387
- #### Gap C (CRITICAL) — Escalation now actually fires
388
-
389
- The score formula's `noProgress` and `deviations` inputs were hardcoded as `false` and `[]` in the plugin. This meant the `no-progress-detector` (weight 0.20) and `deviation-detector` (weight 0.20) signals always contributed 0. Combined with default thresholds, the maximum possible score was -0.55 — never reaching `escalateThreshold: 0.6` or `stopThreshold: 0.8`.
390
-
391
- **Fix:**
392
- 1. **Derive `noProgress`** from the recent tool call window. If the last 5 tool calls contain no `write`/`edit`/`task` (i.e. the agent is only reading/grepping without producing artifacts), `noProgress = true`.
393
- 2. **Derive `deviations`** from accumulated protocol violations. The audit hook now stores violations in `state.accumulatedDeviations` (capped at 5 per session); the orchestrator input reads them.
394
- 3. **Lower default thresholds** to match the new worst-case math:
395
- - `escalateThreshold`: 0.6 → 0.45
396
- - `stopThreshold`: 0.8 → 0.55
397
-
398
- Now worst-case state (no oracle, no progress, 2 grave deviations, iteration at limit, stop-advice lessons) produces score ≈ -0.55 → `stop` action fires.
399
-
400
- #### Gap Q (HIGH) — File paths threaded through pipeline
401
-
402
- `orchestrator.ts` was hardcoding `filesChanged: []` instead of `input.filePaths`. This meant lesson extraction never saw the actual changed files, so F7.5's file-basename FTS indexing was empty.
403
-
404
- **Fix:**
405
- 1. Track `recentWriteFilePaths` in AuditState (alongside existing `recentWriteContents`).
406
- 2. Capture `filePath` from `toolInput.args` on write/edit tool calls.
407
- 3. New `MetaGovernorInput.filePaths?: readonly string[]` passed through to `observeAndLearn`.
408
-
409
- #### Gap D (HIGH) — Three config fields now actually do something
410
-
411
- Three fields were in the schema and config projection but NEVER consulted by the logic:
412
-
413
- - `closedLoop.saveLessons` — parallel to `saveDecisions`. When `false`, lessons are skipped (decision records still save).
414
- - `intervention.includeDecisionHistory` — when `true`, `messages.transform` prepends recent intervention texts (capped at `maxHistoryMessages`) so the LLM sees its history of decisions.
415
- - `intervention.maxHistoryMessages` — limit for the above (default 5).
416
-
417
- **Fix:** All three fields now control behavior. Track `recentInterventionTexts` in AuditState, format them into the injection text.
418
-
419
- #### Bonus — iteration-budget signal wired (Oracle finding)
420
-
421
- Oracle flagged a pre-existing gap alongside Gap C: `iteration` was hardcoded `0` in the orchestrator input, making the `iteration-budget` signal (weight 0.15) effectively dead.
422
-
423
- Fix:
424
- - Added `iteration: number` to `AuditState`, incremented per tool call.
425
- - Threaded `iteration: sessionState?.iteration ?? 0` into `MetaGovernorInput`.
426
- - `maxIterations` now reads from config instead of being hardcoded.
427
-
428
- Worst-case score math updated: with iteration at 100% (-0.12), all signals bad, no oracle → score = -0.65 → `stop` action fires.
429
-
430
- #### Gap I (MEDIUM) — `verifyDelivery` return type includes "expired"
431
-
432
- The TypeScript signature was `Promise<"delivered" | "pending">` but the registry could return `"expired"`. The expired case leaked through as `"pending"` silently.
433
-
434
- **Fix:** Signature updated to `Promise<"delivered" | "pending" | "expired">`. Bridge tools now distinguish: `"delivered"` (verified), `"pending"` (still polling), `"expired"` (TTL elapsed).
435
-
436
- ### Test & build status
437
-
438
- - **521/521 tests pass** (up from 518 — 3 new tests for the v0.17.2 fixes).
439
- - `bun run typecheck` clean.
440
- - `bun build.ts` clean (0.34 MB dist).
441
- - `npm pack --dry-run` validated.
442
-
443
- ### Migration
444
-
445
- No user action required. Two behavior changes:
446
-
447
- 1. **Escalation now fires more aggressively.** If your agent has been producing violations and not making progress, expect to see escalate → Oracle prompts more often. This is the intended behavior; v0.17.0 was incorrectly silent.
448
- 2. **`includeDecisionHistory` and `maxHistoryMessages` are now functional.** If you set them in v0.17.0 expecting them to work, they will now actually take effect.
449
-
450
- ### Audit roadmap (status as of v0.17.2)
451
-
452
- | Release | Status | Scope |
453
- |---------|--------|-------|
454
- | v0.15.1 (F0) | ✅ | Hotfix self-dep |
455
- | v0.16.0 (F1-F7) | ✅ | Memory hygiene, dead code, tool coverage, CI |
456
- | v0.17.0 | ✅ | F5.1 escalate, F5.4 cap, F3.6 delivery verify |
457
- | v0.17.1 | ✅ | Audit args fix |
458
- | v0.17.2 | ✅ | Gap C (escalation live), Q (file paths), D (config fields), I (delivery expired) |
459
-
460
-
461
-
462
- ## v0.17.3 — Fix Gap I properly (patch)
463
-
464
- v0.17.3 is a single-bug patch. During live verification of v0.17.2, Gap I was found to be incompletely fixed.
465
-
466
- ### The bug (v0.17.2 cosmetic fix)
467
-
468
- The `verifyDelivery` export signature was widened to include `"expired"` in v0.17.2, and the bridge tools' title/output text was updated to handle it. **BUT the underlying `pollForDelivery` helper was still collapsing `"expired"` → `"pending"` silently:**
469
-
470
- ```ts
471
- // v0.17.2 (BUG):
472
- return status === "delivered" ? "delivered" : "pending"
473
- ```
474
-
475
- So bridge tools could never report `"expired"` to the user, even though the registry correctly tracked it. Live verification confirmed: `deliveryStatus` always showed `"pending"`.
476
-
477
- ### The fix (v0.17.3)
478
-
479
- ```ts
480
- // v0.17.3:
481
- return await pendingRegistryRef.awaitDelivery({ sessionID, mcpTool, timeoutMs })
482
- ```
483
-
484
- Now the actual status from the registry propagates through. `"expired"` flows end-to-end to the bridge tool's `metadata.deliveryStatus` and title.
485
-
486
- ### Tests
487
-
488
- Added 3 RED tests in `src/custom-tools.test.ts`:
489
- - Returns `"expired"` when registry entry exists past timeout (real registry instance)
490
- - Returns `"delivered"` when `markDelivered` fires before timeout
491
- - Returns `"pending"` when no registry is configured
492
-
493
- ### Test & build status
494
-
495
- - **525/525 tests pass** (up from 522 in v0.17.2 — 3 new tests for pollForDelivery).
496
- - `bun run typecheck` clean.
497
- - `bun build.ts` clean (0.34 MB dist).
498
-
499
- ### Migration
500
-
501
- No user action required. Bridge tools will now correctly distinguish all three delivery states:
502
- - `"delivered"` — LLM's MCP tool call was observed within 1.5s
503
- - `"expired"` — TTL elapsed without delivery (entry expires after 10s, but bridge tool sees this immediately as "expired" when polling times out at 1.5s)
504
- - `"pending"` — no registry configured (graceful degradation for tests/mocks)
505
-
506
- ### Audit roadmap (status as of v0.17.3)
507
-
508
- | Release | Status | Scope |
509
- |---------|--------|-------|
510
- | v0.15.1 → v0.17.2 | ✅ | All audit findings + 5 gap fixes |
511
- | v0.17.3 | ✅ | Gap I real fix (pollForDelivery returns "expired") |
512
-
513
- Two remaining gaps documented but require SDK support to fix:
514
- - `recentTurnTokens: []` — token-predictor signal dead (10% of score); needs per-turn token counts from OpenCode SDK
515
- - `agentName` defaults to `"unknown"` — cosmetic, no functional impact
516
-
517
-
518
-
519
- ## v0.18.0 — Audit remediation: 7+ silent config drops + circular ref crash
520
-
521
- v0.18.0 is a thorough-audit patch release. Each fix addresses a bug found by testing every public function with edge cases and adversarial inputs.
522
-
523
- ### Highlights
524
-
525
- | # | Bug | Severity | Fix |
526
- |---|-----|----------|-----|
527
- | 1 | `file-logger.redactData` crashed on circular references with stack overflow | 🔴 CRITICAL | `WeakSet` guard + `try/catch` fallback |
528
- | 2 | `loadOrchestratorConfig` only projected `closedLoop.saveDecisions` — `enabled`, `minSeverityToLearn`, `maxLessonsPerSession`, `saveLessons` were silently dropped | 🔴 CRITICAL | Project all 5 fields |
529
- | 3 | `loadOrchestratorConfig` didn't project `decision.warnMessageTemplate`, `escalateMessageTemplate`, `stopMessageTemplate` | 🟠 HIGH | Project all 3 templates |
530
- | 4 | `loadOrchestratorConfig` didn't project `scoring.paralysisThreshold`, `defaultEscalationTarget` | 🟠 HIGH | Project all fields |
531
- | 5 | `loadOrchestratorConfig` had `memory.timeoutMs` field name mismatch (schema said `agentmemoryTimeoutMs`) | 🟠 HIGH | Accept both names |
532
- | 6 | `isMetaGovernorEnabled` only checked top-level `enabled`, not `meta_governor.enabled` (wrapped shape from `opencode.jsonc`) | 🟠 HIGH | Check both shapes |
533
- | 7 | `createMetricsCollector` crashed when called without config (`config.version` on `undefined`) | 🟠 HIGH | Accept `Partial<MetricsCollectorConfig>` |
534
- | 8 | `metrics.inc` crashed on unknown event names (`bucket.count++` on `undefined`) | 🟠 HIGH | Guard `if (!bucket) return` |
535
- | 9 | `isNewerVersion` returned `false` for `installed=null` (no upgrade triggered for fresh installs) | 🟡 MEDIUM | Return `true` when installed is null AND latest is valid |
536
-
537
- ### Test & build status
538
-
539
- - **557/557 tests pass** (up from 530 in v0.17.3 — 27 new tests for the audit fixes).
540
- - `bun run typecheck` clean.
541
- - `bun build.ts` clean (0.34 MB dist).
542
- - `npm pack --dry-run` validated.
543
-
544
- ### Migration
545
-
546
- No user action required. The fix to `loadOrchestratorConfig` means **users who were setting `closedLoop.maxLessonsPerSession` or other previously-dropped fields will now see those values actually take effect**. If you had a config like `{ "closedLoop": { "maxLessonsPerSession": 50 } }` before v0.18.0, it was silently being overridden to 20. Starting v0.18.0, the value 50 is now respected.
547
-
548
- ### Audit roadmap (status as of v0.18.0)
549
-
550
- | Release | Status | Scope |
551
- |---------|--------|-------|
552
- | v0.15.1 → v0.17.3 | ✅ | All audit findings + deferred items + audit args fix + gap fixes |
553
- | v0.18.0 | ✅ | 7 silent config drops + circular ref crash + metrics crashes + upgrade trigger |
554
-
555
- This release closes the final round of gaps found by a thorough function-by-function audit. The plugin now correctly projects **all** user configuration, handles **all** circular reference cases, and fails safely on **all** missing-input scenarios.
556
-
557
- ## Auto-upgrade (v0.12.0)
558
-
559
- On plugin load, queries npm/pip registries to check whether newer versions
560
- of **codegraph** or **graphify** exist. Config: `graphSync.autoUpgrade` (default `true`),
561
- `graphSync.upgradeCheckTtlMs` (default `86400000`).
562
-
563
- ## License
564
-
565
- MIT
1
+ # @herjarsa/omo-meta-governor
2
+
3
+ Self-judging agent orchestration layer for OpenCode. Observes tool executions,
4
+ reads session state, scores progress, and dispatches decisions. Includes **15 custom tools**
5
+ that the agent can invoke across CodeGraph, Graphify, AFT, AgentMemory, Magic Context, and SQLite.
6
+
7
+ ## Install
8
+
9
+ ```bash
10
+ npm install @herjarsa/omo-meta-governor
11
+ ```
12
+
13
+ ## Usage
14
+
15
+ Add as a plugin in your OpenCode config:
16
+
17
+ ```jsonc
18
+ {
19
+ "plugins": ["@herjarsa/omo-meta-governor"]
20
+ }
21
+ ```
22
+
23
+ The 15 custom tools register automatically (even without setting enabled:true).
24
+ To also enable the governance pipeline (intervention, protocol enforcement):
25
+
26
+ ```jsonc
27
+ {
28
+ "meta_governor": {
29
+ "enabled": true,
30
+ "intervention": {
31
+ "mode": "message",
32
+ "minActionForMessage": "warn"
33
+ }
34
+ }
35
+ }
36
+ ```
37
+
38
+ ## 15 Custom Tools
39
+
40
+ The plugin registers 15 tools the LLM can invoke. All available immediately on install.
41
+
42
+ ### Code Search & Navigation
43
+
44
+ | Tool | What it does | Use case |
45
+ |------|-------------|----------|
46
+ | `omo_search` | Semantic code search via codegraph/graphify with AFT fallback | Architecture questions, finding features — USE THIS FIRST |
47
+ | `omo_find` | Exact symbol lookup (definition + direct callers) via codegraph node | "Find the function `validateToken`" |
48
+ | `omo_impact` | Impact analysis: callers, transitive callers, test files, doc files | Run BEFORE modifying a function |
49
+ | `omo_path` | Shortest conceptual path between two concepts via graphify | "How does auth connect to database?" |
50
+ | `omo_explain` | Plain-language explanation of a concept via graphify | "What is the SwinTransformer?" |
51
+ | `omo_outline` | Structural outline of files/directories via AFT | Understanding a new file's structure |
52
+
53
+ ### Lesson & Memory
54
+
55
+ | Tool | What it does | Use case |
56
+ |------|-------------|----------|
57
+ | `omo_recall` | Search past lessons via local SQLite FTS5 (fast, always available) | "How did we set up auth before?" |
58
+ | `omo_recall_mcp` | Search cross-session memory via AgentMemory | "What did we learn about X in previous sessions?" |
59
+ | `omo_remember` | Save a fact/observation to cross-session AgentMemory | "Remember this bug pattern for next time" |
60
+
61
+ ### Rules & Notes
62
+
63
+ | Tool | What it does | Use case |
64
+ |------|-------------|----------|
65
+ | `omo_rule` | Save a durable rule to Magic Context (ctx_memory) | "Always use bun:sqlite, not better-sqlite3" |
66
+ | `omo_history` | Search git history + past messages via ctx_search | "When did we add this feature?" |
67
+ | `omo_note` | Write ephemeral session note via ctx_note | "Currently debugging auth in module X" |
68
+
69
+ ### Safety & Status
70
+
71
+ | Tool | What it does | Use case |
72
+ |------|-------------|----------|
73
+ | `omo_checkpoint` | Create a named AFT snapshot before risky changes | Undo protection before refactoring |
74
+ | `omo_undo` | Revert to most recent AFT checkpoint | "That broke things, revert it" |
75
+ | `omo_health` | Show plugin runtime status: metrics, decisions, errors | "Is the plugin working?" |
76
+
77
+ ## Health & Observability
78
+
79
+ The plugin exposes a health JSON file at `~/.config/opencode/meta-governor-health.json`:
80
+
81
+ ```bash
82
+ cat ~/.config/opencode/meta-governor-health.json
83
+ ```
84
+
85
+ Or the agent can call `omo_health` directly to get a formatted report.
86
+
87
+ Structured JSONL logs at `~/.config/opencode/meta-governor.log` with size-based rotation
88
+ (10MB max, 5 rotated files).
89
+
90
+ ## Persistence
91
+
92
+ Lessons learned by the plugin persist in **SQLite** at `~/.omo-meta-governor/meta-governor.db`
93
+ with full-text search (FTS5) for fast recall. Zero dependencies needed — uses Bun's built-in
94
+ `bun:sqlite`.
95
+
96
+ Optionally, the Opción A tools (`omo_remember`, `omo_recall_mcp`, `omo_rule`, `omo_history`,
97
+ `omo_note`) can bridge to AgentMemory and Magic Context via `session.prompt()` — the LLM
98
+ receives a structured instruction to call the appropriate MCP tool.
99
+
100
+ ## Graph Sync (v0.11.0)
101
+
102
+ MetaGovernor wires the plugin into the native git hooks of **codegraph** and
103
+ **graphify** so each commit automatically reindexes both graphs.
104
+
105
+ ### What it does on first load in a project
106
+
107
+ 1. **Auto-install** codegraph via `npm i -D @colbymchenry/codegraph` and
108
+ graphify via `pip install graphifyy` (falls back to `uv tool install
109
+ graphifyy`) if they're not already on PATH.
110
+ 2. **Run `codegraph init`** + **`graphify . --no-viz`** to build the initial
111
+ indexes for the project.
112
+ 3. **Run `graphify hook install`** to wire up the native `post-commit` and
113
+ `post-checkout` git hooks.
114
+
115
+ ### What it does on each `git commit`
116
+
117
+ - **Primary path** (native git hook): `graphify update` runs in background.
118
+ - **Backup path** (plugin's `tool.execute.after`): detects `git commit` in
119
+ bash commands and runs `codegraph sync -q [path]`.
120
+
121
+ ## Intervention
122
+
123
+ MetaGovernor can inject governance decisions into the agent's context.
124
+ Enabled when `meta_governor.enabled: true` in config.
125
+
126
+ ### Modes
127
+
128
+ | Mode | Mechanism | Effect |
129
+ |------|-----------|--------|
130
+ | `silent` | (none) | Decision is logged only |
131
+ | `message` | `experimental.chat.messages.transform` | Injects a synthetic user message visible to the LLM |
132
+ | `system` | `experimental.chat.system.transform` | Appends guidance to the system prompt |
133
+
134
+ ### Configuration
135
+
136
+ ```jsonc
137
+ {
138
+ "meta_governor": {
139
+ "enabled": true,
140
+ "intervention": {
141
+ "mode": "message",
142
+ "minActionForMessage": "warn",
143
+ "maxInterventionsPerSession": 3,
144
+ "respectDoneSignal": true,
145
+ "phaseAwareDoneSignal": true // v0.15.0: multi-phase plan support
146
+ }
147
+ }
148
+ }
149
+ ```
150
+
151
+ ### Fields
152
+
153
+ | Field | Default | Description |
154
+ |-------|---------|-------------|
155
+ | `mode` | `"message"` | How to inject: `"silent"`, `"message"`, or `"system"` |
156
+ | `minActionForMessage` | `"warn"` | Minimum action: `"warn"`, `"escalate"`, or `"stop"` |
157
+ | `maxInterventionsPerSession` | `3` | Hard cap on injections per session |
158
+ | `respectDoneSignal` | `true` | Stop injecting after terminal signal + Oracle verified |
159
+ | `phaseAwareDoneSignal` | `false` | **v0.15.0**: when `true`, only `<promise>PLAN-COMPLETE</promise>` latches intervention. DONE/PHASE-N-COMPLETE are per-phase hints. Recommended for multi-phase plans. |
160
+
161
+ ### Multi-phase plans (v0.15.0)
162
+
163
+ For work plans with multiple phases (e.g. Sisyphus/Prometheus work plans),
164
+ configure `phaseAwareDoneSignal: true` and emit `<promise>PLAN-COMPLETE</promise>`
165
+ only when the **entire** plan is verified done by Oracle. The new markers:
166
+
167
+ | Marker | Effect |
168
+ |--------|--------|
169
+ | `<promise>DONE</promise>` | Per-phase hint. Logged but does NOT latch intervention (when `phaseAwareDoneSignal: true`). |
170
+ | `<promise>PHASE-N-COMPLETE</promise>` | Per-phase hint (e.g. `<promise>PHASE-1-COMPLETE</promise>`). Same as DONE — logged, does NOT latch. |
171
+ | `<promise>PLAN-COMPLETE</promise>` | Terminal. Latches intervention when Oracle has verified. |
172
+
173
+ **Migration**: existing v0.10.0–v0.14.x users keep working without changes (default
174
+ `phaseAwareDoneSignal: false` preserves the legacy single-task behavior). Set the
175
+ flag to `true` and switch your terminal marker to `PLAN-COMPLETE` to enable
176
+ multi-phase governance.
177
+
178
+ ## v0.16.0 — Audit remediation: memory hygiene, dead code, tool coverage, CI
179
+
180
+ v0.16.0 closes the 50+ findings from the multi-front audit at `.omo/ulw-research/20260727-000530/plan-audit-v0.15.0.md`. The release is **additive in behavior, no breaking API changes** for users — only internal cleanup, dead code removal, and CI hardening.
181
+
182
+ ### Highlights
183
+
184
+ #### Memory hygiene (F1)
185
+
186
+ - **`AuditStateCache`** (`src/audit-state-cache.ts`) — TTL+LRU bounded cache (100 entries, 1h TTL) replaces the bare `Map` that accumulated audit state without bounds. Stale sessions are evicted automatically.
187
+ - **`TTLQueue`** (`src/ttl-queue.ts`) — TTL-based expiration for `pendingBotFeedback` and `pendingViolations` queues. Previously unbounded.
188
+ - Removed dynamic `require("node:fs")` inside `shouldInjectPlanReminder` — replaced with static ESM imports (no more runtime module resolution failures).
189
+
190
+ #### Dead code elimination (F2)
191
+
192
+ - `takeAnyDecision()` — deprecated; removed from the active governance pipeline.
193
+ - `systemInjection` — now awaited eagerly instead of fire-and-forget, eliminating a silent failure route.
194
+ - `logToFile` in `graph-sync.ts` — wired to the real JSONL file logger (was a no-op stub).
195
+ - Plugin version — derived from `package.json` at runtime instead of hardcoded "0.13.0" (closes the version-drift bug where `omo_health` reported stale versions).
196
+
197
+ #### Tool bug fixes (F3)
198
+
199
+ - **AFT checkpoint/undo**: args split on whitespace broke names with spaces. Rewrote arg construction with proper quoting.
200
+ - **AFT subcommand**: now uses `options.projectDir` instead of `process.cwd()`.
201
+ - **graphify binary override**: `omo_path` / `omo_explain` honored the `graphifyBin` option (was hardcoded).
202
+ - **`as never` cast** on `setClient` → proper runtime guard that validates client shape.
203
+ - **`session-bridge`**: replaced module-level `_client` with `AsyncLocalStorage` for per-request isolation. Concurrent sessions no longer race on the same client reference.
204
+
205
+ #### Test coverage (F4)
206
+
207
+ - 22 tests covering all 15 custom tools (`src/custom-tools.test.ts`). Previously the entire public tool surface had zero test coverage.
208
+ - 12 tests for `decision-store` (previously untested).
209
+
210
+ #### Type/token pipeline (F5)
211
+
212
+ - `token-predictor` refactor: dead code (`delegate`/`switch-model`) removed; output is now informational-only as designed.
213
+ - Type alignment across `types.ts`, `token-predictor.ts`, `orchestrator.ts`.
214
+
215
+ #### CI matrix (F6)
216
+
217
+ - `bun run typecheck` now runs on **macos-latest** and **windows-latest** (was Ubuntu-only).
218
+ - Removed `package-lock.json` (bun project — canonical is `bun.lock`).
219
+ - Secret redaction layer in `logToFile` (JWT, OpenAI keys, Bearer tokens, GitHub PATs, generic key:value patterns).
220
+ - Implementation plan renamed `IMPLEMENTATION_PLAN.md` → `ARCHITECTURE.md`.
221
+
222
+ #### Final refactors (F7)
223
+
224
+ - Score formula documented (header doc with full formula spec).
225
+ - **NaN guard** in `score()` — defaults to neutral continue when `iterationRatio` or `ambient.iteration/maxIterations` produce NaN.
226
+ - `ACTION_SEVERITY` keyed by `DecisionHandlerOutput["action"]` union literal (was bare `Record<string, number>`).
227
+ - `projectHasCodegraph` / `projectHasGraphify` IIFE booleans replaced with lookup-time calls to `graphRetrieval.hasCodegraphDir(cwd)`.
228
+ - `extractConcepts` includes file basename for FTS lookup by tool/file name.
229
+ - Backup graph-sync uses `triggerReindex` (was `triggerCodegraphSync`) — reindexes both codegraph AND graphify backends.
230
+
231
+ ### Test & build status
232
+
233
+ - **495/495 tests pass** (up from 487 in v0.15.0/0.15.1).
234
+ - `bun run typecheck` clean.
235
+ - `bun build.ts` clean (0.34 MB dist).
236
+ - `npm pack --dry-run` validated (no forbidden artifacts).
237
+
238
+ ### Migration
239
+
240
+ No user action required. All changes are internal. The default `phaseAwareDoneSignal` is still `false` for backward compatibility; the v0.15.0 multi-phase behavior is preserved when explicitly enabled.
241
+
242
+ ### Deferred to v0.17.0
243
+
244
+ - F5.1 — wiring `escalate` action to a real dispatcher (Oracle is recommended but not yet wired).
245
+ - F5.4 — `maxLessonsPerSession` enforcement (config field exists but is not enforced).
246
+ - F3.6 — Bridge tools lying about delivery (5 tools still return "dispatched" without polling). Recommend the user explicitly request this if delivery verification is critical.
247
+
248
+
249
+
250
+
251
+ ## v0.17.0 — Wire escalate to Oracle, enforce lesson cap, verify bridge delivery
252
+
253
+ v0.17.0 closes the 3 deferred items from the v0.16.0 audit: **F5.1** (escalate → Oracle), **F5.4** (`maxLessonsPerSession` enforcement), and **F3.6** (bridge tool delivery verification).
254
+
255
+ ### Highlights
256
+
257
+ #### F5.1 — Escalate action now fires Oracle (v0.17.0)
258
+
259
+ When the scoring engine produces an `escalate` action with target `oracle`, the plugin's `tool.execute.after` hook now fires a `session.prompt()` instructing the LLM to invoke `task(subagent_type=oracle)`. The prompt includes the decision reasoning, evidence count, and a verification pass directive. New `buildEscalationPrompt()` function in `session-bridge.ts` is the pure prompt builder (testable in isolation). User-targeted escalations get a separate prompt asking the LLM to summarize for human input.
260
+
261
+ ```ts
262
+ // Decision flow when score lands in escalate band:
263
+ score ≤ -escalateThreshold (default -0.6)
264
+ → decision.action = "escalate"
265
+ → decision.shouldEscalateTo = "oracle" (or "user" for grave deviations)
266
+ → plugin fires session.prompt with buildEscalationPrompt(...)
267
+ → LLM invokes Oracle (or summarizes for user)
268
+ → Oracle verifies → oracleInvoked=true → governance continues
269
+ ```
270
+
271
+ #### F5.4 — `maxLessonsPerSession` is now enforced
272
+
273
+ The cap (default 20) was a config field that was never enforced. v0.17.0 adds:
274
+ - `currentLessonCount` on `LearnFromOutcomeInput` and `MetaGovernorInput`
275
+ - `lessonCount` tracked in per-session `AuditState`
276
+ - `observeAndLearn()` short-circuits when `currentLessonCount >= maxLessonsPerSession`
277
+ - The orchestrator increments `sessionState.lessonCount` after each successful save
278
+ - **Cap semantics: inclusive** — when count equals cap, no more lessons are saved
279
+
280
+ #### F3.6 — Bridge tool delivery verification
281
+
282
+ The 5 bridge tools (`omo_remember`, `omo_recall_mcp`, `omo_rule`, `omo_history`, `omo_note`) previously returned "dispatched" after the `session.prompt()` was queued — without verifying the LLM actually called the MCP tool. v0.17.0 adds:
283
+
284
+ - **New `PendingDeliveryRegistry` module** (`src/delivery-registry.ts`) — tracks pending dispatches per session with TTL-based cleanup.
285
+ - **`tool.execute.after` hook** marks deliveries when a matching MCP tool call is observed.
286
+ - **All 5 bridge tools** now report `deliveryStatus: "delivered" | "pending"` in their tool result and metadata, and briefly poll (1.5s) for fast deliveries.
287
+ - When the LLM follows the prompt, the tool returns immediately with `"delivered"`. When it doesn't, the tool returns `"pending"` and the entry expires silently after 10s.
288
+
289
+ ```ts
290
+ // Bridge tool result metadata now includes:
291
+ {
292
+ tool: "omo_remember",
293
+ ok: true,
294
+ deliveryStatus: "delivered" | "pending",
295
+ messageID: "...",
296
+ durationMs: 1234,
297
+ contentLength: 256
298
+ }
299
+ ```
300
+
301
+ ### Test & build status
302
+
303
+ - **514/514 tests pass** (up from 495 in v0.16.0 — 5 + 4 + 10 new tests across F5.4, F5.1, F3.6).
304
+ - `bun run typecheck` clean.
305
+ - `bun build.ts` clean (0.34 MB dist).
306
+ - `npm pack --dry-run` validated.
307
+
308
+ ### Migration
309
+
310
+ No user action required. All changes are internal or additive:
311
+ - `deliveryStatus` is an additive metadata field — existing consumers ignore it.
312
+ - `maxLessonsPerSession` is now actually enforced — if you have sessions that previously saved more than 20 lessons (e.g. from before the cap was added), this may surprise you. Bump the cap in your config if needed.
313
+ - `escalate` action now actively fires Oracle — this is the first version where Oracle is auto-invoked, not just manually invoked by the LLM.
314
+
315
+ ### Audit roadmap (status as of v0.17.0)
316
+
317
+ | Release | Status | Scope |
318
+ |---------|--------|-------|
319
+ | v0.15.1 (F0) | ✅ Shipped | Hotfix self-dep + npm pack gate |
320
+ | v0.16.0 (F1-F7) | ✅ Shipped | Memory hygiene, dead code, tool coverage, CI |
321
+ | v0.17.0 (deferred) | ✅ Shipped | F5.1 escalate, F5.4 cap, F3.6 delivery verify |
322
+
323
+ All audit findings are now closed. Future work focuses on new features and user-driven feedback.
324
+
325
+
326
+
327
+ ## v0.17.1 — Audit args fix (patch)
328
+
329
+ v0.17.1 is a single-bug patch release. The fix addresses an issue discovered during v0.17.0 verification:
330
+
331
+ ### The bug
332
+
333
+ The `tool.execute.before` hook passed an empty `{}` object as the second argument to `auditToolCall()`. This meant the audit function never saw the tool's args (e.g. file content for write tools) and could never detect:
334
+
335
+ - `@ts-ignore` / `@ts-expect-error` directives
336
+ - `as any` type assertions
337
+ - `catch(e) {}` empty catch blocks
338
+
339
+ The hook signature was also incomplete — it didn't receive the `output` parameter that contains the mutable args, even though the SDK provides it.
340
+
341
+ ### The fix
342
+
343
+ Two changes in `src/plugin.ts`:
344
+
345
+ 1. **Hook signature updated** to receive the `output` parameter:
346
+ ```ts
347
+ "tool.execute.before": async (
348
+ toolInput: { tool: string; sessionID: string; callID: string },
349
+ _output: { args: unknown },
350
+ ): Promise<void> => {
351
+ ```
352
+
353
+ 2. **Audit call** now passes `_output.args` instead of `{}`:
354
+ ```ts
355
+ const violations = auditToolCall(toolInput.tool, _output.args, { ... })
356
+ ```
357
+
358
+ ### Tests
359
+
360
+ Added 4 new tests in `src/plugin.test.ts`:
361
+ - `@ts-ignore + as any` in args → `no-type-suppression` violation detected and injected
362
+ - `catch(e) {}` in args → `no-empty-catch` violation detected and injected
363
+ - Clean code → no violation injected (false-positive guard)
364
+ - `auditToolCalls: false` → audit short-circuits (regression check)
365
+
366
+ ### Test & build status
367
+
368
+ - **518/518 tests pass** (up from 514 in v0.17.0 — 4 new audit tests).
369
+ - `bun run typecheck` clean.
370
+ - `bun build.ts` clean (0.34 MB dist).
371
+
372
+ ### Migration
373
+
374
+ No user action required. The audit detection now correctly fires when the agent writes forbidden patterns. This means:
375
+
376
+ - **If your agent previously wrote `@ts-ignore` without being flagged**: it will now be flagged with `[GRAVE] no-type-suppression: ...` injected as a synthetic user message.
377
+ - **If you want to disable the audit**: set `protocolEnforcement.auditToolCalls: false` (already supported).
378
+
379
+
380
+
381
+ ## v0.17.2 — Fix escalation dead code + 4 audit gaps
382
+
383
+ v0.17.2 closes 4 gaps discovered during live verification of v0.17.0/v0.17.1. The most important: F5.1 (escalate → Oracle) was effectively dead in production due to two compounding bugs.
384
+
385
+ ### Highlights
386
+
387
+ #### Gap C (CRITICAL) — Escalation now actually fires
388
+
389
+ The score formula's `noProgress` and `deviations` inputs were hardcoded as `false` and `[]` in the plugin. This meant the `no-progress-detector` (weight 0.20) and `deviation-detector` (weight 0.20) signals always contributed 0. Combined with default thresholds, the maximum possible score was -0.55 — never reaching `escalateThreshold: 0.6` or `stopThreshold: 0.8`.
390
+
391
+ **Fix:**
392
+ 1. **Derive `noProgress`** from the recent tool call window. If the last 5 tool calls contain no `write`/`edit`/`task` (i.e. the agent is only reading/grepping without producing artifacts), `noProgress = true`.
393
+ 2. **Derive `deviations`** from accumulated protocol violations. The audit hook now stores violations in `state.accumulatedDeviations` (capped at 5 per session); the orchestrator input reads them.
394
+ 3. **Lower default thresholds** to match the new worst-case math:
395
+ - `escalateThreshold`: 0.6 → 0.45
396
+ - `stopThreshold`: 0.8 → 0.55
397
+
398
+ Now worst-case state (no oracle, no progress, 2 grave deviations, iteration at limit, stop-advice lessons) produces score ≈ -0.55 → `stop` action fires.
399
+
400
+ #### Gap Q (HIGH) — File paths threaded through pipeline
401
+
402
+ `orchestrator.ts` was hardcoding `filesChanged: []` instead of `input.filePaths`. This meant lesson extraction never saw the actual changed files, so F7.5's file-basename FTS indexing was empty.
403
+
404
+ **Fix:**
405
+ 1. Track `recentWriteFilePaths` in AuditState (alongside existing `recentWriteContents`).
406
+ 2. Capture `filePath` from `toolInput.args` on write/edit tool calls.
407
+ 3. New `MetaGovernorInput.filePaths?: readonly string[]` passed through to `observeAndLearn`.
408
+
409
+ #### Gap D (HIGH) — Three config fields now actually do something
410
+
411
+ Three fields were in the schema and config projection but NEVER consulted by the logic:
412
+
413
+ - `closedLoop.saveLessons` — parallel to `saveDecisions`. When `false`, lessons are skipped (decision records still save).
414
+ - `intervention.includeDecisionHistory` — when `true`, `messages.transform` prepends recent intervention texts (capped at `maxHistoryMessages`) so the LLM sees its history of decisions.
415
+ - `intervention.maxHistoryMessages` — limit for the above (default 5).
416
+
417
+ **Fix:** All three fields now control behavior. Track `recentInterventionTexts` in AuditState, format them into the injection text.
418
+
419
+ #### Bonus — iteration-budget signal wired (Oracle finding)
420
+
421
+ Oracle flagged a pre-existing gap alongside Gap C: `iteration` was hardcoded `0` in the orchestrator input, making the `iteration-budget` signal (weight 0.15) effectively dead.
422
+
423
+ Fix:
424
+ - Added `iteration: number` to `AuditState`, incremented per tool call.
425
+ - Threaded `iteration: sessionState?.iteration ?? 0` into `MetaGovernorInput`.
426
+ - `maxIterations` now reads from config instead of being hardcoded.
427
+
428
+ Worst-case score math updated: with iteration at 100% (-0.12), all signals bad, no oracle → score = -0.65 → `stop` action fires.
429
+
430
+ #### Gap I (MEDIUM) — `verifyDelivery` return type includes "expired"
431
+
432
+ The TypeScript signature was `Promise<"delivered" | "pending">` but the registry could return `"expired"`. The expired case leaked through as `"pending"` silently.
433
+
434
+ **Fix:** Signature updated to `Promise<"delivered" | "pending" | "expired">`. Bridge tools now distinguish: `"delivered"` (verified), `"pending"` (still polling), `"expired"` (TTL elapsed).
435
+
436
+ ### Test & build status
437
+
438
+ - **521/521 tests pass** (up from 518 — 3 new tests for the v0.17.2 fixes).
439
+ - `bun run typecheck` clean.
440
+ - `bun build.ts` clean (0.34 MB dist).
441
+ - `npm pack --dry-run` validated.
442
+
443
+ ### Migration
444
+
445
+ No user action required. Two behavior changes:
446
+
447
+ 1. **Escalation now fires more aggressively.** If your agent has been producing violations and not making progress, expect to see escalate → Oracle prompts more often. This is the intended behavior; v0.17.0 was incorrectly silent.
448
+ 2. **`includeDecisionHistory` and `maxHistoryMessages` are now functional.** If you set them in v0.17.0 expecting them to work, they will now actually take effect.
449
+
450
+ ### Audit roadmap (status as of v0.17.2)
451
+
452
+ | Release | Status | Scope |
453
+ |---------|--------|-------|
454
+ | v0.15.1 (F0) | ✅ | Hotfix self-dep |
455
+ | v0.16.0 (F1-F7) | ✅ | Memory hygiene, dead code, tool coverage, CI |
456
+ | v0.17.0 | ✅ | F5.1 escalate, F5.4 cap, F3.6 delivery verify |
457
+ | v0.17.1 | ✅ | Audit args fix |
458
+ | v0.17.2 | ✅ | Gap C (escalation live), Q (file paths), D (config fields), I (delivery expired) |
459
+
460
+
461
+
462
+ ## v0.17.3 — Fix Gap I properly (patch)
463
+
464
+ v0.17.3 is a single-bug patch. During live verification of v0.17.2, Gap I was found to be incompletely fixed.
465
+
466
+ ### The bug (v0.17.2 cosmetic fix)
467
+
468
+ The `verifyDelivery` export signature was widened to include `"expired"` in v0.17.2, and the bridge tools' title/output text was updated to handle it. **BUT the underlying `pollForDelivery` helper was still collapsing `"expired"` → `"pending"` silently:**
469
+
470
+ ```ts
471
+ // v0.17.2 (BUG):
472
+ return status === "delivered" ? "delivered" : "pending"
473
+ ```
474
+
475
+ So bridge tools could never report `"expired"` to the user, even though the registry correctly tracked it. Live verification confirmed: `deliveryStatus` always showed `"pending"`.
476
+
477
+ ### The fix (v0.17.3)
478
+
479
+ ```ts
480
+ // v0.17.3:
481
+ return await pendingRegistryRef.awaitDelivery({ sessionID, mcpTool, timeoutMs })
482
+ ```
483
+
484
+ Now the actual status from the registry propagates through. `"expired"` flows end-to-end to the bridge tool's `metadata.deliveryStatus` and title.
485
+
486
+ ### Tests
487
+
488
+ Added 3 RED tests in `src/custom-tools.test.ts`:
489
+ - Returns `"expired"` when registry entry exists past timeout (real registry instance)
490
+ - Returns `"delivered"` when `markDelivered` fires before timeout
491
+ - Returns `"pending"` when no registry is configured
492
+
493
+ ### Test & build status
494
+
495
+ - **525/525 tests pass** (up from 522 in v0.17.2 — 3 new tests for pollForDelivery).
496
+ - `bun run typecheck` clean.
497
+ - `bun build.ts` clean (0.34 MB dist).
498
+
499
+ ### Migration
500
+
501
+ No user action required. Bridge tools will now correctly distinguish all three delivery states:
502
+ - `"delivered"` — LLM's MCP tool call was observed within 1.5s
503
+ - `"expired"` — TTL elapsed without delivery (entry expires after 10s, but bridge tool sees this immediately as "expired" when polling times out at 1.5s)
504
+ - `"pending"` — no registry configured (graceful degradation for tests/mocks)
505
+
506
+ ### Audit roadmap (status as of v0.17.3)
507
+
508
+ | Release | Status | Scope |
509
+ |---------|--------|-------|
510
+ | v0.15.1 → v0.17.2 | ✅ | All audit findings + 5 gap fixes |
511
+ | v0.17.3 | ✅ | Gap I real fix (pollForDelivery returns "expired") |
512
+
513
+ Two remaining gaps documented but require SDK support to fix:
514
+ - `recentTurnTokens: []` — token-predictor signal dead (10% of score); needs per-turn token counts from OpenCode SDK
515
+ - `agentName` defaults to `"unknown"` — cosmetic, no functional impact
516
+
517
+
518
+
519
+ ## v0.18.0 — Audit remediation: 7+ silent config drops + circular ref crash
520
+
521
+ v0.18.0 is a thorough-audit patch release. Each fix addresses a bug found by testing every public function with edge cases and adversarial inputs.
522
+
523
+ ### Highlights
524
+
525
+ | # | Bug | Severity | Fix |
526
+ |---|-----|----------|-----|
527
+ | 1 | `file-logger.redactData` crashed on circular references with stack overflow | 🔴 CRITICAL | `WeakSet` guard + `try/catch` fallback |
528
+ | 2 | `loadOrchestratorConfig` only projected `closedLoop.saveDecisions` — `enabled`, `minSeverityToLearn`, `maxLessonsPerSession`, `saveLessons` were silently dropped | 🔴 CRITICAL | Project all 5 fields |
529
+ | 3 | `loadOrchestratorConfig` didn't project `decision.warnMessageTemplate`, `escalateMessageTemplate`, `stopMessageTemplate` | 🟠 HIGH | Project all 3 templates |
530
+ | 4 | `loadOrchestratorConfig` didn't project `scoring.paralysisThreshold`, `defaultEscalationTarget` | 🟠 HIGH | Project all fields |
531
+ | 5 | `loadOrchestratorConfig` had `memory.timeoutMs` field name mismatch (schema said `agentmemoryTimeoutMs`) | 🟠 HIGH | Accept both names |
532
+ | 6 | `isMetaGovernorEnabled` only checked top-level `enabled`, not `meta_governor.enabled` (wrapped shape from `opencode.jsonc`) | 🟠 HIGH | Check both shapes |
533
+ | 7 | `createMetricsCollector` crashed when called without config (`config.version` on `undefined`) | 🟠 HIGH | Accept `Partial<MetricsCollectorConfig>` |
534
+ | 8 | `metrics.inc` crashed on unknown event names (`bucket.count++` on `undefined`) | 🟠 HIGH | Guard `if (!bucket) return` |
535
+ | 9 | `isNewerVersion` returned `false` for `installed=null` (no upgrade triggered for fresh installs) | 🟡 MEDIUM | Return `true` when installed is null AND latest is valid |
536
+
537
+ ### Test & build status
538
+
539
+ - **557/557 tests pass** (up from 530 in v0.17.3 — 27 new tests for the audit fixes).
540
+ - `bun run typecheck` clean.
541
+ - `bun build.ts` clean (0.34 MB dist).
542
+ - `npm pack --dry-run` validated.
543
+
544
+ ### Migration
545
+
546
+ No user action required. The fix to `loadOrchestratorConfig` means **users who were setting `closedLoop.maxLessonsPerSession` or other previously-dropped fields will now see those values actually take effect**. If you had a config like `{ "closedLoop": { "maxLessonsPerSession": 50 } }` before v0.18.0, it was silently being overridden to 20. Starting v0.18.0, the value 50 is now respected.
547
+
548
+ ### Audit roadmap (status as of v0.18.0)
549
+
550
+ | Release | Status | Scope |
551
+ |---------|--------|-------|
552
+ | v0.15.1 → v0.17.3 | ✅ | All audit findings + deferred items + audit args fix + gap fixes |
553
+ | v0.18.0 | ✅ | 7 silent config drops + circular ref crash + metrics crashes + upgrade trigger |
554
+
555
+ This release closes the final round of gaps found by a thorough function-by-function audit. The plugin now correctly projects **all** user configuration, handles **all** circular reference cases, and fails safely on **all** missing-input scenarios.
556
+
557
+ ## Auto-upgrade (v0.12.0)
558
+
559
+ On plugin load, queries npm/pip registries to check whether newer versions
560
+ of **codegraph** or **graphify** exist. Config: `graphSync.autoUpgrade` (default `true`),
561
+ `graphSync.upgradeCheckTtlMs` (default `86400000`).
562
+
563
+ ## License
564
+
565
+ MIT