@herjarsa/omo-meta-governor 0.17.0 → 0.17.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -322,6 +322,141 @@ No user action required. All changes are internal or additive:
322
322
 
323
323
  All audit findings are now closed. Future work focuses on new features and user-driven feedback.
324
324
 
325
+
326
+
327
+ ## v0.17.1 — Audit args fix (patch)
328
+
329
+ v0.17.1 is a single-bug patch release. The fix addresses an issue discovered during v0.17.0 verification:
330
+
331
+ ### The bug
332
+
333
+ The `tool.execute.before` hook passed an empty `{}` object as the second argument to `auditToolCall()`. This meant the audit function never saw the tool's args (e.g. file content for write tools) and could never detect:
334
+
335
+ - `@ts-ignore` / `@ts-expect-error` directives
336
+ - `as any` type assertions
337
+ - `catch(e) {}` empty catch blocks
338
+
339
+ The hook signature was also incomplete — it didn't receive the `output` parameter that contains the mutable args, even though the SDK provides it.
340
+
341
+ ### The fix
342
+
343
+ Two changes in `src/plugin.ts`:
344
+
345
+ 1. **Hook signature updated** to receive the `output` parameter:
346
+ ```ts
347
+ "tool.execute.before": async (
348
+ toolInput: { tool: string; sessionID: string; callID: string },
349
+ _output: { args: unknown },
350
+ ): Promise<void> => {
351
+ ```
352
+
353
+ 2. **Audit call** now passes `_output.args` instead of `{}`:
354
+ ```ts
355
+ const violations = auditToolCall(toolInput.tool, _output.args, { ... })
356
+ ```
357
+
358
+ ### Tests
359
+
360
+ Added 4 new tests in `src/plugin.test.ts`:
361
+ - `@ts-ignore + as any` in args → `no-type-suppression` violation detected and injected
362
+ - `catch(e) {}` in args → `no-empty-catch` violation detected and injected
363
+ - Clean code → no violation injected (false-positive guard)
364
+ - `auditToolCalls: false` → audit short-circuits (regression check)
365
+
366
+ ### Test & build status
367
+
368
+ - **518/518 tests pass** (up from 514 in v0.17.0 — 4 new audit tests).
369
+ - `bun run typecheck` clean.
370
+ - `bun build.ts` clean (0.34 MB dist).
371
+
372
+ ### Migration
373
+
374
+ No user action required. The audit detection now correctly fires when the agent writes forbidden patterns. This means:
375
+
376
+ - **If your agent previously wrote `@ts-ignore` without being flagged**: it will now be flagged with `[GRAVE] no-type-suppression: ...` injected as a synthetic user message.
377
+ - **If you want to disable the audit**: set `protocolEnforcement.auditToolCalls: false` (already supported).
378
+
379
+
380
+
381
+ ## v0.17.2 — Fix escalation dead code + 4 audit gaps
382
+
383
+ v0.17.2 closes 4 gaps discovered during live verification of v0.17.0/v0.17.1. The most important: F5.1 (escalate → Oracle) was effectively dead in production due to two compounding bugs.
384
+
385
+ ### Highlights
386
+
387
+ #### Gap C (CRITICAL) — Escalation now actually fires
388
+
389
+ The score formula's `noProgress` and `deviations` inputs were hardcoded as `false` and `[]` in the plugin. This meant the `no-progress-detector` (weight 0.20) and `deviation-detector` (weight 0.20) signals always contributed 0. Combined with default thresholds, the maximum possible score was -0.55 — never reaching `escalateThreshold: 0.6` or `stopThreshold: 0.8`.
390
+
391
+ **Fix:**
392
+ 1. **Derive `noProgress`** from the recent tool call window. If the last 5 tool calls contain no `write`/`edit`/`task` (i.e. the agent is only reading/grepping without producing artifacts), `noProgress = true`.
393
+ 2. **Derive `deviations`** from accumulated protocol violations. The audit hook now stores violations in `state.accumulatedDeviations` (capped at 5 per session); the orchestrator input reads them.
394
+ 3. **Lower default thresholds** to match the new worst-case math:
395
+ - `escalateThreshold`: 0.6 → 0.45
396
+ - `stopThreshold`: 0.8 → 0.55
397
+
398
+ Now worst-case state (no oracle, no progress, 2 grave deviations, iteration at limit, stop-advice lessons) produces score ≈ -0.55 → `stop` action fires.
399
+
400
+ #### Gap Q (HIGH) — File paths threaded through pipeline
401
+
402
+ `orchestrator.ts` was hardcoding `filesChanged: []` instead of `input.filePaths`. This meant lesson extraction never saw the actual changed files, so F7.5's file-basename FTS indexing was empty.
403
+
404
+ **Fix:**
405
+ 1. Track `recentWriteFilePaths` in AuditState (alongside existing `recentWriteContents`).
406
+ 2. Capture `filePath` from `toolInput.args` on write/edit tool calls.
407
+ 3. New `MetaGovernorInput.filePaths?: readonly string[]` passed through to `observeAndLearn`.
408
+
409
+ #### Gap D (HIGH) — Three config fields now actually do something
410
+
411
+ Three fields were in the schema and config projection but NEVER consulted by the logic:
412
+
413
+ - `closedLoop.saveLessons` — parallel to `saveDecisions`. When `false`, lessons are skipped (decision records still save).
414
+ - `intervention.includeDecisionHistory` — when `true`, `messages.transform` prepends recent intervention texts (capped at `maxHistoryMessages`) so the LLM sees its history of decisions.
415
+ - `intervention.maxHistoryMessages` — limit for the above (default 5).
416
+
417
+ **Fix:** All three fields now control behavior. Track `recentInterventionTexts` in AuditState, format them into the injection text.
418
+
419
+ #### Bonus — iteration-budget signal wired (Oracle finding)
420
+
421
+ Oracle flagged a pre-existing gap alongside Gap C: `iteration` was hardcoded `0` in the orchestrator input, making the `iteration-budget` signal (weight 0.15) effectively dead.
422
+
423
+ Fix:
424
+ - Added `iteration: number` to `AuditState`, incremented per tool call.
425
+ - Threaded `iteration: sessionState?.iteration ?? 0` into `MetaGovernorInput`.
426
+ - `maxIterations` now reads from config instead of being hardcoded.
427
+
428
+ Worst-case score math updated: with iteration at 100% (-0.12), all signals bad, no oracle → score = -0.65 → `stop` action fires.
429
+
430
+ #### Gap I (MEDIUM) — `verifyDelivery` return type includes "expired"
431
+
432
+ The TypeScript signature was `Promise<"delivered" | "pending">` but the registry could return `"expired"`. The expired case leaked through as `"pending"` silently.
433
+
434
+ **Fix:** Signature updated to `Promise<"delivered" | "pending" | "expired">`. Bridge tools now distinguish: `"delivered"` (verified), `"pending"` (still polling), `"expired"` (TTL elapsed).
435
+
436
+ ### Test & build status
437
+
438
+ - **521/521 tests pass** (up from 518 — 3 new tests for the v0.17.2 fixes).
439
+ - `bun run typecheck` clean.
440
+ - `bun build.ts` clean (0.34 MB dist).
441
+ - `npm pack --dry-run` validated.
442
+
443
+ ### Migration
444
+
445
+ No user action required. Two behavior changes:
446
+
447
+ 1. **Escalation now fires more aggressively.** If your agent has been producing violations and not making progress, expect to see escalate → Oracle prompts more often. This is the intended behavior; v0.17.0 was incorrectly silent.
448
+ 2. **`includeDecisionHistory` and `maxHistoryMessages` are now functional.** If you set them in v0.17.0 expecting them to work, they will now actually take effect.
449
+
450
+ ### Audit roadmap (status as of v0.17.2)
451
+
452
+ | Release | Status | Scope |
453
+ |---------|--------|-------|
454
+ | v0.15.1 (F0) | ✅ | Hotfix self-dep |
455
+ | v0.16.0 (F1-F7) | ✅ | Memory hygiene, dead code, tool coverage, CI |
456
+ | v0.17.0 | ✅ | F5.1 escalate, F5.4 cap, F3.6 delivery verify |
457
+ | v0.17.1 | ✅ | Audit args fix |
458
+ | v0.17.2 | ✅ | Gap C (escalation live), Q (file paths), D (config fields), I (delivery expired) |
459
+
325
460
  ## Auto-upgrade (v0.12.0)
326
461
 
327
462
  On plugin load, queries npm/pip registries to check whether newer versions
@@ -46,6 +46,16 @@ declare let pendingRegistryRef: {
46
46
  * Exposed as a setter so we don't need to thread it through every tool deps.
47
47
  */
48
48
  export declare function setPendingDeliveryRegistry(registry: typeof pendingRegistryRef): void;
49
+ /**
50
+ * v0.17.2 (Gap I): Updated return type to include "expired" so the bridge
51
+ * tools can distinguish between "LLM hasn't called yet" (pending) and
52
+ * "TTL elapsed without delivery" (expired).
53
+ *
54
+ * Returns "delivered" if the LLM's MCP tool call was observed within the
55
+ * timeout, "expired" if the pending entry's TTL elapsed without delivery,
56
+ * "pending" otherwise (poll still active, no result yet).
57
+ */
58
+ export declare function verifyDelivery(sessionID: string, mcpTool: string): Promise<"delivered" | "pending" | "expired">;
49
59
  export interface OmoSearchDeps {
50
60
  graphRetrieval: GraphRetrieval;
51
61
  cwd: string;