@herjarsa/omo-meta-governor 0.17.0 → 0.17.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +135 -0
- package/dist/custom-tools.d.ts +10 -0
- package/dist/index.js +52 -46
- package/dist/index.js.map +7 -7
- package/dist/types.d.ts +7 -0
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -322,6 +322,141 @@ No user action required. All changes are internal or additive:
|
|
|
322
322
|
|
|
323
323
|
All audit findings are now closed. Future work focuses on new features and user-driven feedback.
|
|
324
324
|
|
|
325
|
+
|
|
326
|
+
|
|
327
|
+
## v0.17.1 — Audit args fix (patch)
|
|
328
|
+
|
|
329
|
+
v0.17.1 is a single-bug patch release. The fix addresses an issue discovered during v0.17.0 verification:
|
|
330
|
+
|
|
331
|
+
### The bug
|
|
332
|
+
|
|
333
|
+
The `tool.execute.before` hook passed an empty `{}` object as the second argument to `auditToolCall()`. This meant the audit function never saw the tool's args (e.g. file content for write tools) and could never detect:
|
|
334
|
+
|
|
335
|
+
- `@ts-ignore` / `@ts-expect-error` directives
|
|
336
|
+
- `as any` type assertions
|
|
337
|
+
- `catch(e) {}` empty catch blocks
|
|
338
|
+
|
|
339
|
+
The hook signature was also incomplete — it didn't receive the `output` parameter that contains the mutable args, even though the SDK provides it.
|
|
340
|
+
|
|
341
|
+
### The fix
|
|
342
|
+
|
|
343
|
+
Two changes in `src/plugin.ts`:
|
|
344
|
+
|
|
345
|
+
1. **Hook signature updated** to receive the `output` parameter:
|
|
346
|
+
```ts
|
|
347
|
+
"tool.execute.before": async (
|
|
348
|
+
toolInput: { tool: string; sessionID: string; callID: string },
|
|
349
|
+
_output: { args: unknown },
|
|
350
|
+
): Promise<void> => {
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
2. **Audit call** now passes `_output.args` instead of `{}`:
|
|
354
|
+
```ts
|
|
355
|
+
const violations = auditToolCall(toolInput.tool, _output.args, { ... })
|
|
356
|
+
```
|
|
357
|
+
|
|
358
|
+
### Tests
|
|
359
|
+
|
|
360
|
+
Added 4 new tests in `src/plugin.test.ts`:
|
|
361
|
+
- `@ts-ignore + as any` in args → `no-type-suppression` violation detected and injected
|
|
362
|
+
- `catch(e) {}` in args → `no-empty-catch` violation detected and injected
|
|
363
|
+
- Clean code → no violation injected (false-positive guard)
|
|
364
|
+
- `auditToolCalls: false` → audit short-circuits (regression check)
|
|
365
|
+
|
|
366
|
+
### Test & build status
|
|
367
|
+
|
|
368
|
+
- **518/518 tests pass** (up from 514 in v0.17.0 — 4 new audit tests).
|
|
369
|
+
- `bun run typecheck` clean.
|
|
370
|
+
- `bun build.ts` clean (0.34 MB dist).
|
|
371
|
+
|
|
372
|
+
### Migration
|
|
373
|
+
|
|
374
|
+
No user action required. The audit detection now correctly fires when the agent writes forbidden patterns. This means:
|
|
375
|
+
|
|
376
|
+
- **If your agent previously wrote `@ts-ignore` without being flagged**: it will now be flagged with `[GRAVE] no-type-suppression: ...` injected as a synthetic user message.
|
|
377
|
+
- **If you want to disable the audit**: set `protocolEnforcement.auditToolCalls: false` (already supported).
|
|
378
|
+
|
|
379
|
+
|
|
380
|
+
|
|
381
|
+
## v0.17.2 — Fix escalation dead code + 4 audit gaps
|
|
382
|
+
|
|
383
|
+
v0.17.2 closes 4 gaps discovered during live verification of v0.17.0/v0.17.1. The most important: F5.1 (escalate → Oracle) was effectively dead in production due to two compounding bugs.
|
|
384
|
+
|
|
385
|
+
### Highlights
|
|
386
|
+
|
|
387
|
+
#### Gap C (CRITICAL) — Escalation now actually fires
|
|
388
|
+
|
|
389
|
+
The score formula's `noProgress` and `deviations` inputs were hardcoded as `false` and `[]` in the plugin. This meant the `no-progress-detector` (weight 0.20) and `deviation-detector` (weight 0.20) signals always contributed 0. Combined with default thresholds, the maximum possible score was -0.55 — never reaching `escalateThreshold: 0.6` or `stopThreshold: 0.8`.
|
|
390
|
+
|
|
391
|
+
**Fix:**
|
|
392
|
+
1. **Derive `noProgress`** from the recent tool call window. If the last 5 tool calls contain no `write`/`edit`/`task` (i.e. the agent is only reading/grepping without producing artifacts), `noProgress = true`.
|
|
393
|
+
2. **Derive `deviations`** from accumulated protocol violations. The audit hook now stores violations in `state.accumulatedDeviations` (capped at 5 per session); the orchestrator input reads them.
|
|
394
|
+
3. **Lower default thresholds** to match the new worst-case math:
|
|
395
|
+
- `escalateThreshold`: 0.6 → 0.45
|
|
396
|
+
- `stopThreshold`: 0.8 → 0.55
|
|
397
|
+
|
|
398
|
+
Now worst-case state (no oracle, no progress, 2 grave deviations, iteration at limit, stop-advice lessons) produces score ≈ -0.55 → `stop` action fires.
|
|
399
|
+
|
|
400
|
+
#### Gap Q (HIGH) — File paths threaded through pipeline
|
|
401
|
+
|
|
402
|
+
`orchestrator.ts` was hardcoding `filesChanged: []` instead of `input.filePaths`. This meant lesson extraction never saw the actual changed files, so F7.5's file-basename FTS indexing was empty.
|
|
403
|
+
|
|
404
|
+
**Fix:**
|
|
405
|
+
1. Track `recentWriteFilePaths` in AuditState (alongside existing `recentWriteContents`).
|
|
406
|
+
2. Capture `filePath` from `toolInput.args` on write/edit tool calls.
|
|
407
|
+
3. New `MetaGovernorInput.filePaths?: readonly string[]` passed through to `observeAndLearn`.
|
|
408
|
+
|
|
409
|
+
#### Gap D (HIGH) — Three config fields now actually do something
|
|
410
|
+
|
|
411
|
+
Three fields were in the schema and config projection but NEVER consulted by the logic:
|
|
412
|
+
|
|
413
|
+
- `closedLoop.saveLessons` — parallel to `saveDecisions`. When `false`, lessons are skipped (decision records still save).
|
|
414
|
+
- `intervention.includeDecisionHistory` — when `true`, `messages.transform` prepends recent intervention texts (capped at `maxHistoryMessages`) so the LLM sees its history of decisions.
|
|
415
|
+
- `intervention.maxHistoryMessages` — limit for the above (default 5).
|
|
416
|
+
|
|
417
|
+
**Fix:** All three fields now control behavior. Track `recentInterventionTexts` in AuditState, format them into the injection text.
|
|
418
|
+
|
|
419
|
+
#### Bonus — iteration-budget signal wired (Oracle finding)
|
|
420
|
+
|
|
421
|
+
Oracle flagged a pre-existing gap alongside Gap C: `iteration` was hardcoded `0` in the orchestrator input, making the `iteration-budget` signal (weight 0.15) effectively dead.
|
|
422
|
+
|
|
423
|
+
Fix:
|
|
424
|
+
- Added `iteration: number` to `AuditState`, incremented per tool call.
|
|
425
|
+
- Threaded `iteration: sessionState?.iteration ?? 0` into `MetaGovernorInput`.
|
|
426
|
+
- `maxIterations` now reads from config instead of being hardcoded.
|
|
427
|
+
|
|
428
|
+
Worst-case score math updated: with iteration at 100% (-0.12), all signals bad, no oracle → score = -0.65 → `stop` action fires.
|
|
429
|
+
|
|
430
|
+
#### Gap I (MEDIUM) — `verifyDelivery` return type includes "expired"
|
|
431
|
+
|
|
432
|
+
The TypeScript signature was `Promise<"delivered" | "pending">` but the registry could return `"expired"`. The expired case leaked through as `"pending"` silently.
|
|
433
|
+
|
|
434
|
+
**Fix:** Signature updated to `Promise<"delivered" | "pending" | "expired">`. Bridge tools now distinguish: `"delivered"` (verified), `"pending"` (still polling), `"expired"` (TTL elapsed).
|
|
435
|
+
|
|
436
|
+
### Test & build status
|
|
437
|
+
|
|
438
|
+
- **521/521 tests pass** (up from 518 — 3 new tests for the v0.17.2 fixes).
|
|
439
|
+
- `bun run typecheck` clean.
|
|
440
|
+
- `bun build.ts` clean (0.34 MB dist).
|
|
441
|
+
- `npm pack --dry-run` validated.
|
|
442
|
+
|
|
443
|
+
### Migration
|
|
444
|
+
|
|
445
|
+
No user action required. Two behavior changes:
|
|
446
|
+
|
|
447
|
+
1. **Escalation now fires more aggressively.** If your agent has been producing violations and not making progress, expect to see escalate → Oracle prompts more often. This is the intended behavior; v0.17.0 was incorrectly silent.
|
|
448
|
+
2. **`includeDecisionHistory` and `maxHistoryMessages` are now functional.** If you set them in v0.17.0 expecting them to work, they will now actually take effect.
|
|
449
|
+
|
|
450
|
+
### Audit roadmap (status as of v0.17.2)
|
|
451
|
+
|
|
452
|
+
| Release | Status | Scope |
|
|
453
|
+
|---------|--------|-------|
|
|
454
|
+
| v0.15.1 (F0) | ✅ | Hotfix self-dep |
|
|
455
|
+
| v0.16.0 (F1-F7) | ✅ | Memory hygiene, dead code, tool coverage, CI |
|
|
456
|
+
| v0.17.0 | ✅ | F5.1 escalate, F5.4 cap, F3.6 delivery verify |
|
|
457
|
+
| v0.17.1 | ✅ | Audit args fix |
|
|
458
|
+
| v0.17.2 | ✅ | Gap C (escalation live), Q (file paths), D (config fields), I (delivery expired) |
|
|
459
|
+
|
|
325
460
|
## Auto-upgrade (v0.12.0)
|
|
326
461
|
|
|
327
462
|
On plugin load, queries npm/pip registries to check whether newer versions
|
package/dist/custom-tools.d.ts
CHANGED
|
@@ -46,6 +46,16 @@ declare let pendingRegistryRef: {
|
|
|
46
46
|
* Exposed as a setter so we don't need to thread it through every tool deps.
|
|
47
47
|
*/
|
|
48
48
|
export declare function setPendingDeliveryRegistry(registry: typeof pendingRegistryRef): void;
|
|
49
|
+
/**
|
|
50
|
+
* v0.17.2 (Gap I): Updated return type to include "expired" so the bridge
|
|
51
|
+
* tools can distinguish between "LLM hasn't called yet" (pending) and
|
|
52
|
+
* "TTL elapsed without delivery" (expired).
|
|
53
|
+
*
|
|
54
|
+
* Returns "delivered" if the LLM's MCP tool call was observed within the
|
|
55
|
+
* timeout, "expired" if the pending entry's TTL elapsed without delivery,
|
|
56
|
+
* "pending" otherwise (poll still active, no result yet).
|
|
57
|
+
*/
|
|
58
|
+
export declare function verifyDelivery(sessionID: string, mcpTool: string): Promise<"delivered" | "pending" | "expired">;
|
|
49
59
|
export interface OmoSearchDeps {
|
|
50
60
|
graphRetrieval: GraphRetrieval;
|
|
51
61
|
cwd: string;
|