agentgate-runtime-control 2.13.14 → 2.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -283,6 +283,48 @@ policies: {
283
283
 
284
284
  `agentgate doctor` warns loudly whenever the effective setting is `'allow'`, so this is never a silent gap in a project that runs `doctor` as part of its setup. `examples/protect-first-tool.mjs` demonstrates the gap and the fix side by side with a `grant_admin` call.
285
285
 
286
+ ## Closing the action-name evasion gap further (2.14)
287
+
288
+ A follow-up adversarial test fed the policy engine action names it had never been designed to see: `delete\u0000all` (embedded NUL), `delete‮` (right-to-left override — makes the name *render* differently than it reads), `%export` (fullwidth lookalikes of "export"), and `../delete` (path-traversal-shaped). None crashed anything, but all of them ALLOWed under the historical default, which is the real finding: a string-based classifier can always be fed a string it wasn't expecting.
289
+
290
+ Three independent changes close this, and they compose with `unknownActionPolicy` above rather than replace it:
291
+
292
+ **1. Unsafe action names are always rejected, unconditionally.** Control characters, bidi-override/zero-width characters, the fullwidth-forms block, and `../`-style path segments have no legitimate reason to appear in an action name, so this check is **not** an opt-in policy — it runs for every request, for every existing user, with no config needed:
293
+
294
+ ```text
295
+ delete\u0000all → BLOCK (winningRule: 'unsafe-action-name')
296
+ delete‮ → BLOCK
297
+ %export → BLOCK
298
+ ../delete → BLOCK
299
+ ```
300
+
301
+ **2. `policies.strictActionNames: true` (opt-in)** additionally requires every action name to match `^[a-z][a-z0-9_.:-]{0,63}$` — a plain lowercase identifier. This is off by default because it's a real behavior change for any system with mixed-case or unusually-shaped action names already in production; turn it on once you've audited your own action-name list.
302
+
303
+ **3. Tool registration metadata is authoritative over the action name a caller sends.** The real fix for "call the dangerous tool under a name the policy engine doesn't recognize" isn't a better regex — it's not trusting the name at all for tools you control. `registerTool()` now accepts classification that lives with the tool, not with whatever the caller's request claims:
304
+
305
+ ```js
306
+ gateway.registerTool({
307
+ name: 'remove_customer',
308
+ actionClass: 'destructive', // or 'readOnly'
309
+ requiresApproval: true,
310
+ environments: ['development', 'staging'], // tool refuses to run outside these
311
+ handler: async (args) => { /* ... */ }
312
+ });
313
+ ```
314
+
315
+ A caller cannot edit another tool's registration, so `grant_admin` registered with `requiresApproval: true` still requires approval even if a request calls it with `action: 'totally_unrecognized_name'`. Metadata sits below your own explicit `blockActions`/`productionBlock`/amount/`approvalActions` rules (those still win) but above name-based classification and `unknownActionPolicy` — see `test/action-hardening-2.14.test.js` for the exact priority ordering.
316
+
317
+ ## Idempotent approval resolution (2.14)
318
+
319
+ `POST /api/approvals/approve` and `POST /api/approvals/deny` accept an `Idempotency-Key` header. A retried request with the same key (same tenant, same path) gets back the **exact same response** instead of re-executing the handler or hitting an "already approved" error:
320
+
321
+ ```text
322
+ POST /api/approvals/approve
323
+ Idempotency-Key: approval_123:resolve:v1
324
+ ```
325
+
326
+ Replayed responses carry an `Idempotency-Replayed: true` header. Keys are cached in memory per control-plane process for 24h by default (`idempotencyTtlMs` to override); a different key against the same approval is **not** treated as a replay and goes through normal single-use resolution. This matters for exactly the failure mode a retrying client hits in practice: a network blip after the first approve succeeded, followed by an automatic retry that must not double-execute a refund or a production change.
327
+
286
328
  ## Approval Flow
287
329
 
288
330
  Sensitive tool calls that evaluate to `ASK` enter a pending approval state and are never executed automatically in enforce mode.
@@ -496,6 +538,28 @@ AgentGate supports deterministic RBAC and ABAC authorization using agent/user id
496
538
 
497
539
  Runs, approvals, agents, and policy versions can be persisted locally through the built-in storage adapters. For production multi-tenant deployments, use the Postgres/Supabase adapters with tenant-scoped sessions and RLS; local JSON persistence is intended for development or single-process deployments. Storage is provider-neutral so a database adapter can be introduced without changing the policy API.
498
540
 
541
+ ### Persistence corruption is fail-closed, not silent (2.13.15+)
542
+
543
+ An independent stress test of 2.13.14 found that writing garbage into `runs.json` or `approvals.json` and restarting made the server come up normally with an **empty collection** — no error, no alert. For `runs.json` that's a silently erased audit trail; for `approvals.json` it means a pending approval can vanish with no trace at all. That was a bug, not a resilience feature.
544
+
545
+ As of 2.13.15, `PersistentCollectionStore` distinguishes three cases:
546
+
547
+ - **File doesn't exist** (first run) — starts empty, as always. Not an error.
548
+ - **File exists but isn't valid JSON, or isn't an array** — this is corruption. By default, the store throws `PersistenceCorruptionError` (`code: 'AGENTGATE_PERSISTENCE_CORRUPT'`) and refuses to start, so a corrupted security-relevant file is loud at startup instead of discovered later as a gap in the audit log.
549
+ - **Any other read failure** (permission denied, disk I/O error) — also propagates as an error rather than being treated as "no data yet."
550
+
551
+ Recovery is opt-in only, because silently discarding the operator's decision here is exactly the bug being fixed:
552
+
553
+ ```js
554
+ createMCPGateway({
555
+ persistence: '.agentgate',
556
+ recoverFromCorruption: true, // off by default — corruption throws unless you opt in
557
+ onPersistenceCorruption: (event) => { /* event.quarantinePath, event.filePath, ... */ }
558
+ });
559
+ ```
560
+
561
+ With `recoverFromCorruption: true`, the corrupt file is renamed to `<file>.corrupt.<timestamp>` (never deleted) so it can be inspected, the collection starts from an empty/seed state, and the event is logged loudly to stderr and handed to `onPersistenceCorruption` if provided — it is never silently absorbed. `gateway.persistenceHealth()` and the Control Plane's `GET /api/ready` (which now returns `503` with `{ ready: false, persistence: { degraded: true, corruptions: [...] } }` in this state) let you alert on it rather than assume health.
562
+
499
563
  ## Multi-Tenant Control Plane (v1.8)
500
564
  AgentGate v1.8 adds tenant isolation, scoped API keys, key rotation/revocation, and tenant-scoped webhook registrations.
501
565
 
@@ -543,8 +607,14 @@ For first-tenant bootstrap over HTTP, set `AGENTGATE_BOOTSTRAP_TOKEN` and send `
543
607
  - Deterministic authorization remains the security authority
544
608
 
545
609
 
610
+ ### Oversized request bodies no longer stall the next request (2.13.15+)
611
+
612
+ A stress test found that a request body larger than `maxBodySize` correctly returned `413`, but the server stopped reading the body as soon as the limit was crossed without draining the rest of what the client was sending. On a keep-alive connection, those unread bytes were still arriving and got misread as the start of the *next* request, so a completely unrelated follow-up request (even an auth check with a fake key) could hang for the full request timeout instead of returning `401` quickly — a cheap way to tie up connections.
613
+
614
+ 2.13.15 fixes this by always fully draining the request body (even once it's known to be oversized — the rest is discarded, not buffered) before responding, and by sending `Connection: close` on `413`/`400` body-parsing errors so the client doesn't attempt to reuse a connection that was cut short either way. Regression tests in `test/oversized-body-recovery.test.js` repeat the exact sequence from the report (oversized request, then a fake-key request) and assert the follow-up never exceeds ~2s.
615
+
546
616
  ### Production Operations v2.1
547
- - `/api/health` and `/api/ready` health/readiness probes
617
+ - `/api/health` and `/api/ready` health/readiness probes (`/api/ready` reports `503` and `persistence.degraded: true` if persistence was recovered from corruption — see Persistence above)
548
618
  - `/api/metrics` deterministic runtime counters and latency telemetry
549
619
  - `/api/events` Server-Sent Events stream for runtime events
550
620
  - Sliding-window HTTP rate limiting with 429 responses
@@ -0,0 +1,60 @@
1
+ # 2.13.15 / 2.14 Hardening — Status and Remaining Infrastructure Plan
2
+
3
+ This tracks the fix plan that came out of the independent 2.13.14 stress-test report. Items are split into what shipped as code in this repo (2.13.15 hotfix + 2.14 action-hardening) versus what requires real infrastructure or a third party and cannot be completed by changing source files alone.
4
+
5
+ ## Shipped as code
6
+
7
+ | Item | Release | Where |
8
+ |---|---|---|
9
+ | Oversized request body no longer stalls the next request | 2.13.15 | `src/control-plane.js` (`readJsonBody`), `test/oversized-body-recovery.test.js` |
10
+ | Persistence corruption fails closed by default (was silent data loss) | 2.13.15 | `src/persistent-store.js` (`PersistenceCorruptionError`), `test/persistence-corruption.test.js` |
11
+ | `/api/ready` reports `503`/`degraded` on recovered corruption | 2.13.15 | `src/control-plane.js`, `src/mcp-gateway.js` (`persistenceHealth()`) |
12
+ | Unsafe action names (control/bidi-override/fullwidth/path-traversal) blocked unconditionally | 2.14 | `src/policy-engine.js` (`UNSAFE_ACTION_NAME`) |
13
+ | Opt-in strict action-name allowlist (`policies.strictActionNames`) | 2.14 | `src/policy-engine.js` (`STRICT_ACTION_NAME`) |
14
+ | Tool registration metadata (`actionClass`, `requiresApproval`, `environments`) as the source of truth, overriding caller-supplied action names | 2.14 | `src/mcp-gateway.js` (`registerTool`), `src/policy-engine.js` |
15
+ | Idempotency-Key support for `/api/approvals/approve` and `/deny` | 2.14 | `src/control-plane.js` |
16
+ | `test/action-hardening-2.14.test.js`, `test/idempotency-approve-deny.test.js` | 2.14 | regression coverage for all of the above |
17
+
18
+ `policies.unknownActionPolicy` (deny-by-default for action names the engine still doesn't recognize) and the Deep Attack Lab shipped in 2.13.14; this round only adds to it rather than replacing it. None of the above change default behavior in a way that breaks an existing deployment, except `strictActionNames`, which is opt-in.
19
+
20
+ ## Still requires real infrastructure or a third party
21
+
22
+ These cannot be "sent as files" — they need an actual environment, actual load, or an actual outside reviewer. Two already have a runnable test pack in this repo; the other two are net new and described below.
23
+
24
+ ### Already have a test pack — just needs an environment to run it in
25
+
26
+ - **PostgreSQL/Supabase + RLS acceptance test** — `docs/managed-postgres-acceptance-test.md`. Run against a real managed Postgres staging instance; records RPO/RTO and backup/restore verification.
27
+ - **External security review** — `docs/external-security-review.md` and `docs/external-security-review-test-pack.md`. Needs an independent reviewer to actually execute and sign off; AgentGate's own team cannot self-certify this gate.
28
+
29
+ ### Net new — no existing doc, plan below
30
+
31
+ **Disk-full test.** The 2.13.14 report deliberately avoided filling a real disk (`EACCES` via a read-only directory was used as a safe proxy, and it did surface a clear exception — see `src/persistent-store.js:save()`). A real disk-full test needs:
32
+ 1. A disposable staging host or container with a **quota-limited volume** (e.g. a small loop-mounted filesystem, or a container volume with a byte quota) — never the host's real root disk.
33
+ 2. Fill the quota while `PersistentCollectionStore.save()` is mid-write (temp-file + `renameSync`), and separately while idle, and confirm in both cases: the process surfaces a clear, loud error (not a silent swallow — this is the same class of bug 2.13.15 just fixed for corruption, so it needs its own explicit check); the `.tmp` file never replaces a good `filePath` via a partial `renameSync`; the control plane's `/api/ready` reflects the failure.
34
+ 3. Record: error surfaced, state of `filePath` vs `filePath.tmp` after the failure, and whether a subsequent write (after freeing space) recovers cleanly.
35
+
36
+ **Multi-instance / concurrent-writer test.** Local JSON persistence (`PersistentCollectionStore`) is explicitly documented as single-process (see README "Persistence"); this test is about *proving* that boundary rather than discovering it in production:
37
+ 1. Point two separate `createMCPGateway({ persistence: sameDir })` processes at the same directory.
38
+ 2. Drive concurrent approvals/writes from both.
39
+ 3. Expected (and acceptable) outcome: last-writer-wins / lost updates — document this explicitly as the known limitation of local JSON mode, with the recommendation to use the PostgreSQL/Supabase adapter (which has real RLS and a real database handling concurrency) for any multi-instance deployment. The point of this test is a documented, demonstrated boundary — not a fix to local JSON mode itself, which isn't meant to be multi-writer-safe.
40
+
41
+ ## Acceptance criteria before using "enterprise-grade" language
42
+
43
+ Carried over from the stress-test report, now with status:
44
+
45
+ | Gate | Target | Status |
46
+ |---|---:|---|
47
+ | Oversized request followed by normal request | 0 timeouts | ✅ shipped 2.13.15, regression-tested |
48
+ | Corrupted persistence startup | fail-closed or documented recovery | ✅ shipped 2.13.15, regression-tested |
49
+ | Pending approvals lost silently | 0 | ✅ shipped 2.13.15 |
50
+ | Unknown sensitive action allowed (name evasion) | 0 | ✅ shipped 2.13.14/2.14 (`unknownActionPolicy` + metadata + unsafe-name block) |
51
+ | Pre-approval handler execution | 0 | ✅ already enforced, regression-tested pre-2.13 |
52
+ | Duplicate approved execution | 0 | ✅ already enforced; 2.14 adds idempotent retries on top |
53
+ | Cross-tenant leak | 0 | ✅ regression-tested (`test/approval-tenant-isolation.test.js`) |
54
+ | Backup/restore | PASS | ⬜ needs a real Postgres staging run — see `docs/managed-postgres-acceptance-test.md` |
55
+ | PostgreSQL RLS | PASS | ⬜ needs a real Postgres staging run — same doc |
56
+ | Disk-full behavior | documented and tested | ⬜ plan above, needs a quota-limited environment |
57
+ | Multi-instance/concurrent writer | documented and tested | ⬜ plan above, needs two real processes |
58
+ | External review | completed or claim explicitly absent | ⬜ needs an actual independent reviewer |
59
+
60
+ Until the ⬜ rows are run against real infrastructure, marketing/sales copy should keep saying what's actually true: the engine, policy, and single-process persistence layer are hardened and regression-tested; multi-instance/Postgres/external-review claims are not yet backed by an executed test run.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "agentgate-runtime-control",
3
- "version": "2.13.14",
3
+ "version": "2.14.0",
4
4
  "type": "module",
5
5
  "description": "Runtime control plane and SaaS governance layer for AI agent tool execution",
6
6
  "exports": {
@@ -18,10 +18,42 @@ import { createRequire } from 'node:module';
18
18
  const require = createRequire(import.meta.url);
19
19
  const PACKAGE_VERSION = require('../package.json').version;
20
20
 
21
+ // Reads and parses a POST body with a hard size cap. Earlier versions bailed
22
+ // out of the read loop as soon as the cap was exceeded without consuming the
23
+ // rest of the incoming request — on a keep-alive connection, those unread
24
+ // bytes are still in flight from the client and get misinterpreted as the
25
+ // start of the next request, which is what caused an unrelated follow-up
26
+ // request on the same socket to hang until timeout instead of failing fast.
27
+ // Fix: once oversized, keep draining every chunk (so the socket/HTTP parser
28
+ // ends this request cleanly) but stop buffering it into memory, then raise a
29
+ // typed error with the right status for the caller to respond with.
30
+ async function readJsonBody(req, maxBodySize) {
31
+ let raw = '';
32
+ let bytes = 0;
33
+ let oversized = false;
34
+ for await (const chunk of req) {
35
+ bytes += chunk.length;
36
+ if (bytes > maxBodySize) { oversized = true; continue; }
37
+ raw += chunk;
38
+ }
39
+ if (oversized) {
40
+ const error = new Error('Payload too large');
41
+ error.status = 413;
42
+ throw error;
43
+ }
44
+ try {
45
+ return raw ? JSON.parse(raw) : {};
46
+ } catch {
47
+ const error = new Error('Invalid JSON');
48
+ error.status = 400;
49
+ throw error;
50
+ }
51
+ }
52
+
21
53
  export function createControlPlane(options = {}) {
22
54
  const gateway = options.gateway || createMCPGateway(options);
23
55
  const staticHtml = options.html;
24
- const agentStore = options.agentStore || (options.persistence ? createPersistentAgentStore({ filePath: options.agentPersistence || `${options.persistence}/agents.json` }) : null);
56
+ const agentStore = options.agentStore || (options.persistence ? createPersistentAgentStore({ filePath: options.agentPersistence || `${options.persistence}/agents.json`, recoverFromCorruption: options.recoverFromCorruption, onCorruption: options.onPersistenceCorruption }) : null);
25
57
  const tenantRegistry = options.tenantRegistry || new TenantRegistry({ filePath: `${options.persistence || '.agentgate'}/tenants.json`, keyPath: `${options.persistence || '.agentgate'}/api-keys.json` });
26
58
  const webhookRegistry = options.webhookRegistry || new WebhookRegistry({ filePath: `${options.persistence || '.agentgate'}/webhooks.json`, deliveryPath: `${options.persistence || '.agentgate'}/webhook-deliveries.json` });
27
59
  const authRequired = options.authRequired !== false;
@@ -35,6 +67,25 @@ export function createControlPlane(options = {}) {
35
67
  const limiter = options.rateLimiter || new SlidingWindowLimiter({ limit: options.rateLimit || 120, windowMs: options.rateWindowMs || 60000 });
36
68
  const tenantLimiter = options.tenantRateLimiter || new SlidingWindowLimiter({ limit: options.tenantRateLimit || options.rateLimit || 120, windowMs: options.rateWindowMs || 60000 });
37
69
  const startedAt = Date.now();
70
+ // Idempotency: a client that retries POST /api/approvals/approve|deny
71
+ // (timeout, connection reset, at-least-once delivery) must not trigger a
72
+ // second side effect. Scoped by tenant + path + the client-supplied
73
+ // Idempotency-Key so a retry with the same key gets back the exact same
74
+ // response instead of re-executing (or hitting "already approved").
75
+ const idempotencyTtlMs = Number(options.idempotencyTtlMs || 24 * 60 * 60 * 1000);
76
+ const idempotencyCache = new Map(); // scopedKey -> { status, result, expiresAt }
77
+ const IDEMPOTENT_PATHS = new Set(['/api/approvals/approve', '/api/approvals/deny']);
78
+ function idempotencyGet(key) {
79
+ if (!key) return null;
80
+ const entry = idempotencyCache.get(key);
81
+ if (!entry) return null;
82
+ if (entry.expiresAt < Date.now()) { idempotencyCache.delete(key); return null; }
83
+ return entry;
84
+ }
85
+ function idempotencyPut(key, status, result) {
86
+ if (!key) return;
87
+ idempotencyCache.set(key, { status, result, expiresAt: Date.now() + idempotencyTtlMs });
88
+ }
38
89
  const policyRegistry = options.policyRegistry || createPolicyRegistry({ filePath: options.policyPersistence || `${options.persistence || '.agentgate'}/policies.json` });
39
90
  const bundleRegistry = options.policyBundleRegistry || createPolicyBundleRegistry({ filePath: options.policyBundlePersistence || `${options.persistence || '.agentgate'}/policy-bundles.json` });
40
91
  const agents = agentStore || { items: [], add(a){ this.items.unshift(a); return a; }, list(filter){ const all=[...this.items]; return typeof filter==='function' ? all.filter(filter) : all; }, get(id){ return this.items.find(a=>a.id===id || a.name===id) || null; }, save(){} };
@@ -69,7 +120,11 @@ export function createControlPlane(options = {}) {
69
120
  const scopedRuns = () => gateway.runs(tenantId);
70
121
  const scopedApprovals = (status) => gateway.approvals(status, tenantId);
71
122
  if (path === '/api/health') return { ok: true, mode: gateway.mode, version: gateway?.serverInfo?.version || PACKAGE_VERSION };
72
- if (path === '/api/ready') return { ok: true, ready: true, mode: gateway.mode, uptimeSeconds: Math.floor((Date.now() - startedAt) / 1000) };
123
+ if (path === '/api/ready') {
124
+ const persistence = gateway.persistenceHealth?.() || { persistent: false, degraded: false, corruptions: [] };
125
+ const ready = !persistence.degraded;
126
+ return { ok: ready, ready, mode: gateway.mode, uptimeSeconds: Math.floor((Date.now() - startedAt) / 1000), persistence };
127
+ }
73
128
  if (path === '/api/kill-switch' && method === 'GET') return gateway.killStatus?.() || {killed:false};
74
129
  if (path === '/api/kill-switch' && method === 'POST') { if (body.enabled === false) return gateway.unkill?.(); return gateway.kill?.(body.reason); }
75
130
  if (path === '/api/metrics') return gateway.telemetry?.snapshot?.() || { uptimeSeconds: Math.floor((Date.now() - startedAt) / 1000), counters: {}, latency: {} };
@@ -171,11 +226,18 @@ export function createControlPlane(options = {}) {
171
226
  if (url.pathname.startsWith('/api/')) {
172
227
  let body = {};
173
228
  if (req.method === 'POST') {
174
- let raw = '';
175
- let bodyBytes = 0;
176
229
  const maxBodySize = Number(options.maxBodySize || 1024 * 1024);
177
- for await (const chunk of req) { bodyBytes += chunk.length; if (bodyBytes > maxBodySize) { res.writeHead(413, {'content-type':'application/json'}); res.end(JSON.stringify({error:'Payload too large'})); return; } raw += chunk; }
178
- try { body = raw ? JSON.parse(raw) : {}; } catch { res.writeHead(400, {'content-type':'application/json'}); res.end(JSON.stringify({error:'Invalid JSON'})); return; }
230
+ try {
231
+ body = await readJsonBody(req, maxBodySize);
232
+ } catch (error) {
233
+ // 'connection: close' on top of fully draining the body (above) is
234
+ // belt-and-suspenders: it tells the client itself not to reuse this
235
+ // socket, so even a client we haven't fully drained yet won't have
236
+ // its next request misread on this connection.
237
+ res.writeHead(error.status || 400, { 'content-type': 'application/json', 'cache-control': 'no-store', 'connection': 'close' });
238
+ res.end(JSON.stringify({ error: error.message }));
239
+ return;
240
+ }
179
241
  }
180
242
  const limitKey = req.headers['x-agentgate-key'] || req.socket.remoteAddress || 'anonymous';
181
243
  const rate = limiter.check(String(limitKey));
@@ -202,9 +264,23 @@ export function createControlPlane(options = {}) {
202
264
  if (!tenantRate.allowed) { res.writeHead(429, {'content-type':'application/json','cache-control':'no-store','retry-after':String(Math.ceil((Date.parse(tenantRate.resetAt)-Date.now())/1000))}); res.end(JSON.stringify({error:'Tenant rate limit exceeded', ...tenantRate})); return; }
203
265
  }
204
266
  }
267
+ const idempotencyHeader = req.headers['idempotency-key'];
268
+ const idempotencyKey = idempotencyHeader && IDEMPOTENT_PATHS.has(url.pathname)
269
+ ? `${authContext.tenantId || 'no-tenant'}:${url.pathname}:${idempotencyHeader}`
270
+ : null;
271
+ if (idempotencyKey) {
272
+ const cached = idempotencyGet(idempotencyKey);
273
+ if (cached) {
274
+ res.writeHead(cached.status, { 'content-type': 'application/json', 'cache-control': 'no-store', 'idempotency-replayed': 'true' });
275
+ res.end(JSON.stringify(cached.result));
276
+ return;
277
+ }
278
+ }
205
279
  const result = await api(url.pathname, req.method, body, Object.fromEntries(url.searchParams.entries()), authContext);
206
280
  if (result._html) { res.writeHead(200, { 'content-type': 'text/html; charset=utf-8', 'cache-control': 'no-store' }); res.end(result.html); return; }
207
- res.writeHead(result.error ? 404 : 200, { 'content-type': 'application/json', 'cache-control': 'no-store' });
281
+ const status = result.error ? 404 : (url.pathname === '/api/ready' && result.ready === false ? 503 : 200);
282
+ if (idempotencyKey) idempotencyPut(idempotencyKey, status, result);
283
+ res.writeHead(status, { 'content-type': 'application/json', 'cache-control': 'no-store' });
208
284
  res.end(JSON.stringify(result));
209
285
  return;
210
286
  }
@@ -32,9 +32,11 @@ export function createMCPGateway(options = {}) {
32
32
  const killReason = { value: options.killReason || 'Emergency security lock' };
33
33
  const maxRuns = Number(options.maxRuns || 25000);
34
34
  let retentionWarned = false;
35
- const runStore = options.runStore || (options.persistence ? createPersistentRunStore({ filePath: options.runPersistence || `${options.persistence}/runs.json`, limit: maxRuns }) : null);
35
+ const recoverFromCorruption = Boolean(options.recoverFromCorruption);
36
+ const onPersistenceCorruption = options.onPersistenceCorruption;
37
+ const runStore = options.runStore || (options.persistence ? createPersistentRunStore({ filePath: options.runPersistence || `${options.persistence}/runs.json`, limit: maxRuns, recoverFromCorruption, onCorruption: onPersistenceCorruption }) : null);
36
38
  const runs = runStore ? null : [];
37
- const approvals = createApprovalStore({ store: options.approvalStore || (options.persistence ? createPersistentApprovalStore({ filePath: options.approvalPersistence || `${options.persistence}/approvals.json`, limit: options.maxApprovals || maxRuns }) : undefined), limit: options.maxApprovals || maxRuns });
39
+ const approvals = createApprovalStore({ store: options.approvalStore || (options.persistence ? createPersistentApprovalStore({ filePath: options.approvalPersistence || `${options.persistence}/approvals.json`, limit: options.maxApprovals || maxRuns, recoverFromCorruption, onCorruption: onPersistenceCorruption }) : undefined), limit: options.maxApprovals || maxRuns });
38
40
  const eventBus = options.eventBus || createEventBus();
39
41
  const telemetry = options.telemetry || createTelemetry();
40
42
  const egressGuard = options.egressGuard || (options.egress ? createEgressGuard(options.egress === true ? {} : options.egress) : null);
@@ -73,6 +75,15 @@ export function createMCPGateway(options = {}) {
73
75
  const count = runStore ? runStore.list().length : runs.length;
74
76
  const nearLimit = count >= Math.max(1, Math.ceil(maxRuns * 0.9));
75
77
  return { count, limit: maxRuns, nearLimit, persistent: Boolean(runStore), truncated: count >= maxRuns };
78
+ },
79
+ // Reports whether persistence had to recover from corruption at startup
80
+ // (only possible when recoverFromCorruption: true was passed — otherwise
81
+ // corruption makes createMCPGateway() throw instead of starting up in a
82
+ // degraded state). Control planes/readiness probes should surface this
83
+ // rather than reporting healthy while sitting on recovered/lost data.
84
+ persistenceHealth: () => {
85
+ const corruptions = [runStore?.corruption, approvals?.corruption].filter(Boolean);
86
+ return { persistent: Boolean(runStore), degraded: corruptions.length > 0, corruptions };
76
87
  }
77
88
  };
78
89
 
@@ -100,7 +111,17 @@ export function createMCPGateway(options = {}) {
100
111
  name: tool.name,
101
112
  description: tool.description || '',
102
113
  inputSchema: tool.inputSchema || { type: 'object' },
103
- handler: tool.handler
114
+ handler: tool.handler,
115
+ // Optional classification recorded at registration time, i.e. by
116
+ // whoever controls the tool list — not by whatever action name a
117
+ // caller's request happens to send. This is what makes it trustworthy:
118
+ // a caller can rename "remove_customer" to "delete" or to a brand-new
119
+ // string the policy engine has never seen, but it cannot edit this
120
+ // tool's registered metadata. See policy-engine.js's handling of
121
+ // request.toolMetadata for how this takes priority over name matching.
122
+ actionClass: tool.actionClass || null, // 'destructive' | 'readOnly' | null
123
+ requiresApproval: tool.requiresApproval === true,
124
+ environments: Array.isArray(tool.environments) ? tool.environments : null
104
125
  });
105
126
  return api;
106
127
  }
@@ -146,7 +167,12 @@ export function createMCPGateway(options = {}) {
146
167
  model: params.model || params.context?.model || input.model,
147
168
  usage: params.usage || params.context?.usage || input.usage,
148
169
  estimatedCost: params.estimatedCost ?? params.context?.estimatedCost ?? input.estimatedCost,
149
- context: params.context || {}
170
+ context: params.context || {},
171
+ // The tool's own registered metadata (never caller-suppliable) — see
172
+ // registerTool() above and policy-engine.js's toolMetadata handling.
173
+ toolMetadata: (tool.actionClass || tool.requiresApproval || tool.environments)
174
+ ? { actionClass: tool.actionClass, requiresApproval: tool.requiresApproval, environments: tool.environments }
175
+ : null
150
176
  };
151
177
  if (killSwitch && mode === 'enforce') {
152
178
  const run = { id: randomUUID(), createdAt:new Date().toISOString(), tenantId:request.tenantId, request, decision:'BLOCK', risk:100, reason:killReason.value, mode, killed:true };
@@ -1,11 +1,29 @@
1
1
  import fs from 'node:fs';
2
2
  import path from 'node:path';
3
3
 
4
+ // Thrown when a persisted collection file exists but cannot be trusted:
5
+ // invalid JSON, or valid JSON that isn't the array shape this store expects.
6
+ // This is distinct from "file does not exist yet" (ENOENT), which is the
7
+ // normal first-run case and is NOT an error.
8
+ export class PersistenceCorruptionError extends Error {
9
+ constructor(message, options) {
10
+ super(message, options);
11
+ this.name = 'PersistenceCorruptionError';
12
+ this.code = 'AGENTGATE_PERSISTENCE_CORRUPT';
13
+ }
14
+ }
15
+
4
16
  export class PersistentCollectionStore {
5
- constructor(filePath, { key = 'id', limit = 25000, seed = [] } = {}) {
17
+ constructor(filePath, { key = 'id', limit = 25000, seed = [], recoverFromCorruption = false, onCorruption } = {}) {
6
18
  this.filePath = path.resolve(filePath);
7
19
  this.key = key;
8
20
  this.limit = limit;
21
+ this.recoverFromCorruption = Boolean(recoverFromCorruption);
22
+ this.onCorruption = typeof onCorruption === 'function' ? onCorruption : null;
23
+ // Set only if a corruption event was handled (requires recoverFromCorruption:
24
+ // true). When this is non-null, the in-memory collection was reset and
25
+ // whatever was in the corrupt file was NOT recovered automatically.
26
+ this.corruption = null;
9
27
  this.items = this.#load(seed);
10
28
  }
11
29
 
@@ -35,22 +53,94 @@ export class PersistentCollectionStore {
35
53
  fs.writeFileSync(temp, JSON.stringify(this.items, null, 2), 'utf8');
36
54
  fs.renameSync(temp, this.filePath);
37
55
  }
56
+
38
57
  #load(seed) {
58
+ let text;
59
+ try {
60
+ text = fs.readFileSync(this.filePath, 'utf8');
61
+ } catch (err) {
62
+ // No file yet is the normal first-run case. Any other read failure
63
+ // (permission denied, I/O error, ...) is an operational failure, not
64
+ // "no data yet" — let it fail closed by propagating the error instead
65
+ // of silently returning an empty collection.
66
+ if (err.code === 'ENOENT') return Array.isArray(seed) ? seed.slice(0, this.limit) : [];
67
+ throw err;
68
+ }
69
+ let parsed;
39
70
  try {
40
- if (!fs.existsSync(this.filePath)) return Array.isArray(seed) ? seed.slice(0, this.limit) : [];
41
- const parsed = JSON.parse(fs.readFileSync(this.filePath, 'utf8'));
42
- return Array.isArray(parsed) ? parsed.slice(0, this.limit) : [];
43
- } catch { return Array.isArray(seed) ? seed.slice(0, this.limit) : []; }
71
+ parsed = JSON.parse(text);
72
+ } catch (err) {
73
+ return this.#handleCorruption(text, new PersistenceCorruptionError(
74
+ `Corrupt persistence file (invalid JSON): ${this.filePath}`,
75
+ { cause: err }
76
+ ), seed);
77
+ }
78
+ if (!Array.isArray(parsed)) {
79
+ return this.#handleCorruption(text, new PersistenceCorruptionError(
80
+ `Corrupt persistence file (expected an array, got ${parsed === null ? 'null' : typeof parsed}): ${this.filePath}`
81
+ ), seed);
82
+ }
83
+ return parsed.slice(0, this.limit);
44
84
  }
85
+
86
+ // Corruption means the file exists but its contents cannot be trusted —
87
+ // tampering, a crash mid-write on an older version, disk corruption, etc.
88
+ // The previous behavior here was to swallow the error and silently return
89
+ // an empty array, which quietly erases audit history and, worse, makes
90
+ // pending approvals vanish with no trace. We now fail closed by default:
91
+ // the store refuses to come up at all, so the operator finds out at
92
+ // startup instead of discovering a gap in the audit trail later.
93
+ //
94
+ // Recovery is opt-in only (recoverFromCorruption: true): the corrupt file
95
+ // is quarantined (renamed, never deleted) and the collection starts from
96
+ // `seed` (normally empty) — but this is explicitly NOT the same as
97
+ // recovering the lost data, and is logged loudly as such.
98
+ #handleCorruption(rawText, error, seed) {
99
+ if (!this.recoverFromCorruption) throw error;
100
+ const quarantinePath = `${this.filePath}.corrupt.${Date.now()}`;
101
+ try {
102
+ fs.renameSync(this.filePath, quarantinePath);
103
+ } catch {
104
+ try { fs.writeFileSync(quarantinePath, rawText ?? '', 'utf8'); } catch { /* best effort */ }
105
+ }
106
+ this.corruption = {
107
+ filePath: this.filePath,
108
+ quarantinePath,
109
+ message: error.message,
110
+ recoveredAt: new Date().toISOString(),
111
+ recoveredCount: Array.isArray(seed) ? seed.length : 0
112
+ };
113
+ console.error(
114
+ `\n🚨 AgentGate: persistence corruption recovered for ${this.filePath}\n` +
115
+ ` Corrupt file quarantined to: ${quarantinePath}\n` +
116
+ ` Started with ${this.corruption.recoveredCount} seed record(s) — the original contents were NOT recovered.\n` +
117
+ ` Investigate the quarantined file before trusting this collection again.\n`
118
+ );
119
+ this.onCorruption?.(this.corruption);
120
+ return Array.isArray(seed) ? seed.slice(0, this.limit) : [];
121
+ }
122
+
45
123
  #trim() { if (this.items.length > this.limit) this.items.splice(this.limit); }
46
124
  }
47
125
 
48
126
  export function createPersistentRunStore(options = {}) {
49
- return new PersistentCollectionStore(options.filePath || '.agentgate/runs.json', { limit: options.limit || 25000 });
127
+ return new PersistentCollectionStore(options.filePath || '.agentgate/runs.json', {
128
+ limit: options.limit || 25000,
129
+ recoverFromCorruption: options.recoverFromCorruption,
130
+ onCorruption: options.onCorruption
131
+ });
50
132
  }
51
133
  export function createPersistentApprovalStore(options = {}) {
52
- return new PersistentCollectionStore(options.filePath || '.agentgate/approvals.json', { limit: options.limit || 25000 });
134
+ return new PersistentCollectionStore(options.filePath || '.agentgate/approvals.json', {
135
+ limit: options.limit || 25000,
136
+ recoverFromCorruption: options.recoverFromCorruption,
137
+ onCorruption: options.onCorruption
138
+ });
53
139
  }
54
140
  export function createPersistentAgentStore(options = {}) {
55
- return new PersistentCollectionStore(options.filePath || '.agentgate/agents.json', { limit: options.limit || 1000 });
141
+ return new PersistentCollectionStore(options.filePath || '.agentgate/agents.json', {
142
+ limit: options.limit || 1000,
143
+ recoverFromCorruption: options.recoverFromCorruption,
144
+ onCorruption: options.onCorruption
145
+ });
56
146
  }
@@ -5,11 +5,38 @@ export const DECISIONS = Object.freeze({ ALLOW: 'ALLOW', ASK: 'ASK', BLOCK: 'BLO
5
5
  const destructive = new Set(['delete','refund','publish','deploy','export_all','update_production']);
6
6
  const readOnly = new Set(['read','search','list','get','fetch']);
7
7
 
8
+ // Characters that have no legitimate reason to appear in an action name and
9
+ // exist only to confuse logs/matching or to smuggle a control sequence past
10
+ // string-equality checks: NUL and other control bytes, zero-width/bidi
11
+ // override characters (can make "delete" *look* like something else when
12
+ // rendered, or hide characters entirely), the fullwidth-forms block (an
13
+ // adversarial test used "%export" — fullwidth lookalikes of ASCII — to
14
+ // probe whether it matched "export"), and "../"-style path segments. This
15
+ // check is unconditional and always on (not an opt-in policy) because no
16
+ // real tool/action name is ever affected by it — unlike a strict allowlist
17
+ // regex, which IS opt-in below to avoid breaking existing naming schemes.
18
+ const UNSAFE_ACTION_NAME = new RegExp(
19
+ '[\\x00-\\x1F\\x7F\\u200B-\\u200F\\u202A-\\u202E\\u2066-\\u2069\\uFEFF\\uFF00-\\uFFEF]|\\.\\.[\\/\\\\]'
20
+ );
21
+ // Opt-in, stricter allowlist for teams that control every action name they
22
+ // register: policies.strictActionNames: true requires lowercase
23
+ // ascii-identifier-shaped names. Off by default so existing naming schemes
24
+ // (mixed case, spaces, unicode brand/customer names, etc.) don't suddenly
25
+ // start failing on upgrade.
26
+ const STRICT_ACTION_NAME = /^[a-z][a-z0-9_.:-]{0,63}$/;
27
+
8
28
  export function evaluate(input = {}, policies = {}) {
9
29
  const auth = authorize(input, policies.authorization);
10
30
  const trace = [{ id: 'authorization', matched: auth.decision === 'BLOCK', decision: auth.decision, reason: auth.reason, matchedRule: auth.matchedRule || null }];
11
31
  if (auth.decision === 'BLOCK') return decision('BLOCK', auth.reason, 96, { authorization: auth, ruleTrace: trace, winningRule: auth.matchedRule || 'authorization' });
12
- const action = String(input.action || '').toLowerCase();
32
+ const rawAction = String(input.action || '');
33
+ if (UNSAFE_ACTION_NAME.test(rawAction)) {
34
+ return decision('BLOCK', 'Action name contains unsafe characters (control, bidi-override, fullwidth-form, or path-traversal)', 99, { ruleTrace: [...trace, { id: 'unsafe-action-name', matched: true, decision: 'BLOCK', reason: 'Action name contains unsafe characters' }], winningRule: 'unsafe-action-name' });
35
+ }
36
+ const action = rawAction.toLowerCase();
37
+ if (policies.strictActionNames && !STRICT_ACTION_NAME.test(action)) {
38
+ return decision('BLOCK', 'Action name does not match the required format under strictActionNames', 97, { ruleTrace: [...trace, { id: 'strict-action-name', matched: true, decision: 'BLOCK', reason: 'Action name does not match the required format under strictActionNames' }], winningRule: 'strict-action-name' });
39
+ }
13
40
  const hasAmount = Object.prototype.hasOwnProperty.call(input, 'amount') && input.amount !== undefined && input.amount !== null;
14
41
  if (hasAmount && (typeof input.amount !== 'number' || !Number.isFinite(input.amount) || input.amount < 0)) {
15
42
  return decision('BLOCK', 'Amount must be a finite non-negative number', 98, { ruleTrace: [...trace, { id: 'invalid-amount', matched: true, decision: 'BLOCK', reason: 'Amount must be a finite non-negative number' }], winningRule: 'invalid-amount' });
@@ -29,6 +56,30 @@ export function evaluate(input = {}, policies = {}) {
29
56
  if (policies.approvalActions?.includes(action)) {
30
57
  return decision('ASK', 'Approval required for sensitive action', 70, { ruleTrace: [...trace, { id: 'approval-action', matched: true, decision: 'ASK', reason: 'Approval required for sensitive action' }], winningRule: 'approval-action' });
31
58
  }
59
+ // A tool's own registered metadata (set by whoever controls the tool list,
60
+ // via registerTool({ actionClass, requiresApproval, environments })) is a
61
+ // more trustworthy signal than the action *name* a caller's request
62
+ // happens to send, because a caller can pick any action string it wants
63
+ // but cannot edit another tool's registration. It sits below the admin's
64
+ // own explicit blockActions/productionBlock/amount/approvalActions rules
65
+ // above (those still win), but above name-based classification and
66
+ // unknownActionPolicy below — this is what closes the "call the dangerous
67
+ // tool under an action name the policy engine doesn't recognize" evasion.
68
+ const toolMeta = input.toolMetadata;
69
+ if (toolMeta) {
70
+ if (Array.isArray(toolMeta.environments) && toolMeta.environments.length && !toolMeta.environments.includes(env)) {
71
+ return decision('BLOCK', 'Tool is not permitted in this environment per its registered metadata', 93, { ruleTrace: [...trace, { id: 'tool-metadata-environment', matched: true, decision: 'BLOCK', reason: 'Tool is not permitted in this environment per its registered metadata' }], winningRule: 'tool-metadata-environment' });
72
+ }
73
+ if (toolMeta.actionClass === 'destructive' && policies.productionBlock && env === 'production') {
74
+ return decision('BLOCK', 'Destructive production action is blocked (per tool registration metadata)', 90, { ruleTrace: [...trace, { id: 'tool-metadata-production-block', matched: true, decision: 'BLOCK', reason: 'Destructive production action is blocked (per tool registration metadata)' }], winningRule: 'tool-metadata-production-block' });
75
+ }
76
+ if (toolMeta.requiresApproval) {
77
+ return decision('ASK', 'Approval required per tool registration metadata', 74, { ruleTrace: [...trace, { id: 'tool-metadata-approval', matched: true, decision: 'ASK', reason: 'Approval required per tool registration metadata' }], winningRule: 'tool-metadata-approval' });
78
+ }
79
+ if (toolMeta.actionClass === 'readOnly') {
80
+ return decision('ALLOW', 'Read-only per tool registration metadata', 5, { ruleTrace: [...trace, { id: 'tool-metadata-readonly', matched: true, decision: 'ALLOW', reason: 'Read-only per tool registration metadata' }], winningRule: 'tool-metadata-readonly' });
81
+ }
82
+ }
32
83
  if (amount > 0 && amount > (policies.autoApproveAmount ?? 500)) {
33
84
  return decision('ASK', 'Amount exceeds the automatic approval threshold', 65, { ruleTrace: [...trace, { id: 'auto-approval-threshold', matched: true, decision: 'ASK', reason: 'Amount exceeds the automatic approval threshold' }], winningRule: 'auto-approval-threshold' });
34
85
  }