agentgate-runtime-control 2.13.15 → 2.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +146 -1
- package/docs/2.14-hardening-plan.md +60 -0
- package/package.json +1 -1
- package/src/control-plane.js +32 -0
- package/src/mcp-gateway.js +17 -2
- package/src/policy-engine.js +52 -1
package/README.md
CHANGED
|
@@ -24,6 +24,18 @@ Useful control-plane endpoints include `/api/observability`, `/api/trace?runId=.
|
|
|
24
24
|
|
|
25
25
|
**Observe → Attack → Enforce → Replay → Report → Govern**
|
|
26
26
|
|
|
27
|
+
### Quickstart
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
npm install agentgate-runtime-control
|
|
31
|
+
npx agentgate init # writes agentgate.config.mjs — tells you to run doctor next
|
|
32
|
+
npx agentgate doctor # checks your config, warns loudly if you're still in observe mode,
|
|
33
|
+
# then tells you to open examples/protect-first-tool.mjs next
|
|
34
|
+
node examples/protect-first-tool.mjs # see a real tool protected end-to-end
|
|
35
|
+
npx agentgate attack --config ./agentgate.config.mjs # attack-test YOUR policy, not the defaults
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
Each command prints what to run next, so you don't have to remember this sequence.
|
|
27
39
|
|
|
28
40
|
## Design Partner Edition
|
|
29
41
|
|
|
@@ -70,6 +82,10 @@ The config file must export an `agentgate` object created with `createAgentGate(
|
|
|
70
82
|
|
|
71
83
|
See [`docs/production-readiness.md`](docs/production-readiness.md), [`docs/production-deployment.md`](docs/production-deployment.md), and [`docs/release-checklist.md`](docs/release-checklist.md) for deployment, operational, performance, and release gates.
|
|
72
84
|
|
|
85
|
+
### About `npm test` on the installed package
|
|
86
|
+
|
|
87
|
+
Running `npm test` inside an **installed** copy of `agentgate-runtime-control` (i.e. from `node_modules`) reports `0 tests` — that's expected, not a bug: the `test/` directory is intentionally not published to npm (see `files` in `package.json`), the same way most published packages don't ship their own test suite to consumers. The real suite (150+ cases, covering policy decisions, the approval lifecycle — including concurrent approve/deny and TTL expiry — attack-lab scenarios, egress guarding, multi-tenant isolation, and more) lives in and runs from the [source repository](https://github.com/walid-agentgate/agentgate) via `node --test`.
|
|
88
|
+
|
|
73
89
|
## Security
|
|
74
90
|
|
|
75
91
|
See [`SECURITY.md`](SECURITY.md) for the security model and vulnerability-reporting guidance. AgentGate provides a deterministic control layer; it does not replace application-level identity, secret management, network isolation, or threat-model testing.
|
|
@@ -159,11 +175,30 @@ console.table(runAttackLab({ productionBlock: true }));
|
|
|
159
175
|
|
|
160
176
|
The built-in lab covers prompt injection, privilege escalation, destructive actions, high-value refunds, and unsafe tool chaining. It is a testing aid, not a guarantee of security.
|
|
161
177
|
|
|
178
|
+
### Deep Attack Lab — the unrecognized-action-name gap
|
|
179
|
+
|
|
180
|
+
The 5 built-in cases above all use action names the policy engine already classifies (`export_all`, `update_production`, `delete`, `refund`, `publish`). A second, larger set specifically attacks action names it does **not** classify — the `unknownActionPolicy` gap described above — plus two "name evasion" cases (the same dangerous action called under a name that isn't in your `blockActions`/`approvalActions`):
|
|
181
|
+
|
|
182
|
+
```bash
|
|
183
|
+
agentgate attack --deep # against built-in default policies
|
|
184
|
+
agentgate attack --deep --config ./agentgate.config.mjs # against YOUR policy
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
or programmatically:
|
|
188
|
+
|
|
189
|
+
```js
|
|
190
|
+
import { runDeepAttackLab, DEEP_ATTACK_CASES } from 'agentgate-runtime-control';
|
|
191
|
+
const results = await runDeepAttackLab(gateway);
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
Under the historical default (`unknownActionPolicy: 'allow'`), most of these legitimately ALLOW — that's the point, and CI should treat that as a finding rather than a passing baseline for anything reachable in production. Set `unknownActionPolicy: 'ask'` or `'block'` and re-run to confirm the gap is closed for your own policy.
|
|
195
|
+
|
|
162
196
|
## CLI
|
|
163
197
|
|
|
164
198
|
```bash
|
|
165
199
|
agentgate test refund 1200
|
|
166
200
|
agentgate attack
|
|
201
|
+
agentgate attack --deep
|
|
167
202
|
```
|
|
168
203
|
|
|
169
204
|
## MCP Gateway
|
|
@@ -227,6 +262,68 @@ const nextPolicy = mergePolicies(currentPolicy, generated.policy);
|
|
|
227
262
|
|
|
228
263
|
Policy generation is deterministic and reviewable. Generated suggestions do not automatically authorize or block traffic until the resulting policy is explicitly applied to a gateway.
|
|
229
264
|
|
|
265
|
+
### `unknownActionPolicy` — what happens to action names AgentGate doesn't recognize
|
|
266
|
+
|
|
267
|
+
The policy engine only classifies a small built-in set of action names as `destructive` (`delete`, `refund`, `publish`, `deploy`, `export_all`, `update_production`) or `readOnly` (`read`, `search`, `list`, `get`, `fetch`). **Any other action name — a typo, a new tool, a third-party integration using its own naming, or something that sounds obviously dangerous like `grant_admin` or `drop_database` — does not match any rule and falls through to `ALLOW` by default.** This is a real gap, not a corner case: it means adding a new tool with an unrecognized action name silently gets no protection at all unless you've explicitly listed it in `approvalActions`/`blockActions`.
|
|
268
|
+
|
|
269
|
+
`policies.unknownActionPolicy` controls that fallback:
|
|
270
|
+
|
|
271
|
+
```js
|
|
272
|
+
policies: {
|
|
273
|
+
// 'allow' (default, kept for backward compatibility): unrecognized actions
|
|
274
|
+
// pass through untouched, exactly as AgentGate has always done.
|
|
275
|
+
// 'ask': unrecognized actions require human approval — the recommended
|
|
276
|
+
// starting point; `agentgate init` sets this for new projects.
|
|
277
|
+
// 'block': unrecognized actions are refused outright — the strictest,
|
|
278
|
+
// deny-by-default option, once every legitimate action name in
|
|
279
|
+
// your system has been classified.
|
|
280
|
+
unknownActionPolicy: 'ask'
|
|
281
|
+
}
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
`agentgate doctor` warns loudly whenever the effective setting is `'allow'`, so this is never a silent gap in a project that runs `doctor` as part of its setup. `examples/protect-first-tool.mjs` demonstrates the gap and the fix side by side with a `grant_admin` call.
|
|
285
|
+
|
|
286
|
+
## Closing the action-name evasion gap further (2.14)
|
|
287
|
+
|
|
288
|
+
A follow-up adversarial test fed the policy engine action names it had never been designed to see: `delete\u0000all` (embedded NUL), `delete` (right-to-left override — makes the name *render* differently than it reads), `%export` (fullwidth lookalikes of "export"), and `../delete` (path-traversal-shaped). None crashed anything, but all of them ALLOWed under the historical default, which is the real finding: a string-based classifier can always be fed a string it wasn't expecting.
|
|
289
|
+
|
|
290
|
+
Three independent changes close this, and they compose with `unknownActionPolicy` above rather than replace it:
|
|
291
|
+
|
|
292
|
+
**1. Unsafe action names are always rejected, unconditionally.** Control characters, bidi-override/zero-width characters, the fullwidth-forms block, and `../`-style path segments have no legitimate reason to appear in an action name, so this check is **not** an opt-in policy — it runs for every request, for every existing user, with no config needed:
|
|
293
|
+
|
|
294
|
+
```text
|
|
295
|
+
delete\u0000all → BLOCK (winningRule: 'unsafe-action-name')
|
|
296
|
+
delete → BLOCK
|
|
297
|
+
%export → BLOCK
|
|
298
|
+
../delete → BLOCK
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
**2. `policies.strictActionNames: true` (opt-in)** additionally requires every action name to match `^[a-z][a-z0-9_.:-]{0,63}$` — a plain lowercase identifier. This is off by default because it's a real behavior change for any system with mixed-case or unusually-shaped action names already in production; turn it on once you've audited your own action-name list.
|
|
302
|
+
|
|
303
|
+
**3. Tool registration metadata is authoritative over the action name a caller sends.** The real fix for "call the dangerous tool under a name the policy engine doesn't recognize" isn't a better regex — it's not trusting the name at all for tools you control. `registerTool()` now accepts classification that lives with the tool, not with whatever the caller's request claims:
|
|
304
|
+
|
|
305
|
+
```js
|
|
306
|
+
gateway.registerTool({
|
|
307
|
+
name: 'remove_customer',
|
|
308
|
+
actionClass: 'destructive', // or 'readOnly'
|
|
309
|
+
requiresApproval: true,
|
|
310
|
+
environments: ['development', 'staging'], // tool refuses to run outside these
|
|
311
|
+
handler: async (args) => { /* ... */ }
|
|
312
|
+
});
|
|
313
|
+
```
|
|
314
|
+
|
|
315
|
+
A caller cannot edit another tool's registration, so `grant_admin` registered with `requiresApproval: true` still requires approval even if a request calls it with `action: 'totally_unrecognized_name'`. Metadata sits below your own explicit `blockActions`/`productionBlock`/amount/`approvalActions` rules (those still win) but above name-based classification and `unknownActionPolicy` — see `test/action-hardening-2.14.test.js` for the exact priority ordering.
|
|
316
|
+
|
|
317
|
+
## Idempotent approval resolution (2.14)
|
|
318
|
+
|
|
319
|
+
`POST /api/approvals/approve` and `POST /api/approvals/deny` accept an `Idempotency-Key` header. A retried request with the same key (same tenant, same path) gets back the **exact same response** instead of re-executing the handler or hitting an "already approved" error:
|
|
320
|
+
|
|
321
|
+
```text
|
|
322
|
+
POST /api/approvals/approve
|
|
323
|
+
Idempotency-Key: approval_123:resolve:v1
|
|
324
|
+
```
|
|
325
|
+
|
|
326
|
+
Replayed responses carry an `Idempotency-Replayed: true` header. Keys are cached in memory per control-plane process for 24h by default (`idempotencyTtlMs` to override); a different key against the same approval is **not** treated as a replay and goes through normal single-use resolution. This matters for exactly the failure mode a retrying client hits in practice: a network blip after the first approve succeeded, followed by an automatic retry that must not double-execute a refund or a production change.
|
|
230
327
|
|
|
231
328
|
## Approval Flow
|
|
232
329
|
|
|
@@ -258,6 +355,26 @@ Approval state is queryable through `gateway.approvals()` and JSON-RPC methods:
|
|
|
258
355
|
|
|
259
356
|
The approval layer is intentionally separate from policy evaluation: policy decides `ALLOW`, `ASK`, or `BLOCK`; approval resolves only the `ASK` path.
|
|
260
357
|
|
|
358
|
+
### Approval lifecycle — who, when, expiry, single-use, revocation
|
|
359
|
+
|
|
360
|
+
- **Who approved / denied, and when**: every approval record carries `createdAt`, `resolvedAt`, and (for a deny) a `resolutionReason`. The run record (`gateway.replay(runId)`) links back to the approval via `approvalId` and stores the same `approval` block for audit export (`gateway.replay()` / `/api/audit/export`). AgentGate itself doesn't have a user identity system, so "who" is whatever identity your own auth layer attaches to the request that calls `approve()`/`deny()` — log that at your call site if you need a named approver.
|
|
361
|
+
- **Expiry (TTL)**: a pending approval expires automatically after **15 minutes** by default (`DEFAULT_APPROVAL_TTL_MS` in `src/approval.js`). Pass `approvalTTLMs` to `createRuntime`/`createAgentGate`/`createMCPGateway` to change it, or `ttlMs: null` on a specific request to disable expiry. Once `expiresAt` passes, the approval flips to `status: 'expired'` the next time it's looked at (list/get/approve/deny), and the original tool call can never be executed late.
|
|
362
|
+
- **Single-use guarantee**: `approve()`/`deny()` are synchronous up to the point where they flip `status` away from `pending` — there is no `await` in between the status check and the status write. Because Node runs JS on a single thread, two calls racing to resolve the same approval (concurrent HTTP requests, a double click, a retried request) can never both see `pending`: the second call always sees the already-resolved status and is rejected with `Approval is already <status>`. There is nothing else to configure for this — it's guaranteed by construction, not by a lock.
|
|
363
|
+
- **Revocation**: there's no separate "revoke" verb — deny a still-pending approval with `gateway.deny(approvalId, reason)` (or `agentgate approval deny <id> <reason>` from the CLI) to take it off the table before anyone acts on it.
|
|
364
|
+
- **Duplicate requests**: each call to a protected tool creates its own approval with its own id — AgentGate does not de-duplicate identical-looking requests. If your agent might retry the same call, treat that as your integration's concern (e.g. an idempotency key on your own tool handler).
|
|
365
|
+
|
|
366
|
+
### Approval CLI
|
|
367
|
+
|
|
368
|
+
Once a gateway or control plane is running (for example via `agentgate dev`), you can list and resolve approvals from the command line instead of writing HTTP calls by hand:
|
|
369
|
+
|
|
370
|
+
```
|
|
371
|
+
agentgate approval list [--status pending|approved|denied|expired] [--url <url>]
|
|
372
|
+
agentgate approval approve <approvalId> [--url <url>] [--key <apiKey>]
|
|
373
|
+
agentgate approval deny <approvalId> [reason] [--url <url>] [--key <apiKey>]
|
|
374
|
+
```
|
|
375
|
+
|
|
376
|
+
By default it talks to `http://localhost:8787` (what `agentgate dev` uses) and, if no `--key`/`AGENTGATE_API_KEY` is given, it automatically picks up the local dev session the same way opening the dashboard in a browser would — no extra setup needed for local testing. Point `--url` at a different host/port for a control plane running elsewhere, and pass `--key` (or set `AGENTGATE_API_KEY`) when auth is required outside local dev.
|
|
377
|
+
|
|
261
378
|
## v1.1 — Developer Integration
|
|
262
379
|
|
|
263
380
|
AgentGate now exposes a single developer-facing runtime:
|
|
@@ -421,6 +538,28 @@ AgentGate supports deterministic RBAC and ABAC authorization using agent/user id
|
|
|
421
538
|
|
|
422
539
|
Runs, approvals, agents, and policy versions can be persisted locally through the built-in storage adapters. For production multi-tenant deployments, use the Postgres/Supabase adapters with tenant-scoped sessions and RLS; local JSON persistence is intended for development or single-process deployments. Storage is provider-neutral so a database adapter can be introduced without changing the policy API.
|
|
423
540
|
|
|
541
|
+
### Persistence corruption is fail-closed, not silent (2.13.15+)
|
|
542
|
+
|
|
543
|
+
An independent stress test of 2.13.14 found that writing garbage into `runs.json` or `approvals.json` and restarting made the server come up normally with an **empty collection** — no error, no alert. For `runs.json` that's a silently erased audit trail; for `approvals.json` it means a pending approval can vanish with no trace at all. That was a bug, not a resilience feature.
|
|
544
|
+
|
|
545
|
+
As of 2.13.15, `PersistentCollectionStore` distinguishes three cases:
|
|
546
|
+
|
|
547
|
+
- **File doesn't exist** (first run) — starts empty, as always. Not an error.
|
|
548
|
+
- **File exists but isn't valid JSON, or isn't an array** — this is corruption. By default, the store throws `PersistenceCorruptionError` (`code: 'AGENTGATE_PERSISTENCE_CORRUPT'`) and refuses to start, so a corrupted security-relevant file is loud at startup instead of discovered later as a gap in the audit log.
|
|
549
|
+
- **Any other read failure** (permission denied, disk I/O error) — also propagates as an error rather than being treated as "no data yet."
|
|
550
|
+
|
|
551
|
+
Recovery is opt-in only, because silently discarding the operator's decision here is exactly the bug being fixed:
|
|
552
|
+
|
|
553
|
+
```js
|
|
554
|
+
createMCPGateway({
|
|
555
|
+
persistence: '.agentgate',
|
|
556
|
+
recoverFromCorruption: true, // off by default — corruption throws unless you opt in
|
|
557
|
+
onPersistenceCorruption: (event) => { /* event.quarantinePath, event.filePath, ... */ }
|
|
558
|
+
});
|
|
559
|
+
```
|
|
560
|
+
|
|
561
|
+
With `recoverFromCorruption: true`, the corrupt file is renamed to `<file>.corrupt.<timestamp>` (never deleted) so it can be inspected, the collection starts from an empty/seed state, and the event is logged loudly to stderr and handed to `onPersistenceCorruption` if provided — it is never silently absorbed. `gateway.persistenceHealth()` and the Control Plane's `GET /api/ready` (which now returns `503` with `{ ready: false, persistence: { degraded: true, corruptions: [...] } }` in this state) let you alert on it rather than assume health.
|
|
562
|
+
|
|
424
563
|
## Multi-Tenant Control Plane (v1.8)
|
|
425
564
|
AgentGate v1.8 adds tenant isolation, scoped API keys, key rotation/revocation, and tenant-scoped webhook registrations.
|
|
426
565
|
|
|
@@ -468,8 +607,14 @@ For first-tenant bootstrap over HTTP, set `AGENTGATE_BOOTSTRAP_TOKEN` and send `
|
|
|
468
607
|
- Deterministic authorization remains the security authority
|
|
469
608
|
|
|
470
609
|
|
|
610
|
+
### Oversized request bodies no longer stall the next request (2.13.15+)
|
|
611
|
+
|
|
612
|
+
A stress test found that a request body larger than `maxBodySize` correctly returned `413`, but the server stopped reading the body as soon as the limit was crossed without draining the rest of what the client was sending. On a keep-alive connection, those unread bytes were still arriving and got misread as the start of the *next* request, so a completely unrelated follow-up request (even an auth check with a fake key) could hang for the full request timeout instead of returning `401` quickly — a cheap way to tie up connections.
|
|
613
|
+
|
|
614
|
+
2.13.15 fixes this by always fully draining the request body (even once it's known to be oversized — the rest is discarded, not buffered) before responding, and by sending `Connection: close` on `413`/`400` body-parsing errors so the client doesn't attempt to reuse a connection that was cut short either way. Regression tests in `test/oversized-body-recovery.test.js` repeat the exact sequence from the report (oversized request, then a fake-key request) and assert the follow-up never exceeds ~2s.
|
|
615
|
+
|
|
471
616
|
### Production Operations v2.1
|
|
472
|
-
- `/api/health` and `/api/ready` health/readiness probes
|
|
617
|
+
- `/api/health` and `/api/ready` health/readiness probes (`/api/ready` reports `503` and `persistence.degraded: true` if persistence was recovered from corruption — see Persistence above)
|
|
473
618
|
- `/api/metrics` deterministic runtime counters and latency telemetry
|
|
474
619
|
- `/api/events` Server-Sent Events stream for runtime events
|
|
475
620
|
- Sliding-window HTTP rate limiting with 429 responses
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# 2.13.15 / 2.14 Hardening — Status and Remaining Infrastructure Plan
|
|
2
|
+
|
|
3
|
+
This tracks the fix plan that came out of the independent 2.13.14 stress-test report. Items are split into what shipped as code in this repo (2.13.15 hotfix + 2.14 action-hardening) versus what requires real infrastructure or a third party and cannot be completed by changing source files alone.
|
|
4
|
+
|
|
5
|
+
## Shipped as code
|
|
6
|
+
|
|
7
|
+
| Item | Release | Where |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| Oversized request body no longer stalls the next request | 2.13.15 | `src/control-plane.js` (`readJsonBody`), `test/oversized-body-recovery.test.js` |
|
|
10
|
+
| Persistence corruption fails closed by default (was silent data loss) | 2.13.15 | `src/persistent-store.js` (`PersistenceCorruptionError`), `test/persistence-corruption.test.js` |
|
|
11
|
+
| `/api/ready` reports `503`/`degraded` on recovered corruption | 2.13.15 | `src/control-plane.js`, `src/mcp-gateway.js` (`persistenceHealth()`) |
|
|
12
|
+
| Unsafe action names (control/bidi-override/fullwidth/path-traversal) blocked unconditionally | 2.14 | `src/policy-engine.js` (`UNSAFE_ACTION_NAME`) |
|
|
13
|
+
| Opt-in strict action-name allowlist (`policies.strictActionNames`) | 2.14 | `src/policy-engine.js` (`STRICT_ACTION_NAME`) |
|
|
14
|
+
| Tool registration metadata (`actionClass`, `requiresApproval`, `environments`) as the source of truth, overriding caller-supplied action names | 2.14 | `src/mcp-gateway.js` (`registerTool`), `src/policy-engine.js` |
|
|
15
|
+
| Idempotency-Key support for `/api/approvals/approve` and `/deny` | 2.14 | `src/control-plane.js` |
|
|
16
|
+
| `test/action-hardening-2.14.test.js`, `test/idempotency-approve-deny.test.js` | 2.14 | regression coverage for all of the above |
|
|
17
|
+
|
|
18
|
+
`policies.unknownActionPolicy` (deny-by-default for action names the engine still doesn't recognize) and the Deep Attack Lab shipped in 2.13.14; this round only adds to it rather than replacing it. None of the above change default behavior in a way that breaks an existing deployment, except `strictActionNames`, which is opt-in.
|
|
19
|
+
|
|
20
|
+
## Still requires real infrastructure or a third party
|
|
21
|
+
|
|
22
|
+
These cannot be "sent as files" — they need an actual environment, actual load, or an actual outside reviewer. Two already have a runnable test pack in this repo; the other two are net new and described below.
|
|
23
|
+
|
|
24
|
+
### Already have a test pack — just needs an environment to run it in
|
|
25
|
+
|
|
26
|
+
- **PostgreSQL/Supabase + RLS acceptance test** — `docs/managed-postgres-acceptance-test.md`. Run against a real managed Postgres staging instance; records RPO/RTO and backup/restore verification.
|
|
27
|
+
- **External security review** — `docs/external-security-review.md` and `docs/external-security-review-test-pack.md`. Needs an independent reviewer to actually execute and sign off; AgentGate's own team cannot self-certify this gate.
|
|
28
|
+
|
|
29
|
+
### Net new — no existing doc, plan below
|
|
30
|
+
|
|
31
|
+
**Disk-full test.** The 2.13.14 report deliberately avoided filling a real disk (`EACCES` via a read-only directory was used as a safe proxy, and it did surface a clear exception — see `src/persistent-store.js:save()`). A real disk-full test needs:
|
|
32
|
+
1. A disposable staging host or container with a **quota-limited volume** (e.g. a small loop-mounted filesystem, or a container volume with a byte quota) — never the host's real root disk.
|
|
33
|
+
2. Fill the quota while `PersistentCollectionStore.save()` is mid-write (temp-file + `renameSync`), and separately while idle, and confirm in both cases: the process surfaces a clear, loud error (not a silent swallow — this is the same class of bug 2.13.15 just fixed for corruption, so it needs its own explicit check); the `.tmp` file never replaces a good `filePath` via a partial `renameSync`; the control plane's `/api/ready` reflects the failure.
|
|
34
|
+
3. Record: error surfaced, state of `filePath` vs `filePath.tmp` after the failure, and whether a subsequent write (after freeing space) recovers cleanly.
|
|
35
|
+
|
|
36
|
+
**Multi-instance / concurrent-writer test.** Local JSON persistence (`PersistentCollectionStore`) is explicitly documented as single-process (see README "Persistence"); this test is about *proving* that boundary rather than discovering it in production:
|
|
37
|
+
1. Point two separate `createMCPGateway({ persistence: sameDir })` processes at the same directory.
|
|
38
|
+
2. Drive concurrent approvals/writes from both.
|
|
39
|
+
3. Expected (and acceptable) outcome: last-writer-wins / lost updates — document this explicitly as the known limitation of local JSON mode, with the recommendation to use the PostgreSQL/Supabase adapter (which has real RLS and a real database handling concurrency) for any multi-instance deployment. The point of this test is a documented, demonstrated boundary — not a fix to local JSON mode itself, which isn't meant to be multi-writer-safe.
|
|
40
|
+
|
|
41
|
+
## Acceptance criteria before using "enterprise-grade" language
|
|
42
|
+
|
|
43
|
+
Carried over from the stress-test report, now with status:
|
|
44
|
+
|
|
45
|
+
| Gate | Target | Status |
|
|
46
|
+
|---|---:|---|
|
|
47
|
+
| Oversized request followed by normal request | 0 timeouts | ✅ shipped 2.13.15, regression-tested |
|
|
48
|
+
| Corrupted persistence startup | fail-closed or documented recovery | ✅ shipped 2.13.15, regression-tested |
|
|
49
|
+
| Pending approvals lost silently | 0 | ✅ shipped 2.13.15 |
|
|
50
|
+
| Unknown sensitive action allowed (name evasion) | 0 | ✅ shipped 2.13.14/2.14 (`unknownActionPolicy` + metadata + unsafe-name block) |
|
|
51
|
+
| Pre-approval handler execution | 0 | ✅ already enforced, regression-tested pre-2.13 |
|
|
52
|
+
| Duplicate approved execution | 0 | ✅ already enforced; 2.14 adds idempotent retries on top |
|
|
53
|
+
| Cross-tenant leak | 0 | ✅ regression-tested (`test/approval-tenant-isolation.test.js`) |
|
|
54
|
+
| Backup/restore | PASS | ⬜ needs a real Postgres staging run — see `docs/managed-postgres-acceptance-test.md` |
|
|
55
|
+
| PostgreSQL RLS | PASS | ⬜ needs a real Postgres staging run — same doc |
|
|
56
|
+
| Disk-full behavior | documented and tested | ⬜ plan above, needs a quota-limited environment |
|
|
57
|
+
| Multi-instance/concurrent writer | documented and tested | ⬜ plan above, needs two real processes |
|
|
58
|
+
| External review | completed or claim explicitly absent | ⬜ needs an actual independent reviewer |
|
|
59
|
+
|
|
60
|
+
Until the ⬜ rows are run against real infrastructure, marketing/sales copy should keep saying what's actually true: the engine, policy, and single-process persistence layer are hardened and regression-tested; multi-instance/Postgres/external-review claims are not yet backed by an executed test run.
|
package/package.json
CHANGED
package/src/control-plane.js
CHANGED
|
@@ -67,6 +67,25 @@ export function createControlPlane(options = {}) {
|
|
|
67
67
|
const limiter = options.rateLimiter || new SlidingWindowLimiter({ limit: options.rateLimit || 120, windowMs: options.rateWindowMs || 60000 });
|
|
68
68
|
const tenantLimiter = options.tenantRateLimiter || new SlidingWindowLimiter({ limit: options.tenantRateLimit || options.rateLimit || 120, windowMs: options.rateWindowMs || 60000 });
|
|
69
69
|
const startedAt = Date.now();
|
|
70
|
+
// Idempotency: a client that retries POST /api/approvals/approve|deny
|
|
71
|
+
// (timeout, connection reset, at-least-once delivery) must not trigger a
|
|
72
|
+
// second side effect. Scoped by tenant + path + the client-supplied
|
|
73
|
+
// Idempotency-Key so a retry with the same key gets back the exact same
|
|
74
|
+
// response instead of re-executing (or hitting "already approved").
|
|
75
|
+
const idempotencyTtlMs = Number(options.idempotencyTtlMs || 24 * 60 * 60 * 1000);
|
|
76
|
+
const idempotencyCache = new Map(); // scopedKey -> { status, result, expiresAt }
|
|
77
|
+
const IDEMPOTENT_PATHS = new Set(['/api/approvals/approve', '/api/approvals/deny']);
|
|
78
|
+
function idempotencyGet(key) {
|
|
79
|
+
if (!key) return null;
|
|
80
|
+
const entry = idempotencyCache.get(key);
|
|
81
|
+
if (!entry) return null;
|
|
82
|
+
if (entry.expiresAt < Date.now()) { idempotencyCache.delete(key); return null; }
|
|
83
|
+
return entry;
|
|
84
|
+
}
|
|
85
|
+
function idempotencyPut(key, status, result) {
|
|
86
|
+
if (!key) return;
|
|
87
|
+
idempotencyCache.set(key, { status, result, expiresAt: Date.now() + idempotencyTtlMs });
|
|
88
|
+
}
|
|
70
89
|
const policyRegistry = options.policyRegistry || createPolicyRegistry({ filePath: options.policyPersistence || `${options.persistence || '.agentgate'}/policies.json` });
|
|
71
90
|
const bundleRegistry = options.policyBundleRegistry || createPolicyBundleRegistry({ filePath: options.policyBundlePersistence || `${options.persistence || '.agentgate'}/policy-bundles.json` });
|
|
72
91
|
const agents = agentStore || { items: [], add(a){ this.items.unshift(a); return a; }, list(filter){ const all=[...this.items]; return typeof filter==='function' ? all.filter(filter) : all; }, get(id){ return this.items.find(a=>a.id===id || a.name===id) || null; }, save(){} };
|
|
@@ -245,9 +264,22 @@ export function createControlPlane(options = {}) {
|
|
|
245
264
|
if (!tenantRate.allowed) { res.writeHead(429, {'content-type':'application/json','cache-control':'no-store','retry-after':String(Math.ceil((Date.parse(tenantRate.resetAt)-Date.now())/1000))}); res.end(JSON.stringify({error:'Tenant rate limit exceeded', ...tenantRate})); return; }
|
|
246
265
|
}
|
|
247
266
|
}
|
|
267
|
+
const idempotencyHeader = req.headers['idempotency-key'];
|
|
268
|
+
const idempotencyKey = idempotencyHeader && IDEMPOTENT_PATHS.has(url.pathname)
|
|
269
|
+
? `${authContext.tenantId || 'no-tenant'}:${url.pathname}:${idempotencyHeader}`
|
|
270
|
+
: null;
|
|
271
|
+
if (idempotencyKey) {
|
|
272
|
+
const cached = idempotencyGet(idempotencyKey);
|
|
273
|
+
if (cached) {
|
|
274
|
+
res.writeHead(cached.status, { 'content-type': 'application/json', 'cache-control': 'no-store', 'idempotency-replayed': 'true' });
|
|
275
|
+
res.end(JSON.stringify(cached.result));
|
|
276
|
+
return;
|
|
277
|
+
}
|
|
278
|
+
}
|
|
248
279
|
const result = await api(url.pathname, req.method, body, Object.fromEntries(url.searchParams.entries()), authContext);
|
|
249
280
|
if (result._html) { res.writeHead(200, { 'content-type': 'text/html; charset=utf-8', 'cache-control': 'no-store' }); res.end(result.html); return; }
|
|
250
281
|
const status = result.error ? 404 : (url.pathname === '/api/ready' && result.ready === false ? 503 : 200);
|
|
282
|
+
if (idempotencyKey) idempotencyPut(idempotencyKey, status, result);
|
|
251
283
|
res.writeHead(status, { 'content-type': 'application/json', 'cache-control': 'no-store' });
|
|
252
284
|
res.end(JSON.stringify(result));
|
|
253
285
|
return;
|
package/src/mcp-gateway.js
CHANGED
|
@@ -111,7 +111,17 @@ export function createMCPGateway(options = {}) {
|
|
|
111
111
|
name: tool.name,
|
|
112
112
|
description: tool.description || '',
|
|
113
113
|
inputSchema: tool.inputSchema || { type: 'object' },
|
|
114
|
-
handler: tool.handler
|
|
114
|
+
handler: tool.handler,
|
|
115
|
+
// Optional classification recorded at registration time, i.e. by
|
|
116
|
+
// whoever controls the tool list — not by whatever action name a
|
|
117
|
+
// caller's request happens to send. This is what makes it trustworthy:
|
|
118
|
+
// a caller can rename "remove_customer" to "delete" or to a brand-new
|
|
119
|
+
// string the policy engine has never seen, but it cannot edit this
|
|
120
|
+
// tool's registered metadata. See policy-engine.js's handling of
|
|
121
|
+
// request.toolMetadata for how this takes priority over name matching.
|
|
122
|
+
actionClass: tool.actionClass || null, // 'destructive' | 'readOnly' | null
|
|
123
|
+
requiresApproval: tool.requiresApproval === true,
|
|
124
|
+
environments: Array.isArray(tool.environments) ? tool.environments : null
|
|
115
125
|
});
|
|
116
126
|
return api;
|
|
117
127
|
}
|
|
@@ -157,7 +167,12 @@ export function createMCPGateway(options = {}) {
|
|
|
157
167
|
model: params.model || params.context?.model || input.model,
|
|
158
168
|
usage: params.usage || params.context?.usage || input.usage,
|
|
159
169
|
estimatedCost: params.estimatedCost ?? params.context?.estimatedCost ?? input.estimatedCost,
|
|
160
|
-
context: params.context || {}
|
|
170
|
+
context: params.context || {},
|
|
171
|
+
// The tool's own registered metadata (never caller-suppliable) — see
|
|
172
|
+
// registerTool() above and policy-engine.js's toolMetadata handling.
|
|
173
|
+
toolMetadata: (tool.actionClass || tool.requiresApproval || tool.environments)
|
|
174
|
+
? { actionClass: tool.actionClass, requiresApproval: tool.requiresApproval, environments: tool.environments }
|
|
175
|
+
: null
|
|
161
176
|
};
|
|
162
177
|
if (killSwitch && mode === 'enforce') {
|
|
163
178
|
const run = { id: randomUUID(), createdAt:new Date().toISOString(), tenantId:request.tenantId, request, decision:'BLOCK', risk:100, reason:killReason.value, mode, killed:true };
|
package/src/policy-engine.js
CHANGED
|
@@ -5,11 +5,38 @@ export const DECISIONS = Object.freeze({ ALLOW: 'ALLOW', ASK: 'ASK', BLOCK: 'BLO
|
|
|
5
5
|
const destructive = new Set(['delete','refund','publish','deploy','export_all','update_production']);
|
|
6
6
|
const readOnly = new Set(['read','search','list','get','fetch']);
|
|
7
7
|
|
|
8
|
+
// Characters that have no legitimate reason to appear in an action name and
|
|
9
|
+
// exist only to confuse logs/matching or to smuggle a control sequence past
|
|
10
|
+
// string-equality checks: NUL and other control bytes, zero-width/bidi
|
|
11
|
+
// override characters (can make "delete" *look* like something else when
|
|
12
|
+
// rendered, or hide characters entirely), the fullwidth-forms block (an
|
|
13
|
+
// adversarial test used "%export" — fullwidth lookalikes of ASCII — to
|
|
14
|
+
// probe whether it matched "export"), and "../"-style path segments. This
|
|
15
|
+
// check is unconditional and always on (not an opt-in policy) because no
|
|
16
|
+
// real tool/action name is ever affected by it — unlike a strict allowlist
|
|
17
|
+
// regex, which IS opt-in below to avoid breaking existing naming schemes.
|
|
18
|
+
const UNSAFE_ACTION_NAME = new RegExp(
|
|
19
|
+
'[\\x00-\\x1F\\x7F\\u200B-\\u200F\\u202A-\\u202E\\u2066-\\u2069\\uFEFF\\uFF00-\\uFFEF]|\\.\\.[\\/\\\\]'
|
|
20
|
+
);
|
|
21
|
+
// Opt-in, stricter allowlist for teams that control every action name they
|
|
22
|
+
// register: policies.strictActionNames: true requires lowercase
|
|
23
|
+
// ascii-identifier-shaped names. Off by default so existing naming schemes
|
|
24
|
+
// (mixed case, spaces, unicode brand/customer names, etc.) don't suddenly
|
|
25
|
+
// start failing on upgrade.
|
|
26
|
+
const STRICT_ACTION_NAME = /^[a-z][a-z0-9_.:-]{0,63}$/;
|
|
27
|
+
|
|
8
28
|
export function evaluate(input = {}, policies = {}) {
|
|
9
29
|
const auth = authorize(input, policies.authorization);
|
|
10
30
|
const trace = [{ id: 'authorization', matched: auth.decision === 'BLOCK', decision: auth.decision, reason: auth.reason, matchedRule: auth.matchedRule || null }];
|
|
11
31
|
if (auth.decision === 'BLOCK') return decision('BLOCK', auth.reason, 96, { authorization: auth, ruleTrace: trace, winningRule: auth.matchedRule || 'authorization' });
|
|
12
|
-
const
|
|
32
|
+
const rawAction = String(input.action || '');
|
|
33
|
+
if (UNSAFE_ACTION_NAME.test(rawAction)) {
|
|
34
|
+
return decision('BLOCK', 'Action name contains unsafe characters (control, bidi-override, fullwidth-form, or path-traversal)', 99, { ruleTrace: [...trace, { id: 'unsafe-action-name', matched: true, decision: 'BLOCK', reason: 'Action name contains unsafe characters' }], winningRule: 'unsafe-action-name' });
|
|
35
|
+
}
|
|
36
|
+
const action = rawAction.toLowerCase();
|
|
37
|
+
if (policies.strictActionNames && !STRICT_ACTION_NAME.test(action)) {
|
|
38
|
+
return decision('BLOCK', 'Action name does not match the required format under strictActionNames', 97, { ruleTrace: [...trace, { id: 'strict-action-name', matched: true, decision: 'BLOCK', reason: 'Action name does not match the required format under strictActionNames' }], winningRule: 'strict-action-name' });
|
|
39
|
+
}
|
|
13
40
|
const hasAmount = Object.prototype.hasOwnProperty.call(input, 'amount') && input.amount !== undefined && input.amount !== null;
|
|
14
41
|
if (hasAmount && (typeof input.amount !== 'number' || !Number.isFinite(input.amount) || input.amount < 0)) {
|
|
15
42
|
return decision('BLOCK', 'Amount must be a finite non-negative number', 98, { ruleTrace: [...trace, { id: 'invalid-amount', matched: true, decision: 'BLOCK', reason: 'Amount must be a finite non-negative number' }], winningRule: 'invalid-amount' });
|
|
@@ -29,6 +56,30 @@ export function evaluate(input = {}, policies = {}) {
|
|
|
29
56
|
if (policies.approvalActions?.includes(action)) {
|
|
30
57
|
return decision('ASK', 'Approval required for sensitive action', 70, { ruleTrace: [...trace, { id: 'approval-action', matched: true, decision: 'ASK', reason: 'Approval required for sensitive action' }], winningRule: 'approval-action' });
|
|
31
58
|
}
|
|
59
|
+
// A tool's own registered metadata (set by whoever controls the tool list,
|
|
60
|
+
// via registerTool({ actionClass, requiresApproval, environments })) is a
|
|
61
|
+
// more trustworthy signal than the action *name* a caller's request
|
|
62
|
+
// happens to send, because a caller can pick any action string it wants
|
|
63
|
+
// but cannot edit another tool's registration. It sits below the admin's
|
|
64
|
+
// own explicit blockActions/productionBlock/amount/approvalActions rules
|
|
65
|
+
// above (those still win), but above name-based classification and
|
|
66
|
+
// unknownActionPolicy below — this is what closes the "call the dangerous
|
|
67
|
+
// tool under an action name the policy engine doesn't recognize" evasion.
|
|
68
|
+
const toolMeta = input.toolMetadata;
|
|
69
|
+
if (toolMeta) {
|
|
70
|
+
if (Array.isArray(toolMeta.environments) && toolMeta.environments.length && !toolMeta.environments.includes(env)) {
|
|
71
|
+
return decision('BLOCK', 'Tool is not permitted in this environment per its registered metadata', 93, { ruleTrace: [...trace, { id: 'tool-metadata-environment', matched: true, decision: 'BLOCK', reason: 'Tool is not permitted in this environment per its registered metadata' }], winningRule: 'tool-metadata-environment' });
|
|
72
|
+
}
|
|
73
|
+
if (toolMeta.actionClass === 'destructive' && policies.productionBlock && env === 'production') {
|
|
74
|
+
return decision('BLOCK', 'Destructive production action is blocked (per tool registration metadata)', 90, { ruleTrace: [...trace, { id: 'tool-metadata-production-block', matched: true, decision: 'BLOCK', reason: 'Destructive production action is blocked (per tool registration metadata)' }], winningRule: 'tool-metadata-production-block' });
|
|
75
|
+
}
|
|
76
|
+
if (toolMeta.requiresApproval) {
|
|
77
|
+
return decision('ASK', 'Approval required per tool registration metadata', 74, { ruleTrace: [...trace, { id: 'tool-metadata-approval', matched: true, decision: 'ASK', reason: 'Approval required per tool registration metadata' }], winningRule: 'tool-metadata-approval' });
|
|
78
|
+
}
|
|
79
|
+
if (toolMeta.actionClass === 'readOnly') {
|
|
80
|
+
return decision('ALLOW', 'Read-only per tool registration metadata', 5, { ruleTrace: [...trace, { id: 'tool-metadata-readonly', matched: true, decision: 'ALLOW', reason: 'Read-only per tool registration metadata' }], winningRule: 'tool-metadata-readonly' });
|
|
81
|
+
}
|
|
82
|
+
}
|
|
32
83
|
if (amount > 0 && amount > (policies.autoApproveAmount ?? 500)) {
|
|
33
84
|
return decision('ASK', 'Amount exceeds the automatic approval threshold', 65, { ruleTrace: [...trace, { id: 'auto-approval-threshold', matched: true, decision: 'ASK', reason: 'Amount exceeds the automatic approval threshold' }], winningRule: 'auto-approval-threshold' });
|
|
34
85
|
}
|