agentgate-runtime-control 2.13.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +229 -0
- package/LICENSE +21 -0
- package/README.md +534 -0
- package/SECURITY.md +19 -0
- package/bin/agentgate.js +303 -0
- package/docs/case-study-technical-validation.md +25 -0
- package/docs/case-study-template.md +37 -0
- package/docs/data-protection.md +49 -0
- package/docs/design-partner-checklist.md +32 -0
- package/docs/design-partner-kit.md +51 -0
- package/docs/design-partner-rollout.md +35 -0
- package/docs/design-partner.md +67 -0
- package/docs/external-security-review-test-pack.md +132 -0
- package/docs/external-security-review.md +33 -0
- package/docs/incident-response.md +54 -0
- package/docs/integration-matrix.md +17 -0
- package/docs/managed-postgres-acceptance-test.md +138 -0
- package/docs/marketing-plan.md +33 -0
- package/docs/observability-alerting.md +44 -0
- package/docs/outreach.md +26 -0
- package/docs/partner-intake-template.md +26 -0
- package/docs/performance-baseline.md +23 -0
- package/docs/performance.md +27 -0
- package/docs/pricing.md +53 -0
- package/docs/production-deployment.md +70 -0
- package/docs/production-quickstart.md +58 -0
- package/docs/production-readiness.md +29 -0
- package/docs/quickstart.md +115 -0
- package/docs/release-checklist.md +33 -0
- package/docs/security-hardening-release-report.md +69 -0
- package/docs/threat-model.md +47 -0
- package/docs/website-copy.md +44 -0
- package/examples/basic.mjs +14 -0
- package/examples/control-plane.mjs +17 -0
- package/examples/design-partner-refund.mjs +33 -0
- package/examples/design-partner-shadow.mjs +27 -0
- package/examples/mcp-gateway.mjs +26 -0
- package/examples/policy-bundle.mjs +18 -0
- package/examples/refund-agent.mjs +20 -0
- package/examples/runtime.mjs +12 -0
- package/package.json +49 -0
- package/schema/postgres.sql +17 -0
- package/src/admin-rbac.js +3 -0
- package/src/agentgate.js +85 -0
- package/src/approval.js +30 -0
- package/src/attack-lab.js +94 -0
- package/src/auth.js +27 -0
- package/src/behavior.js +146 -0
- package/src/control-plane.js +215 -0
- package/src/egress-guard.js +132 -0
- package/src/event-bus.js +10 -0
- package/src/identity.js +109 -0
- package/src/index.js +44 -0
- package/src/local-experience.js +46 -0
- package/src/mcp-gateway.js +383 -0
- package/src/mcp-scanner.js +45 -0
- package/src/middleware.js +17 -0
- package/src/multi-tenant.js +29 -0
- package/src/observability.js +395 -0
- package/src/oidc.js +38 -0
- package/src/persistent-store.js +56 -0
- package/src/policy-builder.js +74 -0
- package/src/policy-bundles.js +17 -0
- package/src/policy-engine.js +67 -0
- package/src/policy-packs.js +115 -0
- package/src/policy-registry.js +58 -0
- package/src/postgres-adapter.js +76 -0
- package/src/runtime.js +138 -0
- package/src/saas.js +67 -0
- package/src/security-report.js +42 -0
- package/src/security-validation.js +92 -0
- package/src/shadow-mode.js +47 -0
- package/src/telemetry.js +28 -0
- package/src/webhook-delivery.js +70 -0
- package/standalone.html +86 -0
|
@@ -0,0 +1,132 @@
|
|
|
1
|
+
# AgentGate — External Security Review Test Pack
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
This is a test package for an **independent reviewer**. It is not an internal audit and must not be described as an external security review until an independent reviewer returns a dated report.
|
|
6
|
+
|
|
7
|
+
## Evidence baseline
|
|
8
|
+
|
|
9
|
+
Reviewer should record:
|
|
10
|
+
|
|
11
|
+
- AgentGate version and SHA-256 of the release ZIP.
|
|
12
|
+
- Node.js version and OS.
|
|
13
|
+
- `npm test` result.
|
|
14
|
+
- `npm audit --omit=dev` result.
|
|
15
|
+
- `node bin/agentgate.js doctor`.
|
|
16
|
+
- `node bin/agentgate.js validate-security`.
|
|
17
|
+
- `node bin/agentgate.js attack-ci`.
|
|
18
|
+
- HTTP Control Plane authentication mode.
|
|
19
|
+
- PostgreSQL version and RLS configuration when persistence is in scope.
|
|
20
|
+
|
|
21
|
+
## Required security tests
|
|
22
|
+
|
|
23
|
+
### A. Authentication / Authorization
|
|
24
|
+
|
|
25
|
+
1. Access every protected Control Plane endpoint without credentials.
|
|
26
|
+
2. Use an invalid API key.
|
|
27
|
+
3. Use a valid key belonging to Tenant A against Tenant B resources.
|
|
28
|
+
4. Attempt approval with a key that cannot approve that tenant/run.
|
|
29
|
+
5. Attempt replay with a key from another tenant.
|
|
30
|
+
6. Verify failures are fail-closed (`401`/`403`) and do not leak data.
|
|
31
|
+
|
|
32
|
+
Expected: no unauthorized action or cross-tenant data exposure.
|
|
33
|
+
|
|
34
|
+
### B. Runtime enforcement
|
|
35
|
+
|
|
36
|
+
Run the built-in Attack Lab and verify:
|
|
37
|
+
|
|
38
|
+
- Prompt injection → BLOCK, execution 0.
|
|
39
|
+
- Privilege escalation → BLOCK, execution 0.
|
|
40
|
+
- Destructive tool → BLOCK, execution 0.
|
|
41
|
+
- High-value refund → BLOCK, execution 0.
|
|
42
|
+
- Unsafe tool chaining → BLOCK, execution 0.
|
|
43
|
+
|
|
44
|
+
### C. Approval security
|
|
45
|
+
|
|
46
|
+
1. Create an ASK request.
|
|
47
|
+
2. Verify execution is 0 before approval.
|
|
48
|
+
3. Approve once and verify exactly one execution.
|
|
49
|
+
4. Replay the same approval request concurrently 10 times.
|
|
50
|
+
5. Verify only one approval wins and only one execution occurs.
|
|
51
|
+
6. Try approval from another tenant.
|
|
52
|
+
|
|
53
|
+
Expected: one successful resolution, no duplicate execution, no cross-tenant approval.
|
|
54
|
+
|
|
55
|
+
### D. Egress / secret leakage
|
|
56
|
+
|
|
57
|
+
Send representative outputs containing:
|
|
58
|
+
|
|
59
|
+
- `secret`
|
|
60
|
+
- `password`
|
|
61
|
+
- `passphrase`
|
|
62
|
+
- `private_key`
|
|
63
|
+
- `access_token`
|
|
64
|
+
- `refresh_token`
|
|
65
|
+
- `authorization`
|
|
66
|
+
- `client_secret`
|
|
67
|
+
- JWT
|
|
68
|
+
- API keys
|
|
69
|
+
- cloud access keys
|
|
70
|
+
- PEM private keys
|
|
71
|
+
- email and phone PII
|
|
72
|
+
|
|
73
|
+
Expected: sensitive values are REDACTED or BLOCKED according to policy and never appear in the returned output/audit preview where policy says they must not.
|
|
74
|
+
|
|
75
|
+
### E. Input robustness
|
|
76
|
+
|
|
77
|
+
Try malformed JSON, missing fields, wrong types, arrays where objects are expected, extremely long strings, nulls, duplicated fields, unexpected action/tool names, and invalid approval IDs.
|
|
78
|
+
|
|
79
|
+
Expected: controlled error; no process crash; no bypass to tool execution.
|
|
80
|
+
|
|
81
|
+
### F. MCP / tool boundary
|
|
82
|
+
|
|
83
|
+
1. Call an unregistered tool.
|
|
84
|
+
2. Call a registered destructive tool with a policy bypass attempt.
|
|
85
|
+
3. Attempt tool chaining where the second action is privileged.
|
|
86
|
+
4. Attempt to alter tool identity/action through user-controlled arguments.
|
|
87
|
+
|
|
88
|
+
Expected: authorization is based on trusted runtime context/policy, not merely model-provided text.
|
|
89
|
+
|
|
90
|
+
### G. Replay / audit integrity
|
|
91
|
+
|
|
92
|
+
Verify a blocked run contains:
|
|
93
|
+
|
|
94
|
+
- runId
|
|
95
|
+
- decision
|
|
96
|
+
- tool/action
|
|
97
|
+
- request/arguments as permitted by logging policy
|
|
98
|
+
- winningRule
|
|
99
|
+
- ruleTrace
|
|
100
|
+
- execution outcome
|
|
101
|
+
- approval state when applicable
|
|
102
|
+
|
|
103
|
+
Attempt to read or modify another tenant's run.
|
|
104
|
+
|
|
105
|
+
Expected: complete evidence for authorized users and no unauthorized access/modification.
|
|
106
|
+
|
|
107
|
+
### H. Persistence / PostgreSQL
|
|
108
|
+
|
|
109
|
+
Use the managed PostgreSQL acceptance pack in `docs/managed-postgres-acceptance-test.md`.
|
|
110
|
+
|
|
111
|
+
### I. Availability / abuse controls
|
|
112
|
+
|
|
113
|
+
Test rate limits, oversized requests, repeated authentication failures, approval floods, and run-retention pressure in a production-like environment.
|
|
114
|
+
|
|
115
|
+
Expected: bounded resource usage, controlled errors, and no security boundary bypass.
|
|
116
|
+
|
|
117
|
+
## Reviewer report format
|
|
118
|
+
|
|
119
|
+
The independent reviewer should return:
|
|
120
|
+
|
|
121
|
+
1. Scope and exclusions.
|
|
122
|
+
2. Environment and versions.
|
|
123
|
+
3. Methodology.
|
|
124
|
+
4. Test cases executed.
|
|
125
|
+
5. Findings with severity (Critical/High/Medium/Low/Informational).
|
|
126
|
+
6. Reproduction steps for each finding.
|
|
127
|
+
7. Evidence.
|
|
128
|
+
8. Remediation status.
|
|
129
|
+
9. Retest results.
|
|
130
|
+
10. Date and reviewer identity/organization.
|
|
131
|
+
|
|
132
|
+
Only after this report exists should AgentGate state that an external security review was completed.
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# External Security Review Package
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
**Not completed.** This repository contains a preparation package only. No third-party security audit, penetration test, SOC 2 assessment, or certification is claimed by AgentGate.
|
|
6
|
+
|
|
7
|
+
## Scope for an external reviewer
|
|
8
|
+
|
|
9
|
+
1. Runtime policy enforcement and pre-execution boundary.
|
|
10
|
+
2. Approval lifecycle and concurrent approval race.
|
|
11
|
+
3. MCP Gateway and tool registration.
|
|
12
|
+
4. HTTP Control Plane authentication and authorization.
|
|
13
|
+
5. Tenant isolation and replay access control.
|
|
14
|
+
6. Egress inspection/redaction/blocking.
|
|
15
|
+
7. Persistence and Postgres/Supabase RLS assumptions.
|
|
16
|
+
8. Audit integrity and export.
|
|
17
|
+
9. Local/dev authentication boundaries.
|
|
18
|
+
10. Dependency and supply-chain exposure.
|
|
19
|
+
|
|
20
|
+
## Evidence to provide
|
|
21
|
+
|
|
22
|
+
- Release ZIP and SHA-256.
|
|
23
|
+
- `npm pack --dry-run` output.
|
|
24
|
+
- Automated test report.
|
|
25
|
+
- Threat model.
|
|
26
|
+
- Production deployment architecture.
|
|
27
|
+
- Database schema/RLS policy.
|
|
28
|
+
- Incident-response runbook.
|
|
29
|
+
- Backup/restore evidence.
|
|
30
|
+
|
|
31
|
+
## Acceptance
|
|
32
|
+
|
|
33
|
+
The review is considered complete only when an independent reviewer supplies a dated report with scope, methodology, findings, severity, and remediation status.
|
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
# Incident Response Runbook
|
|
2
|
+
|
|
3
|
+
This runbook is an operational template for AgentGate deployments. It does not replace the customer's incident-response policy.
|
|
4
|
+
|
|
5
|
+
## Severity
|
|
6
|
+
|
|
7
|
+
### SEV-1
|
|
8
|
+
Confirmed or suspected unauthorized side effect, cross-tenant data exposure, credential compromise, or bypass of an enforced security boundary.
|
|
9
|
+
|
|
10
|
+
### SEV-2
|
|
11
|
+
Security control degradation without confirmed unauthorized side effect, persistent service failure, or material audit/replay loss.
|
|
12
|
+
|
|
13
|
+
### SEV-3
|
|
14
|
+
Non-security defect, degraded dashboard, reporting issue, or recoverable integration problem.
|
|
15
|
+
|
|
16
|
+
## Immediate containment
|
|
17
|
+
|
|
18
|
+
1. Preserve relevant run IDs, approval IDs, policy versions, timestamps, and request metadata.
|
|
19
|
+
2. Activate the kill switch or move affected tools to Block/Observe as appropriate.
|
|
20
|
+
3. Revoke or rotate compromised credentials.
|
|
21
|
+
4. Isolate the affected tenant/tool/integration.
|
|
22
|
+
5. Preserve database and application logs before cleanup.
|
|
23
|
+
|
|
24
|
+
## Investigation
|
|
25
|
+
|
|
26
|
+
Correlate:
|
|
27
|
+
|
|
28
|
+
```text
|
|
29
|
+
runId
|
|
30
|
+
approvalId
|
|
31
|
+
tenantId
|
|
32
|
+
agent
|
|
33
|
+
policy version
|
|
34
|
+
winningRule
|
|
35
|
+
ruleTrace
|
|
36
|
+
tool/action
|
|
37
|
+
time window
|
|
38
|
+
egress findings
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Use Replay and Audit Export before modifying historical records.
|
|
42
|
+
|
|
43
|
+
## Recovery
|
|
44
|
+
|
|
45
|
+
1. Patch or roll back the affected runtime/policy.
|
|
46
|
+
2. Re-run the relevant Attack Lab and customer-specific abuse cases.
|
|
47
|
+
3. Verify tenant isolation.
|
|
48
|
+
4. Verify approval single-execution behavior.
|
|
49
|
+
5. Restore from backup if data integrity is affected.
|
|
50
|
+
6. Document the exact recovery point and customer impact.
|
|
51
|
+
|
|
52
|
+
## Post-incident
|
|
53
|
+
|
|
54
|
+
Record root cause, affected versions, scope, evidence, containment, customer notification decision, corrective action, and regression tests. Add a regression test for every confirmed bypass or operational failure.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# Integration Matrix
|
|
2
|
+
|
|
3
|
+
| Integration | Status | Primary path | Notes |
|
|
4
|
+
|---|---|---|---|
|
|
5
|
+
| Node.js SDK | Supported | `createAgentGate`, `protect`, `createRuntime` | Primary developer experience |
|
|
6
|
+
| TypeScript | Supported via Node API | SDK exports | Types can be layered by the consuming project |
|
|
7
|
+
| MCP | Supported | MCP Gateway | Runtime enforcement + approval |
|
|
8
|
+
| HTTP Control Plane | Supported | `/api/*` | Auth/tenant scoping required in production |
|
|
9
|
+
| PostgreSQL | Supported | `PostgresStoreAdapter` | RLS + tenant scope |
|
|
10
|
+
| Supabase | Supported | `createSupabaseAdapter` | Use RLS and tenant predicates |
|
|
11
|
+
| OpenAI Agents | Integration pattern | Protect the side-effecting tool | No vendor lock-in claim |
|
|
12
|
+
| Anthropic agents | Integration pattern | Protect the side-effecting tool/MCP path | Validate SDK/tool adapter in pilot |
|
|
13
|
+
| LangChain | Integration pattern | Protect the tool boundary | Validate in pilot |
|
|
14
|
+
| Vercel AI | Integration pattern | Protect server-side tool handlers | Do not expose server secrets to browser |
|
|
15
|
+
| Python/.NET/Go | Not first-class SDKs | HTTP/MCP boundary | Use Control Plane/MCP until native SDK exists |
|
|
16
|
+
|
|
17
|
+
AgentGate should not claim native framework support unless the integration has a maintained example and release test.
|
|
@@ -0,0 +1,138 @@
|
|
|
1
|
+
# AgentGate — Managed PostgreSQL Production Acceptance Test
|
|
2
|
+
|
|
3
|
+
## Purpose
|
|
4
|
+
|
|
5
|
+
Run this against a **managed PostgreSQL test/staging database** (for example a managed PostgreSQL provider), not a local developer database. Do not use production customer data for destructive restore testing.
|
|
6
|
+
|
|
7
|
+
## Preconditions
|
|
8
|
+
|
|
9
|
+
- PostgreSQL endpoint and database credentials.
|
|
10
|
+
- Application role with only the permissions required by AgentGate.
|
|
11
|
+
- Separate migration/admin role.
|
|
12
|
+
- TLS enabled for database connections.
|
|
13
|
+
- `schema/postgres.sql` applied.
|
|
14
|
+
- RLS enabled and forced on `agentgate_records`.
|
|
15
|
+
- At least two tenants: `tenant-a` and `tenant-b`.
|
|
16
|
+
|
|
17
|
+
## Test 1 — Schema and RLS
|
|
18
|
+
|
|
19
|
+
Verify:
|
|
20
|
+
|
|
21
|
+
```sql
|
|
22
|
+
SELECT relrowsecurity, relforcerowsecurity
|
|
23
|
+
FROM pg_class
|
|
24
|
+
WHERE oid = 'agentgate_records'::regclass;
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Expected: both values are `true`.
|
|
28
|
+
|
|
29
|
+
Verify the policy:
|
|
30
|
+
|
|
31
|
+
```sql
|
|
32
|
+
SELECT policyname, permissive, roles, cmd
|
|
33
|
+
FROM pg_policies
|
|
34
|
+
WHERE tablename = 'agentgate_records';
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Expected: tenant isolation policy exists.
|
|
38
|
+
|
|
39
|
+
## Test 2 — Tenant A isolation
|
|
40
|
+
|
|
41
|
+
In an application connection:
|
|
42
|
+
|
|
43
|
+
```sql
|
|
44
|
+
BEGIN;
|
|
45
|
+
SELECT set_config('agentgate.tenant_id', 'tenant-a', true);
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
Create/read Tenant A records through `PostgresStoreAdapter`.
|
|
49
|
+
|
|
50
|
+
Then attempt to read Tenant B by ID/list query.
|
|
51
|
+
|
|
52
|
+
Expected: zero Tenant B rows.
|
|
53
|
+
|
|
54
|
+
## Test 3 — Cross-tenant write
|
|
55
|
+
|
|
56
|
+
With `agentgate.tenant_id = tenant-a`, attempt to insert a row with `tenant_id = tenant-b`.
|
|
57
|
+
|
|
58
|
+
Expected: PostgreSQL rejects it with an RLS policy violation.
|
|
59
|
+
|
|
60
|
+
## Test 4 — Backup
|
|
61
|
+
|
|
62
|
+
Create known test records:
|
|
63
|
+
|
|
64
|
+
- Tenant A Run: BLOCK, execution 0.
|
|
65
|
+
- Tenant A Approval: pending.
|
|
66
|
+
- Tenant B Run: ASK.
|
|
67
|
+
- Tenant B Approval: approved.
|
|
68
|
+
|
|
69
|
+
Run a provider-native backup or `pg_dump` where permitted.
|
|
70
|
+
|
|
71
|
+
Record backup identifier, timestamp, size, and checksum if available.
|
|
72
|
+
|
|
73
|
+
## Test 5 — Restore drill
|
|
74
|
+
|
|
75
|
+
Restore into an isolated staging database/clone.
|
|
76
|
+
|
|
77
|
+
Verify:
|
|
78
|
+
|
|
79
|
+
- Tenant A Run exists.
|
|
80
|
+
- Tenant A Approval exists.
|
|
81
|
+
- Tenant B Run exists.
|
|
82
|
+
- Tenant B Approval exists.
|
|
83
|
+
- Replay fields survive.
|
|
84
|
+
- `winningRule` survives.
|
|
85
|
+
- `ruleTrace` survives.
|
|
86
|
+
- execution count/state survives.
|
|
87
|
+
- RLS remains enabled and forced.
|
|
88
|
+
- cross-tenant read remains blocked.
|
|
89
|
+
|
|
90
|
+
## Test 6 — Application role boundaries
|
|
91
|
+
|
|
92
|
+
Verify the application role cannot:
|
|
93
|
+
|
|
94
|
+
- `DROP DATABASE`
|
|
95
|
+
- create/drop arbitrary databases
|
|
96
|
+
- become superuser
|
|
97
|
+
- disable RLS
|
|
98
|
+
- alter security policies
|
|
99
|
+
- perform unrestricted cross-tenant reads
|
|
100
|
+
|
|
101
|
+
Migration/admin credentials must be separate.
|
|
102
|
+
|
|
103
|
+
## Test 7 — TLS
|
|
104
|
+
|
|
105
|
+
Verify the application connection uses TLS and certificate verification appropriate to the provider.
|
|
106
|
+
|
|
107
|
+
Record the provider, TLS mode, and certificate verification configuration.
|
|
108
|
+
|
|
109
|
+
## Test 8 — Backup recovery objective
|
|
110
|
+
|
|
111
|
+
Record:
|
|
112
|
+
|
|
113
|
+
- backup frequency
|
|
114
|
+
- retention
|
|
115
|
+
- last successful backup timestamp
|
|
116
|
+
- restore start/end time
|
|
117
|
+
- RPO achieved
|
|
118
|
+
- RTO achieved
|
|
119
|
+
- whether restore verification is automated
|
|
120
|
+
|
|
121
|
+
## Acceptance criteria
|
|
122
|
+
|
|
123
|
+
| Gate | Required result |
|
|
124
|
+
|---|---|
|
|
125
|
+
| RLS enabled | PASS |
|
|
126
|
+
| RLS forced | PASS |
|
|
127
|
+
| Cross-tenant read | 0 rows / denied |
|
|
128
|
+
| Cross-tenant write | denied |
|
|
129
|
+
| Backup | successful |
|
|
130
|
+
| Restore | successful |
|
|
131
|
+
| Runs restored | PASS |
|
|
132
|
+
| Approvals restored | PASS |
|
|
133
|
+
| Replay restored | PASS |
|
|
134
|
+
| Tenant isolation after restore | PASS |
|
|
135
|
+
| App role least privilege | PASS |
|
|
136
|
+
| TLS | PASS |
|
|
137
|
+
|
|
138
|
+
A single failure in tenant isolation or unauthorized application privilege is a security blocker and should be remediated before production use.
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# AgentGate Marketing Plan — After Design Partner Readiness
|
|
2
|
+
|
|
3
|
+
## Phase 1 — Evidence
|
|
4
|
+
Secure 3–5 design partners and publish only partner-approved results.
|
|
5
|
+
|
|
6
|
+
## Phase 2 — Positioning
|
|
7
|
+
Use one category statement:
|
|
8
|
+
|
|
9
|
+
> AgentGate is the runtime control plane for AI agent actions.
|
|
10
|
+
|
|
11
|
+
Core promise:
|
|
12
|
+
|
|
13
|
+
> Stop dangerous agent actions before they execute — and prove what happened.
|
|
14
|
+
|
|
15
|
+
## Phase 3 — Content
|
|
16
|
+
1. Technical demo: refund agent.
|
|
17
|
+
2. Attack Lab demonstration.
|
|
18
|
+
3. Shadow → Enforce walkthrough.
|
|
19
|
+
4. Replay evidence walkthrough.
|
|
20
|
+
5. Partner case study when approved.
|
|
21
|
+
|
|
22
|
+
## Phase 4 — Distribution
|
|
23
|
+
- GitHub / npm developer funnel.
|
|
24
|
+
- Founder-led technical outreach.
|
|
25
|
+
- AI engineering/security communities.
|
|
26
|
+
- Short technical videos demonstrating real decisions.
|
|
27
|
+
- Documentation and integration examples.
|
|
28
|
+
|
|
29
|
+
## Phase 5 — Conversion funnel
|
|
30
|
+
Content → Quickstart → protected tool → Attack Lab → Shadow → Design Partner → Production → Paid plan.
|
|
31
|
+
|
|
32
|
+
## Claims policy
|
|
33
|
+
Prefer measured statements. Avoid absolute claims such as “makes agents safe,” “prevents all prompt injection,” or “replaces IAM/SIEM.”
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Observability and Alerting Runbook
|
|
2
|
+
|
|
3
|
+
## Health
|
|
4
|
+
|
|
5
|
+
Probe:
|
|
6
|
+
|
|
7
|
+
```text
|
|
8
|
+
GET /api/health
|
|
9
|
+
GET /api/ready
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
## Security alerts
|
|
13
|
+
|
|
14
|
+
Alert on:
|
|
15
|
+
|
|
16
|
+
- unexpected increase in BLOCK/ASK volume for a production tool;
|
|
17
|
+
- cross-tenant authorization failures;
|
|
18
|
+
- repeated egress secret findings;
|
|
19
|
+
- approval failures or abnormal approval volume;
|
|
20
|
+
- kill-switch activation;
|
|
21
|
+
- policy activation/rollback events;
|
|
22
|
+
- audit persistence failures.
|
|
23
|
+
|
|
24
|
+
## Performance alerts
|
|
25
|
+
|
|
26
|
+
Alert on sustained P95/P99 latency above the deployment's tested baseline, error rate, and database connection failures. Do not use the development benchmark as an SLA.
|
|
27
|
+
|
|
28
|
+
## Operational dashboards
|
|
29
|
+
|
|
30
|
+
At minimum track:
|
|
31
|
+
|
|
32
|
+
```text
|
|
33
|
+
requests
|
|
34
|
+
ALLOW / ASK / BLOCK
|
|
35
|
+
handler executions
|
|
36
|
+
approval resolution latency
|
|
37
|
+
policy version
|
|
38
|
+
run persistence failures
|
|
39
|
+
API errors
|
|
40
|
+
P50 / P95 / P99
|
|
41
|
+
active tenants
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Use `/api/metrics`, `/api/observability`, `/api/security-report`, and `/api/audit/export` as evidence sources where appropriate.
|
package/docs/outreach.md
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# AgentGate Outreach — Design Partner Draft
|
|
2
|
+
|
|
3
|
+
## Target
|
|
4
|
+
AI SaaS/startup teams with agents that can execute real side effects: refunds, cancellations, production changes, exports, or CRM mutations.
|
|
5
|
+
|
|
6
|
+
## Short message
|
|
7
|
+
I built AgentGate, a runtime control plane for AI agent actions. It sits between an agent and its tools and can deterministically ALLOW, ASK, or BLOCK sensitive actions, with replayable audit evidence.
|
|
8
|
+
|
|
9
|
+
I’m looking for 3–5 design partners. We start in sandbox/replay or Shadow Mode, then protect one sensitive tool. The goal is to prove one thing: a blocked action produces zero side effects.
|
|
10
|
+
|
|
11
|
+
Would you be open to testing one real agent/tool workflow with us?
|
|
12
|
+
|
|
13
|
+
## Technical follow-up
|
|
14
|
+
The integration is designed to start with one tool and one Policy Pack. We provide the pilot checklist, acceptance gates, Attack Lab scenarios, and evidence/replay workflow.
|
|
15
|
+
|
|
16
|
+
## Follow-up 1
|
|
17
|
+
The pilot does not require production enforcement at the start. We can replay representative traffic first, show would-ALLOW/ASK/BLOCK decisions, review mismatches, then enforce one sensitive tool.
|
|
18
|
+
|
|
19
|
+
## Follow-up 2
|
|
20
|
+
If useful, I can run the first test with your refund/export/deployment tool and leave the decision evidence with your team.
|
|
21
|
+
|
|
22
|
+
## Outreach rules
|
|
23
|
+
- Do not claim customer results before a customer exists.
|
|
24
|
+
- Do not claim prevention of every prompt injection.
|
|
25
|
+
- Do not claim compliance certification.
|
|
26
|
+
- Lead with a concrete side-effecting tool, not generic AI security.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Partner Intake Template
|
|
2
|
+
|
|
3
|
+
- Partner:
|
|
4
|
+
- Agent:
|
|
5
|
+
- Sensitive tool:
|
|
6
|
+
- Environment:
|
|
7
|
+
- Tool side effects:
|
|
8
|
+
- Identity model:
|
|
9
|
+
- Tenant model:
|
|
10
|
+
- Existing authorization:
|
|
11
|
+
- Existing approval path:
|
|
12
|
+
- Existing audit evidence:
|
|
13
|
+
- Sandbox/replay source:
|
|
14
|
+
- Technical owner:
|
|
15
|
+
- Pilot start:
|
|
16
|
+
- Pilot target end:
|
|
17
|
+
|
|
18
|
+
## Success criteria
|
|
19
|
+
|
|
20
|
+
- BLOCK -> 0 handler execution
|
|
21
|
+
- ASK -> 0 pre-approval execution
|
|
22
|
+
- Approved ASK -> exactly 1 execution
|
|
23
|
+
- Cross-tenant access -> 0 leaks
|
|
24
|
+
- Undetected egress secrets -> 0
|
|
25
|
+
- Audit mismatch -> 0
|
|
26
|
+
- Replay after restart -> PASS
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# Performance Baseline — 2.13.8
|
|
2
|
+
|
|
3
|
+
Measured with `npm run benchmark` on the release build.
|
|
4
|
+
|
|
5
|
+
Environment:
|
|
6
|
+
|
|
7
|
+
```text
|
|
8
|
+
Node: v22.16.0
|
|
9
|
+
OS/arch: linux-x64
|
|
10
|
+
CPU: Intel(R) Xeon(R) Platinum 8370C CPU @ 2.80GHz
|
|
11
|
+
Memory: 5.81 GB
|
|
12
|
+
Samples: 5,000 per workload
|
|
13
|
+
Concurrency: 100
|
|
14
|
+
Policy: productionBlock=true, approvalAmount=5000, autoApproveAmount=500
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
| Workload | P50 | P95 | P99 | Throughput |
|
|
18
|
+
|---|---:|---:|---:|---:|
|
|
19
|
+
| ALLOW check | 2.323 ms | 5.225 ms | 7.570 ms | 35,915.98 ops/s |
|
|
20
|
+
| BLOCK check | 2.239 ms | 5.168 ms | 8.393 ms | 40,660.09 ops/s |
|
|
21
|
+
| ALLOW execute | 1.577 ms | 6.573 ms | 28.485 ms | 39,524.91 ops/s |
|
|
22
|
+
|
|
23
|
+
These figures are an environment-specific baseline, not an SLA. Database persistence, network calls, model latency, egress inspection configuration, policy size and deployment topology can materially change end-to-end latency.
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Performance Benchmark
|
|
2
|
+
|
|
3
|
+
AgentGate security tests and performance benchmarks are separate concerns. A concurrency/security test proves correctness under load; it is not a latency SLA.
|
|
4
|
+
|
|
5
|
+
Run:
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
npm run benchmark
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
The benchmark reports P50/P95/P99 latency and throughput for deterministic `check()` decisions and an execution path. Results are environment-specific and must be rerun on the customer's target infrastructure before publishing an SLA.
|
|
12
|
+
|
|
13
|
+
## What to record
|
|
14
|
+
|
|
15
|
+
- Node version
|
|
16
|
+
- CPU / memory
|
|
17
|
+
- AgentGate version
|
|
18
|
+
- policy size/type
|
|
19
|
+
- concurrency
|
|
20
|
+
- sample count
|
|
21
|
+
- P50/P95/P99
|
|
22
|
+
- throughput
|
|
23
|
+
- error count
|
|
24
|
+
|
|
25
|
+
## Release policy
|
|
26
|
+
|
|
27
|
+
No universal latency number is promised by AgentGate. A customer-facing performance claim must cite the exact benchmark environment and policy workload.
|
package/docs/pricing.md
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
# AgentGate Pricing — Initial Test Model
|
|
2
|
+
|
|
3
|
+
This is a proposed pricing experiment, not a claim of validated market pricing. Do not publish these numbers as final until at least 3 design-partner conversations test willingness to pay.
|
|
4
|
+
|
|
5
|
+
## Free / Developer
|
|
6
|
+
**$0**
|
|
7
|
+
|
|
8
|
+
For local evaluation and open-source adoption.
|
|
9
|
+
|
|
10
|
+
- SDK / runtime
|
|
11
|
+
- Policy Packs
|
|
12
|
+
- Simulate + Attack Lab
|
|
13
|
+
- Local Replay
|
|
14
|
+
- Shadow Mode
|
|
15
|
+
- Community support
|
|
16
|
+
|
|
17
|
+
## Pro
|
|
18
|
+
**$99/month** starting point
|
|
19
|
+
|
|
20
|
+
For one production team that needs centralized controls.
|
|
21
|
+
|
|
22
|
+
- Everything in Developer
|
|
23
|
+
- Hosted Control Plane
|
|
24
|
+
- Team access
|
|
25
|
+
- Central policy management
|
|
26
|
+
- Approval workflows
|
|
27
|
+
- Replayable audit evidence
|
|
28
|
+
- Runtime monitoring
|
|
29
|
+
- Usage limits to be finalized from partner data
|
|
30
|
+
|
|
31
|
+
## Team / Production
|
|
32
|
+
**$499/month** starting point
|
|
33
|
+
|
|
34
|
+
For teams operating multiple agents or tenants.
|
|
35
|
+
|
|
36
|
+
- Everything in Pro
|
|
37
|
+
- Multi-tenant controls
|
|
38
|
+
- Advanced policy bundles
|
|
39
|
+
- Egress protection
|
|
40
|
+
- Security reporting
|
|
41
|
+
- Higher usage limits
|
|
42
|
+
- Priority support
|
|
43
|
+
|
|
44
|
+
## Enterprise
|
|
45
|
+
**Custom**
|
|
46
|
+
|
|
47
|
+
For organizations requiring deployment, identity, retention, support, or contractual requirements beyond the standard plans.
|
|
48
|
+
|
|
49
|
+
### Pricing principles
|
|
50
|
+
- Do not charge for a dashboard alone.
|
|
51
|
+
- Price around protected runtime usage and control-plane value.
|
|
52
|
+
- Keep local evaluation friction low.
|
|
53
|
+
- Revisit price after design-partner evidence.
|
|
@@ -0,0 +1,70 @@
|
|
|
1
|
+
# Production Deployment Reference
|
|
2
|
+
|
|
3
|
+
## Recommended topology
|
|
4
|
+
|
|
5
|
+
```text
|
|
6
|
+
Agent / Worker
|
|
7
|
+
|
|
|
8
|
+
v
|
|
9
|
+
AgentGate Runtime / MCP Gateway
|
|
10
|
+
|
|
|
11
|
+
+---- PostgreSQL (durable runs, approvals, policy metadata)
|
|
12
|
+
|
|
|
13
|
+
+---- Control Plane (private network)
|
|
14
|
+
|
|
|
15
|
+
+---- HTTPS / OIDC / reverse proxy
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
### Development vs production
|
|
19
|
+
|
|
20
|
+
- Development may use the built-in file persistence and local browser session.
|
|
21
|
+
- Production should use a durable PostgreSQL/Supabase-backed store, tenant-scoped authentication, HTTPS, and a private Control Plane network.
|
|
22
|
+
- Do not expose the service-role database credential or other server secrets to the browser.
|
|
23
|
+
- Keep the Control Plane separate from untrusted agent traffic where practical.
|
|
24
|
+
|
|
25
|
+
## PostgreSQL
|
|
26
|
+
|
|
27
|
+
Apply `schema/postgres.sql` to a dedicated database. The schema enables RLS and requires a tenant scope for reads/writes.
|
|
28
|
+
|
|
29
|
+
For pooled PostgreSQL clients, AgentGate scopes each transaction with `SET LOCAL agentgate.tenant_id`. The adapter also keeps explicit `tenant_id` predicates as defense in depth.
|
|
30
|
+
|
|
31
|
+
### Connection requirements
|
|
32
|
+
|
|
33
|
+
- TLS enabled between AgentGate and PostgreSQL.
|
|
34
|
+
- Least-privilege database role.
|
|
35
|
+
- Connection pool sized for the actual workload.
|
|
36
|
+
- Statement and connection timeouts configured by the deployment platform.
|
|
37
|
+
- Backups enabled at the database layer.
|
|
38
|
+
|
|
39
|
+
### Readiness
|
|
40
|
+
|
|
41
|
+
Use:
|
|
42
|
+
|
|
43
|
+
```text
|
|
44
|
+
GET /api/health
|
|
45
|
+
GET /api/ready
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
`/api/ready` should be used by orchestrators for readiness, while liveness should remain lightweight.
|
|
49
|
+
|
|
50
|
+
## Container deployment
|
|
51
|
+
|
|
52
|
+
The repository includes a Node 20 Alpine `Dockerfile`. Build and run it behind an HTTPS reverse proxy. Set `NODE_ENV=production` and a durable persistence configuration; do not use local JSON persistence as the authoritative store for a multi-instance production deployment.
|
|
53
|
+
|
|
54
|
+
## Rollout
|
|
55
|
+
|
|
56
|
+
1. Install the exact package version and retain the lockfile.
|
|
57
|
+
2. Run `npm test` and `node scripts/release-check.mjs`.
|
|
58
|
+
3. Apply database migrations/schema and verify RLS.
|
|
59
|
+
4. Start one canary instance in Observe/Shadow mode.
|
|
60
|
+
5. Run Attack Lab and application-specific abuse cases.
|
|
61
|
+
6. Verify audit export, replay, approval, and backup/restore procedures.
|
|
62
|
+
7. Switch only the reviewed tools to Enforce.
|
|
63
|
+
8. Expand tool-by-tool after acceptance criteria are met.
|
|
64
|
+
|
|
65
|
+
## Rollback
|
|
66
|
+
|
|
67
|
+
- Keep the previous application image/package available.
|
|
68
|
+
- Never roll back the database blindly: check schema compatibility first.
|
|
69
|
+
- Disable newly activated policies before reverting runtime code when policy/runtime compatibility is uncertain.
|
|
70
|
+
- Preserve exported audit evidence before destructive rollback operations.
|