insika 0.0.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +361 -0
- data/LICENSE +21 -0
- data/README.md +136 -2
- data/bin/insika +366 -0
- data/docs/AGENTS.md +618 -0
- data/docs/ARCHITECTURE.md +333 -0
- data/docs/BENCHMARK.md +114 -0
- data/docs/CHANNELS.md +453 -0
- data/docs/CONTEXT.md +117 -0
- data/docs/DEPLOY.md +354 -0
- data/docs/EMBEDDING.md +198 -0
- data/docs/EVALS.md +273 -0
- data/docs/LOADTEST.md +232 -0
- data/docs/OBSERVABILITY.md +374 -0
- data/docs/PLUGINS.md +211 -0
- data/docs/REFINEMENT.md +477 -0
- data/docs/RELEASING.md +70 -0
- data/docs/RUNNING-LOCAL.md +153 -0
- data/docs/SANDBOX.md +114 -0
- data/docs/SECURITY.md +375 -0
- data/docs/SKILLS.md +284 -0
- data/docs/TOOLS.md +302 -0
- data/docs/WHY.md +137 -0
- data/docs/WORKFLOWS.md +225 -0
- data/docs/build.md +14 -0
- data/docs/index.md +68 -0
- data/docs/onboarding/start.md +126 -0
- data/docs/operate.md +12 -0
- data/docs/ship.md +10 -0
- data/docs/understand.md +10 -0
- data/lib/insika/agent_file_store.rb +125 -0
- data/lib/insika/agent_profile.rb +255 -0
- data/lib/insika/alert_dispatcher.rb +139 -0
- data/lib/insika/allowlist.rb +28 -0
- data/lib/insika/baseline_store.rb +74 -0
- data/lib/insika/budget_ledger.rb +135 -0
- data/lib/insika/capability/resolved_tool.rb +34 -0
- data/lib/insika/capability_registry.rb +112 -0
- data/lib/insika/channel_delivery.rb +153 -0
- data/lib/insika/channel_registry.rb +30 -0
- data/lib/insika/channels/relay.rb +178 -0
- data/lib/insika/channels/web/widget.js +283 -0
- data/lib/insika/channels/web.rb +211 -0
- data/lib/insika/channels/webhook.rb +58 -0
- data/lib/insika/chat_builder.rb +303 -0
- data/lib/insika/checkpoint.rb +13 -0
- data/lib/insika/checkpoint_store.rb +153 -0
- data/lib/insika/circuit_state.rb +114 -0
- data/lib/insika/coercion.rb +58 -0
- data/lib/insika/command.rb +32 -0
- data/lib/insika/command_bus.rb +39 -0
- data/lib/insika/commands/agent_payload.rb +43 -0
- data/lib/insika/commands/approve_action.rb +46 -0
- data/lib/insika/commands/cancel_task.rb +33 -0
- data/lib/insika/commands/create_agent.rb +54 -0
- data/lib/insika/commands/create_session.rb +67 -0
- data/lib/insika/commands/delete_agent.rb +33 -0
- data/lib/insika/commands/delete_agent_file.rb +50 -0
- data/lib/insika/commands/delete_data_tool.rb +33 -0
- data/lib/insika/commands/delete_llm_provider.rb +36 -0
- data/lib/insika/commands/delete_mcp.rb +30 -0
- data/lib/insika/commands/delete_skill.rb +43 -0
- data/lib/insika/commands/delete_system_file.rb +29 -0
- data/lib/insika/commands/gate_refinement.rb +245 -0
- data/lib/insika/commands/import_mcp_tools.rb +48 -0
- data/lib/insika/commands/import_tools.rb +81 -0
- data/lib/insika/commands/issue_tenant_token.rb +41 -0
- data/lib/insika/commands/memory_add_note.rb +32 -0
- data/lib/insika/commands/memory_forget_fact.rb +32 -0
- data/lib/insika/commands/memory_put_fact.rb +35 -0
- data/lib/insika/commands/pause_task.rb +29 -0
- data/lib/insika/commands/resolve_refinement.rb +126 -0
- data/lib/insika/commands/restore_agent_file.rb +36 -0
- data/lib/insika/commands/restore_data_tool.rb +34 -0
- data/lib/insika/commands/restore_system_file.rb +31 -0
- data/lib/insika/commands/resume_task.rb +85 -0
- data/lib/insika/commands/revoke_token.rb +39 -0
- data/lib/insika/commands/rotate_tenant_token.rb +43 -0
- data/lib/insika/commands/run_refinement.rb +133 -0
- data/lib/insika/commands/send_message.rb +150 -0
- data/lib/insika/commands/set_agent_tools.rb +39 -0
- data/lib/insika/commands/set_skill_agents.rb +112 -0
- data/lib/insika/commands/trigger_workflow.rb +80 -0
- data/lib/insika/commands/update_agent.rb +49 -0
- data/lib/insika/commands/update_settings.rb +33 -0
- data/lib/insika/commands/upsert_llm_provider.rb +34 -0
- data/lib/insika/commands/upsert_mcp.rb +32 -0
- data/lib/insika/commands/write_agent_file.rb +57 -0
- data/lib/insika/commands/write_data_tool.rb +43 -0
- data/lib/insika/commands/write_golden.rb +58 -0
- data/lib/insika/commands/write_skill.rb +60 -0
- data/lib/insika/commands/write_system_file.rb +31 -0
- data/lib/insika/config_store.rb +89 -0
- data/lib/insika/context/builder.rb +166 -0
- data/lib/insika/context/catalog_provider.rb +23 -0
- data/lib/insika/context/fragment.rb +43 -0
- data/lib/insika/context/priority.rb +30 -0
- data/lib/insika/context/provider.rb +19 -0
- data/lib/insika/context/providers/memory.rb +60 -0
- data/lib/insika/context/providers/prompt.rb +105 -0
- data/lib/insika/context/providers/request.rb +32 -0
- data/lib/insika/context/providers/session.rb +123 -0
- data/lib/insika/context/providers/skill.rb +24 -0
- data/lib/insika/context/providers/skill_trigger.rb +128 -0
- data/lib/insika/context/providers/tool_search.rb +20 -0
- data/lib/insika/context_trace_store.rb +92 -0
- data/lib/insika/delegation_store.rb +153 -0
- data/lib/insika/doctor.rb +539 -0
- data/lib/insika/dsl/definition.rb +55 -0
- data/lib/insika/dsl/runtime.rb +382 -0
- data/lib/insika/dsl/server_boot.rb +98 -0
- data/lib/insika/dsl/system.rb +93 -0
- data/lib/insika/dsl/workflow_adapter.rb +59 -0
- data/lib/insika/dsl.rb +364 -0
- data/lib/insika/edge_limiter.rb +268 -0
- data/lib/insika/egress_guard.rb +75 -0
- data/lib/insika/env_schema.rb +249 -0
- data/lib/insika/errors.rb +201 -0
- data/lib/insika/evals/assertions.rb +247 -0
- data/lib/insika/evals/baseline.rb +69 -0
- data/lib/insika/evals/golden.rb +172 -0
- data/lib/insika/evals/judge.rb +225 -0
- data/lib/insika/evals/pairwise.rb +178 -0
- data/lib/insika/evals/report.rb +115 -0
- data/lib/insika/evals/runner.rb +141 -0
- data/lib/insika/evals/transport.rb +178 -0
- data/lib/insika/event.rb +18 -0
- data/lib/insika/event_stream.rb +132 -0
- data/lib/insika/executor.rb +1995 -0
- data/lib/insika/frontmatter.rb +42 -0
- data/lib/insika/golden_store.rb +145 -0
- data/lib/insika/hooks.rb +48 -0
- data/lib/insika/http_client.rb +63 -0
- data/lib/insika/inbound_log.rb +84 -0
- data/lib/insika/llm_configurator.rb +99 -0
- data/lib/insika/llm_provider_store.rb +83 -0
- data/lib/insika/loop_detector.rb +143 -0
- data/lib/insika/mcp_http_client.rb +67 -0
- data/lib/insika/mcp_store.rb +115 -0
- data/lib/insika/mcp_tool_ingestor.rb +143 -0
- data/lib/insika/memory_store.rb +93 -0
- data/lib/insika/message_origin.rb +76 -0
- data/lib/insika/middleware.rb +36 -0
- data/lib/insika/model_policy.rb +52 -0
- data/lib/insika/model_resolver.rb +176 -0
- data/lib/insika/model_selection.rb +115 -0
- data/lib/insika/onboarding.rb +208 -0
- data/lib/insika/outbox_store.rb +166 -0
- data/lib/insika/overlay_tool_registry.rb +102 -0
- data/lib/insika/pack.rb +102 -0
- data/lib/insika/pack_importer.rb +123 -0
- data/lib/insika/pending_action_store.rb +120 -0
- data/lib/insika/plugin/loader.rb +356 -0
- data/lib/insika/plugin.rb +35 -0
- data/lib/insika/policy/engine.rb +83 -0
- data/lib/insika/policy/policy.rb +120 -0
- data/lib/insika/policy_registry.rb +23 -0
- data/lib/insika/profile_source.rb +143 -0
- data/lib/insika/prompt_catalog.rb +61 -0
- data/lib/insika/provider_error_classifier.rb +160 -0
- data/lib/insika/queue_policy.rb +167 -0
- data/lib/insika/recovery.rb +168 -0
- data/lib/insika/refinement/candidate.rb +159 -0
- data/lib/insika/refinement/evidence_collector.rb +371 -0
- data/lib/insika/refinement/gate.rb +234 -0
- data/lib/insika/refinement/panel.rb +222 -0
- data/lib/insika/refinement/proposer.rb +262 -0
- data/lib/insika/refinement_store.rb +295 -0
- data/lib/insika/registry.rb +59 -0
- data/lib/insika/reliability.rb +185 -0
- data/lib/insika/safety/config.rb +109 -0
- data/lib/insika/safety/detectors.rb +176 -0
- data/lib/insika/safety/factory.rb +102 -0
- data/lib/insika/safety/input_guardrail.rb +102 -0
- data/lib/insika/safety/moderator.rb +94 -0
- data/lib/insika/safety/output_filter.rb +79 -0
- data/lib/insika/safety/output_validator.rb +101 -0
- data/lib/insika/safety/safe_responses.rb +47 -0
- data/lib/insika/sandbox/boundary.rb +93 -0
- data/lib/insika/sandbox/docker.rb +74 -0
- data/lib/insika/sandbox/local.rb +33 -0
- data/lib/insika/sandbox/runner.rb +80 -0
- data/lib/insika/sandbox.rb +85 -0
- data/lib/insika/schema_guard.rb +147 -0
- data/lib/insika/secret_masking.rb +34 -0
- data/lib/insika/server/a2a/agent_card.rb +27 -0
- data/lib/insika/server/a2a/app.rb +112 -0
- data/lib/insika/server/a2a/client.rb +101 -0
- data/lib/insika/server/a2a/errors.rb +32 -0
- data/lib/insika/server/a2a/http.rb +42 -0
- data/lib/insika/server/a2a/message.rb +27 -0
- data/lib/insika/server/a2a/protocol.rb +45 -0
- data/lib/insika/server/a2a/remotes.rb +25 -0
- data/lib/insika/server/a2a/task_projection.rb +40 -0
- data/lib/insika/server/app.rb +1022 -0
- data/lib/insika/server/boot.rb +119 -0
- data/lib/insika/server/rack_app.rb +118 -0
- data/lib/insika/server/responses.rb +165 -0
- data/lib/insika/server/sse_body.rb +96 -0
- data/lib/insika/server/tenant_auth.rb +61 -0
- data/lib/insika/session_actor.rb +162 -0
- data/lib/insika/session_store.rb +143 -0
- data/lib/insika/settings_store.rb +154 -0
- data/lib/insika/shutdown.rb +125 -0
- data/lib/insika/skill_catalog.rb +220 -0
- data/lib/insika/skill_store.rb +127 -0
- data/lib/insika/steer_injector.rb +110 -0
- data/lib/insika/store.rb +52 -0
- data/lib/insika/stores/memory.rb +123 -0
- data/lib/insika/stores/sqlite.rb +183 -0
- data/lib/insika/studio/app.rb +1693 -0
- data/lib/insika/studio/assets/dist/application.css +1 -0
- data/lib/insika/studio/assets/dist/application.js +70 -0
- data/lib/insika/studio/forms.rb +335 -0
- data/lib/insika/studio/nav_icons.rb +31 -0
- data/lib/insika/studio/views/_message.erb +44 -0
- data/lib/insika/studio/views/agent_detail.erb +285 -0
- data/lib/insika/studio/views/agents.erb +63 -0
- data/lib/insika/studio/views/approvals.erb +41 -0
- data/lib/insika/studio/views/chats.erb +34 -0
- data/lib/insika/studio/views/evals.erb +83 -0
- data/lib/insika/studio/views/home.erb +72 -0
- data/lib/insika/studio/views/layout.erb +94 -0
- data/lib/insika/studio/views/login.erb +17 -0
- data/lib/insika/studio/views/mcp.erb +91 -0
- data/lib/insika/studio/views/not_found.erb +5 -0
- data/lib/insika/studio/views/playground.erb +47 -0
- data/lib/insika/studio/views/refinement.erb +234 -0
- data/lib/insika/studio/views/session.erb +137 -0
- data/lib/insika/studio/views/settings.erb +168 -0
- data/lib/insika/studio/views/skills.erb +141 -0
- data/lib/insika/studio/views/system_files.erb +65 -0
- data/lib/insika/studio/views/task.erb +105 -0
- data/lib/insika/studio/views/tasks.erb +33 -0
- data/lib/insika/studio/views/tool_edit.erb +107 -0
- data/lib/insika/studio/views/tools.erb +89 -0
- data/lib/insika/subagent_graph.rb +96 -0
- data/lib/insika/system_file_store.rb +96 -0
- data/lib/insika/task_actor.rb +128 -0
- data/lib/insika/task_store.rb +250 -0
- data/lib/insika/telemetry/pricing.rb +104 -0
- data/lib/insika/telemetry/recorder.rb +228 -0
- data/lib/insika/telemetry.rb +127 -0
- data/lib/insika/testing/store_contract.rb +270 -0
- data/lib/insika/tick.rb +122 -0
- data/lib/insika/token_estimator.rb +16 -0
- data/lib/insika/token_store.rb +168 -0
- data/lib/insika/tool_assembly.rb +140 -0
- data/lib/insika/tool_catalog.rb +89 -0
- data/lib/insika/tool_definition.rb +518 -0
- data/lib/insika/tool_envelope.rb +140 -0
- data/lib/insika/tool_manifest.rb +218 -0
- data/lib/insika/tool_output_compressor.rb +100 -0
- data/lib/insika/tool_registry.rb +21 -0
- data/lib/insika/tool_store.rb +135 -0
- data/lib/insika/tool_trace_store.rb +92 -0
- data/lib/insika/tools/a2a_remote.rb +48 -0
- data/lib/insika/tools/agent_enum.rb +68 -0
- data/lib/insika/tools/concurrency.rb +54 -0
- data/lib/insika/tools/data_defined_tool.rb +219 -0
- data/lib/insika/tools/load_skill.rb +99 -0
- data/lib/insika/tools/remember.rb +53 -0
- data/lib/insika/tools/stuck_signal.rb +44 -0
- data/lib/insika/tools/subagent.rb +75 -0
- data/lib/insika/tools/subagents.rb +77 -0
- data/lib/insika/tools/tool_search.rb +94 -0
- data/lib/insika/turn_output.rb +139 -0
- data/lib/insika/turn_state.rb +162 -0
- data/lib/insika/turn_timing.rb +56 -0
- data/lib/insika/usage_ledger.rb +47 -0
- data/lib/insika/version.rb +3 -1
- data/lib/insika/wiring/graph.rb +249 -0
- data/lib/insika/workflow.rb +185 -0
- data/lib/insika/workflow_registry.rb +33 -0
- data/lib/insika.rb +220 -4
- metadata +412 -8
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Running locally
|
|
3
|
+
parent: Build an agent
|
|
4
|
+
nav_order: 8
|
|
5
|
+
permalink: /running-local/
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Running the Insika locally
|
|
9
|
+
|
|
10
|
+
Boots the engine single-process, serving `/studio` and `/v1/*` against a demo
|
|
11
|
+
agent (the `bia` persona on DeepSeek). Every message runs the **same**
|
|
12
|
+
`send_message` the API runs — real tools, skills, and memory.
|
|
13
|
+
|
|
14
|
+
## Boot
|
|
15
|
+
|
|
16
|
+
```bash
|
|
17
|
+
cd insika
|
|
18
|
+
DEEPSEEK_API_KEY=sk-... bundle exec ruby scripts/serve_real.rb
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
> Use **`bundle exec`** (bundler isolation matters — the optional OpenTelemetry gem
|
|
22
|
+
> is in the Gemfile). It is single-process: `Ctrl-C` frees the port immediately.
|
|
23
|
+
|
|
24
|
+
Open `http://localhost:9292`:
|
|
25
|
+
|
|
26
|
+
| URL | What |
|
|
27
|
+
|-----|------|
|
|
28
|
+
| `/studio` | management UI (log in with the token; default `local-demo`) |
|
|
29
|
+
| `/studio/chats` | chat with the demo agent (`agent: bia`, `session_id: web`, multi-turn ready) |
|
|
30
|
+
| `/studio/tasks` | tasks / approvals console |
|
|
31
|
+
| `/v1/responses` | OpenAI-Responses ingress (Bearer) — the drop-in API contract |
|
|
32
|
+
| `/v1/agents` | provisioning by definition/pack (Bearer) — `POST` imports, `DELETE /:id` removes |
|
|
33
|
+
| `/v1/messages` | `send_message` sugar (Bearer; SSE when `?stream` is set) |
|
|
34
|
+
|
|
35
|
+
`POST /v1/messages?stream=false` answers the aggregated turn as JSON. It is also
|
|
36
|
+
the only HTTP surface that can report a **coalesced** message, so it is the only
|
|
37
|
+
one on which an agent configured with `queue_mode: "collect"`
|
|
38
|
+
([Agents](AGENTS.md#queue_mode--when-a-message-arrives-while-the-agent-is-busy))
|
|
39
|
+
actually coalesces:
|
|
40
|
+
|
|
41
|
+
```jsonc
|
|
42
|
+
// this call owns the reply
|
|
43
|
+
{ "task_id": "9f3c…", "content": "…" }
|
|
44
|
+
|
|
45
|
+
// this one joined a turn already waiting — deliver NOTHING, no stream is opened
|
|
46
|
+
{ "task_id": "9f3c…", "merged": true }
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
`/v1/responses` and any open stream never coalesce: their bodies have nowhere to
|
|
50
|
+
put that verdict, and a caller that cannot hear it would send the same answer once
|
|
51
|
+
per fragment. Those requests fall back to one turn per message.
|
|
52
|
+
|
|
53
|
+
Every `/v1` and `/a2a` route needs the Bearer. The exceptions are `/up` and, when
|
|
54
|
+
`INSIKA_ONBOARDING` is on, `/start.md`, `/models.json` and `/docs*` — a route is closed
|
|
55
|
+
unless it is on the allowlist in `lib/insika/server/app.rb`. With no token configured at all the
|
|
56
|
+
whole surface answers `503`, never open by omission.
|
|
57
|
+
|
|
58
|
+
## Variables (all optional)
|
|
59
|
+
|
|
60
|
+
| Env | Default | Effect |
|
|
61
|
+
|-----|---------|--------|
|
|
62
|
+
| `INSIKA_DB` | — (ephemeral memory) | SQLite path → config + execution survive a restart |
|
|
63
|
+
| `BIND` | `http://localhost:9292` | host:port |
|
|
64
|
+
| `ADMIN_TOKEN` | `local-demo` | token for `/studio` |
|
|
65
|
+
| `OPENCLAW_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN` | Bearer for the whole `/v1` + `/a2a` surface |
|
|
66
|
+
| `DEEPSEEK_MODEL` | `deepseek-v4-flash` | model |
|
|
67
|
+
|
|
68
|
+
With persistence:
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
DEEPSEEK_API_KEY=sk-... INSIKA_DB=./insika.db bundle exec ruby scripts/serve_real.rb
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
## Pointing a Responses client at the local engine
|
|
75
|
+
|
|
76
|
+
Anything that speaks the OpenAI Responses contract can drive the local engine —
|
|
77
|
+
that is the whole point of the `/v1/responses` drop-in. Point your client's base
|
|
78
|
+
URL at `http://localhost:9292`, send the API Bearer, and address an agent by id as
|
|
79
|
+
the `model`:
|
|
80
|
+
|
|
81
|
+
```bash
|
|
82
|
+
curl -N http://localhost:9292/v1/responses \
|
|
83
|
+
-H "Authorization: Bearer local-demo" \
|
|
84
|
+
-H "Content-Type: application/json" \
|
|
85
|
+
-d '{ "model": "bia", "user": "web", "stream": true, "input": "hello" }'
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
`user` is the session id (any stable id for a multi-turn conversation).
|
|
89
|
+
|
|
90
|
+
## Wiring an agent's tools back to your backend
|
|
91
|
+
|
|
92
|
+
If the agent has **data-tools** that call your own HTTP backend (see
|
|
93
|
+
[Tools](TOOLS.md)), and both the engine and that backend run on your machine, the
|
|
94
|
+
tools target `http://localhost:<port>/…` — plain `http` on a loopback address,
|
|
95
|
+
which the egress guard blocks by default (SSRF defense). Enable the opt-in
|
|
96
|
+
**pinned to your backend's host** when you boot:
|
|
97
|
+
|
|
98
|
+
```bash
|
|
99
|
+
INSIKA_EGRESS_ALLOW_HTTP=1 INSIKA_EGRESS_ALLOW_PRIVATE=1 \
|
|
100
|
+
INSIKA_EGRESS_HOSTS=localhost,127.0.0.1 \
|
|
101
|
+
DEEPSEEK_API_KEY=sk-... bundle exec ruby scripts/serve_real.rb
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
> `INSIKA_EGRESS_HOSTS` restricts the opened egress to just the internal host
|
|
105
|
+
> (defense-in-depth) — without it, `ALLOW_PRIVATE` opens *any* private destination.
|
|
106
|
+
> These `ALLOW_*` vars are for the fully-local loop only; never set them in the
|
|
107
|
+
> cloud. See [Security](SECURITY.md#egress-the-ssrf-boundary).
|
|
108
|
+
|
|
109
|
+
## Provisioning an agent
|
|
110
|
+
|
|
111
|
+
An agent is created from a **definition** — a folder ("pack") with an agent config,
|
|
112
|
+
prompt files, skills, and one data-tool per file:
|
|
113
|
+
|
|
114
|
+
```
|
|
115
|
+
<pack>/
|
|
116
|
+
agent.config.json # { id, model, provider, memory, metadata }
|
|
117
|
+
*.md # prompt files (identity, tools notes, …)
|
|
118
|
+
skills/<name>/SKILL.md
|
|
119
|
+
tools/<tool>.json # one data-tool per file
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
> **Data-tool URLs must be literal on the pack path.** The pack import does not
|
|
123
|
+
> resolve `{{env.*}}` — bake the backend base URL into each `tools/*.json` at
|
|
124
|
+
> generation time. (Only the *manifest* path resolves `{{env.*}}`.) See
|
|
125
|
+
> [Tools](TOOLS.md#the-one-gotcha-env-templating-is-manifest-only).
|
|
126
|
+
|
|
127
|
+
Provision it (runs as a client against the live server; the internal token comes
|
|
128
|
+
from the environment, never disk):
|
|
129
|
+
|
|
130
|
+
```bash
|
|
131
|
+
INSIKA_URL=http://localhost:9292 OPENCLAW_GATEWAY_TOKEN=local-demo \
|
|
132
|
+
bundle exec ruby scripts/import_pack.rb /path/to/pack
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
…or `POST /v1/agents` directly, or build the agent by hand in the `/studio`. All
|
|
136
|
+
paths land on the same import. See [Agents](AGENTS.md) for the from-scratch flow
|
|
137
|
+
and the `Insika.agent { … }` DSL.
|
|
138
|
+
|
|
139
|
+
## Observability (OpenTelemetry, opt-in)
|
|
140
|
+
|
|
141
|
+
OpenTelemetry is **off by default** (the gems do not even load). To turn it on, see
|
|
142
|
+
traces in a local collector (a one-line Jaeger), and read the attribute convention
|
|
143
|
+
and the production config, see [OBSERVABILITY.md](OBSERVABILITY.md). In short:
|
|
144
|
+
`INSIKA_OTEL=1` + `OTEL_EXPORTER_OTLP_ENDPOINT=…`, and every turn becomes a
|
|
145
|
+
`insika.turn` trace with `insika.tool` / `insika.data_tool` children, plus counters
|
|
146
|
+
and histograms (`insika.turns`, `insika.turn.duration`, `insika.tokens`,
|
|
147
|
+
`insika.cost`) you can chart without aggregating spans.
|
|
148
|
+
|
|
149
|
+
## See also
|
|
150
|
+
|
|
151
|
+
- [Agents](AGENTS.md) — create and configure an agent.
|
|
152
|
+
- [Tools](TOOLS.md) — define data-tools and troubleshoot egress.
|
|
153
|
+
- [Deploy](DEPLOY.md) — running the same image durably in a container.
|
data/docs/SANDBOX.md
ADDED
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Sandbox
|
|
3
|
+
parent: Ship it
|
|
4
|
+
nav_order: 2
|
|
5
|
+
permalink: /sandbox/
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Sandbox — confined execution
|
|
9
|
+
|
|
10
|
+
`Insika::Sandbox` is the engine primitive for **confined execution**: a single,
|
|
11
|
+
pluggable interface a tool holds to touch the filesystem and run commands inside
|
|
12
|
+
a bounded environment. It follows the principle of the *narrowest sandbox that
|
|
13
|
+
supports the task* — cheap in-process confinement by default, real container
|
|
14
|
+
isolation when the task warrants it, chosen by configuration rather than code.
|
|
15
|
+
|
|
16
|
+
It has two halves:
|
|
17
|
+
|
|
18
|
+
- **FS confinement (`Boundary`) — always on, always host-side.** Every path a
|
|
19
|
+
tool resolves is proven to live inside a single root *before any IO happens*.
|
|
20
|
+
- **Command exec — via a swappable provider.** `local` (in-process, the default)
|
|
21
|
+
or `docker` (isolated container). The provider is selected by data on the agent
|
|
22
|
+
profile, not by branching in tool code.
|
|
23
|
+
|
|
24
|
+
## The FS boundary
|
|
25
|
+
|
|
26
|
+
`Insika::Sandbox::Boundary` confines every path to one root directory. It
|
|
27
|
+
rejects, before touching disk:
|
|
28
|
+
|
|
29
|
+
- `..` traversal that escapes the root (path normalized, then string-contained on
|
|
30
|
+
a separator boundary — so `/ws-evil` does not pass for root `/ws`);
|
|
31
|
+
- absolute paths outside the root;
|
|
32
|
+
- symlinks: the final path component may never be a symlink (an `lstat` check that
|
|
33
|
+
also catches a *broken* symlink whose target does not yet exist), and any
|
|
34
|
+
existing target's real path (`File.realpath`) must also be contained — a symlink
|
|
35
|
+
inside the root pointing outside is refused rather than followed.
|
|
36
|
+
|
|
37
|
+
An escape raises `Insika::Sandbox::Escape`; tools rescue it and return a
|
|
38
|
+
structured `{ error: ... }` to the model, never a crashed turn.
|
|
39
|
+
|
|
40
|
+
The boundary is **always host-side**, for both providers. With `docker`, the FS
|
|
41
|
+
tools still read and write host files (confined by the boundary); the container
|
|
42
|
+
sees the same bytes through a bind mount. Only the risky part — shell exec — is
|
|
43
|
+
isolated. That is the "narrowest sandbox" line: don't containerize a file read.
|
|
44
|
+
|
|
45
|
+
## Exec providers
|
|
46
|
+
|
|
47
|
+
`sandbox.exec(command)` returns a `Insika::Sandbox::Result`
|
|
48
|
+
(`exit_status`, `output`, `timed_out`). Both providers enforce a **hard-kill
|
|
49
|
+
wall-clock timeout**: on the deadline the process — and, via its own process
|
|
50
|
+
group, any children — is force-killed and the partial output is returned with
|
|
51
|
+
`timed_out: true`. (The engine's fiber contract forbids `Timeout.timeout`; the
|
|
52
|
+
deadline is enforced with a bounded `Thread#join` and a process-group kill.)
|
|
53
|
+
|
|
54
|
+
### `local` (default)
|
|
55
|
+
|
|
56
|
+
Runs the command in-process (`/bin/bash -c`) with the working directory pinned to
|
|
57
|
+
the root. Cheap, no daemon, no image pull. It is **not** an isolation boundary for
|
|
58
|
+
a shell (a command can still read absolute paths or `cd ..`); shell tools stay
|
|
59
|
+
approval-gated, and untrusted execution should use `docker`.
|
|
60
|
+
|
|
61
|
+
### `docker`
|
|
62
|
+
|
|
63
|
+
Runs the command inside a throwaway container (`docker run --rm`) with the root
|
|
64
|
+
bind-mounted at a fixed workdir. Conservative defaults: `--network none`, a memory
|
|
65
|
+
cap, a cpu cap, and a minimal image — all overridable. The container is named so
|
|
66
|
+
the timeout teardown can `docker kill` the exact container.
|
|
67
|
+
|
|
68
|
+
## Declaring a sandbox (config-over-code)
|
|
69
|
+
|
|
70
|
+
The provider and its policy are **data on the agent profile**, under the
|
|
71
|
+
`sandbox` key — never a branch in tool code:
|
|
72
|
+
|
|
73
|
+
```ruby
|
|
74
|
+
Insika::AgentProfile.build(
|
|
75
|
+
id: "coder",
|
|
76
|
+
# ...
|
|
77
|
+
sandbox: {
|
|
78
|
+
provider: "docker", # "local" (default) | "docker"
|
|
79
|
+
root: "/srv/project", # confinement root (default: cwd)
|
|
80
|
+
timeout: 120, # per-exec wall-clock seconds
|
|
81
|
+
image: "ruby:3.3", # docker only
|
|
82
|
+
network: "none", # docker only (default: none)
|
|
83
|
+
memory: "512m", # docker only
|
|
84
|
+
cpus: "1.0" # docker only
|
|
85
|
+
}
|
|
86
|
+
)
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
A deployment turns that config into the object every tool holds:
|
|
90
|
+
|
|
91
|
+
```ruby
|
|
92
|
+
sandbox = Insika::Sandbox.build(profile.sandbox) # => Insika::Sandbox::Env
|
|
93
|
+
|
|
94
|
+
sandbox.resolve("src/app.rb") # host path, guaranteed inside the root
|
|
95
|
+
sandbox.exec("bundle exec rspec") # => Result(exit_status:, output:, timed_out:)
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
The config round-trips through the JSON store (Studio edits), so profile and
|
|
99
|
+
runtime never drift. Absent (`nil`) config means a deployment builds a `local`
|
|
100
|
+
sandbox by default.
|
|
101
|
+
|
|
102
|
+
## Reference deployment
|
|
103
|
+
|
|
104
|
+
`plugins/insika-code` (a tier-2 code plugin — see [Plugins](PLUGINS.md)) is the
|
|
105
|
+
reference consumer: the FS/shell toolset
|
|
106
|
+
(`read_file`, `list_dir`, `grep`, `write_file`, `edit_file`, `bash`) is built on
|
|
107
|
+
this primitive, and `examples/insika-code/boot.rb` declares the `sandbox` block
|
|
108
|
+
on its profile while the plugin builds the matching `Sandbox` from the same
|
|
109
|
+
config. See [`examples/insika-code/README.md`](https://github.com/guizaols/insika/blob/main/examples/insika-code/README.md).
|
|
110
|
+
|
|
111
|
+
The sandbox is one of two independent controls on the high-risk tools; the other
|
|
112
|
+
is the engine's human-approval gate (`approvals_required` + the `ToolEnvelope`
|
|
113
|
+
suspend/resume path). They compose: confinement bounds *where* a tool can act;
|
|
114
|
+
approval bounds *whether* it acts at all.
|
data/docs/SECURITY.md
ADDED
|
@@ -0,0 +1,375 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Security
|
|
3
|
+
parent: Ship it
|
|
4
|
+
nav_order: 1
|
|
5
|
+
permalink: /security/
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Security
|
|
9
|
+
|
|
10
|
+
An agent runtime runs untrusted input through a model that can call tools and
|
|
11
|
+
touch the outside world. Insika treats that as the core problem, not an add-on.
|
|
12
|
+
Every control below is **built into the engine**, **configured as data** (not
|
|
13
|
+
hand-rolled per agent), and composes with the others. This page is the map;
|
|
14
|
+
each section links to the deeper guide.
|
|
15
|
+
|
|
16
|
+
The layers, from the edge inward:
|
|
17
|
+
|
|
18
|
+
0. **[The Bearer gate](#the-bearer-gate)** — nothing but the health probe answers without a token.
|
|
19
|
+
1. **[Edge limits](#edge-limits)** — stop a flood before it costs anything.
|
|
20
|
+
2. **[Input guardrails](#guardrails)** — refuse injection/abuse without a model turn.
|
|
21
|
+
3. **[Human approval](#human-approval)** — gate high-risk tool calls on an operator.
|
|
22
|
+
4. **[Egress guard](#egress-the-ssrf-boundary)** — bound where a tool can reach.
|
|
23
|
+
5. **[Sandbox](#sandbox-confined-execution)** — bound where code can run.
|
|
24
|
+
6. **[Output guardrails](#guardrails)** — moderate and redact what streams back.
|
|
25
|
+
7. **[Secrets](#secrets-live-only-in-the-environment)** — never on disk, never in the model.
|
|
26
|
+
|
|
27
|
+
## The Bearer gate
|
|
28
|
+
|
|
29
|
+
The `/v1` and `/a2a` surface answers only with
|
|
30
|
+
`Authorization: Bearer <OPENCLAW_GATEWAY_TOKEN>` (which falls back to `ADMIN_TOKEN`).
|
|
31
|
+
The check runs in the router, **before** any dispatch, against an **allowlist** of
|
|
32
|
+
public routes — so a route added later is closed until someone deliberately publishes
|
|
33
|
+
it. Only these answer without a token:
|
|
34
|
+
|
|
35
|
+
| Route | Why |
|
|
36
|
+
|---|---|
|
|
37
|
+
| `GET /up` | health/readiness probe; touches no store |
|
|
38
|
+
| `GET /start.md`, `/models.json`, `/docs`, `/docs/<name>.md` | the onboarding surface, opt-in via `INSIKA_ONBOARDING`: it exists to be read by a coding agent that has no credential yet |
|
|
39
|
+
| `GET /.well-known/agent-card.json` | A2A discovery — the card is the advertisement |
|
|
40
|
+
|
|
41
|
+
**With no token configured, the surface is not open — it is `503`.** Fail-closed by
|
|
42
|
+
construction, the same posture as `/studio`, which denies login without `ADMIN_TOKEN`.
|
|
43
|
+
`insika doctor` warns when neither is set.
|
|
44
|
+
|
|
45
|
+
This matters most for `POST /v1/commands/<type>`, the generic Command ingress: it can
|
|
46
|
+
dispatch **any** registered authoring Command (`write_agent_file`, `write_data_tool`,
|
|
47
|
+
`upsert_llm_provider`, `update_settings`, `delete_agent`). Treat that token as
|
|
48
|
+
operator-grade — it is not a read key, and a leak is agent takeover, not just usage.
|
|
49
|
+
|
|
50
|
+
### Multi-tenant mode (`INSIKA_TENANCY=multi_tenant`)
|
|
51
|
+
|
|
52
|
+
A single operator-grade token cannot host N stores. With `INSIKA_TENANCY=multi_tenant`
|
|
53
|
+
the Bearer is resolved to a **principal** before the routes:
|
|
54
|
+
|
|
55
|
+
- **Per-tenant tokens** (`POST /v1/commands/issue_tenant_token`) scope a caller to
|
|
56
|
+
one tenant; **operator** tokens (and the legacy gateway token — an existing
|
|
57
|
+
deployment switching modes keeps its credential) have the run of the deployment.
|
|
58
|
+
- Tokens are stored **only as SHA-256 hashes**; the plaintext is shown exactly once,
|
|
59
|
+
at issue time. Rotation (`rotate_tenant_token`) and revocation (`revoke_token`)
|
|
60
|
+
are operator commands — revoking one tenant's token never touches another's.
|
|
61
|
+
- A tenant principal reaches only its **own runtime surfaces** (`/v1/sessions`,
|
|
62
|
+
`/v1/messages`, `/v1/responses`, workflow runs, and its own session/task/event
|
|
63
|
+
reads). Every authoring/provisioning/config surface answers `403` to a tenant.
|
|
64
|
+
- Isolation is the key, not a convention: a tenant's sessions live under
|
|
65
|
+
`<tenant>:<session-id>`, its commands carry `meta.tenant` (memory scoping,
|
|
66
|
+
event tagging), and reading another tenant's session/task reads as `404`.
|
|
67
|
+
|
|
68
|
+
`single_tenant` (the default) is exactly the classic behavior above — one
|
|
69
|
+
operator credential, no principal, no stamping.
|
|
70
|
+
|
|
71
|
+
## The `/v1` contract is versioned by date
|
|
72
|
+
|
|
73
|
+
Every `/v1` route reads an optional `Insika-Version: YYYY-MM-DD` header, checked
|
|
74
|
+
before the Bearer gate above. Absent header means today's (only) behaviour; an
|
|
75
|
+
unknown value is a `400`, not a silent fallback — a caller that pins a version
|
|
76
|
+
finds out immediately that it does not exist, rather than being served whatever
|
|
77
|
+
happens to be current. `/a2a` and `/channels/<id>/…` are versioned by their own
|
|
78
|
+
contracts (JSON-RPC, the platform's own shape) and never read this header.
|
|
79
|
+
|
|
80
|
+
## Channels authenticate themselves
|
|
81
|
+
|
|
82
|
+
`/channels/<id>/…` is the one route family that does **not** answer to the gateway
|
|
83
|
+
token — and it is not an exception to the rule above, it is the same rule with a
|
|
84
|
+
different credential. A messaging platform has no way to send your gateway token,
|
|
85
|
+
and neither has a visitor's browser; what they can send is their own scheme (a
|
|
86
|
+
shared secret for the [relay](CHANNELS.md), an origin for the widget, an HMAC
|
|
87
|
+
signature for a channel you write yourself). So the channel does the check, and the router refuses before
|
|
88
|
+
parsing anything:
|
|
89
|
+
|
|
90
|
+
| The channel says | The route answers |
|
|
91
|
+
|---|---|
|
|
92
|
+
| `:ok` | the turn is dispatched |
|
|
93
|
+
| `:unauthorized` | `401` |
|
|
94
|
+
| `:disabled` — no credential configured | `503` |
|
|
95
|
+
| the channel has no `authenticate` at all | `503` |
|
|
96
|
+
|
|
97
|
+
There is no path to an open channel route. A relay with no `INSIKA_RELAY_TOKEN` is
|
|
98
|
+
not mounted at all (`404`); one that is mounted always has a secret. That is
|
|
99
|
+
deliberate: a public inbound route with an LLM behind it is a money faucet, and
|
|
100
|
+
[edge limits](#edge-limits) are the second line, not the first.
|
|
101
|
+
|
|
102
|
+
The routes are enumerated in the router, not prefix-matched, so a channel route
|
|
103
|
+
added tomorrow is gated by default and publishing it is a deliberate edit.
|
|
104
|
+
|
|
105
|
+
Two more things a relay operator owns:
|
|
106
|
+
|
|
107
|
+
- **The callback URL is egress.** The delivery POST goes through the
|
|
108
|
+
[egress guard](#egress-the-ssrf-boundary) on every call, not once at boot — a
|
|
109
|
+
hostname that resolved publicly yesterday can resolve to `169.254.169.254`
|
|
110
|
+
today, and that POST carries a customer's conversation.
|
|
111
|
+
- **`event_id` is a safety property, not an optimization.** Without it a retried
|
|
112
|
+
webhook is a second turn you pay for and a second message the customer reads.
|
|
113
|
+
|
|
114
|
+
### A public channel: the web widget
|
|
115
|
+
|
|
116
|
+
The [widget](CHANNELS.md#the-web-widget) is different from every other surface here
|
|
117
|
+
in one way that changes the whole posture: **the caller is an anonymous browser, so
|
|
118
|
+
there is no secret to check.** Three controls stand in for the missing credential,
|
|
119
|
+
and it is worth being precise about which of them is actually load-bearing.
|
|
120
|
+
|
|
121
|
+
- **The rate limit is the real defense, and it is mandatory.** The widget answers
|
|
122
|
+
`503` until a chat rate limit exists — the platform's `edge.chat_rate_limit` or a
|
|
123
|
+
per-agent `limits.chat_rate_limit` on every published agent. This is the only
|
|
124
|
+
place in the engine that refuses to serve rather than warn, because the failure
|
|
125
|
+
mode is a bill rather than an error. The bucket is the minted session id, and it
|
|
126
|
+
is checked *before* the input guardrail so a flood cannot even spend the
|
|
127
|
+
moderator. See [edge limits](#edge-limits).
|
|
128
|
+
- **The agent allowlist is a real boundary.** `INSIKA_WIDGET_AGENTS` is what an
|
|
129
|
+
anonymous visitor may address. Editing `data-agent` in devtools to name an
|
|
130
|
+
internal agent gets a `422`, not that agent.
|
|
131
|
+
- **The origin allowlist is a browser courtesy, not a control.** `Access-Control-
|
|
132
|
+
Allow-Origin` is enforced by the browser, and curl sends whatever origin it
|
|
133
|
+
likes. It stops another *site* from embedding your widget; it does not stop a
|
|
134
|
+
script. Configure it (exact match, no wildcards — `https://shop.example` does not
|
|
135
|
+
admit `https://a.shop.example`) and then do not count on it.
|
|
136
|
+
|
|
137
|
+
Two more properties worth knowing:
|
|
138
|
+
|
|
139
|
+
- **The engine issues session ids; the client never proposes one.** `POST
|
|
140
|
+
/channels/web/messages` with an unminted id is a `404`. Create-on-write on an
|
|
141
|
+
anonymous endpoint means anyone who guesses an id joins someone else's
|
|
142
|
+
conversation, and the ids are 128 random bits for the same reason. A session also
|
|
143
|
+
belongs to exactly one channel — a widget visitor cannot stream a relay
|
|
144
|
+
customer's conversation by pasting its id.
|
|
145
|
+
- **What the visitor types is data, at the most cuttable priority.** Untrusted
|
|
146
|
+
input from a public channel enters at `REQUEST` (40) like any other turn content,
|
|
147
|
+
and the [guardrails](#guardrails) run before the model. A channel may refuse a
|
|
148
|
+
request; it can never widen what the agent is allowed to do.
|
|
149
|
+
|
|
150
|
+
## Edge limits
|
|
151
|
+
|
|
152
|
+
The next gate. The edge limiter wraps a turn **before** the input guardrail,
|
|
153
|
+
so a flood cannot even spend an LLM moderator call. Two independent, opt-in
|
|
154
|
+
ceilings (nil/0 = off):
|
|
155
|
+
|
|
156
|
+
- **`chat_rate_limit`** — turn attempts per session per `chat_rate_window`.
|
|
157
|
+
Counted on entry (blocked attempts still count). Keyed by session id, so id
|
|
158
|
+
rotation defeats the per-session limit — the per-agent token ceiling is the
|
|
159
|
+
backstop.
|
|
160
|
+
- **`agent_token_ceiling`** — total tokens per agent per `agent_token_window`.
|
|
161
|
+
Checked on entry against a ledger, recorded after the turn. Advisory under
|
|
162
|
+
concurrency (overshoot ≈ in-flight turns).
|
|
163
|
+
|
|
164
|
+
On breach: a graceful halt returning a configurable `limit_response` with **zero
|
|
165
|
+
LLM calls**. A *resumed* turn is never re-counted.
|
|
166
|
+
|
|
167
|
+
> ⚠️ Windows live at the platform level; the token window defaults to **86400
|
|
168
|
+
> (daily)**. "500k tokens per hour" means `agent_token_ceiling = 500000` **and**
|
|
169
|
+
> `agent_token_window = 3600`. A per-agent ceiling that is *present but nil* reads
|
|
170
|
+
> as OFF for that agent — leave the key **absent** to inherit the platform value,
|
|
171
|
+
> `0` to explicitly disable. A malformed value in the Studio raises a validation
|
|
172
|
+
> error rather than silently disabling a production limit.
|
|
173
|
+
|
|
174
|
+
See [Agents §Layer 4](AGENTS.md#layer-4-edge-limits-flood-and-spend-control).
|
|
175
|
+
|
|
176
|
+
## Guardrails
|
|
177
|
+
|
|
178
|
+
Content safety runs on both sides of a turn, **opt-in per agent**. An agent that
|
|
179
|
+
configures nothing gets a conservative default: deterministic detectors on, LLM
|
|
180
|
+
moderator off. See [`examples/guardrails/`](https://github.com/guizaols/insika/tree/main/examples/guardrails/).
|
|
181
|
+
|
|
182
|
+
- **Input** — deterministic detectors (prompt-injection, abuse) run *before* the
|
|
183
|
+
model. A flagged input gets a **safe refusal without burning a model turn** — an
|
|
184
|
+
injection or a flood never reaches the provider. An LLM moderator can be layered
|
|
185
|
+
on top. The moderator is **fail-open**: an error or an unparseable reply never
|
|
186
|
+
blocks a legitimate customer — but silence is not a negative. That third state
|
|
187
|
+
surfaces as a `:guardrail_flagged` event with category `moderator_unavailable`,
|
|
188
|
+
so a degraded tier is distinguishable from a healthy one in the audit stream.
|
|
189
|
+
- **Output** — moderation plus PII/secret redaction on the streamed response, and
|
|
190
|
+
a post-turn validator.
|
|
191
|
+
|
|
192
|
+
```jsonc
|
|
193
|
+
{
|
|
194
|
+
"input": true,
|
|
195
|
+
"output": true,
|
|
196
|
+
"moderator": "provider/model", // optional LLM moderator; omit for detectors only
|
|
197
|
+
"strictness": "low | medium | high",
|
|
198
|
+
"responses": { "injection": "safe reply…", "default": "…" }
|
|
199
|
+
}
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
Strictness selects the detector categories (`low` = injection only; `medium`
|
|
203
|
+
(default) and `high` add sexual and abuse). Safe-reply lookup falls back per
|
|
204
|
+
category: the agent's category reply → the agent's default → the builtin
|
|
205
|
+
category → the builtin default. All of it is editable in the Studio Configuration
|
|
206
|
+
form. See [Agents §Layer 3](AGENTS.md#layer-3-guardrails-content-safety).
|
|
207
|
+
|
|
208
|
+
## Human approval
|
|
209
|
+
|
|
210
|
+
Some tool calls should not happen unattended. Mark them with
|
|
211
|
+
`approvals_required: [tool names]`: the approval policy *tags* those tools, and
|
|
212
|
+
when the model tries to call one, the turn **suspends** and waits for an operator
|
|
213
|
+
to approve or reject it in the Studio. Approval and confinement are independent
|
|
214
|
+
and compose — approval bounds *whether* a tool acts; the sandbox bounds *where* it
|
|
215
|
+
can act.
|
|
216
|
+
|
|
217
|
+
The wait is bounded by `approval_timeout` (default ~1h), and a turn that is
|
|
218
|
+
waiting on approval is *not* killed by the ordinary turn timeout. The suspend/
|
|
219
|
+
resume path is durable: an approval that arrives after a restart still resumes the
|
|
220
|
+
turn from its checkpoint (see
|
|
221
|
+
[Architecture](ARCHITECTURE.md#durability-checkpoints-and-resume)).
|
|
222
|
+
|
|
223
|
+
## Egress: the SSRF boundary
|
|
224
|
+
|
|
225
|
+
Every outbound HTTP call from a data tool passes through the **EgressGuard**, a
|
|
226
|
+
Server-Side Request Forgery defense. The default posture is **strict: public
|
|
227
|
+
`https` only** — private and loopback addresses and plain `http` are refused
|
|
228
|
+
unless explicitly opted in.
|
|
229
|
+
|
|
230
|
+
| Env | Effect |
|
|
231
|
+
|-----|--------|
|
|
232
|
+
| `INSIKA_EGRESS_HOSTS` | host allowlist (CSV) — the safe, specific way to permit a backend |
|
|
233
|
+
| `INSIKA_EGRESS_ALLOW_HTTP=1` | permit plain `http` — **loopback dev only** |
|
|
234
|
+
| `INSIKA_EGRESS_ALLOW_PRIVATE=1` | permit private/loopback IPs — **dev only** |
|
|
235
|
+
|
|
236
|
+
Restricting `INSIKA_EGRESS_HOSTS` to exactly the hosts a tool needs is
|
|
237
|
+
defense-in-depth: without it, `ALLOW_PRIVATE` opens *any* private destination.
|
|
238
|
+
**Never set the `ALLOW_*` vars in a cloud deployment** — a public backend over
|
|
239
|
+
`https` already passes the strict default.
|
|
240
|
+
|
|
241
|
+
> ⚠️ A blocked egress **fails silently** — the tool returns an error to the model,
|
|
242
|
+
> the request never leaves the process, and the conversation *looks* fine. Verify
|
|
243
|
+
> tool health by the Studio session **trace** (a healthy call shows the backend's
|
|
244
|
+
> `200`), never by the model's reply. Full detail in
|
|
245
|
+
> [Tools §Egress](TOOLS.md#egress-the-ssrf-guard-and-its-silent-failure).
|
|
246
|
+
|
|
247
|
+
## Sandbox: confined execution
|
|
248
|
+
|
|
249
|
+
Code tools that touch the filesystem or run commands do so through a **sandbox** —
|
|
250
|
+
a single, pluggable primitive chosen by config, following the principle of the
|
|
251
|
+
*narrowest sandbox that supports the task*:
|
|
252
|
+
|
|
253
|
+
- **Filesystem confinement is always on, host-side.** Every path is proven to
|
|
254
|
+
live inside one root *before any IO happens* — `..` escapes, absolute paths
|
|
255
|
+
outside the root, and symlinks pointing out are all refused. An escape returns a
|
|
256
|
+
structured error, never a crashed turn.
|
|
257
|
+
- **Command execution is via a swappable provider.** `local` (in-process, the
|
|
258
|
+
default — cheap, but *not* an isolation boundary for a shell) or `docker` (a
|
|
259
|
+
throwaway container with `--network none`, memory/cpu caps, and a minimal image).
|
|
260
|
+
Untrusted execution should use `docker`.
|
|
261
|
+
|
|
262
|
+
Both providers enforce a hard-kill wall-clock timeout that also reaps child
|
|
263
|
+
processes. The provider is **data on the agent profile** (a `sandbox` block), not
|
|
264
|
+
a branch in tool code. Full detail — including the config shape — in
|
|
265
|
+
[Sandbox](SANDBOX.md).
|
|
266
|
+
|
|
267
|
+
## Secrets live only in the environment
|
|
268
|
+
|
|
269
|
+
Secrets never live on disk in agent or tool definitions, and are never exposed to
|
|
270
|
+
the model:
|
|
271
|
+
|
|
272
|
+
- A data tool references a secret with `{{secret.*}}`, allowed **only** inside a
|
|
273
|
+
header named in `secret_headers`; the real value is injected at provision time
|
|
274
|
+
and stored masked. A stray `{{secret.*}}` anywhere else is rejected. See
|
|
275
|
+
[Tools](TOOLS.md#data-tools-a-tool-is-a-row).
|
|
276
|
+
- Full tool results are stored **masked** in the trace store; the copy persisted
|
|
277
|
+
into the transcript is capped.
|
|
278
|
+
- Provider keys and the API bearer token come from the environment (see
|
|
279
|
+
[Deploy](DEPLOY.md)). Rotating the API bearer requires updating both the runtime
|
|
280
|
+
and every consumer in the same step.
|
|
281
|
+
|
|
282
|
+
## Grading sends text to a judge (evals)
|
|
283
|
+
|
|
284
|
+
A rubric is scored by a model, so running an eval sends that case's **user turns and
|
|
285
|
+
the assistant reply** to every judge in the panel ([Evals](EVALS.md)). Two
|
|
286
|
+
consequences worth stating out loud:
|
|
287
|
+
|
|
288
|
+
- **The judges are a second provider surface.** Configure them deliberately: a panel
|
|
289
|
+
of three models is three vendors seeing those conversations. Judges are opt-in and
|
|
290
|
+
empty by default — with none configured, only the deterministic assertions run and
|
|
291
|
+
nothing leaves the deployment.
|
|
292
|
+
- **A case is curated text, not live traffic.** The corpus is authored from real
|
|
293
|
+
conversations with PII removed at curation time. That masking is a human step, not
|
|
294
|
+
an automatic one: treat a golden case as something that WILL be read by an external
|
|
295
|
+
model and reviewed in a pull request.
|
|
296
|
+
|
|
297
|
+
## Reading traffic back (refinement)
|
|
298
|
+
|
|
299
|
+
A refinement run reads an agent's own transcripts and tool traces to report what
|
|
300
|
+
broke ([Refinement](REFINEMENT.md)). Two properties keep that from becoming a
|
|
301
|
+
second copy of your customers' data:
|
|
302
|
+
|
|
303
|
+
- **It quotes as little as possible, redacted.** Only the `repetition` finding
|
|
304
|
+
carries customer words, and every snippet goes through the same detectors that
|
|
305
|
+
redact a customer-facing turn (formatted CPF/CNPJ, API secrets). Tool arguments
|
|
306
|
+
and results are never copied — only the normalized error signature is. That is
|
|
307
|
+
not a general PII scrubber: a phone number typed into a chat can survive into a
|
|
308
|
+
snippet, so the page sits behind the Studio login like the transcripts do.
|
|
309
|
+
- **Provenance is ids.** A run record stores session ids, never their contents, and
|
|
310
|
+
the events it emits carry counts only.
|
|
311
|
+
|
|
312
|
+
### Editing an agent from its traffic
|
|
313
|
+
|
|
314
|
+
A run can also propose a change to the agent's instructions, and that path is opt-in
|
|
315
|
+
per agent (`refinement.mode`), off by default, and bounded by construction rather
|
|
316
|
+
than by instruction:
|
|
317
|
+
|
|
318
|
+
| Surface | Reachable? | Why |
|
|
319
|
+
|---|---|---|
|
|
320
|
+
| the files listed in `refinement.files` | **yes** | text you already edit by hand, versioned, one-click restore |
|
|
321
|
+
| skill bodies and descriptions | **yes** | same trust level, same history |
|
|
322
|
+
| guardrails, the safety corpus | **no** | a constrained thing does not edit its own constraints |
|
|
323
|
+
| tool definitions and schemas | **no** | tools are authored by a person; a wrong schema is theirs to fix |
|
|
324
|
+
| policies, approvals, denied tools | **no** | authorization is not a prompt concern |
|
|
325
|
+
| model pins, limits, edge config | **no** | cost and latency are the operator's decisions |
|
|
326
|
+
| the system preamble the engine assembles | **no** | the fixed frame of a turn |
|
|
327
|
+
|
|
328
|
+
There is no code path to the "no" rows — not a rule in a prompt. Three more
|
|
329
|
+
properties are worth stating because each one is a way this could have gone wrong:
|
|
330
|
+
|
|
331
|
+
- **An edit is verified by running it, not by asking a model.** The candidate is
|
|
332
|
+
applied to a throwaway clone of the agent and the golden set is replayed against
|
|
333
|
+
it; any regression disqualifies it. An agent with no cases, with no recorded
|
|
334
|
+
baseline, or with a baseline in which nothing passes, **cannot be edited at all** —
|
|
335
|
+
the gate refuses instead of passing vacuously. That last case is the subtle one: a
|
|
336
|
+
regression is measured against a case that was passing, so an all-red baseline
|
|
337
|
+
cannot produce one and would wave everything through.
|
|
338
|
+
- **A human approves.** A gate pass parks the proposal for review; nothing applies
|
|
339
|
+
itself. Approving writes through the versioned file store, so undo is the Restore
|
|
340
|
+
button that was already there.
|
|
341
|
+
- **Prompt injection buys nothing.** Evidence reaches a proposer as quoted, masked
|
|
342
|
+
data, and whatever comes back is validated against the candidate schema and the
|
|
343
|
+
file allowlist. The worst an injected instruction can achieve is a proposal that
|
|
344
|
+
gets dropped or fails the gate.
|
|
345
|
+
- **The proposer reads the allowlisted files and nothing else.** Those it gets
|
|
346
|
+
verbatim and unmasked, which sends a model nothing it was not already sent on every
|
|
347
|
+
turn — they are the agent's own instructions. Masking them would break anchoring
|
|
348
|
+
(a `before` copied from a masked view never matches the real file) and protect
|
|
349
|
+
nothing. Files outside the allowlist are not shown at all.
|
|
350
|
+
|
|
351
|
+
**Know what the gate does not measure.** It catches an edit that breaks a case you
|
|
352
|
+
wrote; it cannot catch one that breaks something no case covers. Two of its blind
|
|
353
|
+
spots are this engine working correctly rather than gaps — an edit cannot remove a
|
|
354
|
+
tool (availability is `tools_allow`, not prose) and it cannot make a reply leak PII
|
|
355
|
+
(the output guardrail redacts first) — but the general point stands, and it is
|
|
356
|
+
[stated plainly in Refinement](REFINEMENT.md#what-the-gate-can-and-cannot-catch).
|
|
357
|
+
The structural limits above hold regardless; the gate is the layer that has to be
|
|
358
|
+
earned with cases.
|
|
359
|
+
|
|
360
|
+
## Config discipline
|
|
361
|
+
|
|
362
|
+
Configuration is validated against a schema of known keys at boot. An unknown key
|
|
363
|
+
in the `INSIKA_` namespace (a typo the runtime would otherwise ignore) or a
|
|
364
|
+
wrong-typed value is **surfaced**, and `INSIKA_CONFIG_STRICT=1` turns findings
|
|
365
|
+
into a boot refusal. By default the engine warns and boots on last-known-good — a
|
|
366
|
+
rotated key or a typo never takes the whole service down. The `insika doctor`
|
|
367
|
+
command runs the same checks on demand against a live database. See
|
|
368
|
+
[Deploy](DEPLOY.md#strict-config-and-insika-doctor).
|
|
369
|
+
|
|
370
|
+
## See also
|
|
371
|
+
|
|
372
|
+
- [Agents](AGENTS.md) — the five access layers per agent.
|
|
373
|
+
- [Tools](TOOLS.md) — egress and secret placeholders in depth.
|
|
374
|
+
- [Sandbox](SANDBOX.md) — the confinement primitive.
|
|
375
|
+
- [Deploy](DEPLOY.md) — tokens, rotation, and strict config.
|