insika 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (100) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +191 -0
  3. data/README.md +9 -6
  4. data/bin/insika +44 -3
  5. data/docs/AGENTS.md +74 -14
  6. data/docs/API.md +73 -0
  7. data/docs/ARCHITECTURE.md +45 -44
  8. data/docs/ARTIFACTS.md +42 -0
  9. data/docs/CHANNELS.md +19 -2
  10. data/docs/CONTEXT.md +86 -38
  11. data/docs/DEPLOY.md +27 -7
  12. data/docs/EVALS.md +98 -8
  13. data/docs/FACTS.md +4 -0
  14. data/docs/KNOWLEDGE.md +7 -0
  15. data/docs/LOADTEST.md +15 -27
  16. data/docs/MEDIA.md +1 -1
  17. data/docs/OBSERVABILITY.md +48 -4
  18. data/docs/POLICY.md +14 -5
  19. data/docs/RELEASING.md +4 -0
  20. data/docs/RUNNING-LOCAL.md +2 -2
  21. data/docs/SECURITY.md +28 -2
  22. data/docs/SOAK.md +1 -1
  23. data/docs/TOOLS.md +150 -32
  24. data/docs/prompts/ADD-TOOL.md +12 -2
  25. data/docs/prompts/DIAGNOSE-TURN.md +3 -0
  26. data/docs/prompts/GO-LIVE.md +6 -4
  27. data/lib/insika/agent_profile.rb +47 -10
  28. data/lib/insika/channels/web/widget.js +33 -0
  29. data/lib/insika/channels/web.rb +5 -2
  30. data/lib/insika/chat_builder.rb +90 -37
  31. data/lib/insika/commands/agent_payload.rb +1 -1
  32. data/lib/insika/commands/run_distillation.rb +5 -8
  33. data/lib/insika/commands/seed_session.rb +118 -0
  34. data/lib/insika/compaction.rb +196 -0
  35. data/lib/insika/context/builder.rb +35 -11
  36. data/lib/insika/context/fragment.rb +4 -1
  37. data/lib/insika/context/priority.rb +8 -0
  38. data/lib/insika/context/provider.rb +5 -0
  39. data/lib/insika/context/providers/briefing.rb +61 -29
  40. data/lib/insika/context/providers/fence_notice.rb +27 -0
  41. data/lib/insika/context/providers/knowledge.rb +7 -4
  42. data/lib/insika/context/providers/memory.rb +8 -4
  43. data/lib/insika/context/providers/session.rb +50 -10
  44. data/lib/insika/context_trace_store.rb +11 -1
  45. data/lib/insika/doctor.rb +213 -10
  46. data/lib/insika/dsl/runtime.rb +5 -0
  47. data/lib/insika/dsl.rb +6 -0
  48. data/lib/insika/edge_limiter.rb +4 -1
  49. data/lib/insika/env_schema.rb +5 -6
  50. data/lib/insika/errors.rb +1 -0
  51. data/lib/insika/evals/assertions.rb +92 -6
  52. data/lib/insika/evals/golden.rb +91 -2
  53. data/lib/insika/evals/runner.rb +20 -0
  54. data/lib/insika/evals/simulator.rb +11 -2
  55. data/lib/insika/evals/transport.rb +118 -16
  56. data/lib/insika/evidence.rb +79 -12
  57. data/lib/insika/executor.rb +94 -23
  58. data/lib/insika/fence.rb +96 -0
  59. data/lib/insika/golden_store.rb +3 -0
  60. data/lib/insika/loop_detector.rb +5 -34
  61. data/lib/insika/mcp_store.rb +5 -2
  62. data/lib/insika/mcp_tool_registry.rb +8 -1
  63. data/lib/insika/memory_store.rb +12 -0
  64. data/lib/insika/overlay_tool_registry.rb +5 -0
  65. data/lib/insika/prefix_fingerprint.rb +32 -27
  66. data/lib/insika/profile_source.rb +8 -0
  67. data/lib/insika/server/app.rb +43 -1
  68. data/lib/insika/server/rack_app.rb +2 -0
  69. data/lib/insika/server/responses.rb +35 -8
  70. data/lib/insika/session_store.rb +38 -5
  71. data/lib/insika/settings_store.rb +18 -2
  72. data/lib/insika/soak/runner.rb +4 -4
  73. data/lib/insika/spoken_transcript.rb +31 -0
  74. data/lib/insika/studio/app.rb +34 -8
  75. data/lib/insika/studio/forms.rb +29 -3
  76. data/lib/insika/studio/views/_agent_tab_config.erb +5 -1
  77. data/lib/insika/studio/views/session.erb +1 -1
  78. data/lib/insika/studio/views/settings.erb +11 -0
  79. data/lib/insika/studio/views/tool_edit.erb +6 -2
  80. data/lib/insika/telemetry/recorder.rb +61 -1
  81. data/lib/insika/templates/daily-digest/README.md +9 -0
  82. data/lib/insika/templates/research-analyst/agent.rb +10 -0
  83. data/lib/insika/tool_assembly.rb +21 -13
  84. data/lib/insika/tool_batch.rb +67 -0
  85. data/lib/insika/tool_definition.rb +73 -10
  86. data/lib/insika/tool_envelope.rb +102 -2
  87. data/lib/insika/tool_store.rb +9 -4
  88. data/lib/insika/tool_trace_store.rb +1 -1
  89. data/lib/insika/tool_usage_report.rb +172 -0
  90. data/lib/insika/tools/data_defined_tool.rb +1 -0
  91. data/lib/insika/tools/present.rb +122 -0
  92. data/lib/insika/tools/run_persona_eval.rb +6 -1
  93. data/lib/insika/tools/tool_search.rb +4 -2
  94. data/lib/insika/turn_budget.rb +91 -0
  95. data/lib/insika/turn_state.rb +13 -1
  96. data/lib/insika/version.rb +1 -1
  97. data/lib/insika/wiring/graph.rb +7 -0
  98. data/lib/insika/wiring/graph_chat.rb +4 -0
  99. data/lib/insika.rb +11 -0
  100. metadata +10 -1
data/docs/FACTS.md CHANGED
@@ -38,6 +38,10 @@ decides; the engine never applies its own proposal.
38
38
  5. Approved facts join the customer's memory cell and are injected by the
39
39
  Memory provider on the next turn of any session of that customer.
40
40
 
41
+ The extraction transcript contains only nonblank `user` and `assistant` text,
42
+ with original message indexes preserved. Tool payloads are excluded regardless
43
+ of `fencing`; an assistant's repetition of tool text can still be included.
44
+
41
45
  ## Enabling it — the `distill:` block
42
46
 
43
47
  Distillation is pack data on the agent, exactly like `refinement:` or
data/docs/KNOWLEDGE.md CHANGED
@@ -27,6 +27,13 @@ and an operator can see, edit and resolve all of it in the Studio. Only the
27
27
  optional FTS5 index remains, deferred with a measured trigger — see
28
28
  [What's not here yet](#whats-not-here-yet).
29
29
 
30
+ Extraction reads only nonblank `user` and `assistant` text, retaining original
31
+ message indexes. Direct tool payloads are excluded regardless of `fencing`;
32
+ assistant paraphrases can still reach the extractor. The names and descriptions
33
+ in the injected `<knowledge>` block are sanitized when
34
+ [fencing](AGENTS.md#fencing--third-party-text-is-data-never-instructions) is on.
35
+ The full body returned by `load_knowledge` is not fenced.
36
+
30
37
  ## The concept format
31
38
 
32
39
  One concept is one record — a markdown document with a YAML frontmatter
data/docs/LOADTEST.md CHANGED
@@ -67,7 +67,7 @@ frame that carries it.
67
67
 
68
68
  ```bash
69
69
  INSIKA_URL=http://localhost:9292 \
70
- OPENCLAW_GATEWAY_TOKEN=xxx \
70
+ INSIKA_GATEWAY_TOKEN=xxx \
71
71
  bundle exec ruby scripts/loadtest.rb \
72
72
  --agents demo,my-store --concurrency 16 --iterations 3 \
73
73
  --message "hi, how are you?"
@@ -80,7 +80,7 @@ Runs against a local server **or** a remote one (e.g. Railway) — just point
80
80
 
81
81
  | Flag | Default | Meaning |
82
82
  |------|---------|---------|
83
- | `--agents a,b,c` | `demo` | comma-separated agent ids (mapped to `model: openclaw:<agent>`) |
83
+ | `--agents a,b,c` | `demo` | comma-separated agent ids (mapped to `model: insika:<agent>`) |
84
84
  | `--concurrency N` | `8` | concurrent turns per wave |
85
85
  | `--iterations N` | `1` | number of waves per agent |
86
86
  | `--message TEXT` | greeting | user message sent every turn |
@@ -95,14 +95,14 @@ Runs against a local server **or** a remote one (e.g. Railway) — just point
95
95
  | Env | Default | Meaning |
96
96
  |-----|---------|---------|
97
97
  | `INSIKA_URL` | `http://localhost:9292` | base URL of the engine |
98
- | `OPENCLAW_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN`, then `local-demo` | Bearer for `/v1/responses` |
98
+ | `INSIKA_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN`, then `local-demo` | Bearer for `/v1/responses` |
99
99
  | `DEEPSEEK_API_KEY` | — | must be configured **on the server** for real turns (not read by the client) |
100
100
 
101
101
  Use `--dry-run` to sanity-check your flags/URL/token before firing real traffic
102
102
  (and to confirm the request body without needing a running server):
103
103
 
104
104
  ```bash
105
- INSIKA_URL=http://localhost:9292 OPENCLAW_GATEWAY_TOKEN=xxx \
105
+ INSIKA_URL=http://localhost:9292 INSIKA_GATEWAY_TOKEN=xxx \
106
106
  bundle exec ruby scripts/loadtest.rb --agents demo --concurrency 16 --dry-run
107
107
  ```
108
108
 
@@ -133,7 +133,7 @@ DEEPSEEK_API_KEY=sk-... ./scripts/loadtest-local.sh [WORKERS] [CONCURRENCY]
133
133
  | Env | Default | Meaning |
134
134
  |-----|---------|---------|
135
135
  | `DEEPSEEK_API_KEY` | — (required) | real turns hit the provider; also auto-sourced from `.env.local` |
136
- | `OPENCLAW_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN`, then `local-demo` | Bearer for the sweep |
136
+ | `INSIKA_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN`, then `local-demo` | Bearer for the sweep |
137
137
  | `PORT` | `9299` | bind port for the local Falcon |
138
138
  | `AGENT` | `demo` | agent id to load |
139
139
 
@@ -158,30 +158,18 @@ the engine, once at the gateway (its `/v1/responses` speaks the same protocol).
158
158
  Keep `--agents`, `--concurrency`, `--iterations` and `--message` identical, and use
159
159
  matching agents on both sides. Compare the printed TTFB/total/cache/error lines.
160
160
 
161
- ### 4b. Reuse OpenClaw's `loadtest-gateway.mjs` unmodified
161
+ ### 4b. Comparing against an OpenClaw gateway
162
162
 
163
- `loadtest.rb` is the Ruby port of OpenClaw's `loadtest-gateway.mjs`. You do **not**
164
- need to change that script to point it at the engine — because the engine is a
165
- drop-in for the gateway, you only change **where it points**:
163
+ The engine speaks its own wire names (`model: insika:<agent>`, header
164
+ `X-Insika-Agent`), so OpenClaw's `loadtest-gateway.mjs` cannot be pointed at it
165
+ unmodified. For a shadow comparison, run `loadtest.rb` (section 4a) against each
166
+ side with identical `--agents`, `--concurrency`, `--iterations` and `--message`,
167
+ and diff the reports. What you still need in hand:
166
168
 
167
- ```bash
168
- # In the OpenClaw checkout, run its gateway loadtest against the HARNESS:
169
- OPENCLAW_GATEWAY_URL=http://localhost:9292 \
170
- OPENCLAW_GATEWAY_TOKEN=<same bearer the engine accepts> \
171
- node scripts/loadtest-gateway.mjs --agents demo --concurrency 16 --iterations 3
172
- ```
173
-
174
- Then run the exact same command with `OPENCLAW_GATEWAY_URL` pointing at the real
175
- gateway, and diff the two reports. This is the shadow comparison the pilot needs.
176
-
177
- **What the operator must have in hand** (this repo does not vendor OpenClaw):
178
-
179
- - The OpenClaw checkout containing `scripts/loadtest-gateway.mjs` and Node installed.
180
- - A **bearer token accepted by both** sides. For the engine that is
181
- `OPENCLAW_GATEWAY_TOKEN` (see DEPLOY.md); point the gateway run at its own token.
182
- - **The same agent id provisioned on both** sides (e.g. `demo`) so `model:
183
- openclaw:<agent>` resolves on each. On the engine, provision via
184
- `scripts/import_pack.rb`.
169
+ - A bearer token accepted by each side — `INSIKA_GATEWAY_TOKEN` for the engine
170
+ (see DEPLOY.md), the gateway's own token for the gateway run.
171
+ - **The same agent id provisioned on both** sides (e.g. `demo`). On the engine,
172
+ provision via `scripts/import_pack.rb`.
185
173
  - The **same provider** (or an equivalent-latency one) behind each, otherwise you
186
174
  are comparing providers, not engines.
187
175
  - Both endpoints reachable from where you run the client, warmed up (hit `/up` on
data/docs/MEDIA.md CHANGED
@@ -87,7 +87,7 @@ receive it (`channel.capabilities` on the request):
87
87
 
88
88
  ```bash
89
89
  curl -X POST /v1/responses -H "Authorization: Bearer $TOKEN" -d '{
90
- "model": "openclaw:store-support", "user": "chat-7",
90
+ "model": "insika:store-support", "user": "chat-7",
91
91
  "input": "manda a foto do sofá da promoção",
92
92
  "channel": { "capabilities": ["image_output", "audio_output"] }
93
93
  }'
@@ -25,13 +25,17 @@ authoring writes (`:golden_written`, `:agent_file_written`, …), queue bookkeep
25
25
  channel delivery (`:channel_delivered` — see [Channels](CHANNELS.md))
26
26
  travel the same stream and are **ignored** by the bridge: they open no span and
27
27
  touch no instrument, because they are not part of a turn's latency or cost. Any
28
- other subscriber still sees them.
28
+ other subscriber still sees them. Additional events feed counters without opening
29
+ spans: `:tool_loop_intervened`, `:tool_blocked` and `:context_compacted`.
29
30
 
30
31
  They are worth subscribing to even so, because each is the ONLY record of
31
32
  something that left no task of its own behind:
32
33
 
33
34
  | Event | Data | What it answers |
34
35
  |---|---|---|
36
+ | `:tool_blocked` | `name`, `gate`, `param` | a provenance gate refused a call before execution; no rejected value in this event |
37
+ | `:ui` | `component`, `title`, `items`, `count`, `dropped` | the presentation selection, including customer-facing card content; see [API](API.md#tool-and-presentation-sse-events) |
38
+ | `:session_seeded` | `session_id`, `tenant`, `keys` | an eval snapshot was loaded; no snapshot contents |
35
39
  | `:turn_coalesced` | `task_id`, `merged`, `arrivals[]` | the fragments a customer typed in a row arrived as separate messages, and when |
36
40
  | `:turn_steered` | `task_id`, `count`, `total` | a message arrived mid-run and was appended to the turn in flight |
37
41
  | `:turn_steer_released` | `task_id`, `released_as`, `count` | the run could not absorb it, so it became the turn `released_as` |
@@ -68,8 +72,9 @@ and correct while the customer got nothing, because delivery is a separate,
68
72
  retried, out-of-band step. `status: "failed"` means the reply is sitting in the
69
73
  outbox and the customer is still waiting.
70
74
 
71
- Counts, ids and times only — never message content. The text lives in the
72
- transcript, which is the surface that is allowed to carry it.
75
+ Operational audit events use metadata rather than transcript bodies. Presentation
76
+ `:ui` events intentionally contain customer-facing card content; tool-call events
77
+ can contain arguments. Do not treat the whole event stream as content-free.
73
78
 
74
79
  The bridge speaks the standard the market already runs on: point any OTLP backend
75
80
  at Insika and a real turn shows up as a full trace, next to counters and histograms
@@ -80,6 +85,14 @@ backend config, no vendor file. It ships a stable set of attribute and instrumen
80
85
  names, and the recipes below tell you what to chart against them — in whatever you
81
86
  already run.
82
87
 
88
+ ### Blocked tool calls
89
+
90
+ `tool_blocked` carries `name`, `gate`, and `param`, with task/session correlation
91
+ in event metadata. It never carries the rejected value. The `insika.tool.blocked`
92
+ counter (unit `{call}`) uses the turn's agent/tenant/command labels plus
93
+ `insika.tool` and `insika.gate`. Session traces keep the `gate` field alongside the
94
+ masked result; `insika tools:report` lists blocked calls separately from errors.
95
+
83
96
  ## Contents
84
97
 
85
98
  - [Turning it on](#turning-it-on-opt-in-parity-when-off)
@@ -143,14 +156,35 @@ knows its outcome.
143
156
  | `insika.turn.duration` | histogram | `s` | same, when both timestamps are known |
144
157
  | `insika.tokens` | counter | `{token}` | the turn reported usage |
145
158
  | `insika.cost` | counter | `{USD}` | the turn's model is priced (see below) |
159
+ | `insika.tool.blocked` | counter | `{call}` | a `tool_blocked` gate refusal |
146
160
  | `insika.tool.calls` | counter | `{call}` | a tool call completes |
147
161
  | `insika.tool.duration` | histogram | `s` | a `tool_call`/`tool_result` pair completes |
162
+ | `insika.cache.hit_rate` | histogram | `%` | a turn reported billed prompt tokens (see below) |
163
+ | `insika.tool.loop_intervened` | counter | `{intervention}` | the loop detector delivered its one-shot warning |
164
+ | `insika.context.compacted` | counter | `{compaction}` | an in-session compaction was persisted (RFC-0044) |
148
165
 
149
166
  `insika.tool.duration` is deliberately **not** recorded for data-tools: those are a
150
167
  single point-in-time event, so there is no measured duration to report. A tool left
151
168
  open by a mid-turn failure is not counted as a completed call either — its span is
152
169
  closed, but a failed call must not inflate the success histogram.
153
170
 
171
+ `insika.cache.hit_rate` is the same arithmetic the Studio's per-agent series uses:
172
+ cache **reads** over the whole **billed** prompt (fresh input + cache reads + cache
173
+ writes), always in `[0,100]`. A turn with no billed prompt tokens records nothing —
174
+ absence is not a 0% hit. It is recorded per turn on the terminal event, so it
175
+ carries the turn labels (`insika.agent`, `insika.model`, …); the token-counter
176
+ recipe below still works and answers the fleet-wide version of the question.
177
+
178
+ `insika.tool.loop_intervened` counts deliveries of the loop detector's one
179
+ warning per turn (see [Agents](AGENTS.md) `max_tool_repeat`), labelled with
180
+ `insika.tool`. It counts interventions, not repeats: a turn contributes at most 1.
181
+
182
+ `insika.context.compacted` counts persisted in-session compactions
183
+ (`:context_compacted` — see [Context](CONTEXT.md)), labelled with
184
+ `insika.agent` and `insika.model` (the summarizer's model, not the turn's). It
185
+ fires post-turn, after the turn span already closed, so it deliberately rides
186
+ its own labels rather than the open-turn set.
187
+
154
188
  ## Attribute reference
155
189
 
156
190
  The same names are used on spans and on metrics. **Metrics carry a deliberate
@@ -267,7 +301,17 @@ climbing while `input` stays flat.
267
301
 
268
302
  **Cache hit ratio**
269
303
  `insika.tokens` filtered to `insika.token.type="cached"` over the same counter
270
- filtered to `input`. This is the number that moves your bill.
304
+ filtered to `input`. This is the number that moves your bill. For the per-turn
305
+ distribution (does every turn hit, or do fleet averages hide cold agents?), chart
306
+ `insika.cache.hit_rate` — p50 by `insika.agent`. Compare with that agent's baseline:
307
+ cache eligibility, expiry and changing history affect the ratio. The
308
+ [identity-prefix fingerprint](CONTEXT.md#the-observable-cache-fingerprints-and-the-invalidation-reason)
309
+ explains identity/tool-schema changes; volatile category digests are separate.
310
+
311
+ **Loop interventions**
312
+ `insika.tool.loop_intervened`, rate, grouped by `insika.agent` and `insika.tool`.
313
+ Any sustained non-zero rate means one tool keeps being retried with identical
314
+ arguments — fix the tool's contract or the prompt, not the detector.
271
315
 
272
316
  **Spend per tenant**
273
317
  `insika.cost`, rate (or `increase` over a billing window), grouped by
data/docs/POLICY.md CHANGED
@@ -18,17 +18,26 @@ matters, and editable hot.
18
18
  ## Layer 1: Tools (what it can call)
19
19
 
20
20
  `tools_allow` / `tools_deny` / `tools_allow_groups` decide which tools enter the
21
- turn's tool-loop, enforced by the tool-allowlist policy. See [Tools](TOOLS.md)
22
- for how tools are defined and registered, and [`examples/data-tool/`](https://github.com/guizaols/insika/tree/main/examples/data-tool/).
21
+ turn's tool-loop, enforced by the builtin `tool_allowlist` policy — which the
22
+ engine adds for you the moment any of the three is declared, so you never have to
23
+ name it in `policies` (see [Agents](AGENTS.md#the-allowlist-convention)). See
24
+ [Tools](TOOLS.md) for how tools are defined and registered, and
25
+ [`examples/data-tool/`](https://github.com/guizaols/insika/tree/main/examples/data-tool/).
23
26
 
24
27
  ## Layer 2: Policies and approvals
25
28
 
26
- Policies are named entries evaluated before the turn runs. Builtins cover
27
- tool-, skill-, and workflow-allowlisting, plus **`ApprovalRequired`** — which
29
+ Policies are named entries evaluated before the turn runs. **Only the policies a
30
+ profile names run** — which is why `tool_allowlist` is added implicitly by a
31
+ declared tool list; an allowlist nobody applies is worse than no allowlist.
32
+ Builtins cover tool-, skill-, and workflow-allowlisting, plus
33
+ **`ApprovalRequired`** — which
28
34
  does not allow or deny but *tags* a tool as needing human approval. Set
29
35
  `approvals_required: [tool names]`; the gate then fires when the model tries to
30
36
  call that tool, suspending the turn until an operator approves it in the Studio.
31
- See [Security](SECURITY.md#human-approval).
37
+ A data tool's `requires_evidence` check runs before this approval gate; an unknown
38
+ ID is blocked without asking the operator. See
39
+ [Tools](TOOLS.md#provenance-checking-ids-before-a-write) and
40
+ [Security](SECURITY.md#human-approval).
32
41
 
33
42
  ## Layer 3: Guardrails (content safety)
34
43
 
data/docs/RELEASING.md CHANGED
@@ -19,6 +19,10 @@ is invisible to it. Do not publish on rspec alone.
19
19
  3. Every new `lib/` file is **tracked in git**. The gemspec's `files` come from
20
20
  `git ls-files`: an untracked file builds without a warning and the installed
21
21
  gem fails at `require` — this is exactly the failure this proof exists to catch.
22
+ 4. If changing the `fencing` default (currently off), first run the deployment's
23
+ golden cases with it enabled, retain the comparison report, and document the
24
+ behavior change in `CHANGELOG.md`. A default change needs evaluation evidence;
25
+ a scheduled version number alone is not the release gate.
22
26
 
23
27
  ## Cut the gem
24
28
 
@@ -62,7 +62,7 @@ whole surface answers `503`, never open by omission.
62
62
  | `INSIKA_DB` | — (ephemeral memory) | SQLite path → config + execution survive a restart |
63
63
  | `BIND` | `http://localhost:9292` | host:port |
64
64
  | `ADMIN_TOKEN` | `local-demo` | token for `/studio` |
65
- | `OPENCLAW_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN` | Bearer for the whole `/v1` + `/a2a` surface |
65
+ | `INSIKA_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN` | Bearer for the whole `/v1` + `/a2a` surface |
66
66
  | `DEEPSEEK_MODEL` | `deepseek-v4-flash` | model |
67
67
 
68
68
  With persistence:
@@ -128,7 +128,7 @@ Provision it (runs as a client against the live server; the internal token comes
128
128
  from the environment, never disk):
129
129
 
130
130
  ```bash
131
- INSIKA_URL=http://localhost:9292 OPENCLAW_GATEWAY_TOKEN=local-demo \
131
+ INSIKA_URL=http://localhost:9292 INSIKA_GATEWAY_TOKEN=local-demo \
132
132
  bundle exec ruby scripts/import_pack.rb /path/to/pack
133
133
  ```
134
134
 
data/docs/SECURITY.md CHANGED
@@ -27,7 +27,7 @@ The layers, from the edge inward:
27
27
  ## The Bearer gate
28
28
 
29
29
  The `/v1` and `/a2a` surface answers only with
30
- `Authorization: Bearer <OPENCLAW_GATEWAY_TOKEN>` (which falls back to `ADMIN_TOKEN`).
30
+ `Authorization: Bearer <INSIKA_GATEWAY_TOKEN>` (which falls back to `ADMIN_TOKEN`).
31
31
  The check runs in the router, **before** any dispatch, against an **allowlist** of
32
32
  public routes — so a route added later is closed until someone deliberately publishes
33
33
  it. Only these answer without a token:
@@ -205,6 +205,32 @@ category: the agent's category reply → the agent's default → the builtin
205
205
  category → the builtin default. All of it is editable in the Studio Configuration
206
206
  form. See [Agents §Layer 3](POLICY.md#layer-3-guardrails-content-safety).
207
207
 
208
+ ## Third-party text is data (fencing)
209
+
210
+ Input guardrails cover the incoming message. The opt-in `fencing` flag also
211
+ sanitizes ordinary tool result strings, memory text and the injected knowledge
212
+ names/descriptions, and places a
213
+ fixed notice under the identity. It reduces known markup and Unicode tricks;
214
+ it does not make arbitrary third-party instructions safe.
215
+
216
+ The user/assistant-only extraction filter applies even when fencing is off.
217
+ Attachment captions and `load_knowledge` bodies are not fenced. The extraction
218
+ filter excludes direct tool messages,
219
+ but assistant paraphrases can still become extraction input. See
220
+ [Agents](AGENTS.md#fencing--third-party-text-is-data-never-instructions) for the
221
+ exact scope, exceptions, default and size cap.
222
+
223
+ ## Write provenance and serial execution
224
+
225
+ A data tool's `requires_evidence` gate checks IDs against the session ledger
226
+ before approval or backend execution. A known ID proves a prior lookup; it does
227
+ not authorize access or establish current stock, price or quantity limits.
228
+ Those checks remain the backend's responsibility.
229
+
230
+ Tools marked `side_effect` execute one at a time within a session, including in
231
+ parallel batches. This is not a cross-session or distributed backend lock.
232
+ See [Tools](TOOLS.md#provenance-checking-ids-before-a-write).
233
+
208
234
  ## Human approval
209
235
 
210
236
  Some tool calls should not happen unattended. Mark them with
@@ -242,7 +268,7 @@ defense-in-depth: without it, `ALLOW_PRIVATE` opens *any* private destination.
242
268
  > the request never leaves the process, and the conversation *looks* fine. Verify
243
269
  > tool health by the Studio session **trace** (a healthy call shows the backend's
244
270
  > `200`), never by the model's reply. Full detail in
245
- > [Tools §Egress](TOOLS.md#egress-the-ssrf-guard-and-its-silent-failure).
271
+ > [Tools §Egress](TOOLS.md#egress-the-ssrf-guard).
246
272
 
247
273
  ## Sandbox: confined execution
248
274
 
data/docs/SOAK.md CHANGED
@@ -88,7 +88,7 @@ insika soak --dry-run --envelope soak-envelope.md
88
88
  # every precondition, and nothing else
89
89
  insika soak --preflight --envelope soak-envelope.md
90
90
 
91
- # the run itself (INSIKA_URL + OPENCLAW_GATEWAY_TOKEN, like loadtest.rb)
91
+ # the run itself (INSIKA_URL + INSIKA_GATEWAY_TOKEN, like loadtest.rb)
92
92
  INSIKA_URL=https://<target> insika soak --run --envelope soak-envelope.md --out soak-out/
93
93
 
94
94
  # resume after a short outage (the gap is recorded and counts against the window)
data/docs/TOOLS.md CHANGED
@@ -12,10 +12,10 @@ kinds, and the distinction that matters is **who can change one at runtime**:
12
12
 
13
13
  | | **Code tool** | **Data tool** | **MCP tool** |
14
14
  |---|---|---|---|
15
- | What | a Ruby class (`< RubyLLM::Tool`) | an HTTP call described by config, no Ruby | an MCP server's tool, called LIVE |
15
+ | What | a Ruby class (`< RubyLLM::Tool`) | an HTTP call or card presentation described by config | an MCP server's tool, called LIVE |
16
16
  | Lives | in the deployment image | as a row in SQLite | on the MCP server, behind a live client |
17
17
  | Editable at runtime | no (shipped in the image) | **yes** (DSL / API / manifest / Studio) | **yes** — enable/edit the *instance* (DSL / CLI / API / JSON import / Studio); the server owns its own tools |
18
- | Reach for it when | logic must run in-process (file edit, shell, subagent) | calling an external HTTP API | adopting a whole external MCP server's toolset |
18
+ | Reach for it when | logic must run in-process (file edit, shell, subagent) | calling an HTTP API or selecting evidence cards | adopting a whole external MCP server's toolset |
19
19
 
20
20
  **MCP tools are not data tools.** Configuring an enabled MCP **instance** (any
21
21
  surface below) is enough — its tools appear automatically, tagged
@@ -120,7 +120,7 @@ once, at ingestion — see the gotcha below before reaching for it:
120
120
  (`object/array/string/number/integer/boolean`); `oneOf`/`anyOf`/`allOf`/`$ref`/
121
121
  `if`/`then`/`else` are forbidden (not every provider supports them).
122
122
  - `side_effect` defaults from the method (GET/HEAD → false, else true) and drives
123
- checkpoint/replay semantics (a completed side-effecting tool is not re-run on
123
+ serial execution within a session and checkpoint/replay semantics (a completed side-effecting tool is not re-run on
124
124
  resume — see [Architecture](ARCHITECTURE.md#durability-checkpoints-and-resume)).
125
125
 
126
126
  ### `halt_when`: when the answer is already out
@@ -199,12 +199,10 @@ backend, not of whoever calls it. Every agent sharing the tool gets the same val
199
199
 
200
200
  ## Evidence: the lean envelope and grounding
201
201
 
202
- A catalog tool returns products; the model should only ever quote the ones the tool
203
- actually returned — the store dies of a SKU the model invented. `evidence` is the
204
- declaration that makes "no claim without a tool ID" an engine rule instead of a
205
- prompt convention. One declaration does **both** jobs: the engine strips the result
206
- down to what the model sees (the lean envelope) **and** records every returned id on
207
- the session's evidence ledger. There is no "lean but not evidence" mode.
202
+ An `evidence` declaration reshapes a tool result into a lean list and records its
203
+ IDs in the session ledger. The ledger supplies the write gate, presentation tools
204
+ and optional output grounding below. Declaring evidence alone does not prevent
205
+ unsupported claims in the final answer.
208
206
 
209
207
  ```jsonc
210
208
  { "name": "search_products",
@@ -226,14 +224,63 @@ the session's evidence ledger. There is no "lean but not evidence" mode.
226
224
  never a null. A malformed evidence result becomes `{ "error": … }` back to the
227
225
  model — a correctable tool answer, exactly like a malformed call.
228
226
  - **Attachments** are the optional second half: `[{ "type": "card"|"image",
229
- "url": "…", "caption": "…" }]` (≤ 16, url ≤ 500 chars, malformed dropped). They
227
+ "url": "…", "caption": "…", "id": "…" }]` (≤ 16, url ≤ 500 chars, malformed dropped). They
230
228
  **never** reach the model context or the transcript — they ride the channel
231
229
  delivery as an additive `attachments` key on the outbox payload, and the channel
232
- (or its consumer) decides what a card looks like.
230
+ (or its consumer) decides what a card looks like. Supply an explicit `id` when
231
+ cards are not one-to-one with items in the same order. Without one, an attachment
232
+ takes its item's ID at the original position, before malformed cards are dropped.
233
+ - Lean `line` passes through the tool-result sanitizer only when `fencing` is on.
234
+ Attachment captions are normalized to UTF-8 but are not fenced. See
235
+ [Fencing](AGENTS.md#fencing--third-party-text-is-data-never-instructions).
233
236
  - A **code tool** opts in the same way: it either returns `{ items, attachments }`
234
237
  directly and declares `evidence` in its registry metadata, or exposes an
235
238
  `evidence` reader. No declaration = today's tool behavior, byte for byte.
236
239
 
240
+ ### Provenance: checking IDs before a write
241
+
242
+ Declare `requires_evidence` on a data-defined tool to accept only IDs previously
243
+ returned by an `evidence` tool in the same session:
244
+
245
+ ```json
246
+ { "requires_evidence": ["product_id"] }
247
+ ```
248
+
249
+ The full form is `{ "requires_evidence": { "params": ["product_id"] } }`.
250
+ The list must be non-empty and name declared top-level parameters. Scalar values
251
+ and every element of an array are converted to strings and compared exactly:
252
+ `SKU-1` and `sku-1` are different IDs. IDs typed by a customer do not count.
253
+ The ledger includes earlier turns and completed evidence results in the current
254
+ turn, capped at the latest 1,000 distinct IDs. Every declared parameter the call
255
+ carries must pass before any write occurs; a parameter the schema marks optional
256
+ and the model leaves out has nothing to check, a required one left out blocks.
257
+ Search first, then write in a later batch: a search and write in the same parallel
258
+ batch have no dependency ordering guarantee.
259
+
260
+ An unknown ID returns `status: "blocked"`, `gate: "provenance"`, the parameter,
261
+ the value, and an instruction to search or look it up before retrying. The backend
262
+ is never called and no operator approval is requested. A missing ledger blocks the
263
+ call too. Omit the declaration to keep the existing behavior; MCP tools do not
264
+ support this declaration.
265
+
266
+ The Studio tool editor exposes `requires_evidence`. Blocked calls appear in session
267
+ traces and `insika tools:report`, emit `tool_blocked` with name/gate/parameter only,
268
+ and increment `insika.tool.blocked`. `insika doctor` warns when an agent allows a
269
+ gated data tool without an allowed data tool declaring `evidence`.
270
+
271
+ ### Side effects in parallel batches
272
+
273
+ With `limits.tool_concurrency > 1`, tools marked `side_effect` execute one at a time
274
+ within the session's runtime. Unmarked tools still run concurrently, including
275
+ while a write is running. MCP tools are marked `side_effect: true` unless the server
276
+ annotates them `readOnlyHint`; a read-only `POST` data tool needs an explicit
277
+ `"side_effect": false` to keep its concurrency.
278
+ A queued write holds no concurrency slot. Different sessions remain independent;
279
+ backend rules such as quantity limits remain the backend's responsibility.
280
+
281
+ The per-tool timeout starts after both gates are acquired. Trace duration includes
282
+ queueing time, so it measures how long the model waited, not just backend execution.
283
+
237
284
  ### Grounding: policing claims against the ledger
238
285
 
239
286
  With the ledger fed, the pack declares how claims are policed — data on the agent,
@@ -263,11 +310,63 @@ grounding mode: :flag, matcher: { sku: '\b[A-Z]{2,4}\d{4,8}\b' }
263
310
  - Grounding is **independent of the guardrails opt-in**: an agent with guardrails
264
311
  off and `grounding.mode: :flag` still gets the check.
265
312
 
313
+ ## Presentation tools: the model picks ids, the engine shows the cards
314
+
315
+ A **presentation tool** selects which evidence cards to show. Declare it with
316
+ `presentation` instead of `request`; it runs in-process and is always
317
+ `side_effect: false`. The model supplies IDs, never card URLs or captions.
318
+
319
+ ```jsonc
320
+ { "name": "present_products",
321
+ "description": "Show product cards to the customer. Pass only ids a search returned.",
322
+ "parameters": [{ "name": "product_ids", "type": "array:string" },
323
+ { "name": "title", "type": "string", "required": false }],
324
+ "presentation": { "component": "product_cards", // what the channel renders
325
+ "ids": "product_ids", // the array:string parameter
326
+ "max": 8 } } // 1..16 (the attachment cap; default 16)
327
+ ```
328
+
329
+ Exactly one of `request` / `presentation`; `ids` must name a declared `array:string`
330
+ parameter; `component` follows the tool-name rule. Create it through the DSL,
331
+ manifest or API. The Studio form preserves an existing `presentation` declaration
332
+ on save but does not expose its fields. Requested IDs are deduplicated in order
333
+ and compared case-sensitively. The engine then:
334
+
335
+ 1. keeps only ids the session's **evidence ledger** has seen — the rest are dropped
336
+ with reason `unknown` (an id the model invented, or the customer typed);
337
+ 2. joins each kept id to the card an evidence tool returned **this session** — this
338
+ turn's first, then the last 64 the ledger kept (cards carry the `id` of the item
339
+ they stand for); a known id with no card is dropped with reason `no_card`;
340
+ 3. truncates to `max` — the overflow is dropped with reason `max`;
341
+ 4. records the selection on the turn and emits `:ui` on the stream (published as
342
+ `insika.ui` on `/v1/responses`, as the `ui` frame on the web channel);
343
+ 5. answers the model with `shown` (IDs) and `dropped` (objects with `id` and `reason`) — plus one
344
+ instruction when nothing could be shown ("name the products in text or search
345
+ again").
346
+
347
+ **Delivery.** When a turn made a presentation call, the outbox `attachments` are
348
+ *exactly the presented cards*, in call order, each stamped with the call's `component`
349
+ and `title`. A turn with no presentation call delivers every hoarded card, as before —
350
+ a pack that declares no presentation tool sees no change. An empty selection also
351
+ suppresses that fallback: it emits `count: 0` and delivers no cards from that call.
352
+ An earlier turn's ID can be shown without a fresh lookup while its card remains
353
+ in the session's 64-card ledger. IDs and cards have separate caps: a known ID can
354
+ still return `no_card`. Cards are stored snapshots, not a stock or price refresh.
355
+
356
+ A presentation tool needs an evidence source to populate the session's cards.
357
+ `insika doctor` warns when an agent allows presentation without an allowed evidence
358
+ data tool (`presentation-tools`); it cannot verify a code-tool evidence source. The line that tells the model
359
+ *when* to show cards ("show cards with present_products, ids only") is the pack's.
360
+
361
+ Not here: partial rendering while arguments stream, per-component enrichment (price
362
+ today, stock — the card is what the evidence tool returned), and a second component
363
+ such as suggestion chips (same mechanism, when a channel asks for it).
364
+
266
365
  ## Registering a tool
267
366
 
268
367
  A tool appears in the Studio panel and enters an agent's tool-loop when it is
269
368
  **registered** in the catalog **and** allowed by the agent's policy allowlist.
270
- Four ways to write a data tool into the store — all **hot** (registry and catalog
369
+ Three ways to write a data tool into the store — all **hot** (registry and catalog
271
370
  reload, no restart):
272
371
 
273
372
  1. **DSL** — `data_tool(name:, …)` in a `Insika.agent { … }` block.
@@ -275,8 +374,6 @@ reload, no restart):
275
374
  3. **Manifest** — `POST /v1/tools/manifest`. Partial failure is isolated: one
276
375
  malformed tool becomes an `errors[]` entry; only a structural manifest error
277
376
  fails the whole request. The response reports `{ version, created, updated, errors }`.
278
- 4. ~~MCP ingestion~~ — retired. An MCP server's tools are no
279
- longer written into this store at all; see [MCP servers](#mcp-servers).
280
377
 
281
378
  ### The one gotcha: env/secret templating is manifest-only
282
379
 
@@ -298,7 +395,7 @@ at turn time, not ingestion).
298
395
  An MCP **instance** is durable config — transport, target, credentials, an
299
396
  `enabled` flag — held in its own store, separate from data tools. Once an
300
397
  instance is enabled, its tools appear in the catalog automatically (group
301
- `mcp:<instance>`, `side_effect: true`), and every call goes straight to the
398
+ `mcp:<instance>`, `side_effect: true` unless annotated `readOnlyHint`), and every call goes straight to the
302
399
  server through a live, held client — the runtime never converts an MCP tool
303
400
  into a stored data tool.
304
401
 
@@ -405,9 +502,10 @@ held client, which does its own discovery on first use regardless of whether
405
502
  A model can ask for several tools in one step. By default the engine runs them one
406
503
  at a time. Set `limits[:tool_concurrency]` above 1 (see
407
504
  [Agents](AGENTS.md#tool_concurrency--parallel-tool-calls)) and the calls in that
408
- batch run concurrently, **at most N in flight**, on the turn's own reactor — so
409
- the wall-clock of a batch of slow data tools approaches the slowest call rather
410
- than their sum. The cap covers every enveloped tool of the turn, including the
505
+ batch run concurrently, **at most N in flight**, on the turn's own reactor.
506
+ Tools marked `side_effect` still execute one at a time per session; queued writes
507
+ acquire that serial gate before taking a shared concurrency slot. See
508
+ [Side effects](#side-effects-in-parallel-batches). The cap covers every enveloped tool, including the
411
509
  ones `tool_search` promotes mid-turn.
412
510
 
413
511
  It applies only to what the *model* fans out. Two primitives already parallelize
@@ -434,7 +532,7 @@ Turning it on changes three things, all of them worth knowing before you do:
434
532
  Approvals and concurrency are mutually exclusive per turn — the approval gate wins
435
533
  and the turn goes serial. That is a deadlock avoided, not a preference.
436
534
 
437
- ## Egress: the SSRF guard (and its silent failure)
535
+ ## Egress: the SSRF guard
438
536
 
439
537
  Data tools make outbound HTTP, so every call passes through the **EgressGuard**, a
440
538
  Server-Side Request Forgery defense. The default posture is **strict: public
@@ -446,17 +544,10 @@ Server-Side Request Forgery defense. The default posture is **strict: public
446
544
  | `INSIKA_EGRESS_ALLOW_HTTP=1` | permit plain `http` — **loopback dev only** |
447
545
  | `INSIKA_EGRESS_ALLOW_PRIVATE=1` | permit private/loopback IPs — **dev only** |
448
546
 
449
- > ⚠️ **Egress failures are silent.** When a tool targets a blocked host (e.g. a
450
- > plain-`http` localhost backend without the opt-ins), the guard turns the block
451
- > into a `{ error: … }` returned **to the model** — the request never leaves the
452
- > process, yet the stream still emits a tool call, so the model narrates a
453
- > plausible failure and the conversation *looks* like it worked. You will not see
454
- > an exception.
455
- >
456
- > **Always verify by the trace, never by the reply:** open the Studio session
457
- > viewer — a healthy call shows the request, args, and the backend's `200`; a
458
- > missing or errored call is almost always egress (host not in the allowlist, or
459
- > `http`/private without the opt-in).
547
+ An egress rejection returns `{ error: … }` to the model without making the HTTP
548
+ request. Inspect the session trace for the actual error; a plausible reply does
549
+ not prove the backend ran. Provenance refusals instead report `status: "blocked"`
550
+ and `gate: "provenance"`.
460
551
 
461
552
  Egress is **orthogonal** to registration and allowlisting: a tool can be
462
553
  registered, allowed, offered to the model, and still blocked at call time.
@@ -471,8 +562,9 @@ Work down this checklist:
471
562
  `data-tools` check is the only place that says so.
472
563
  2. **Allowed for this agent?** In `tools_allow` (or an allowed group), and not in
473
564
  `tools_deny`?
474
- 3. **Egress?** If it *appears and is called* but "fails", open the trace — a
475
- blocked call is ~99% egress.
565
+ 3. **Call refused or failed?** Inspect the trace. For `gate: "provenance"`, look up
566
+ the ID through an evidence tool before retrying. For an error, check its message
567
+ for schema, egress, timeout or backend failures.
476
568
  4. **URL literal?** For non-manifest tools, an unresolved `{{env.*}}` would have
477
569
  422'd at import — re-check the definition.
478
570
 
@@ -485,6 +577,32 @@ and gets the URL back; the tenant is bound from the turn, never a parameter the
485
577
  model types. See [Artifacts](ARTIFACTS.md) for the tool contract, the serving
486
578
  routes, the signed link and the retention/LGPD reach.
487
579
 
580
+ ## The usage report — `insika tools:report`
581
+
582
+ The per-session trace answers "what did this conversation call"; nothing used to
583
+ answer "what does this agent carry and never use". The report aggregates the
584
+ stored traces per agent (tasks → sessions → `tool_traces`, the same read the
585
+ Studio does) and flags four shapes:
586
+
587
+ - **`never_called`** — in `tools_allow`, zero calls in any stored trace. Dead
588
+ weight: its schema ships on every request and buys nothing.
589
+ - **`error_rate`** — over 30% conventional errors (the trace's `ok` flag) inside
590
+ the window (default 14 days). Either the tool is broken or the model cannot
591
+ hold its contract.
592
+ - **`stale`** — called at some point, but not once inside the window.
593
+ - **`blocked`** — gate refusals in the window, counted by gate. These do not count
594
+ as conventional tool errors.
595
+
596
+ ```bash
597
+ insika tools:report # every stored agent
598
+ insika tools:report --agent store-support # one agent
599
+ insika tools:report --days 30 --json # wider window, machine-readable
600
+ ```
601
+
602
+ Read-only by design: the report names candidates, the **operator** removes — a
603
+ flagged tool may still be the one a rare but critical flow needs. Counts are
604
+ "at least", never exact: the trace keeps a capped tail per session.
605
+
488
606
  ## See also
489
607
 
490
608
  - [Agents](AGENTS.md) — allowlists, groups, and per-agent tool exposure.