insika 0.7.0 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +191 -0
- data/README.md +9 -6
- data/bin/insika +44 -3
- data/docs/AGENTS.md +74 -14
- data/docs/API.md +73 -0
- data/docs/ARCHITECTURE.md +45 -44
- data/docs/ARTIFACTS.md +42 -0
- data/docs/CHANNELS.md +19 -2
- data/docs/CONTEXT.md +86 -38
- data/docs/DEPLOY.md +27 -7
- data/docs/EVALS.md +98 -8
- data/docs/FACTS.md +4 -0
- data/docs/KNOWLEDGE.md +7 -0
- data/docs/LOADTEST.md +15 -27
- data/docs/MEDIA.md +1 -1
- data/docs/OBSERVABILITY.md +48 -4
- data/docs/POLICY.md +14 -5
- data/docs/RELEASING.md +4 -0
- data/docs/RUNNING-LOCAL.md +2 -2
- data/docs/SECURITY.md +28 -2
- data/docs/SOAK.md +1 -1
- data/docs/TOOLS.md +150 -32
- data/docs/prompts/ADD-TOOL.md +12 -2
- data/docs/prompts/DIAGNOSE-TURN.md +3 -0
- data/docs/prompts/GO-LIVE.md +6 -4
- data/lib/insika/agent_profile.rb +47 -10
- data/lib/insika/channels/web/widget.js +33 -0
- data/lib/insika/channels/web.rb +5 -2
- data/lib/insika/chat_builder.rb +90 -37
- data/lib/insika/commands/agent_payload.rb +1 -1
- data/lib/insika/commands/run_distillation.rb +5 -8
- data/lib/insika/commands/seed_session.rb +118 -0
- data/lib/insika/compaction.rb +196 -0
- data/lib/insika/context/builder.rb +35 -11
- data/lib/insika/context/fragment.rb +4 -1
- data/lib/insika/context/priority.rb +8 -0
- data/lib/insika/context/provider.rb +5 -0
- data/lib/insika/context/providers/briefing.rb +61 -29
- data/lib/insika/context/providers/fence_notice.rb +27 -0
- data/lib/insika/context/providers/knowledge.rb +7 -4
- data/lib/insika/context/providers/memory.rb +8 -4
- data/lib/insika/context/providers/session.rb +50 -10
- data/lib/insika/context_trace_store.rb +11 -1
- data/lib/insika/doctor.rb +213 -10
- data/lib/insika/dsl/runtime.rb +5 -0
- data/lib/insika/dsl.rb +6 -0
- data/lib/insika/edge_limiter.rb +4 -1
- data/lib/insika/env_schema.rb +5 -6
- data/lib/insika/errors.rb +1 -0
- data/lib/insika/evals/assertions.rb +92 -6
- data/lib/insika/evals/golden.rb +91 -2
- data/lib/insika/evals/runner.rb +20 -0
- data/lib/insika/evals/simulator.rb +11 -2
- data/lib/insika/evals/transport.rb +118 -16
- data/lib/insika/evidence.rb +79 -12
- data/lib/insika/executor.rb +94 -23
- data/lib/insika/fence.rb +96 -0
- data/lib/insika/golden_store.rb +3 -0
- data/lib/insika/loop_detector.rb +5 -34
- data/lib/insika/mcp_store.rb +5 -2
- data/lib/insika/mcp_tool_registry.rb +8 -1
- data/lib/insika/memory_store.rb +12 -0
- data/lib/insika/overlay_tool_registry.rb +5 -0
- data/lib/insika/prefix_fingerprint.rb +32 -27
- data/lib/insika/profile_source.rb +8 -0
- data/lib/insika/server/app.rb +43 -1
- data/lib/insika/server/rack_app.rb +2 -0
- data/lib/insika/server/responses.rb +35 -8
- data/lib/insika/session_store.rb +38 -5
- data/lib/insika/settings_store.rb +18 -2
- data/lib/insika/soak/runner.rb +4 -4
- data/lib/insika/spoken_transcript.rb +31 -0
- data/lib/insika/studio/app.rb +34 -8
- data/lib/insika/studio/forms.rb +29 -3
- data/lib/insika/studio/views/_agent_tab_config.erb +5 -1
- data/lib/insika/studio/views/session.erb +1 -1
- data/lib/insika/studio/views/settings.erb +11 -0
- data/lib/insika/studio/views/tool_edit.erb +6 -2
- data/lib/insika/telemetry/recorder.rb +61 -1
- data/lib/insika/templates/daily-digest/README.md +9 -0
- data/lib/insika/templates/research-analyst/agent.rb +10 -0
- data/lib/insika/tool_assembly.rb +21 -13
- data/lib/insika/tool_batch.rb +67 -0
- data/lib/insika/tool_definition.rb +73 -10
- data/lib/insika/tool_envelope.rb +102 -2
- data/lib/insika/tool_store.rb +9 -4
- data/lib/insika/tool_trace_store.rb +1 -1
- data/lib/insika/tool_usage_report.rb +172 -0
- data/lib/insika/tools/data_defined_tool.rb +1 -0
- data/lib/insika/tools/present.rb +122 -0
- data/lib/insika/tools/run_persona_eval.rb +6 -1
- data/lib/insika/tools/tool_search.rb +4 -2
- data/lib/insika/turn_budget.rb +91 -0
- data/lib/insika/turn_state.rb +13 -1
- data/lib/insika/version.rb +1 -1
- data/lib/insika/wiring/graph.rb +7 -0
- data/lib/insika/wiring/graph_chat.rb +4 -0
- data/lib/insika.rb +11 -0
- metadata +10 -1
data/docs/FACTS.md
CHANGED
|
@@ -38,6 +38,10 @@ decides; the engine never applies its own proposal.
|
|
|
38
38
|
5. Approved facts join the customer's memory cell and are injected by the
|
|
39
39
|
Memory provider on the next turn of any session of that customer.
|
|
40
40
|
|
|
41
|
+
The extraction transcript contains only nonblank `user` and `assistant` text,
|
|
42
|
+
with original message indexes preserved. Tool payloads are excluded regardless
|
|
43
|
+
of `fencing`; an assistant's repetition of tool text can still be included.
|
|
44
|
+
|
|
41
45
|
## Enabling it — the `distill:` block
|
|
42
46
|
|
|
43
47
|
Distillation is pack data on the agent, exactly like `refinement:` or
|
data/docs/KNOWLEDGE.md
CHANGED
|
@@ -27,6 +27,13 @@ and an operator can see, edit and resolve all of it in the Studio. Only the
|
|
|
27
27
|
optional FTS5 index remains, deferred with a measured trigger — see
|
|
28
28
|
[What's not here yet](#whats-not-here-yet).
|
|
29
29
|
|
|
30
|
+
Extraction reads only nonblank `user` and `assistant` text, retaining original
|
|
31
|
+
message indexes. Direct tool payloads are excluded regardless of `fencing`;
|
|
32
|
+
assistant paraphrases can still reach the extractor. The names and descriptions
|
|
33
|
+
in the injected `<knowledge>` block are sanitized when
|
|
34
|
+
[fencing](AGENTS.md#fencing--third-party-text-is-data-never-instructions) is on.
|
|
35
|
+
The full body returned by `load_knowledge` is not fenced.
|
|
36
|
+
|
|
30
37
|
## The concept format
|
|
31
38
|
|
|
32
39
|
One concept is one record — a markdown document with a YAML frontmatter
|
data/docs/LOADTEST.md
CHANGED
|
@@ -67,7 +67,7 @@ frame that carries it.
|
|
|
67
67
|
|
|
68
68
|
```bash
|
|
69
69
|
INSIKA_URL=http://localhost:9292 \
|
|
70
|
-
|
|
70
|
+
INSIKA_GATEWAY_TOKEN=xxx \
|
|
71
71
|
bundle exec ruby scripts/loadtest.rb \
|
|
72
72
|
--agents demo,my-store --concurrency 16 --iterations 3 \
|
|
73
73
|
--message "hi, how are you?"
|
|
@@ -80,7 +80,7 @@ Runs against a local server **or** a remote one (e.g. Railway) — just point
|
|
|
80
80
|
|
|
81
81
|
| Flag | Default | Meaning |
|
|
82
82
|
|------|---------|---------|
|
|
83
|
-
| `--agents a,b,c` | `demo` | comma-separated agent ids (mapped to `model:
|
|
83
|
+
| `--agents a,b,c` | `demo` | comma-separated agent ids (mapped to `model: insika:<agent>`) |
|
|
84
84
|
| `--concurrency N` | `8` | concurrent turns per wave |
|
|
85
85
|
| `--iterations N` | `1` | number of waves per agent |
|
|
86
86
|
| `--message TEXT` | greeting | user message sent every turn |
|
|
@@ -95,14 +95,14 @@ Runs against a local server **or** a remote one (e.g. Railway) — just point
|
|
|
95
95
|
| Env | Default | Meaning |
|
|
96
96
|
|-----|---------|---------|
|
|
97
97
|
| `INSIKA_URL` | `http://localhost:9292` | base URL of the engine |
|
|
98
|
-
| `
|
|
98
|
+
| `INSIKA_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN`, then `local-demo` | Bearer for `/v1/responses` |
|
|
99
99
|
| `DEEPSEEK_API_KEY` | — | must be configured **on the server** for real turns (not read by the client) |
|
|
100
100
|
|
|
101
101
|
Use `--dry-run` to sanity-check your flags/URL/token before firing real traffic
|
|
102
102
|
(and to confirm the request body without needing a running server):
|
|
103
103
|
|
|
104
104
|
```bash
|
|
105
|
-
INSIKA_URL=http://localhost:9292
|
|
105
|
+
INSIKA_URL=http://localhost:9292 INSIKA_GATEWAY_TOKEN=xxx \
|
|
106
106
|
bundle exec ruby scripts/loadtest.rb --agents demo --concurrency 16 --dry-run
|
|
107
107
|
```
|
|
108
108
|
|
|
@@ -133,7 +133,7 @@ DEEPSEEK_API_KEY=sk-... ./scripts/loadtest-local.sh [WORKERS] [CONCURRENCY]
|
|
|
133
133
|
| Env | Default | Meaning |
|
|
134
134
|
|-----|---------|---------|
|
|
135
135
|
| `DEEPSEEK_API_KEY` | — (required) | real turns hit the provider; also auto-sourced from `.env.local` |
|
|
136
|
-
| `
|
|
136
|
+
| `INSIKA_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN`, then `local-demo` | Bearer for the sweep |
|
|
137
137
|
| `PORT` | `9299` | bind port for the local Falcon |
|
|
138
138
|
| `AGENT` | `demo` | agent id to load |
|
|
139
139
|
|
|
@@ -158,30 +158,18 @@ the engine, once at the gateway (its `/v1/responses` speaks the same protocol).
|
|
|
158
158
|
Keep `--agents`, `--concurrency`, `--iterations` and `--message` identical, and use
|
|
159
159
|
matching agents on both sides. Compare the printed TTFB/total/cache/error lines.
|
|
160
160
|
|
|
161
|
-
### 4b.
|
|
161
|
+
### 4b. Comparing against an OpenClaw gateway
|
|
162
162
|
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
163
|
+
The engine speaks its own wire names (`model: insika:<agent>`, header
|
|
164
|
+
`X-Insika-Agent`), so OpenClaw's `loadtest-gateway.mjs` cannot be pointed at it
|
|
165
|
+
unmodified. For a shadow comparison, run `loadtest.rb` (section 4a) against each
|
|
166
|
+
side with identical `--agents`, `--concurrency`, `--iterations` and `--message`,
|
|
167
|
+
and diff the reports. What you still need in hand:
|
|
166
168
|
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
node scripts/loadtest-gateway.mjs --agents demo --concurrency 16 --iterations 3
|
|
172
|
-
```
|
|
173
|
-
|
|
174
|
-
Then run the exact same command with `OPENCLAW_GATEWAY_URL` pointing at the real
|
|
175
|
-
gateway, and diff the two reports. This is the shadow comparison the pilot needs.
|
|
176
|
-
|
|
177
|
-
**What the operator must have in hand** (this repo does not vendor OpenClaw):
|
|
178
|
-
|
|
179
|
-
- The OpenClaw checkout containing `scripts/loadtest-gateway.mjs` and Node installed.
|
|
180
|
-
- A **bearer token accepted by both** sides. For the engine that is
|
|
181
|
-
`OPENCLAW_GATEWAY_TOKEN` (see DEPLOY.md); point the gateway run at its own token.
|
|
182
|
-
- **The same agent id provisioned on both** sides (e.g. `demo`) so `model:
|
|
183
|
-
openclaw:<agent>` resolves on each. On the engine, provision via
|
|
184
|
-
`scripts/import_pack.rb`.
|
|
169
|
+
- A bearer token accepted by each side — `INSIKA_GATEWAY_TOKEN` for the engine
|
|
170
|
+
(see DEPLOY.md), the gateway's own token for the gateway run.
|
|
171
|
+
- **The same agent id provisioned on both** sides (e.g. `demo`). On the engine,
|
|
172
|
+
provision via `scripts/import_pack.rb`.
|
|
185
173
|
- The **same provider** (or an equivalent-latency one) behind each, otherwise you
|
|
186
174
|
are comparing providers, not engines.
|
|
187
175
|
- Both endpoints reachable from where you run the client, warmed up (hit `/up` on
|
data/docs/MEDIA.md
CHANGED
|
@@ -87,7 +87,7 @@ receive it (`channel.capabilities` on the request):
|
|
|
87
87
|
|
|
88
88
|
```bash
|
|
89
89
|
curl -X POST /v1/responses -H "Authorization: Bearer $TOKEN" -d '{
|
|
90
|
-
"model": "
|
|
90
|
+
"model": "insika:store-support", "user": "chat-7",
|
|
91
91
|
"input": "manda a foto do sofá da promoção",
|
|
92
92
|
"channel": { "capabilities": ["image_output", "audio_output"] }
|
|
93
93
|
}'
|
data/docs/OBSERVABILITY.md
CHANGED
|
@@ -25,13 +25,17 @@ authoring writes (`:golden_written`, `:agent_file_written`, …), queue bookkeep
|
|
|
25
25
|
channel delivery (`:channel_delivered` — see [Channels](CHANNELS.md))
|
|
26
26
|
travel the same stream and are **ignored** by the bridge: they open no span and
|
|
27
27
|
touch no instrument, because they are not part of a turn's latency or cost. Any
|
|
28
|
-
other subscriber still sees them.
|
|
28
|
+
other subscriber still sees them. Additional events feed counters without opening
|
|
29
|
+
spans: `:tool_loop_intervened`, `:tool_blocked` and `:context_compacted`.
|
|
29
30
|
|
|
30
31
|
They are worth subscribing to even so, because each is the ONLY record of
|
|
31
32
|
something that left no task of its own behind:
|
|
32
33
|
|
|
33
34
|
| Event | Data | What it answers |
|
|
34
35
|
|---|---|---|
|
|
36
|
+
| `:tool_blocked` | `name`, `gate`, `param` | a provenance gate refused a call before execution; no rejected value in this event |
|
|
37
|
+
| `:ui` | `component`, `title`, `items`, `count`, `dropped` | the presentation selection, including customer-facing card content; see [API](API.md#tool-and-presentation-sse-events) |
|
|
38
|
+
| `:session_seeded` | `session_id`, `tenant`, `keys` | an eval snapshot was loaded; no snapshot contents |
|
|
35
39
|
| `:turn_coalesced` | `task_id`, `merged`, `arrivals[]` | the fragments a customer typed in a row arrived as separate messages, and when |
|
|
36
40
|
| `:turn_steered` | `task_id`, `count`, `total` | a message arrived mid-run and was appended to the turn in flight |
|
|
37
41
|
| `:turn_steer_released` | `task_id`, `released_as`, `count` | the run could not absorb it, so it became the turn `released_as` |
|
|
@@ -68,8 +72,9 @@ and correct while the customer got nothing, because delivery is a separate,
|
|
|
68
72
|
retried, out-of-band step. `status: "failed"` means the reply is sitting in the
|
|
69
73
|
outbox and the customer is still waiting.
|
|
70
74
|
|
|
71
|
-
|
|
72
|
-
|
|
75
|
+
Operational audit events use metadata rather than transcript bodies. Presentation
|
|
76
|
+
`:ui` events intentionally contain customer-facing card content; tool-call events
|
|
77
|
+
can contain arguments. Do not treat the whole event stream as content-free.
|
|
73
78
|
|
|
74
79
|
The bridge speaks the standard the market already runs on: point any OTLP backend
|
|
75
80
|
at Insika and a real turn shows up as a full trace, next to counters and histograms
|
|
@@ -80,6 +85,14 @@ backend config, no vendor file. It ships a stable set of attribute and instrumen
|
|
|
80
85
|
names, and the recipes below tell you what to chart against them — in whatever you
|
|
81
86
|
already run.
|
|
82
87
|
|
|
88
|
+
### Blocked tool calls
|
|
89
|
+
|
|
90
|
+
`tool_blocked` carries `name`, `gate`, and `param`, with task/session correlation
|
|
91
|
+
in event metadata. It never carries the rejected value. The `insika.tool.blocked`
|
|
92
|
+
counter (unit `{call}`) uses the turn's agent/tenant/command labels plus
|
|
93
|
+
`insika.tool` and `insika.gate`. Session traces keep the `gate` field alongside the
|
|
94
|
+
masked result; `insika tools:report` lists blocked calls separately from errors.
|
|
95
|
+
|
|
83
96
|
## Contents
|
|
84
97
|
|
|
85
98
|
- [Turning it on](#turning-it-on-opt-in-parity-when-off)
|
|
@@ -143,14 +156,35 @@ knows its outcome.
|
|
|
143
156
|
| `insika.turn.duration` | histogram | `s` | same, when both timestamps are known |
|
|
144
157
|
| `insika.tokens` | counter | `{token}` | the turn reported usage |
|
|
145
158
|
| `insika.cost` | counter | `{USD}` | the turn's model is priced (see below) |
|
|
159
|
+
| `insika.tool.blocked` | counter | `{call}` | a `tool_blocked` gate refusal |
|
|
146
160
|
| `insika.tool.calls` | counter | `{call}` | a tool call completes |
|
|
147
161
|
| `insika.tool.duration` | histogram | `s` | a `tool_call`/`tool_result` pair completes |
|
|
162
|
+
| `insika.cache.hit_rate` | histogram | `%` | a turn reported billed prompt tokens (see below) |
|
|
163
|
+
| `insika.tool.loop_intervened` | counter | `{intervention}` | the loop detector delivered its one-shot warning |
|
|
164
|
+
| `insika.context.compacted` | counter | `{compaction}` | an in-session compaction was persisted (RFC-0044) |
|
|
148
165
|
|
|
149
166
|
`insika.tool.duration` is deliberately **not** recorded for data-tools: those are a
|
|
150
167
|
single point-in-time event, so there is no measured duration to report. A tool left
|
|
151
168
|
open by a mid-turn failure is not counted as a completed call either — its span is
|
|
152
169
|
closed, but a failed call must not inflate the success histogram.
|
|
153
170
|
|
|
171
|
+
`insika.cache.hit_rate` is the same arithmetic the Studio's per-agent series uses:
|
|
172
|
+
cache **reads** over the whole **billed** prompt (fresh input + cache reads + cache
|
|
173
|
+
writes), always in `[0,100]`. A turn with no billed prompt tokens records nothing —
|
|
174
|
+
absence is not a 0% hit. It is recorded per turn on the terminal event, so it
|
|
175
|
+
carries the turn labels (`insika.agent`, `insika.model`, …); the token-counter
|
|
176
|
+
recipe below still works and answers the fleet-wide version of the question.
|
|
177
|
+
|
|
178
|
+
`insika.tool.loop_intervened` counts deliveries of the loop detector's one
|
|
179
|
+
warning per turn (see [Agents](AGENTS.md) `max_tool_repeat`), labelled with
|
|
180
|
+
`insika.tool`. It counts interventions, not repeats: a turn contributes at most 1.
|
|
181
|
+
|
|
182
|
+
`insika.context.compacted` counts persisted in-session compactions
|
|
183
|
+
(`:context_compacted` — see [Context](CONTEXT.md)), labelled with
|
|
184
|
+
`insika.agent` and `insika.model` (the summarizer's model, not the turn's). It
|
|
185
|
+
fires post-turn, after the turn span already closed, so it deliberately rides
|
|
186
|
+
its own labels rather than the open-turn set.
|
|
187
|
+
|
|
154
188
|
## Attribute reference
|
|
155
189
|
|
|
156
190
|
The same names are used on spans and on metrics. **Metrics carry a deliberate
|
|
@@ -267,7 +301,17 @@ climbing while `input` stays flat.
|
|
|
267
301
|
|
|
268
302
|
**Cache hit ratio**
|
|
269
303
|
`insika.tokens` filtered to `insika.token.type="cached"` over the same counter
|
|
270
|
-
filtered to `input`. This is the number that moves your bill.
|
|
304
|
+
filtered to `input`. This is the number that moves your bill. For the per-turn
|
|
305
|
+
distribution (does every turn hit, or do fleet averages hide cold agents?), chart
|
|
306
|
+
`insika.cache.hit_rate` — p50 by `insika.agent`. Compare with that agent's baseline:
|
|
307
|
+
cache eligibility, expiry and changing history affect the ratio. The
|
|
308
|
+
[identity-prefix fingerprint](CONTEXT.md#the-observable-cache-fingerprints-and-the-invalidation-reason)
|
|
309
|
+
explains identity/tool-schema changes; volatile category digests are separate.
|
|
310
|
+
|
|
311
|
+
**Loop interventions**
|
|
312
|
+
`insika.tool.loop_intervened`, rate, grouped by `insika.agent` and `insika.tool`.
|
|
313
|
+
Any sustained non-zero rate means one tool keeps being retried with identical
|
|
314
|
+
arguments — fix the tool's contract or the prompt, not the detector.
|
|
271
315
|
|
|
272
316
|
**Spend per tenant**
|
|
273
317
|
`insika.cost`, rate (or `increase` over a billing window), grouped by
|
data/docs/POLICY.md
CHANGED
|
@@ -18,17 +18,26 @@ matters, and editable hot.
|
|
|
18
18
|
## Layer 1: Tools (what it can call)
|
|
19
19
|
|
|
20
20
|
`tools_allow` / `tools_deny` / `tools_allow_groups` decide which tools enter the
|
|
21
|
-
turn's tool-loop, enforced by the
|
|
22
|
-
for
|
|
21
|
+
turn's tool-loop, enforced by the builtin `tool_allowlist` policy — which the
|
|
22
|
+
engine adds for you the moment any of the three is declared, so you never have to
|
|
23
|
+
name it in `policies` (see [Agents](AGENTS.md#the-allowlist-convention)). See
|
|
24
|
+
[Tools](TOOLS.md) for how tools are defined and registered, and
|
|
25
|
+
[`examples/data-tool/`](https://github.com/guizaols/insika/tree/main/examples/data-tool/).
|
|
23
26
|
|
|
24
27
|
## Layer 2: Policies and approvals
|
|
25
28
|
|
|
26
|
-
Policies are named entries evaluated before the turn runs.
|
|
27
|
-
|
|
29
|
+
Policies are named entries evaluated before the turn runs. **Only the policies a
|
|
30
|
+
profile names run** — which is why `tool_allowlist` is added implicitly by a
|
|
31
|
+
declared tool list; an allowlist nobody applies is worse than no allowlist.
|
|
32
|
+
Builtins cover tool-, skill-, and workflow-allowlisting, plus
|
|
33
|
+
**`ApprovalRequired`** — which
|
|
28
34
|
does not allow or deny but *tags* a tool as needing human approval. Set
|
|
29
35
|
`approvals_required: [tool names]`; the gate then fires when the model tries to
|
|
30
36
|
call that tool, suspending the turn until an operator approves it in the Studio.
|
|
31
|
-
|
|
37
|
+
A data tool's `requires_evidence` check runs before this approval gate; an unknown
|
|
38
|
+
ID is blocked without asking the operator. See
|
|
39
|
+
[Tools](TOOLS.md#provenance-checking-ids-before-a-write) and
|
|
40
|
+
[Security](SECURITY.md#human-approval).
|
|
32
41
|
|
|
33
42
|
## Layer 3: Guardrails (content safety)
|
|
34
43
|
|
data/docs/RELEASING.md
CHANGED
|
@@ -19,6 +19,10 @@ is invisible to it. Do not publish on rspec alone.
|
|
|
19
19
|
3. Every new `lib/` file is **tracked in git**. The gemspec's `files` come from
|
|
20
20
|
`git ls-files`: an untracked file builds without a warning and the installed
|
|
21
21
|
gem fails at `require` — this is exactly the failure this proof exists to catch.
|
|
22
|
+
4. If changing the `fencing` default (currently off), first run the deployment's
|
|
23
|
+
golden cases with it enabled, retain the comparison report, and document the
|
|
24
|
+
behavior change in `CHANGELOG.md`. A default change needs evaluation evidence;
|
|
25
|
+
a scheduled version number alone is not the release gate.
|
|
22
26
|
|
|
23
27
|
## Cut the gem
|
|
24
28
|
|
data/docs/RUNNING-LOCAL.md
CHANGED
|
@@ -62,7 +62,7 @@ whole surface answers `503`, never open by omission.
|
|
|
62
62
|
| `INSIKA_DB` | — (ephemeral memory) | SQLite path → config + execution survive a restart |
|
|
63
63
|
| `BIND` | `http://localhost:9292` | host:port |
|
|
64
64
|
| `ADMIN_TOKEN` | `local-demo` | token for `/studio` |
|
|
65
|
-
| `
|
|
65
|
+
| `INSIKA_GATEWAY_TOKEN` | falls back to `ADMIN_TOKEN` | Bearer for the whole `/v1` + `/a2a` surface |
|
|
66
66
|
| `DEEPSEEK_MODEL` | `deepseek-v4-flash` | model |
|
|
67
67
|
|
|
68
68
|
With persistence:
|
|
@@ -128,7 +128,7 @@ Provision it (runs as a client against the live server; the internal token comes
|
|
|
128
128
|
from the environment, never disk):
|
|
129
129
|
|
|
130
130
|
```bash
|
|
131
|
-
INSIKA_URL=http://localhost:9292
|
|
131
|
+
INSIKA_URL=http://localhost:9292 INSIKA_GATEWAY_TOKEN=local-demo \
|
|
132
132
|
bundle exec ruby scripts/import_pack.rb /path/to/pack
|
|
133
133
|
```
|
|
134
134
|
|
data/docs/SECURITY.md
CHANGED
|
@@ -27,7 +27,7 @@ The layers, from the edge inward:
|
|
|
27
27
|
## The Bearer gate
|
|
28
28
|
|
|
29
29
|
The `/v1` and `/a2a` surface answers only with
|
|
30
|
-
`Authorization: Bearer <
|
|
30
|
+
`Authorization: Bearer <INSIKA_GATEWAY_TOKEN>` (which falls back to `ADMIN_TOKEN`).
|
|
31
31
|
The check runs in the router, **before** any dispatch, against an **allowlist** of
|
|
32
32
|
public routes — so a route added later is closed until someone deliberately publishes
|
|
33
33
|
it. Only these answer without a token:
|
|
@@ -205,6 +205,32 @@ category: the agent's category reply → the agent's default → the builtin
|
|
|
205
205
|
category → the builtin default. All of it is editable in the Studio Configuration
|
|
206
206
|
form. See [Agents §Layer 3](POLICY.md#layer-3-guardrails-content-safety).
|
|
207
207
|
|
|
208
|
+
## Third-party text is data (fencing)
|
|
209
|
+
|
|
210
|
+
Input guardrails cover the incoming message. The opt-in `fencing` flag also
|
|
211
|
+
sanitizes ordinary tool result strings, memory text and the injected knowledge
|
|
212
|
+
names/descriptions, and places a
|
|
213
|
+
fixed notice under the identity. It reduces known markup and Unicode tricks;
|
|
214
|
+
it does not make arbitrary third-party instructions safe.
|
|
215
|
+
|
|
216
|
+
The user/assistant-only extraction filter applies even when fencing is off.
|
|
217
|
+
Attachment captions and `load_knowledge` bodies are not fenced. The extraction
|
|
218
|
+
filter excludes direct tool messages,
|
|
219
|
+
but assistant paraphrases can still become extraction input. See
|
|
220
|
+
[Agents](AGENTS.md#fencing--third-party-text-is-data-never-instructions) for the
|
|
221
|
+
exact scope, exceptions, default and size cap.
|
|
222
|
+
|
|
223
|
+
## Write provenance and serial execution
|
|
224
|
+
|
|
225
|
+
A data tool's `requires_evidence` gate checks IDs against the session ledger
|
|
226
|
+
before approval or backend execution. A known ID proves a prior lookup; it does
|
|
227
|
+
not authorize access or establish current stock, price or quantity limits.
|
|
228
|
+
Those checks remain the backend's responsibility.
|
|
229
|
+
|
|
230
|
+
Tools marked `side_effect` execute one at a time within a session, including in
|
|
231
|
+
parallel batches. This is not a cross-session or distributed backend lock.
|
|
232
|
+
See [Tools](TOOLS.md#provenance-checking-ids-before-a-write).
|
|
233
|
+
|
|
208
234
|
## Human approval
|
|
209
235
|
|
|
210
236
|
Some tool calls should not happen unattended. Mark them with
|
|
@@ -242,7 +268,7 @@ defense-in-depth: without it, `ALLOW_PRIVATE` opens *any* private destination.
|
|
|
242
268
|
> the request never leaves the process, and the conversation *looks* fine. Verify
|
|
243
269
|
> tool health by the Studio session **trace** (a healthy call shows the backend's
|
|
244
270
|
> `200`), never by the model's reply. Full detail in
|
|
245
|
-
> [Tools §Egress](TOOLS.md#egress-the-ssrf-guard
|
|
271
|
+
> [Tools §Egress](TOOLS.md#egress-the-ssrf-guard).
|
|
246
272
|
|
|
247
273
|
## Sandbox: confined execution
|
|
248
274
|
|
data/docs/SOAK.md
CHANGED
|
@@ -88,7 +88,7 @@ insika soak --dry-run --envelope soak-envelope.md
|
|
|
88
88
|
# every precondition, and nothing else
|
|
89
89
|
insika soak --preflight --envelope soak-envelope.md
|
|
90
90
|
|
|
91
|
-
# the run itself (INSIKA_URL +
|
|
91
|
+
# the run itself (INSIKA_URL + INSIKA_GATEWAY_TOKEN, like loadtest.rb)
|
|
92
92
|
INSIKA_URL=https://<target> insika soak --run --envelope soak-envelope.md --out soak-out/
|
|
93
93
|
|
|
94
94
|
# resume after a short outage (the gap is recorded and counts against the window)
|
data/docs/TOOLS.md
CHANGED
|
@@ -12,10 +12,10 @@ kinds, and the distinction that matters is **who can change one at runtime**:
|
|
|
12
12
|
|
|
13
13
|
| | **Code tool** | **Data tool** | **MCP tool** |
|
|
14
14
|
|---|---|---|---|
|
|
15
|
-
| What | a Ruby class (`< RubyLLM::Tool`) | an HTTP call described by config
|
|
15
|
+
| What | a Ruby class (`< RubyLLM::Tool`) | an HTTP call or card presentation described by config | an MCP server's tool, called LIVE |
|
|
16
16
|
| Lives | in the deployment image | as a row in SQLite | on the MCP server, behind a live client |
|
|
17
17
|
| Editable at runtime | no (shipped in the image) | **yes** (DSL / API / manifest / Studio) | **yes** — enable/edit the *instance* (DSL / CLI / API / JSON import / Studio); the server owns its own tools |
|
|
18
|
-
| Reach for it when | logic must run in-process (file edit, shell, subagent) | calling an
|
|
18
|
+
| Reach for it when | logic must run in-process (file edit, shell, subagent) | calling an HTTP API or selecting evidence cards | adopting a whole external MCP server's toolset |
|
|
19
19
|
|
|
20
20
|
**MCP tools are not data tools.** Configuring an enabled MCP **instance** (any
|
|
21
21
|
surface below) is enough — its tools appear automatically, tagged
|
|
@@ -120,7 +120,7 @@ once, at ingestion — see the gotcha below before reaching for it:
|
|
|
120
120
|
(`object/array/string/number/integer/boolean`); `oneOf`/`anyOf`/`allOf`/`$ref`/
|
|
121
121
|
`if`/`then`/`else` are forbidden (not every provider supports them).
|
|
122
122
|
- `side_effect` defaults from the method (GET/HEAD → false, else true) and drives
|
|
123
|
-
checkpoint/replay semantics (a completed side-effecting tool is not re-run on
|
|
123
|
+
serial execution within a session and checkpoint/replay semantics (a completed side-effecting tool is not re-run on
|
|
124
124
|
resume — see [Architecture](ARCHITECTURE.md#durability-checkpoints-and-resume)).
|
|
125
125
|
|
|
126
126
|
### `halt_when`: when the answer is already out
|
|
@@ -199,12 +199,10 @@ backend, not of whoever calls it. Every agent sharing the tool gets the same val
|
|
|
199
199
|
|
|
200
200
|
## Evidence: the lean envelope and grounding
|
|
201
201
|
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
down to what the model sees (the lean envelope) **and** records every returned id on
|
|
207
|
-
the session's evidence ledger. There is no "lean but not evidence" mode.
|
|
202
|
+
An `evidence` declaration reshapes a tool result into a lean list and records its
|
|
203
|
+
IDs in the session ledger. The ledger supplies the write gate, presentation tools
|
|
204
|
+
and optional output grounding below. Declaring evidence alone does not prevent
|
|
205
|
+
unsupported claims in the final answer.
|
|
208
206
|
|
|
209
207
|
```jsonc
|
|
210
208
|
{ "name": "search_products",
|
|
@@ -226,14 +224,63 @@ the session's evidence ledger. There is no "lean but not evidence" mode.
|
|
|
226
224
|
never a null. A malformed evidence result becomes `{ "error": … }` back to the
|
|
227
225
|
model — a correctable tool answer, exactly like a malformed call.
|
|
228
226
|
- **Attachments** are the optional second half: `[{ "type": "card"|"image",
|
|
229
|
-
"url": "…", "caption": "…" }]` (≤ 16, url ≤ 500 chars, malformed dropped). They
|
|
227
|
+
"url": "…", "caption": "…", "id": "…" }]` (≤ 16, url ≤ 500 chars, malformed dropped). They
|
|
230
228
|
**never** reach the model context or the transcript — they ride the channel
|
|
231
229
|
delivery as an additive `attachments` key on the outbox payload, and the channel
|
|
232
|
-
(or its consumer) decides what a card looks like.
|
|
230
|
+
(or its consumer) decides what a card looks like. Supply an explicit `id` when
|
|
231
|
+
cards are not one-to-one with items in the same order. Without one, an attachment
|
|
232
|
+
takes its item's ID at the original position, before malformed cards are dropped.
|
|
233
|
+
- Lean `line` passes through the tool-result sanitizer only when `fencing` is on.
|
|
234
|
+
Attachment captions are normalized to UTF-8 but are not fenced. See
|
|
235
|
+
[Fencing](AGENTS.md#fencing--third-party-text-is-data-never-instructions).
|
|
233
236
|
- A **code tool** opts in the same way: it either returns `{ items, attachments }`
|
|
234
237
|
directly and declares `evidence` in its registry metadata, or exposes an
|
|
235
238
|
`evidence` reader. No declaration = today's tool behavior, byte for byte.
|
|
236
239
|
|
|
240
|
+
### Provenance: checking IDs before a write
|
|
241
|
+
|
|
242
|
+
Declare `requires_evidence` on a data-defined tool to accept only IDs previously
|
|
243
|
+
returned by an `evidence` tool in the same session:
|
|
244
|
+
|
|
245
|
+
```json
|
|
246
|
+
{ "requires_evidence": ["product_id"] }
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
The full form is `{ "requires_evidence": { "params": ["product_id"] } }`.
|
|
250
|
+
The list must be non-empty and name declared top-level parameters. Scalar values
|
|
251
|
+
and every element of an array are converted to strings and compared exactly:
|
|
252
|
+
`SKU-1` and `sku-1` are different IDs. IDs typed by a customer do not count.
|
|
253
|
+
The ledger includes earlier turns and completed evidence results in the current
|
|
254
|
+
turn, capped at the latest 1,000 distinct IDs. Every declared parameter the call
|
|
255
|
+
carries must pass before any write occurs; a parameter the schema marks optional
|
|
256
|
+
and the model leaves out has nothing to check, a required one left out blocks.
|
|
257
|
+
Search first, then write in a later batch: a search and write in the same parallel
|
|
258
|
+
batch have no dependency ordering guarantee.
|
|
259
|
+
|
|
260
|
+
An unknown ID returns `status: "blocked"`, `gate: "provenance"`, the parameter,
|
|
261
|
+
the value, and an instruction to search or look it up before retrying. The backend
|
|
262
|
+
is never called and no operator approval is requested. A missing ledger blocks the
|
|
263
|
+
call too. Omit the declaration to keep the existing behavior; MCP tools do not
|
|
264
|
+
support this declaration.
|
|
265
|
+
|
|
266
|
+
The Studio tool editor exposes `requires_evidence`. Blocked calls appear in session
|
|
267
|
+
traces and `insika tools:report`, emit `tool_blocked` with name/gate/parameter only,
|
|
268
|
+
and increment `insika.tool.blocked`. `insika doctor` warns when an agent allows a
|
|
269
|
+
gated data tool without an allowed data tool declaring `evidence`.
|
|
270
|
+
|
|
271
|
+
### Side effects in parallel batches
|
|
272
|
+
|
|
273
|
+
With `limits.tool_concurrency > 1`, tools marked `side_effect` execute one at a time
|
|
274
|
+
within the session's runtime. Unmarked tools still run concurrently, including
|
|
275
|
+
while a write is running. MCP tools are marked `side_effect: true` unless the server
|
|
276
|
+
annotates them `readOnlyHint`; a read-only `POST` data tool needs an explicit
|
|
277
|
+
`"side_effect": false` to keep its concurrency.
|
|
278
|
+
A queued write holds no concurrency slot. Different sessions remain independent;
|
|
279
|
+
backend rules such as quantity limits remain the backend's responsibility.
|
|
280
|
+
|
|
281
|
+
The per-tool timeout starts after both gates are acquired. Trace duration includes
|
|
282
|
+
queueing time, so it measures how long the model waited, not just backend execution.
|
|
283
|
+
|
|
237
284
|
### Grounding: policing claims against the ledger
|
|
238
285
|
|
|
239
286
|
With the ledger fed, the pack declares how claims are policed — data on the agent,
|
|
@@ -263,11 +310,63 @@ grounding mode: :flag, matcher: { sku: '\b[A-Z]{2,4}\d{4,8}\b' }
|
|
|
263
310
|
- Grounding is **independent of the guardrails opt-in**: an agent with guardrails
|
|
264
311
|
off and `grounding.mode: :flag` still gets the check.
|
|
265
312
|
|
|
313
|
+
## Presentation tools: the model picks ids, the engine shows the cards
|
|
314
|
+
|
|
315
|
+
A **presentation tool** selects which evidence cards to show. Declare it with
|
|
316
|
+
`presentation` instead of `request`; it runs in-process and is always
|
|
317
|
+
`side_effect: false`. The model supplies IDs, never card URLs or captions.
|
|
318
|
+
|
|
319
|
+
```jsonc
|
|
320
|
+
{ "name": "present_products",
|
|
321
|
+
"description": "Show product cards to the customer. Pass only ids a search returned.",
|
|
322
|
+
"parameters": [{ "name": "product_ids", "type": "array:string" },
|
|
323
|
+
{ "name": "title", "type": "string", "required": false }],
|
|
324
|
+
"presentation": { "component": "product_cards", // what the channel renders
|
|
325
|
+
"ids": "product_ids", // the array:string parameter
|
|
326
|
+
"max": 8 } } // 1..16 (the attachment cap; default 16)
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
Exactly one of `request` / `presentation`; `ids` must name a declared `array:string`
|
|
330
|
+
parameter; `component` follows the tool-name rule. Create it through the DSL,
|
|
331
|
+
manifest or API. The Studio form preserves an existing `presentation` declaration
|
|
332
|
+
on save but does not expose its fields. Requested IDs are deduplicated in order
|
|
333
|
+
and compared case-sensitively. The engine then:
|
|
334
|
+
|
|
335
|
+
1. keeps only ids the session's **evidence ledger** has seen — the rest are dropped
|
|
336
|
+
with reason `unknown` (an id the model invented, or the customer typed);
|
|
337
|
+
2. joins each kept id to the card an evidence tool returned **this session** — this
|
|
338
|
+
turn's first, then the last 64 the ledger kept (cards carry the `id` of the item
|
|
339
|
+
they stand for); a known id with no card is dropped with reason `no_card`;
|
|
340
|
+
3. truncates to `max` — the overflow is dropped with reason `max`;
|
|
341
|
+
4. records the selection on the turn and emits `:ui` on the stream (published as
|
|
342
|
+
`insika.ui` on `/v1/responses`, as the `ui` frame on the web channel);
|
|
343
|
+
5. answers the model with `shown` (IDs) and `dropped` (objects with `id` and `reason`) — plus one
|
|
344
|
+
instruction when nothing could be shown ("name the products in text or search
|
|
345
|
+
again").
|
|
346
|
+
|
|
347
|
+
**Delivery.** When a turn made a presentation call, the outbox `attachments` are
|
|
348
|
+
*exactly the presented cards*, in call order, each stamped with the call's `component`
|
|
349
|
+
and `title`. A turn with no presentation call delivers every hoarded card, as before —
|
|
350
|
+
a pack that declares no presentation tool sees no change. An empty selection also
|
|
351
|
+
suppresses that fallback: it emits `count: 0` and delivers no cards from that call.
|
|
352
|
+
An earlier turn's ID can be shown without a fresh lookup while its card remains
|
|
353
|
+
in the session's 64-card ledger. IDs and cards have separate caps: a known ID can
|
|
354
|
+
still return `no_card`. Cards are stored snapshots, not a stock or price refresh.
|
|
355
|
+
|
|
356
|
+
A presentation tool needs an evidence source to populate the session's cards.
|
|
357
|
+
`insika doctor` warns when an agent allows presentation without an allowed evidence
|
|
358
|
+
data tool (`presentation-tools`); it cannot verify a code-tool evidence source. The line that tells the model
|
|
359
|
+
*when* to show cards ("show cards with present_products, ids only") is the pack's.
|
|
360
|
+
|
|
361
|
+
Not here: partial rendering while arguments stream, per-component enrichment (price
|
|
362
|
+
today, stock — the card is what the evidence tool returned), and a second component
|
|
363
|
+
such as suggestion chips (same mechanism, when a channel asks for it).
|
|
364
|
+
|
|
266
365
|
## Registering a tool
|
|
267
366
|
|
|
268
367
|
A tool appears in the Studio panel and enters an agent's tool-loop when it is
|
|
269
368
|
**registered** in the catalog **and** allowed by the agent's policy allowlist.
|
|
270
|
-
|
|
369
|
+
Three ways to write a data tool into the store — all **hot** (registry and catalog
|
|
271
370
|
reload, no restart):
|
|
272
371
|
|
|
273
372
|
1. **DSL** — `data_tool(name:, …)` in a `Insika.agent { … }` block.
|
|
@@ -275,8 +374,6 @@ reload, no restart):
|
|
|
275
374
|
3. **Manifest** — `POST /v1/tools/manifest`. Partial failure is isolated: one
|
|
276
375
|
malformed tool becomes an `errors[]` entry; only a structural manifest error
|
|
277
376
|
fails the whole request. The response reports `{ version, created, updated, errors }`.
|
|
278
|
-
4. ~~MCP ingestion~~ — retired. An MCP server's tools are no
|
|
279
|
-
longer written into this store at all; see [MCP servers](#mcp-servers).
|
|
280
377
|
|
|
281
378
|
### The one gotcha: env/secret templating is manifest-only
|
|
282
379
|
|
|
@@ -298,7 +395,7 @@ at turn time, not ingestion).
|
|
|
298
395
|
An MCP **instance** is durable config — transport, target, credentials, an
|
|
299
396
|
`enabled` flag — held in its own store, separate from data tools. Once an
|
|
300
397
|
instance is enabled, its tools appear in the catalog automatically (group
|
|
301
|
-
`mcp:<instance>`, `side_effect: true`), and every call goes straight to the
|
|
398
|
+
`mcp:<instance>`, `side_effect: true` unless annotated `readOnlyHint`), and every call goes straight to the
|
|
302
399
|
server through a live, held client — the runtime never converts an MCP tool
|
|
303
400
|
into a stored data tool.
|
|
304
401
|
|
|
@@ -405,9 +502,10 @@ held client, which does its own discovery on first use regardless of whether
|
|
|
405
502
|
A model can ask for several tools in one step. By default the engine runs them one
|
|
406
503
|
at a time. Set `limits[:tool_concurrency]` above 1 (see
|
|
407
504
|
[Agents](AGENTS.md#tool_concurrency--parallel-tool-calls)) and the calls in that
|
|
408
|
-
batch run concurrently, **at most N in flight**, on the turn's own reactor
|
|
409
|
-
|
|
410
|
-
|
|
505
|
+
batch run concurrently, **at most N in flight**, on the turn's own reactor.
|
|
506
|
+
Tools marked `side_effect` still execute one at a time per session; queued writes
|
|
507
|
+
acquire that serial gate before taking a shared concurrency slot. See
|
|
508
|
+
[Side effects](#side-effects-in-parallel-batches). The cap covers every enveloped tool, including the
|
|
411
509
|
ones `tool_search` promotes mid-turn.
|
|
412
510
|
|
|
413
511
|
It applies only to what the *model* fans out. Two primitives already parallelize
|
|
@@ -434,7 +532,7 @@ Turning it on changes three things, all of them worth knowing before you do:
|
|
|
434
532
|
Approvals and concurrency are mutually exclusive per turn — the approval gate wins
|
|
435
533
|
and the turn goes serial. That is a deadlock avoided, not a preference.
|
|
436
534
|
|
|
437
|
-
## Egress: the SSRF guard
|
|
535
|
+
## Egress: the SSRF guard
|
|
438
536
|
|
|
439
537
|
Data tools make outbound HTTP, so every call passes through the **EgressGuard**, a
|
|
440
538
|
Server-Side Request Forgery defense. The default posture is **strict: public
|
|
@@ -446,17 +544,10 @@ Server-Side Request Forgery defense. The default posture is **strict: public
|
|
|
446
544
|
| `INSIKA_EGRESS_ALLOW_HTTP=1` | permit plain `http` — **loopback dev only** |
|
|
447
545
|
| `INSIKA_EGRESS_ALLOW_PRIVATE=1` | permit private/loopback IPs — **dev only** |
|
|
448
546
|
|
|
449
|
-
|
|
450
|
-
|
|
451
|
-
|
|
452
|
-
|
|
453
|
-
> plausible failure and the conversation *looks* like it worked. You will not see
|
|
454
|
-
> an exception.
|
|
455
|
-
>
|
|
456
|
-
> **Always verify by the trace, never by the reply:** open the Studio session
|
|
457
|
-
> viewer — a healthy call shows the request, args, and the backend's `200`; a
|
|
458
|
-
> missing or errored call is almost always egress (host not in the allowlist, or
|
|
459
|
-
> `http`/private without the opt-in).
|
|
547
|
+
An egress rejection returns `{ error: … }` to the model without making the HTTP
|
|
548
|
+
request. Inspect the session trace for the actual error; a plausible reply does
|
|
549
|
+
not prove the backend ran. Provenance refusals instead report `status: "blocked"`
|
|
550
|
+
and `gate: "provenance"`.
|
|
460
551
|
|
|
461
552
|
Egress is **orthogonal** to registration and allowlisting: a tool can be
|
|
462
553
|
registered, allowed, offered to the model, and still blocked at call time.
|
|
@@ -471,8 +562,9 @@ Work down this checklist:
|
|
|
471
562
|
`data-tools` check is the only place that says so.
|
|
472
563
|
2. **Allowed for this agent?** In `tools_allow` (or an allowed group), and not in
|
|
473
564
|
`tools_deny`?
|
|
474
|
-
3. **
|
|
475
|
-
|
|
565
|
+
3. **Call refused or failed?** Inspect the trace. For `gate: "provenance"`, look up
|
|
566
|
+
the ID through an evidence tool before retrying. For an error, check its message
|
|
567
|
+
for schema, egress, timeout or backend failures.
|
|
476
568
|
4. **URL literal?** For non-manifest tools, an unresolved `{{env.*}}` would have
|
|
477
569
|
422'd at import — re-check the definition.
|
|
478
570
|
|
|
@@ -485,6 +577,32 @@ and gets the URL back; the tenant is bound from the turn, never a parameter the
|
|
|
485
577
|
model types. See [Artifacts](ARTIFACTS.md) for the tool contract, the serving
|
|
486
578
|
routes, the signed link and the retention/LGPD reach.
|
|
487
579
|
|
|
580
|
+
## The usage report — `insika tools:report`
|
|
581
|
+
|
|
582
|
+
The per-session trace answers "what did this conversation call"; nothing used to
|
|
583
|
+
answer "what does this agent carry and never use". The report aggregates the
|
|
584
|
+
stored traces per agent (tasks → sessions → `tool_traces`, the same read the
|
|
585
|
+
Studio does) and flags four shapes:
|
|
586
|
+
|
|
587
|
+
- **`never_called`** — in `tools_allow`, zero calls in any stored trace. Dead
|
|
588
|
+
weight: its schema ships on every request and buys nothing.
|
|
589
|
+
- **`error_rate`** — over 30% conventional errors (the trace's `ok` flag) inside
|
|
590
|
+
the window (default 14 days). Either the tool is broken or the model cannot
|
|
591
|
+
hold its contract.
|
|
592
|
+
- **`stale`** — called at some point, but not once inside the window.
|
|
593
|
+
- **`blocked`** — gate refusals in the window, counted by gate. These do not count
|
|
594
|
+
as conventional tool errors.
|
|
595
|
+
|
|
596
|
+
```bash
|
|
597
|
+
insika tools:report # every stored agent
|
|
598
|
+
insika tools:report --agent store-support # one agent
|
|
599
|
+
insika tools:report --days 30 --json # wider window, machine-readable
|
|
600
|
+
```
|
|
601
|
+
|
|
602
|
+
Read-only by design: the report names candidates, the **operator** removes — a
|
|
603
|
+
flagged tool may still be the one a rare but critical flow needs. Counts are
|
|
604
|
+
"at least", never exact: the trace keeps a capped tail per session.
|
|
605
|
+
|
|
488
606
|
## See also
|
|
489
607
|
|
|
490
608
|
- [Agents](AGENTS.md) — allowlists, groups, and per-agent tool exposure.
|