insika 0.0.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (277) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +361 -0
  3. data/LICENSE +21 -0
  4. data/README.md +136 -2
  5. data/bin/insika +366 -0
  6. data/docs/AGENTS.md +618 -0
  7. data/docs/ARCHITECTURE.md +333 -0
  8. data/docs/BENCHMARK.md +114 -0
  9. data/docs/CHANNELS.md +453 -0
  10. data/docs/CONTEXT.md +117 -0
  11. data/docs/DEPLOY.md +354 -0
  12. data/docs/EMBEDDING.md +198 -0
  13. data/docs/EVALS.md +273 -0
  14. data/docs/LOADTEST.md +232 -0
  15. data/docs/OBSERVABILITY.md +374 -0
  16. data/docs/PLUGINS.md +211 -0
  17. data/docs/REFINEMENT.md +477 -0
  18. data/docs/RELEASING.md +70 -0
  19. data/docs/RUNNING-LOCAL.md +153 -0
  20. data/docs/SANDBOX.md +114 -0
  21. data/docs/SECURITY.md +375 -0
  22. data/docs/SKILLS.md +284 -0
  23. data/docs/TOOLS.md +302 -0
  24. data/docs/WHY.md +137 -0
  25. data/docs/WORKFLOWS.md +225 -0
  26. data/docs/build.md +14 -0
  27. data/docs/index.md +68 -0
  28. data/docs/onboarding/start.md +126 -0
  29. data/docs/operate.md +12 -0
  30. data/docs/ship.md +10 -0
  31. data/docs/understand.md +10 -0
  32. data/lib/insika/agent_file_store.rb +125 -0
  33. data/lib/insika/agent_profile.rb +255 -0
  34. data/lib/insika/alert_dispatcher.rb +139 -0
  35. data/lib/insika/allowlist.rb +28 -0
  36. data/lib/insika/baseline_store.rb +74 -0
  37. data/lib/insika/budget_ledger.rb +135 -0
  38. data/lib/insika/capability/resolved_tool.rb +34 -0
  39. data/lib/insika/capability_registry.rb +112 -0
  40. data/lib/insika/channel_delivery.rb +153 -0
  41. data/lib/insika/channel_registry.rb +30 -0
  42. data/lib/insika/channels/relay.rb +178 -0
  43. data/lib/insika/channels/web/widget.js +283 -0
  44. data/lib/insika/channels/web.rb +211 -0
  45. data/lib/insika/channels/webhook.rb +58 -0
  46. data/lib/insika/chat_builder.rb +303 -0
  47. data/lib/insika/checkpoint.rb +13 -0
  48. data/lib/insika/checkpoint_store.rb +153 -0
  49. data/lib/insika/circuit_state.rb +114 -0
  50. data/lib/insika/coercion.rb +58 -0
  51. data/lib/insika/command.rb +32 -0
  52. data/lib/insika/command_bus.rb +39 -0
  53. data/lib/insika/commands/agent_payload.rb +43 -0
  54. data/lib/insika/commands/approve_action.rb +46 -0
  55. data/lib/insika/commands/cancel_task.rb +33 -0
  56. data/lib/insika/commands/create_agent.rb +54 -0
  57. data/lib/insika/commands/create_session.rb +67 -0
  58. data/lib/insika/commands/delete_agent.rb +33 -0
  59. data/lib/insika/commands/delete_agent_file.rb +50 -0
  60. data/lib/insika/commands/delete_data_tool.rb +33 -0
  61. data/lib/insika/commands/delete_llm_provider.rb +36 -0
  62. data/lib/insika/commands/delete_mcp.rb +30 -0
  63. data/lib/insika/commands/delete_skill.rb +43 -0
  64. data/lib/insika/commands/delete_system_file.rb +29 -0
  65. data/lib/insika/commands/gate_refinement.rb +245 -0
  66. data/lib/insika/commands/import_mcp_tools.rb +48 -0
  67. data/lib/insika/commands/import_tools.rb +81 -0
  68. data/lib/insika/commands/issue_tenant_token.rb +41 -0
  69. data/lib/insika/commands/memory_add_note.rb +32 -0
  70. data/lib/insika/commands/memory_forget_fact.rb +32 -0
  71. data/lib/insika/commands/memory_put_fact.rb +35 -0
  72. data/lib/insika/commands/pause_task.rb +29 -0
  73. data/lib/insika/commands/resolve_refinement.rb +126 -0
  74. data/lib/insika/commands/restore_agent_file.rb +36 -0
  75. data/lib/insika/commands/restore_data_tool.rb +34 -0
  76. data/lib/insika/commands/restore_system_file.rb +31 -0
  77. data/lib/insika/commands/resume_task.rb +85 -0
  78. data/lib/insika/commands/revoke_token.rb +39 -0
  79. data/lib/insika/commands/rotate_tenant_token.rb +43 -0
  80. data/lib/insika/commands/run_refinement.rb +133 -0
  81. data/lib/insika/commands/send_message.rb +150 -0
  82. data/lib/insika/commands/set_agent_tools.rb +39 -0
  83. data/lib/insika/commands/set_skill_agents.rb +112 -0
  84. data/lib/insika/commands/trigger_workflow.rb +80 -0
  85. data/lib/insika/commands/update_agent.rb +49 -0
  86. data/lib/insika/commands/update_settings.rb +33 -0
  87. data/lib/insika/commands/upsert_llm_provider.rb +34 -0
  88. data/lib/insika/commands/upsert_mcp.rb +32 -0
  89. data/lib/insika/commands/write_agent_file.rb +57 -0
  90. data/lib/insika/commands/write_data_tool.rb +43 -0
  91. data/lib/insika/commands/write_golden.rb +58 -0
  92. data/lib/insika/commands/write_skill.rb +60 -0
  93. data/lib/insika/commands/write_system_file.rb +31 -0
  94. data/lib/insika/config_store.rb +89 -0
  95. data/lib/insika/context/builder.rb +166 -0
  96. data/lib/insika/context/catalog_provider.rb +23 -0
  97. data/lib/insika/context/fragment.rb +43 -0
  98. data/lib/insika/context/priority.rb +30 -0
  99. data/lib/insika/context/provider.rb +19 -0
  100. data/lib/insika/context/providers/memory.rb +60 -0
  101. data/lib/insika/context/providers/prompt.rb +105 -0
  102. data/lib/insika/context/providers/request.rb +32 -0
  103. data/lib/insika/context/providers/session.rb +123 -0
  104. data/lib/insika/context/providers/skill.rb +24 -0
  105. data/lib/insika/context/providers/skill_trigger.rb +128 -0
  106. data/lib/insika/context/providers/tool_search.rb +20 -0
  107. data/lib/insika/context_trace_store.rb +92 -0
  108. data/lib/insika/delegation_store.rb +153 -0
  109. data/lib/insika/doctor.rb +539 -0
  110. data/lib/insika/dsl/definition.rb +55 -0
  111. data/lib/insika/dsl/runtime.rb +382 -0
  112. data/lib/insika/dsl/server_boot.rb +98 -0
  113. data/lib/insika/dsl/system.rb +93 -0
  114. data/lib/insika/dsl/workflow_adapter.rb +59 -0
  115. data/lib/insika/dsl.rb +364 -0
  116. data/lib/insika/edge_limiter.rb +268 -0
  117. data/lib/insika/egress_guard.rb +75 -0
  118. data/lib/insika/env_schema.rb +249 -0
  119. data/lib/insika/errors.rb +201 -0
  120. data/lib/insika/evals/assertions.rb +247 -0
  121. data/lib/insika/evals/baseline.rb +69 -0
  122. data/lib/insika/evals/golden.rb +172 -0
  123. data/lib/insika/evals/judge.rb +225 -0
  124. data/lib/insika/evals/pairwise.rb +178 -0
  125. data/lib/insika/evals/report.rb +115 -0
  126. data/lib/insika/evals/runner.rb +141 -0
  127. data/lib/insika/evals/transport.rb +178 -0
  128. data/lib/insika/event.rb +18 -0
  129. data/lib/insika/event_stream.rb +132 -0
  130. data/lib/insika/executor.rb +1995 -0
  131. data/lib/insika/frontmatter.rb +42 -0
  132. data/lib/insika/golden_store.rb +145 -0
  133. data/lib/insika/hooks.rb +48 -0
  134. data/lib/insika/http_client.rb +63 -0
  135. data/lib/insika/inbound_log.rb +84 -0
  136. data/lib/insika/llm_configurator.rb +99 -0
  137. data/lib/insika/llm_provider_store.rb +83 -0
  138. data/lib/insika/loop_detector.rb +143 -0
  139. data/lib/insika/mcp_http_client.rb +67 -0
  140. data/lib/insika/mcp_store.rb +115 -0
  141. data/lib/insika/mcp_tool_ingestor.rb +143 -0
  142. data/lib/insika/memory_store.rb +93 -0
  143. data/lib/insika/message_origin.rb +76 -0
  144. data/lib/insika/middleware.rb +36 -0
  145. data/lib/insika/model_policy.rb +52 -0
  146. data/lib/insika/model_resolver.rb +176 -0
  147. data/lib/insika/model_selection.rb +115 -0
  148. data/lib/insika/onboarding.rb +208 -0
  149. data/lib/insika/outbox_store.rb +166 -0
  150. data/lib/insika/overlay_tool_registry.rb +102 -0
  151. data/lib/insika/pack.rb +102 -0
  152. data/lib/insika/pack_importer.rb +123 -0
  153. data/lib/insika/pending_action_store.rb +120 -0
  154. data/lib/insika/plugin/loader.rb +356 -0
  155. data/lib/insika/plugin.rb +35 -0
  156. data/lib/insika/policy/engine.rb +83 -0
  157. data/lib/insika/policy/policy.rb +120 -0
  158. data/lib/insika/policy_registry.rb +23 -0
  159. data/lib/insika/profile_source.rb +143 -0
  160. data/lib/insika/prompt_catalog.rb +61 -0
  161. data/lib/insika/provider_error_classifier.rb +160 -0
  162. data/lib/insika/queue_policy.rb +167 -0
  163. data/lib/insika/recovery.rb +168 -0
  164. data/lib/insika/refinement/candidate.rb +159 -0
  165. data/lib/insika/refinement/evidence_collector.rb +371 -0
  166. data/lib/insika/refinement/gate.rb +234 -0
  167. data/lib/insika/refinement/panel.rb +222 -0
  168. data/lib/insika/refinement/proposer.rb +262 -0
  169. data/lib/insika/refinement_store.rb +295 -0
  170. data/lib/insika/registry.rb +59 -0
  171. data/lib/insika/reliability.rb +185 -0
  172. data/lib/insika/safety/config.rb +109 -0
  173. data/lib/insika/safety/detectors.rb +176 -0
  174. data/lib/insika/safety/factory.rb +102 -0
  175. data/lib/insika/safety/input_guardrail.rb +102 -0
  176. data/lib/insika/safety/moderator.rb +94 -0
  177. data/lib/insika/safety/output_filter.rb +79 -0
  178. data/lib/insika/safety/output_validator.rb +101 -0
  179. data/lib/insika/safety/safe_responses.rb +47 -0
  180. data/lib/insika/sandbox/boundary.rb +93 -0
  181. data/lib/insika/sandbox/docker.rb +74 -0
  182. data/lib/insika/sandbox/local.rb +33 -0
  183. data/lib/insika/sandbox/runner.rb +80 -0
  184. data/lib/insika/sandbox.rb +85 -0
  185. data/lib/insika/schema_guard.rb +147 -0
  186. data/lib/insika/secret_masking.rb +34 -0
  187. data/lib/insika/server/a2a/agent_card.rb +27 -0
  188. data/lib/insika/server/a2a/app.rb +112 -0
  189. data/lib/insika/server/a2a/client.rb +101 -0
  190. data/lib/insika/server/a2a/errors.rb +32 -0
  191. data/lib/insika/server/a2a/http.rb +42 -0
  192. data/lib/insika/server/a2a/message.rb +27 -0
  193. data/lib/insika/server/a2a/protocol.rb +45 -0
  194. data/lib/insika/server/a2a/remotes.rb +25 -0
  195. data/lib/insika/server/a2a/task_projection.rb +40 -0
  196. data/lib/insika/server/app.rb +1022 -0
  197. data/lib/insika/server/boot.rb +119 -0
  198. data/lib/insika/server/rack_app.rb +118 -0
  199. data/lib/insika/server/responses.rb +165 -0
  200. data/lib/insika/server/sse_body.rb +96 -0
  201. data/lib/insika/server/tenant_auth.rb +61 -0
  202. data/lib/insika/session_actor.rb +162 -0
  203. data/lib/insika/session_store.rb +143 -0
  204. data/lib/insika/settings_store.rb +154 -0
  205. data/lib/insika/shutdown.rb +125 -0
  206. data/lib/insika/skill_catalog.rb +220 -0
  207. data/lib/insika/skill_store.rb +127 -0
  208. data/lib/insika/steer_injector.rb +110 -0
  209. data/lib/insika/store.rb +52 -0
  210. data/lib/insika/stores/memory.rb +123 -0
  211. data/lib/insika/stores/sqlite.rb +183 -0
  212. data/lib/insika/studio/app.rb +1693 -0
  213. data/lib/insika/studio/assets/dist/application.css +1 -0
  214. data/lib/insika/studio/assets/dist/application.js +70 -0
  215. data/lib/insika/studio/forms.rb +335 -0
  216. data/lib/insika/studio/nav_icons.rb +31 -0
  217. data/lib/insika/studio/views/_message.erb +44 -0
  218. data/lib/insika/studio/views/agent_detail.erb +285 -0
  219. data/lib/insika/studio/views/agents.erb +63 -0
  220. data/lib/insika/studio/views/approvals.erb +41 -0
  221. data/lib/insika/studio/views/chats.erb +34 -0
  222. data/lib/insika/studio/views/evals.erb +83 -0
  223. data/lib/insika/studio/views/home.erb +72 -0
  224. data/lib/insika/studio/views/layout.erb +94 -0
  225. data/lib/insika/studio/views/login.erb +17 -0
  226. data/lib/insika/studio/views/mcp.erb +91 -0
  227. data/lib/insika/studio/views/not_found.erb +5 -0
  228. data/lib/insika/studio/views/playground.erb +47 -0
  229. data/lib/insika/studio/views/refinement.erb +234 -0
  230. data/lib/insika/studio/views/session.erb +137 -0
  231. data/lib/insika/studio/views/settings.erb +168 -0
  232. data/lib/insika/studio/views/skills.erb +141 -0
  233. data/lib/insika/studio/views/system_files.erb +65 -0
  234. data/lib/insika/studio/views/task.erb +105 -0
  235. data/lib/insika/studio/views/tasks.erb +33 -0
  236. data/lib/insika/studio/views/tool_edit.erb +107 -0
  237. data/lib/insika/studio/views/tools.erb +89 -0
  238. data/lib/insika/subagent_graph.rb +96 -0
  239. data/lib/insika/system_file_store.rb +96 -0
  240. data/lib/insika/task_actor.rb +128 -0
  241. data/lib/insika/task_store.rb +250 -0
  242. data/lib/insika/telemetry/pricing.rb +104 -0
  243. data/lib/insika/telemetry/recorder.rb +228 -0
  244. data/lib/insika/telemetry.rb +127 -0
  245. data/lib/insika/testing/store_contract.rb +270 -0
  246. data/lib/insika/tick.rb +122 -0
  247. data/lib/insika/token_estimator.rb +16 -0
  248. data/lib/insika/token_store.rb +168 -0
  249. data/lib/insika/tool_assembly.rb +140 -0
  250. data/lib/insika/tool_catalog.rb +89 -0
  251. data/lib/insika/tool_definition.rb +518 -0
  252. data/lib/insika/tool_envelope.rb +140 -0
  253. data/lib/insika/tool_manifest.rb +218 -0
  254. data/lib/insika/tool_output_compressor.rb +100 -0
  255. data/lib/insika/tool_registry.rb +21 -0
  256. data/lib/insika/tool_store.rb +135 -0
  257. data/lib/insika/tool_trace_store.rb +92 -0
  258. data/lib/insika/tools/a2a_remote.rb +48 -0
  259. data/lib/insika/tools/agent_enum.rb +68 -0
  260. data/lib/insika/tools/concurrency.rb +54 -0
  261. data/lib/insika/tools/data_defined_tool.rb +219 -0
  262. data/lib/insika/tools/load_skill.rb +99 -0
  263. data/lib/insika/tools/remember.rb +53 -0
  264. data/lib/insika/tools/stuck_signal.rb +44 -0
  265. data/lib/insika/tools/subagent.rb +75 -0
  266. data/lib/insika/tools/subagents.rb +77 -0
  267. data/lib/insika/tools/tool_search.rb +94 -0
  268. data/lib/insika/turn_output.rb +139 -0
  269. data/lib/insika/turn_state.rb +162 -0
  270. data/lib/insika/turn_timing.rb +56 -0
  271. data/lib/insika/usage_ledger.rb +47 -0
  272. data/lib/insika/version.rb +3 -1
  273. data/lib/insika/wiring/graph.rb +249 -0
  274. data/lib/insika/workflow.rb +185 -0
  275. data/lib/insika/workflow_registry.rb +33 -0
  276. data/lib/insika.rb +220 -4
  277. metadata +412 -8
data/docs/SKILLS.md ADDED
@@ -0,0 +1,284 @@
1
+ ---
2
+ title: Skills
3
+ parent: Build an agent
4
+ nav_order: 3
5
+ permalink: /skills/
6
+ ---
7
+
8
+ # Skills
9
+
10
+ A **skill** is a named playbook an agent loads **on demand**. It is a directory
11
+ with a `SKILL.md` file: YAML frontmatter (`name` + `description`) followed by a
12
+ Markdown body. The agent always sees the *name and description* of each skill it
13
+ is allowed; it pulls the full *body* into context only when a turn actually calls
14
+ for it. That is **progressive loading** — an agent can "know" twenty skills exist
15
+ while paying for the text of only the ones it opens.
16
+
17
+ See [`examples/skills/`](https://github.com/guizaols/insika/tree/main/examples/skills/) for a runnable one.
18
+
19
+ ## Format
20
+
21
+ ```markdown
22
+ ---
23
+ name: refunds # must equal the directory name
24
+ description: When and how to process a refund # the Level-1 trigger text
25
+ triggers: [refund, money back] # optional: deterministic activation (below)
26
+ companions: [refund-policy] # optional: skills this one cannot work without
27
+ ---
28
+
29
+ <the full playbook body — loaded only on demand>
30
+ ```
31
+
32
+ Frontmatter is parsed tolerantly. A skill's canonical name is its directory name,
33
+ which the `name:` field must match.
34
+
35
+ ## Progressive loading: two levels
36
+
37
+ - **Level 1 — metadata only.** A context provider injects an `<available_skills>`
38
+ list into the system prompt — the name, the one-line description and the
39
+ `triggers:` of every allowed skill — telling the model to load a skill before
40
+ acting on it. Cheap, and always present for allowed skills. **This is the routing
41
+ table, and it is generated:** it cannot disagree with the allowlist, so do not
42
+ hand-write one in a prompt file (see [Drift guards](#drift-guards)).
43
+ - **Level 2 — the full body.** A built-in `load_skill` tool returns the skill body
44
+ on demand. It enforces the agent's skill allowlist and is wired **automatically**
45
+ whenever the agent has any allowed skills — you do not add it to `tools_allow`.
46
+ - **Deterministic activation — `triggers:`.** When the user message contains one
47
+ of the skill's `triggers`, the body is injected for that turn — no model
48
+ decision, no `load_skill` call. Only matched skills, only that turn. Use it for
49
+ skills that MUST fire on known phrases; model loading stays as the fallback for
50
+ everything else. Matching is on **whole words**, case-insensitive and
51
+ accent-folded: `presente` fires on *um presente* and *presénte*, never inside
52
+ *apresente*.
53
+
54
+ The Level-1 list is budgeted like any other context fragment
55
+ (see [Context](CONTEXT.md)); the Level-2 body only costs tokens on the turns that
56
+ open it.
57
+
58
+ ### Only trigger a skill that can finish the turn alone
59
+
60
+ A `triggers:` match is not a hint — the body lands in the prompt with the
61
+ authority of an instruction. So put triggers only on a skill that is
62
+ **self-sufficient** for the turn it fires on.
63
+
64
+ The failure mode is counter-intuitive: injecting a skill that is only *part* of
65
+ the answer is **worse than injecting nothing**. Give the model a reference table
66
+ whose procedure lives in a companion skill, and it now holds a plausible
67
+ half-recipe — so it never calls `load_skill` for the other half, and improvises
68
+ the missing part. A precise trigger on the wrong kind of skill still breaks the
69
+ turn.
70
+
71
+ Reference tables, vocabularies and lookup maps are the skills to leave on
72
+ level 1. Whole procedures ("run this journey", "recover from this error") are the
73
+ ones worth triggering.
74
+
75
+ ## Always-on skills: `skills_eager`
76
+
77
+ A skill that every turn needs — output format, the marker vocabulary, how to
78
+ recover from a failed tool — should not depend on the model choosing to load it.
79
+ Name it on the **agent** and its body is in the prompt on every turn:
80
+
81
+ ```ruby
82
+ Insika.agent("consultant") do
83
+ skills_eager "recommendation-formatting", "tool-error-recovery"
84
+ # skills_eager # or: every allowed skill (a corpus that fits the budget)
85
+ # skills_eager false # or: none — the default
86
+ end
87
+ ```
88
+
89
+ An eager skill also **leaves level 1**: it is absent from `<available_skills>` and
90
+ `load_skill` refuses to serve it. There is no level 2 left to fetch, and a catalog
91
+ pointing at a body already in the prompt only invites a call that pays for a
92
+ duplicate.
93
+
94
+ ### Why the agent decides, and not the skill
95
+
96
+ Eagerness used to be an `eager: true` key in the `SKILL.md` frontmatter. That put
97
+ the decision on the wrong object: **skills are shared.** `escalation-to-human`,
98
+ `recommendation-formatting` and `tool-error-recovery` each sit in several agents'
99
+ allowlists, and one flag on the skill forced one decision onto every agent holding
100
+ it — with no way to be always-on for the agent that needs it and discretionary for
101
+ the one that does not.
102
+
103
+ `skills_eager` is a per-agent list, so the same shared skill can be both. The
104
+ frontmatter key is **ignored** — `insika doctor` flags any skill still carrying it,
105
+ and names the agent setting that replaced it.
106
+
107
+ A name that is not in the agent's `skills` allowlist is a no-op (eagerness is
108
+ intersected with what the agent is allowed to see); `doctor` flags that too.
109
+
110
+ ### Keep the discretionary skills on the load path
111
+
112
+ Making everything eager is a trap, and the reason is not the tokens: **it costs you
113
+ the signal**. When every body is present on every turn, "which skills were active"
114
+ is always "all of them", and you can no longer tell which one the model reached for.
115
+ The `load_skill` call is the only record of that choice — it is a persisted tool
116
+ message, so it shows up in the transcript on its own.
117
+
118
+ So the split is: **eager for what the turn always needs, `load_skill` for what the
119
+ turn might need.** The second group is where you want the model's choice on the
120
+ record, because that is the group where a wrong choice is worth seeing.
121
+
122
+ The token trade is real but smaller than it looks: eager bodies sit at a fixed
123
+ position ahead of the history, so they belong to the **cacheable prefix**, and they
124
+ are still evictable under budget pressure, unlike the pinned identity. Conditional
125
+ injection is what breaks that prefix, on exactly the turns it fires.
126
+
127
+ ## Seeing which skills were active, and why
128
+
129
+ The load path is legible for free: `load_skill` is a tool, so the call is a
130
+ persisted message and shows up in the transcript on its own. The deterministic paths
131
+ are not a call, so the engine reports them itself — **with a reason per skill**:
132
+
133
+ | reason | what it means |
134
+ |---|---|
135
+ | `eager` | the agent's `skills_eager` names it, so every turn gets it |
136
+ | `trigger:<phrase>` | this message matched that `triggers:` entry — the phrase as **authored**, so you can find the line to edit |
137
+ | `pack` | a plugin's own context provider supplied the body |
138
+
139
+ Where it shows up, per turn, in the Studio session screen:
140
+
141
+ - an **activation card in the transcript thread**, placed at the top of its turn and
142
+ in the same visual language as a tool result, so a context-injected skill and a
143
+ model-loaded one read the same way;
144
+ - the **Context card**, next to that category's token count, for the after-the-fact
145
+ audit;
146
+ - the `skill_activated` **event** (`skills: [{name, reason}]`, `source: "context"`),
147
+ with full task/session correlation.
148
+
149
+ All three are computed from what actually reached the prompt **after the budget
150
+ cut**: a body the budget evicted is reported as an eviction, never as an activation.
151
+ A turn that mixes both paths is labelled `mixed`, and each line keeps its own reason.
152
+
153
+ ## Where skills live: the store, over a disk seed
154
+
155
+ - Skills live as **rows in SQLite** — one row per skill, holding the entire
156
+ `SKILL.md`, versioned (recent revisions are retained).
157
+ - On-disk `SKILL.md` files in configured roots are loaded as a **seed**, then the
158
+ store is **overlaid on top — the store wins**. A reload swaps the index
159
+ atomically, so edits take effect **without a restart**.
160
+
161
+ > ⚠️ **Committing a `.md` file to the repo does not make a skill show up on a
162
+ > running deployment.** The on-disk file is only a seed for a *fresh* box; a live
163
+ > box serves the store, and a deploy does not rewrite the database. Editing is a
164
+ > runtime operation (Studio / API / DSL), not a commit. See
165
+ > [Context](CONTEXT.md#the-volume).
166
+
167
+ ## Pairs that must not break: `companions:`
168
+
169
+ Injecting *part* of an answer is worse than injecting nothing. Give the model a line
170
+ map whose query-construction rules live in another skill and it holds a plausible
171
+ half-recipe — so it never calls `load_skill` for the other half, and improvises the
172
+ missing part. Measured on a real pack: the map arrived by trigger, the rules did not,
173
+ and the searches came out malformed. Twice.
174
+
175
+ Declare the dependency and it travels with whatever brought it — a trigger match, the
176
+ agent's eager set, or a `load_skill` call (which returns both bodies in the one call):
177
+
178
+ ```yaml
179
+ companions: [query-construction]
180
+ ```
181
+
182
+ Two deliberate limits:
183
+
184
+ - **One level, no transitive walk.** A cycle would be a hang and a chain a budget
185
+ blowout, and "cannot work without" is a direct relationship.
186
+ - **Never widens an allowlist.** A companion the agent is not allowed to load is
187
+ simply absent; `insika doctor` flags the declaration instead.
188
+
189
+ ## Specializing a shared skill for one agent
190
+
191
+ Skills are shared on purpose: `escalation-to-human` belongs in several agents'
192
+ allowlists. But sometimes one agent needs a different version of the same skill —
193
+ its own return policy, its own store name — and forking it under a second name
194
+ throws the sharing away and leaves two things to keep in step.
195
+
196
+ So the store has a second scope, and resolution is a **precedence chain** with one
197
+ more dimension:
198
+
199
+ ```
200
+ for agent A, skill <name>: (A, <name>) in the agent scope,
201
+ then <name> in the shared scope
202
+ ```
203
+
204
+ Three cases fall out of that one rule:
205
+
206
+ | case | what exists in the store |
207
+ |---|---|
208
+ | **shared** | only the shared record — every agent gets the same body |
209
+ | **override** | both — the agent's wins, for that agent only |
210
+ | **agent-private** | only the agent record — invisible elsewhere, and the name may collide freely |
211
+
212
+ The **name never changes.** An override keeps saying `name: escalation-to-human`
213
+ inside, because it *is* that skill, specialized; the allowlist, the `<available_skills>`
214
+ list, `load_skill` and the activation card all keep showing the bare name. What
215
+ decides which record you get is its **position in the store**, never the frontmatter —
216
+ otherwise an override would clobber the shared skill for everybody.
217
+
218
+ Write one with `agent:`, and remove it the same way (which un-specializes, leaving
219
+ the shared skill in place):
220
+
221
+ ```ruby
222
+ dispatch(:write_skill, { name: "escalation-to-human", agent: "store-cacau", content: md })
223
+ dispatch(:delete_skill, { name: "escalation-to-human", agent: "store-cacau" })
224
+ ```
225
+
226
+ In the Studio: **Skills → specialize for this agent**, which seeds the override from
227
+ the shared body.
228
+
229
+ ## Making a new skill "show up"
230
+
231
+ For an agent to actually use a skill, **both** conditions must hold:
232
+
233
+ 1. The skill exists as a row in the store (a `SKILL.md` written into it).
234
+ 2. The skill is in that agent's `skills` allowlist
235
+ (`nil` = all, `[]` = none, `[names]` = those — see
236
+ [Agents](AGENTS.md#the-allowlist-convention)).
237
+
238
+ Miss either and the skill is invisible: not in the store → nothing to load; not
239
+ in the allowlist → the model never sees it in `<available_skills>`.
240
+
241
+ Two ways to satisfy both:
242
+
243
+ - **Via a definition/pack import.** The import writes each skill directory into
244
+ the store and sets the agent's `skills` allowlist **authoritatively** from the
245
+ skills present — so a re-import that drops a skill also removes it. Keep the
246
+ definition complete.
247
+ - **Directly (Studio / API / DSL).** Write the skill (upserts the row and reloads
248
+ the catalog atomically — live immediately), then attach it to the agent(s) by
249
+ adding its name to the `skills` allowlist.
250
+
251
+ ## Verify it showed up
252
+
253
+ - In the Studio agent's Skills section, the skill is listed and allowed.
254
+ - In a turn, the skill appears in the `<available_skills>` list and the model can
255
+ `load_skill` it (the body loads on demand).
256
+ - If the model never mentions it → check the allowlist (condition 2). If
257
+ `load_skill` errors → the store row is missing or misnamed (condition 1); the
258
+ `name:` frontmatter must equal the directory name.
259
+
260
+ ## Drift guards
261
+
262
+ A skill catalog drifts against the prose that routes to it, and every way it happened
263
+ on the pilot was silent — found by reading a customer conversation days later. So the
264
+ routing table is **generated** (above), and `insika doctor` reports the residue the
265
+ generator cannot remove. Every check takes mechanical inputs only — names, allowlists,
266
+ agent identities — because one false positive is enough for an operator to stop
267
+ reading the doctor:
268
+
269
+ | finding | what it means |
270
+ |---|---|
271
+ | a prompt file names a skill outside that agent's allowlist | leftover hand-written routing: the model is told to use something it cannot load |
272
+ | a shared skill's body names one of its own holders | specialized text in shared clothing — the other holders are served that store's policy as their own. Specialize it instead |
273
+ | a body references another catalog skill without declaring it a companion | the pair can still arrive apart |
274
+ | a declared companion is outside an agent's allowlist | the pair cannot travel for that agent, and the engine will not widen the allowlist |
275
+ | a skill still declares `eager:` in its frontmatter | the key is ignored; the decision moved to the agent |
276
+ | an agent marks a skill eager that it does not allow | the name is a no-op |
277
+
278
+ ## See also
279
+
280
+ - [Context](CONTEXT.md) — how the skills list is budgeted into a turn.
281
+ - [Agents](AGENTS.md) — the skills allowlist.
282
+ - [Tools](TOOLS.md) — `load_skill` and deferred-tool progressive disclosure.
283
+ - [Plugins](PLUGINS.md) — shipping skills inside a plugin, and the two extension tiers.
284
+ - [`examples/skills/`](https://github.com/guizaols/insika/tree/main/examples/skills/) — progressive loading, runnable.
data/docs/TOOLS.md ADDED
@@ -0,0 +1,302 @@
1
+ ---
2
+ title: Tools
3
+ parent: Build an agent
4
+ nav_order: 2
5
+ permalink: /tools/
6
+ ---
7
+
8
+ # Tools
9
+
10
+ A **tool** is a function the model can call inside a turn. Insika has three
11
+ kinds, and the distinction that matters is **who can change one at runtime**:
12
+
13
+ | | **Code tool** | **Data tool** | **MCP tool** |
14
+ |---|---|---|---|
15
+ | What | a Ruby class (`< RubyLLM::Tool`) | an HTTP call described by config, no Ruby | an MCP server's tool, ingested |
16
+ | Lives | in the deployment image | as a row in SQLite | as data-tool rows in SQLite |
17
+ | Editable at runtime | no (shipped in the image) | **yes** (DSL / API / manifest / Studio) | **yes** (re-ingest) |
18
+ | Reach for it when | logic must run in-process (file edit, shell, subagent) | calling an external HTTP API | adopting a whole MCP toolset at once |
19
+
20
+ **MCP tools are not a separate runtime type.** An MCP ingestor discovers an MCP
21
+ server's tools and turns each into an HTTP **data tool** that posts a JSON-RPC
22
+ `tools/call`, tagged with a `group` naming the source instance. (Only
23
+ HTTP-transport MCP servers are ingestible; stdio is rejected.)
24
+
25
+ Code tools **win name collisions** — you cannot register a data tool whose name
26
+ shadows a code tool.
27
+
28
+ ## Data tools: a tool is a row
29
+
30
+ A data tool is defined entirely by config — this is the operator-facing kind, and
31
+ the one you create and change without a rebuild. See
32
+ [`examples/data-tool/`](https://github.com/guizaols/insika/tree/main/examples/data-tool/) for a runnable one.
33
+
34
+ ```jsonc
35
+ {
36
+ "name": "search_products", // /\A[a-z][a-z0-9_]*\z/
37
+ "description": "Search the catalog", // required — this is what the model reads
38
+ "parameters": { /* JSON Schema, safe subset */ },
39
+ "request": {
40
+ "method": "POST", // GET | HEAD | POST | PUT | PATCH | DELETE
41
+ "url": "https://api.example.com/search",
42
+ "headers": { "X-Session": "{{ctx.chat_id}}",
43
+ "Authorization": "Bearer {{secret.api_token}}" },
44
+ "query": {}, "body": "…"
45
+ },
46
+ "response": { "extract": "json_path", "path": "$.results" },
47
+ "secret_headers": ["Authorization"],
48
+ "side_effect": true, "timeout": 30, "group": "catalog", "tags": []
49
+ }
50
+ ```
51
+
52
+ ### Parameters: the schema is the contract
53
+
54
+ `parameters` is **JSON Schema**, and it reaches the provider verbatim — it is the only
55
+ thing telling the model what shape to send. The engine never fills a gap in it.
56
+
57
+ For simple params there is a flat sugar (what the Studio's textarea and a hand-written
58
+ manifest accept), one line per param:
59
+
60
+ ```
61
+ cep | string | required | The ZIP code to look up
62
+ tags | array:string | optional | Labels to filter by
63
+ quantity | integer | required | How many
64
+ ```
65
+
66
+ Types are `string`, `number`, `integer`, `boolean`, and `array:<scalar>` for a list.
67
+ There is **no bare `array`**: a list without an item type is an incomplete declaration,
68
+ and it is rejected instead of being guessed at. A list of **objects** — the common
69
+ `[{query, filters}]` shape — cannot be written in the flat form at all; write the JSON
70
+ Schema, which is what the Studio field reads when the text starts with `{`:
71
+
72
+ ```jsonc
73
+ { "type": "object",
74
+ "properties": {
75
+ "query_filter_pairs": {
76
+ "type": "array",
77
+ "items": { "type": "object",
78
+ "properties": { "query": { "type": "string" },
79
+ "filters": { "type": "object", "properties": {} } },
80
+ "required": ["query"] } } },
81
+ "required": ["query_filter_pairs"] }
82
+ ```
83
+
84
+ **Arguments are checked against the schema at call time.** A call the schema does not
85
+ allow never becomes a request: it returns an `{ error: … }` naming the path
86
+ (`query_filter_pairs[0]: expected an object, got a string`), which the model reads and
87
+ retries against. Structure is strict; a scalar may arrive in its lossless string form
88
+ (`"2"`, `"true"`) and is never coerced — what the model sent is what the request carries.
89
+
90
+ **Placeholders** are resolved at turn time:
91
+
92
+ - `{{param}}` — a declared top-level parameter, filled from the model's call.
93
+ - `{{ctx.*}}` — turn context set **server-side, never by the model**: a closed set
94
+ of `chat_id`, `store_id`, `agent_id`, `tenant`. This is how a tool knows *which*
95
+ session/agent it is acting for without trusting the model.
96
+ - `{{secret.*}}` — allowed **only** inside a header named in `secret_headers`.
97
+ A secret placeholder anywhere else is rejected (it would leak unmasked). The
98
+ real secret value is injected at provision time and never lives on disk.
99
+
100
+ **Validation** happens on ingestion. Common rejections:
101
+
102
+ - `url` must be `http`/`https` — anything else is a 422.
103
+ - `parameters` is a **safe subset** of JSON Schema
104
+ (`object/array/string/number/integer/boolean`); `oneOf`/`anyOf`/`allOf`/`$ref`/
105
+ `if`/`then`/`else` are forbidden (not every provider supports them).
106
+ - `side_effect` defaults from the method (GET/HEAD → false, else true) and drives
107
+ checkpoint/replay semantics (a completed side-effecting tool is not re-run on
108
+ resume — see [Architecture](ARCHITECTURE.md#durability-checkpoints-and-resume)).
109
+
110
+ ### `halt_when`: when the answer is already out
111
+
112
+ Some tools do the work **and** deliver the news. A backend that subscribes a customer
113
+ and sends its own confirmation over the channel has already said everything there is to
114
+ say: if the model then writes "all set, you're subscribed!", the person gets the message
115
+ twice. The usual patch is to ask the model to stay quiet in the tool's instructions —
116
+ which works until the turn it doesn't, and the failure lands in front of a customer.
117
+
118
+ `halt_when` moves the decision from the prompt to the engine. It reads the tool's own
119
+ **response**, and when it matches, the turn ends right there — no further provider call:
120
+
121
+ ```jsonc
122
+ { "name": "subscribe_to_learning_path",
123
+ "request": { "method": "POST", "url": "https://app.example/subscribe" },
124
+ "halt_when": { "json_path": "tool_result.status", "equals": ["SUBSCRIBED"] } }
125
+ ```
126
+
127
+ By **result**, not by tool. The same call that goes silent on `SUBSCRIBED` must let the
128
+ model explain a `SUBSCRIPTION_FAILED` ("you are already enrolled") — one tool, two
129
+ endings, decided by what the backend actually returned.
130
+
131
+ - `json_path` is a dotted path into the parsed response body, and `equals` a list of
132
+ values compared **as strings** (a status is a label; JSON types vary by backend).
133
+ - It reads the **body**, independently of `response.extract` — which shapes what the
134
+ *model* sees, not what the engine decides on.
135
+ - It only fires on a **2xx**. An error response that happens to carry the value is a
136
+ failure, and a failure must reach the model.
137
+ - A non-JSON body or a missing path simply does not match: a turn never ends on a guess.
138
+
139
+ A halted turn keeps whatever the model had already streamed *before* the call (usually a
140
+ "let me get that for you") and adds nothing after it.
141
+
142
+ #### `say`: what the customer gets when the model wrote nothing first
143
+
144
+ The model does not always introduce the call. Then the lead-in is empty, and the turn
145
+ used to publish **nothing** — measured on a real store, two escalation turns in a row
146
+ delivered silence to the customer. `say` is the answer for that turn, and only that
147
+ turn: when there **is** a lead-in it still wins, because two messages for one
148
+ escalation is what `halt_when` exists to prevent.
149
+
150
+ It cannot be inferred. `json_path` + `equals` cannot supply it either — the matched
151
+ value is by definition one of the `equals` tokens, so publishing it would ship
152
+ `SUBSCRIBED` to a person as often as it ships a sentence. So you name it, in one of two
153
+ shapes:
154
+
155
+ ```jsonc
156
+ // the sentence the backend itself returned
157
+ "halt_when": { "json_path": "tool_result.status", "equals": ["SUBSCRIBED"],
158
+ "say": { "json_path": "tool_result.message" } }
159
+
160
+ // a literal the CHANNEL knows how to resolve
161
+ "halt_when": { "json_path": "tool_result", "equals": ["…"],
162
+ "say": { "text": "CALL_SUPPORT" } }
163
+ ```
164
+
165
+ The literal form replaces the usual workaround: instructing the model to emit a control
166
+ token and parsing it downstream. The token now comes from the **tool's contract**,
167
+ deterministically, instead of depending on the model complying with a sentence in a
168
+ prompt.
169
+
170
+ - Exactly one of `text` or `json_path` — two answers to "what does the customer get" is
171
+ a configuration nobody can read, so both (or neither) is refused at load.
172
+ - A `json_path` that does not resolve to a **string** publishes nothing: a hash or a
173
+ number reaching a customer as the answer is never what someone meant.
174
+ - Omit `say` and the behaviour is unchanged — a halt with no lead-in completes empty,
175
+ which is what a channel consumer drops.
176
+
177
+ `say` is declared on the **tool**, because what a backend answers is a property of that
178
+ backend, not of whoever calls it. Every agent sharing the tool gets the same value.
179
+
180
+ > The Studio's tool editor does not render this field (nor `group`/`tags`), but a save
181
+ > there **preserves** it — the form carries the stored values through instead of
182
+ > replacing the record with only what it shows.
183
+
184
+ ## Registering a tool
185
+
186
+ A tool appears in the Studio panel and enters an agent's tool-loop when it is
187
+ **registered** in the catalog **and** allowed by the agent's policy allowlist.
188
+ Four ways to write a data tool into the store — all **hot** (registry and catalog
189
+ reload, no restart):
190
+
191
+ 1. **DSL** — `data_tool(name:, …)` in a `Insika.agent { … }` block.
192
+ 2. **Studio** — the Tools panel editor.
193
+ 3. **Manifest** — `POST /v1/tools/manifest`. Partial failure is isolated: one
194
+ malformed tool becomes an `errors[]` entry; only a structural manifest error
195
+ fails the whole request. The response reports `{ version, created, updated, errors }`.
196
+ 4. **MCP ingestion** — import a server; each of its tools becomes a data tool.
197
+
198
+ ### The one gotcha: env templating is manifest-only
199
+
200
+ `{{env.*}}` (and `{{secret.*}}`) are substituted **at ingestion, on the manifest
201
+ path**. Other write paths do **not** resolve `{{env.*}}` — a literal
202
+ `{{env.API_URL}}` there fails the `http`/`https` URL check and 422s. Rule:
203
+ **manifest tools may template the URL with `{{env.*}}`; tools written any other
204
+ way must ship a literal URL.** `{{ctx.*}}` and `{{param}}` work everywhere (they
205
+ resolve at turn time, not ingestion).
206
+
207
+ ## Making it appear — and enter the tool-loop
208
+
209
+ 1. **Panel visibility** = registered in the catalog. Data tools are marked
210
+ editable; code tools are allow/deny only.
211
+ 2. **Per-agent exposure** is set from the same panel, or by the agent's allowlist.
212
+ 3. **Entering the tool-loop** is decided by the **policy allowlist**, not by tool
213
+ type: deny wins, otherwise the agent sees `tools_allow ∪ tools_allow_groups`
214
+ (or all, when both are absent). See [Agents](AGENTS.md#the-allowlist-convention).
215
+ 4. **Deferred tools** (`tools_deferred`) are *not* offered directly — they appear
216
+ as a short "available tools" list and the model must call `tool_search` to
217
+ enable one. This is progressive disclosure for large toolsets — see
218
+ [Context](CONTEXT.md).
219
+
220
+ ## Parallel tool calls
221
+
222
+ A model can ask for several tools in one step. By default the engine runs them one
223
+ at a time. Set `limits[:tool_concurrency]` above 1 (see
224
+ [Agents](AGENTS.md#tool_concurrency--parallel-tool-calls)) and the calls in that
225
+ batch run concurrently, **at most N in flight**, on the turn's own reactor — so
226
+ the wall-clock of a batch of slow data tools approaches the slowest call rather
227
+ than their sum. The cap covers every enveloped tool of the turn, including the
228
+ ones `tool_search` promotes mid-turn.
229
+
230
+ It applies only to what the *model* fans out. Two primitives already parallelize
231
+ deterministically and are unaffected: `spawn_subagents` (capped at 8 children) and
232
+ `Insika::Tools::Concurrency.gather` (fan-out inside one tool). System tools —
233
+ `tool_search`, `load_skill`, `remember`, `spawn_subagent` — are not enveloped and
234
+ so are not gated by the cap; they are trivial or capped on their own.
235
+
236
+ Turning it on changes three things, all of them worth knowing before you do:
237
+
238
+ - **`max_tool_calls` becomes approximate.** The limit is checked per call, but a
239
+ call that trips it does not stop its siblings — the whole batch finishes and the
240
+ turn then fails. With a cap of 4, up to 3 extra tools may have executed. The turn
241
+ still fails at the right boundary; the count is just no longer exact.
242
+ - **The transcript records results in completion order.** Providers key results by
243
+ `tool_call_id`, so the wire stays valid and persistence is faithful to what was
244
+ sent — but a replayed transcript no longer reads in call order.
245
+ - **`turn_timeout` can overrun by up to `tool_timeout`.** A turn deadline does not
246
+ cancel a tool call already in flight in a sibling fiber; it waits for it. Each
247
+ call is still bounded by its own `tool_timeout`, which is what bounds the
248
+ overrun. Serial execution is unaffected (there, the deadline lands directly in
249
+ the fiber running the tool).
250
+
251
+ Approvals and concurrency are mutually exclusive per turn — the approval gate wins
252
+ and the turn goes serial. That is a deadlock avoided, not a preference.
253
+
254
+ ## Egress: the SSRF guard (and its silent failure)
255
+
256
+ Data tools make outbound HTTP, so every call passes through the **EgressGuard**, a
257
+ Server-Side Request Forgery defense. The default posture is **strict: public
258
+ `https` only.** Three env vars widen it:
259
+
260
+ | Env | Effect |
261
+ |-----|--------|
262
+ | `INSIKA_EGRESS_HOSTS` | allowlist of hosts (CSV). The safe way to permit a specific backend. |
263
+ | `INSIKA_EGRESS_ALLOW_HTTP=1` | permit plain `http` — **loopback dev only** |
264
+ | `INSIKA_EGRESS_ALLOW_PRIVATE=1` | permit private/loopback IPs — **dev only** |
265
+
266
+ > ⚠️ **Egress failures are silent.** When a tool targets a blocked host (e.g. a
267
+ > plain-`http` localhost backend without the opt-ins), the guard turns the block
268
+ > into a `{ error: … }` returned **to the model** — the request never leaves the
269
+ > process, yet the stream still emits a tool call, so the model narrates a
270
+ > plausible failure and the conversation *looks* like it worked. You will not see
271
+ > an exception.
272
+ >
273
+ > **Always verify by the trace, never by the reply:** open the Studio session
274
+ > viewer — a healthy call shows the request, args, and the backend's `200`; a
275
+ > missing or errored call is almost always egress (host not in the allowlist, or
276
+ > `http`/private without the opt-in).
277
+
278
+ Egress is **orthogonal** to registration and allowlisting: a tool can be
279
+ registered, allowed, offered to the model, and still blocked at call time.
280
+
281
+ ## Troubleshooting: "the tool is missing"
282
+
283
+ Work down this checklist:
284
+
285
+ 1. **Registered?** Is it in the catalog (Studio Tools panel)? If not, the write
286
+ or import failed — check the manifest `errors[]`, and run `insika doctor`: a stored
287
+ definition that no longer validates is dropped from the catalog, and the
288
+ `data-tools` check is the only place that says so.
289
+ 2. **Allowed for this agent?** In `tools_allow` (or an allowed group), and not in
290
+ `tools_deny`?
291
+ 3. **Egress?** If it *appears and is called* but "fails", open the trace — a
292
+ blocked call is ~99% egress.
293
+ 4. **URL literal?** For non-manifest tools, an unresolved `{{env.*}}` would have
294
+ 422'd at import — re-check the definition.
295
+
296
+ ## See also
297
+
298
+ - [Agents](AGENTS.md) — allowlists, groups, and per-agent tool exposure.
299
+ - [Plugins](PLUGINS.md) — where a code tool comes from, and how to package one.
300
+ - [Security](SECURITY.md) — egress, sandbox, and approval gating together.
301
+ - [Architecture](ARCHITECTURE.md) — the tool-loop and side-effect checkpointing.
302
+ - [`examples/data-tool/`](https://github.com/guizaols/insika/tree/main/examples/data-tool/) — a runnable data tool + the egress note.