insika 0.0.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (277) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +361 -0
  3. data/LICENSE +21 -0
  4. data/README.md +136 -2
  5. data/bin/insika +366 -0
  6. data/docs/AGENTS.md +618 -0
  7. data/docs/ARCHITECTURE.md +333 -0
  8. data/docs/BENCHMARK.md +114 -0
  9. data/docs/CHANNELS.md +453 -0
  10. data/docs/CONTEXT.md +117 -0
  11. data/docs/DEPLOY.md +354 -0
  12. data/docs/EMBEDDING.md +198 -0
  13. data/docs/EVALS.md +273 -0
  14. data/docs/LOADTEST.md +232 -0
  15. data/docs/OBSERVABILITY.md +374 -0
  16. data/docs/PLUGINS.md +211 -0
  17. data/docs/REFINEMENT.md +477 -0
  18. data/docs/RELEASING.md +70 -0
  19. data/docs/RUNNING-LOCAL.md +153 -0
  20. data/docs/SANDBOX.md +114 -0
  21. data/docs/SECURITY.md +375 -0
  22. data/docs/SKILLS.md +284 -0
  23. data/docs/TOOLS.md +302 -0
  24. data/docs/WHY.md +137 -0
  25. data/docs/WORKFLOWS.md +225 -0
  26. data/docs/build.md +14 -0
  27. data/docs/index.md +68 -0
  28. data/docs/onboarding/start.md +126 -0
  29. data/docs/operate.md +12 -0
  30. data/docs/ship.md +10 -0
  31. data/docs/understand.md +10 -0
  32. data/lib/insika/agent_file_store.rb +125 -0
  33. data/lib/insika/agent_profile.rb +255 -0
  34. data/lib/insika/alert_dispatcher.rb +139 -0
  35. data/lib/insika/allowlist.rb +28 -0
  36. data/lib/insika/baseline_store.rb +74 -0
  37. data/lib/insika/budget_ledger.rb +135 -0
  38. data/lib/insika/capability/resolved_tool.rb +34 -0
  39. data/lib/insika/capability_registry.rb +112 -0
  40. data/lib/insika/channel_delivery.rb +153 -0
  41. data/lib/insika/channel_registry.rb +30 -0
  42. data/lib/insika/channels/relay.rb +178 -0
  43. data/lib/insika/channels/web/widget.js +283 -0
  44. data/lib/insika/channels/web.rb +211 -0
  45. data/lib/insika/channels/webhook.rb +58 -0
  46. data/lib/insika/chat_builder.rb +303 -0
  47. data/lib/insika/checkpoint.rb +13 -0
  48. data/lib/insika/checkpoint_store.rb +153 -0
  49. data/lib/insika/circuit_state.rb +114 -0
  50. data/lib/insika/coercion.rb +58 -0
  51. data/lib/insika/command.rb +32 -0
  52. data/lib/insika/command_bus.rb +39 -0
  53. data/lib/insika/commands/agent_payload.rb +43 -0
  54. data/lib/insika/commands/approve_action.rb +46 -0
  55. data/lib/insika/commands/cancel_task.rb +33 -0
  56. data/lib/insika/commands/create_agent.rb +54 -0
  57. data/lib/insika/commands/create_session.rb +67 -0
  58. data/lib/insika/commands/delete_agent.rb +33 -0
  59. data/lib/insika/commands/delete_agent_file.rb +50 -0
  60. data/lib/insika/commands/delete_data_tool.rb +33 -0
  61. data/lib/insika/commands/delete_llm_provider.rb +36 -0
  62. data/lib/insika/commands/delete_mcp.rb +30 -0
  63. data/lib/insika/commands/delete_skill.rb +43 -0
  64. data/lib/insika/commands/delete_system_file.rb +29 -0
  65. data/lib/insika/commands/gate_refinement.rb +245 -0
  66. data/lib/insika/commands/import_mcp_tools.rb +48 -0
  67. data/lib/insika/commands/import_tools.rb +81 -0
  68. data/lib/insika/commands/issue_tenant_token.rb +41 -0
  69. data/lib/insika/commands/memory_add_note.rb +32 -0
  70. data/lib/insika/commands/memory_forget_fact.rb +32 -0
  71. data/lib/insika/commands/memory_put_fact.rb +35 -0
  72. data/lib/insika/commands/pause_task.rb +29 -0
  73. data/lib/insika/commands/resolve_refinement.rb +126 -0
  74. data/lib/insika/commands/restore_agent_file.rb +36 -0
  75. data/lib/insika/commands/restore_data_tool.rb +34 -0
  76. data/lib/insika/commands/restore_system_file.rb +31 -0
  77. data/lib/insika/commands/resume_task.rb +85 -0
  78. data/lib/insika/commands/revoke_token.rb +39 -0
  79. data/lib/insika/commands/rotate_tenant_token.rb +43 -0
  80. data/lib/insika/commands/run_refinement.rb +133 -0
  81. data/lib/insika/commands/send_message.rb +150 -0
  82. data/lib/insika/commands/set_agent_tools.rb +39 -0
  83. data/lib/insika/commands/set_skill_agents.rb +112 -0
  84. data/lib/insika/commands/trigger_workflow.rb +80 -0
  85. data/lib/insika/commands/update_agent.rb +49 -0
  86. data/lib/insika/commands/update_settings.rb +33 -0
  87. data/lib/insika/commands/upsert_llm_provider.rb +34 -0
  88. data/lib/insika/commands/upsert_mcp.rb +32 -0
  89. data/lib/insika/commands/write_agent_file.rb +57 -0
  90. data/lib/insika/commands/write_data_tool.rb +43 -0
  91. data/lib/insika/commands/write_golden.rb +58 -0
  92. data/lib/insika/commands/write_skill.rb +60 -0
  93. data/lib/insika/commands/write_system_file.rb +31 -0
  94. data/lib/insika/config_store.rb +89 -0
  95. data/lib/insika/context/builder.rb +166 -0
  96. data/lib/insika/context/catalog_provider.rb +23 -0
  97. data/lib/insika/context/fragment.rb +43 -0
  98. data/lib/insika/context/priority.rb +30 -0
  99. data/lib/insika/context/provider.rb +19 -0
  100. data/lib/insika/context/providers/memory.rb +60 -0
  101. data/lib/insika/context/providers/prompt.rb +105 -0
  102. data/lib/insika/context/providers/request.rb +32 -0
  103. data/lib/insika/context/providers/session.rb +123 -0
  104. data/lib/insika/context/providers/skill.rb +24 -0
  105. data/lib/insika/context/providers/skill_trigger.rb +128 -0
  106. data/lib/insika/context/providers/tool_search.rb +20 -0
  107. data/lib/insika/context_trace_store.rb +92 -0
  108. data/lib/insika/delegation_store.rb +153 -0
  109. data/lib/insika/doctor.rb +539 -0
  110. data/lib/insika/dsl/definition.rb +55 -0
  111. data/lib/insika/dsl/runtime.rb +382 -0
  112. data/lib/insika/dsl/server_boot.rb +98 -0
  113. data/lib/insika/dsl/system.rb +93 -0
  114. data/lib/insika/dsl/workflow_adapter.rb +59 -0
  115. data/lib/insika/dsl.rb +364 -0
  116. data/lib/insika/edge_limiter.rb +268 -0
  117. data/lib/insika/egress_guard.rb +75 -0
  118. data/lib/insika/env_schema.rb +249 -0
  119. data/lib/insika/errors.rb +201 -0
  120. data/lib/insika/evals/assertions.rb +247 -0
  121. data/lib/insika/evals/baseline.rb +69 -0
  122. data/lib/insika/evals/golden.rb +172 -0
  123. data/lib/insika/evals/judge.rb +225 -0
  124. data/lib/insika/evals/pairwise.rb +178 -0
  125. data/lib/insika/evals/report.rb +115 -0
  126. data/lib/insika/evals/runner.rb +141 -0
  127. data/lib/insika/evals/transport.rb +178 -0
  128. data/lib/insika/event.rb +18 -0
  129. data/lib/insika/event_stream.rb +132 -0
  130. data/lib/insika/executor.rb +1995 -0
  131. data/lib/insika/frontmatter.rb +42 -0
  132. data/lib/insika/golden_store.rb +145 -0
  133. data/lib/insika/hooks.rb +48 -0
  134. data/lib/insika/http_client.rb +63 -0
  135. data/lib/insika/inbound_log.rb +84 -0
  136. data/lib/insika/llm_configurator.rb +99 -0
  137. data/lib/insika/llm_provider_store.rb +83 -0
  138. data/lib/insika/loop_detector.rb +143 -0
  139. data/lib/insika/mcp_http_client.rb +67 -0
  140. data/lib/insika/mcp_store.rb +115 -0
  141. data/lib/insika/mcp_tool_ingestor.rb +143 -0
  142. data/lib/insika/memory_store.rb +93 -0
  143. data/lib/insika/message_origin.rb +76 -0
  144. data/lib/insika/middleware.rb +36 -0
  145. data/lib/insika/model_policy.rb +52 -0
  146. data/lib/insika/model_resolver.rb +176 -0
  147. data/lib/insika/model_selection.rb +115 -0
  148. data/lib/insika/onboarding.rb +208 -0
  149. data/lib/insika/outbox_store.rb +166 -0
  150. data/lib/insika/overlay_tool_registry.rb +102 -0
  151. data/lib/insika/pack.rb +102 -0
  152. data/lib/insika/pack_importer.rb +123 -0
  153. data/lib/insika/pending_action_store.rb +120 -0
  154. data/lib/insika/plugin/loader.rb +356 -0
  155. data/lib/insika/plugin.rb +35 -0
  156. data/lib/insika/policy/engine.rb +83 -0
  157. data/lib/insika/policy/policy.rb +120 -0
  158. data/lib/insika/policy_registry.rb +23 -0
  159. data/lib/insika/profile_source.rb +143 -0
  160. data/lib/insika/prompt_catalog.rb +61 -0
  161. data/lib/insika/provider_error_classifier.rb +160 -0
  162. data/lib/insika/queue_policy.rb +167 -0
  163. data/lib/insika/recovery.rb +168 -0
  164. data/lib/insika/refinement/candidate.rb +159 -0
  165. data/lib/insika/refinement/evidence_collector.rb +371 -0
  166. data/lib/insika/refinement/gate.rb +234 -0
  167. data/lib/insika/refinement/panel.rb +222 -0
  168. data/lib/insika/refinement/proposer.rb +262 -0
  169. data/lib/insika/refinement_store.rb +295 -0
  170. data/lib/insika/registry.rb +59 -0
  171. data/lib/insika/reliability.rb +185 -0
  172. data/lib/insika/safety/config.rb +109 -0
  173. data/lib/insika/safety/detectors.rb +176 -0
  174. data/lib/insika/safety/factory.rb +102 -0
  175. data/lib/insika/safety/input_guardrail.rb +102 -0
  176. data/lib/insika/safety/moderator.rb +94 -0
  177. data/lib/insika/safety/output_filter.rb +79 -0
  178. data/lib/insika/safety/output_validator.rb +101 -0
  179. data/lib/insika/safety/safe_responses.rb +47 -0
  180. data/lib/insika/sandbox/boundary.rb +93 -0
  181. data/lib/insika/sandbox/docker.rb +74 -0
  182. data/lib/insika/sandbox/local.rb +33 -0
  183. data/lib/insika/sandbox/runner.rb +80 -0
  184. data/lib/insika/sandbox.rb +85 -0
  185. data/lib/insika/schema_guard.rb +147 -0
  186. data/lib/insika/secret_masking.rb +34 -0
  187. data/lib/insika/server/a2a/agent_card.rb +27 -0
  188. data/lib/insika/server/a2a/app.rb +112 -0
  189. data/lib/insika/server/a2a/client.rb +101 -0
  190. data/lib/insika/server/a2a/errors.rb +32 -0
  191. data/lib/insika/server/a2a/http.rb +42 -0
  192. data/lib/insika/server/a2a/message.rb +27 -0
  193. data/lib/insika/server/a2a/protocol.rb +45 -0
  194. data/lib/insika/server/a2a/remotes.rb +25 -0
  195. data/lib/insika/server/a2a/task_projection.rb +40 -0
  196. data/lib/insika/server/app.rb +1022 -0
  197. data/lib/insika/server/boot.rb +119 -0
  198. data/lib/insika/server/rack_app.rb +118 -0
  199. data/lib/insika/server/responses.rb +165 -0
  200. data/lib/insika/server/sse_body.rb +96 -0
  201. data/lib/insika/server/tenant_auth.rb +61 -0
  202. data/lib/insika/session_actor.rb +162 -0
  203. data/lib/insika/session_store.rb +143 -0
  204. data/lib/insika/settings_store.rb +154 -0
  205. data/lib/insika/shutdown.rb +125 -0
  206. data/lib/insika/skill_catalog.rb +220 -0
  207. data/lib/insika/skill_store.rb +127 -0
  208. data/lib/insika/steer_injector.rb +110 -0
  209. data/lib/insika/store.rb +52 -0
  210. data/lib/insika/stores/memory.rb +123 -0
  211. data/lib/insika/stores/sqlite.rb +183 -0
  212. data/lib/insika/studio/app.rb +1693 -0
  213. data/lib/insika/studio/assets/dist/application.css +1 -0
  214. data/lib/insika/studio/assets/dist/application.js +70 -0
  215. data/lib/insika/studio/forms.rb +335 -0
  216. data/lib/insika/studio/nav_icons.rb +31 -0
  217. data/lib/insika/studio/views/_message.erb +44 -0
  218. data/lib/insika/studio/views/agent_detail.erb +285 -0
  219. data/lib/insika/studio/views/agents.erb +63 -0
  220. data/lib/insika/studio/views/approvals.erb +41 -0
  221. data/lib/insika/studio/views/chats.erb +34 -0
  222. data/lib/insika/studio/views/evals.erb +83 -0
  223. data/lib/insika/studio/views/home.erb +72 -0
  224. data/lib/insika/studio/views/layout.erb +94 -0
  225. data/lib/insika/studio/views/login.erb +17 -0
  226. data/lib/insika/studio/views/mcp.erb +91 -0
  227. data/lib/insika/studio/views/not_found.erb +5 -0
  228. data/lib/insika/studio/views/playground.erb +47 -0
  229. data/lib/insika/studio/views/refinement.erb +234 -0
  230. data/lib/insika/studio/views/session.erb +137 -0
  231. data/lib/insika/studio/views/settings.erb +168 -0
  232. data/lib/insika/studio/views/skills.erb +141 -0
  233. data/lib/insika/studio/views/system_files.erb +65 -0
  234. data/lib/insika/studio/views/task.erb +105 -0
  235. data/lib/insika/studio/views/tasks.erb +33 -0
  236. data/lib/insika/studio/views/tool_edit.erb +107 -0
  237. data/lib/insika/studio/views/tools.erb +89 -0
  238. data/lib/insika/subagent_graph.rb +96 -0
  239. data/lib/insika/system_file_store.rb +96 -0
  240. data/lib/insika/task_actor.rb +128 -0
  241. data/lib/insika/task_store.rb +250 -0
  242. data/lib/insika/telemetry/pricing.rb +104 -0
  243. data/lib/insika/telemetry/recorder.rb +228 -0
  244. data/lib/insika/telemetry.rb +127 -0
  245. data/lib/insika/testing/store_contract.rb +270 -0
  246. data/lib/insika/tick.rb +122 -0
  247. data/lib/insika/token_estimator.rb +16 -0
  248. data/lib/insika/token_store.rb +168 -0
  249. data/lib/insika/tool_assembly.rb +140 -0
  250. data/lib/insika/tool_catalog.rb +89 -0
  251. data/lib/insika/tool_definition.rb +518 -0
  252. data/lib/insika/tool_envelope.rb +140 -0
  253. data/lib/insika/tool_manifest.rb +218 -0
  254. data/lib/insika/tool_output_compressor.rb +100 -0
  255. data/lib/insika/tool_registry.rb +21 -0
  256. data/lib/insika/tool_store.rb +135 -0
  257. data/lib/insika/tool_trace_store.rb +92 -0
  258. data/lib/insika/tools/a2a_remote.rb +48 -0
  259. data/lib/insika/tools/agent_enum.rb +68 -0
  260. data/lib/insika/tools/concurrency.rb +54 -0
  261. data/lib/insika/tools/data_defined_tool.rb +219 -0
  262. data/lib/insika/tools/load_skill.rb +99 -0
  263. data/lib/insika/tools/remember.rb +53 -0
  264. data/lib/insika/tools/stuck_signal.rb +44 -0
  265. data/lib/insika/tools/subagent.rb +75 -0
  266. data/lib/insika/tools/subagents.rb +77 -0
  267. data/lib/insika/tools/tool_search.rb +94 -0
  268. data/lib/insika/turn_output.rb +139 -0
  269. data/lib/insika/turn_state.rb +162 -0
  270. data/lib/insika/turn_timing.rb +56 -0
  271. data/lib/insika/usage_ledger.rb +47 -0
  272. data/lib/insika/version.rb +3 -1
  273. data/lib/insika/wiring/graph.rb +249 -0
  274. data/lib/insika/workflow.rb +185 -0
  275. data/lib/insika/workflow_registry.rb +33 -0
  276. data/lib/insika.rb +220 -4
  277. metadata +412 -8
data/docs/WHY.md ADDED
@@ -0,0 +1,137 @@
1
+ ---
2
+ title: Why Insika
3
+ parent: Understand the idea
4
+ nav_order: 1
5
+ permalink: /why/
6
+ ---
7
+
8
+ # Why Insika
9
+
10
+ **Ruby is ready for production AI.** The narrative that you must reach for Python
11
+ to ship serious LLM applications is out of date. [RubyLLM](https://rubyllm.com)
12
+ gave the language a first-class, provider-agnostic LLM client — chat, tools,
13
+ streaming, embeddings, moderation — a genuinely valuable foundation that closed
14
+ the "Ruby can't do AI" gap. Insika is the next layer up: it takes those
15
+ primitives and makes an agent *dependable in production*.
16
+
17
+ Because **an agent without a harness is just a chat loop.** The model does the
18
+ reasoning, but everything that makes an agent survive contact with real traffic —
19
+ durable state, tools, guardrails, evals, operations — lives in the engine around
20
+ it. Those harnesses exist, mature, for other ecosystems. Ruby teams have mostly
21
+ been told to glue libraries together. They shouldn't have to.
22
+
23
+ Insika is an **agent runtime for Ruby**: the turn pipeline, the operational
24
+ surface, and the safety layer, in one deployable piece — behind an
25
+ OpenAI-Responses-compatible API. It stands on RubyLLM for the provider layer and
26
+ adds everything between a chat call and a production agent.
27
+
28
+ And it does it *ergonomically*. A complete, running agent is a few lines of Ruby —
29
+ the same speed-to-first-agent story Python teams tell, without leaving Ruby:
30
+
31
+ ```ruby
32
+ require "insika"
33
+
34
+ agent = Insika.agent("assistant") do
35
+ model "deepseek-v4-flash"
36
+ provider :deepseek
37
+ instructions "You are a concise, friendly assistant."
38
+ end
39
+
40
+ puts agent.reply("hi, what can you do?") # one turn, in-process
41
+ agent.serve # ...or a full server: /studio + /v1/responses
42
+ ```
43
+
44
+ Every capability below is reachable the same way — see [`examples/`](https://github.com/guizaols/insika/tree/main/examples/),
45
+ one small runnable project per capability.
46
+
47
+ ## Four ways to run an agent
48
+
49
+ There are, roughly, four ways to put an LLM agent into production. They're not
50
+ wrong — they're different amounts of "build it yourself":
51
+
52
+ - **Raw SDK loop.** Call the provider SDK directly and hand-roll the
53
+ reason→act→observe loop. Minimal to start; everything operational is on you.
54
+ - **Assemble a framework.** Compose library pieces — a tool abstraction here, a
55
+ memory store there, your own persistence and moderation. Flexible, but you own
56
+ the integration and the gaps between the parts.
57
+ - **Hosted agent gateway.** A managed service runs the agent behind an API. You
58
+ get operations for free, but your agents, prompts, and conversation data live in
59
+ someone else's system.
60
+ - **An agent runtime (this project).** The loop, durability, tools, safety, and an
61
+ operations UI as one thing you deploy in your own infrastructure.
62
+
63
+ How they compare on what production actually demands:
64
+
65
+ | | Raw SDK loop | Assemble a framework | Hosted gateway | **Insika (runtime)** |
66
+ |---|:---:|:---:|:---:|:---:|
67
+ | Get a first agent talking | ✅ | ⚠️ | ✅ | ✅ |
68
+ | Add/change a tool without a redeploy | ❌ | ❌ | ⚠️ | ✅ |
69
+ | Durable, resumable turns (crash mid-conversation) | ❌ | ⚠️ | ✅ | ✅ |
70
+ | Content-safety guardrails built in | ❌ | ⚠️ | ⚠️ | ✅ |
71
+ | Evals / regression gating as a primitive | ❌ | ⚠️ | ⚠️ | ✅ |
72
+ | Operations UI included | ❌ | ❌ | ✅ | ✅ |
73
+ | Runs in your infrastructure (you own the data) | ✅ | ✅ | ❌ | ✅ |
74
+ | One container, no orchestrator to start | ✅ | ⚠️ | — | ✅ |
75
+
76
+ ✅ built in · ⚠️ possible, but you build/integrate it · ❌ not addressed · — n/a
77
+
78
+ ## What we do differently
79
+
80
+ - **Tools are data, not code.** Define a tool with a JSON Schema manifest and a
81
+ declarative binding, and the running agent picks it up — no rebuild, no redeploy.
82
+ Import whole toolsets from an MCP server or a manifest at runtime. In most setups
83
+ a new tool is a code change; here it's a row.
84
+ - **Durability is the default.** Every turn checkpoints; sessions and tasks recover
85
+ across restarts; continuous replication is one env var away, with a documented
86
+ restore drill. If the box dies mid-conversation, the conversation doesn't.
87
+ - **Safety is built in, not bolted on.** Input guardrails (prompt-injection and
88
+ abuse handling that fails gracefully *without* burning a model turn), content
89
+ moderation, PII/secret redaction in the output stream, and a post-turn validator
90
+ — configured, not hand-rolled, and opt-in per agent.
91
+ - **Evals are a primitive.** Golden conversations, LLM-as-judge, and baseline
92
+ gating run against the *same* API your users hit — so a prompt or model change
93
+ that regresses behavior fails loudly before it ships.
94
+ - **An operations UI is included.** Studio lets an operator manage agents, tools,
95
+ approvals, and traces — pause a task, approve a sensitive action, inspect a
96
+ tool-call trace — without writing code. Headless when you want it, operable when
97
+ you need it.
98
+ - **One container, one file.** SQLite in WAL mode with streaming replication. No
99
+ queue cluster, no orchestrator, no managed database required to start — and it
100
+ runs a real production workload today.
101
+ - **Fibers, not a thread per request.** An LLM turn is almost all *waiting* on the
102
+ provider. Insika runs on the [Async](https://github.com/socketry/async) fiber
103
+ scheduler and serves under [Falcon](https://github.com/socketry/falcon), so one
104
+ process handles thousands of concurrent turns on a handful of connections instead
105
+ of pinning a heavyweight thread per call. This is exactly the model
106
+ [RubyLLM's async guide](https://rubyllm.com/async/) recommends — *"Falcon is a
107
+ Ruby application server built on fibers; with Falcon, async just works"* — and it
108
+ composes for free, because RubyLLM's HTTP cooperates with Ruby's fiber scheduler.
109
+ It's a big part of why one box goes so far.
110
+ - **The engine gets out of the way.** On top of that concurrency, all the machinery
111
+ above adds **well under a millisecond of overhead per turn**; a single process
112
+ sustains thousands of turns per second of pure engine work. A neutral,
113
+ provider-free benchmark in the repo reproduces it with one command — see
114
+ [BENCHMARK.md](BENCHMARK.md). The rest of a turn's latency is the model, not us.
115
+ *(That number is engine overhead only — end-to-end latency is provider-bound and
116
+ not claimed here.)*
117
+
118
+ ## When a lighter tool is the right call
119
+
120
+ Insika is for when an agent has to run **unattended, durably, and safely, in
121
+ production**. If that's not you yet, reach for less:
122
+
123
+ - If you only need LLM calls with tools in Ruby and nothing operational around
124
+ them, **[RubyLLM](https://rubyllm.com) alone is excellent** — and it's exactly
125
+ what Insika builds on. Reach for Insika when "nothing operational around them"
126
+ stops being true: when the agent has to run unattended, recover from crashes,
127
+ enforce safety, and be operable by someone who isn't you.
128
+ - If your team is committed to another language ecosystem, use a runtime native to
129
+ it — the ideas here travel, the code doesn't.
130
+ - If you want a minimal terminal coding assistant rather than a product platform,
131
+ a single-purpose CLI agent will be simpler.
132
+
133
+ ## See also
134
+
135
+ - [README](https://github.com/guizaols/insika#readme) — quickstart and the drop-in `/v1/responses` API.
136
+ - [BENCHMARK.md](BENCHMARK.md) — the neutral, reproducible engine benchmark.
137
+ - [OBSERVABILITY.md](OBSERVABILITY.md) — OpenTelemetry tracing (opt-in).
data/docs/WORKFLOWS.md ADDED
@@ -0,0 +1,225 @@
1
+ ---
2
+ title: Workflows
3
+ parent: Build an agent
4
+ nav_order: 5
5
+ permalink: /workflows/
6
+ ---
7
+
8
+ # Workflows
9
+
10
+ A single agent's tool-loop is one shape of work: the model decides, calls a tool,
11
+ looks at the result, decides again. Plenty of real work has a shape you already
12
+ know — draft then edit, classify then answer, three reviewers then a summary,
13
+ try until it passes. For those, the choice of "what happens next" belongs in
14
+ **Ruby**, not in a prompt.
15
+
16
+ That is a **workflow**: deterministic orchestration around agent turns.
17
+
18
+ ## The one decision that matters
19
+
20
+ **Who chooses the next step — your code, or the model?**
21
+
22
+ | | **Workflow** | **Delegation (`subagents`)** |
23
+ |---|---|---|
24
+ | Chooses the next step | your Ruby | the model |
25
+ | Cost of deciding | zero | a model call |
26
+ | Repeatable | yes — same path every run | no |
27
+ | Reach for it when | the shape is known in advance | the work depends on what the user said |
28
+ | Surface | `workflow` in a system + `POST /v1/workflows/:name` | `subagents` on a profile + `spawn_subagent(s)` |
29
+
30
+ They compose: a workflow step can `ask` an agent that itself delegates. Start
31
+ with a workflow and hand the choice to the model only where the choice is
32
+ genuinely open — see [Agents](AGENTS.md#delegation-subagents) for delegation.
33
+
34
+ ## Declaring one
35
+
36
+ Workflows live in a **system** (`Insika.system`), because their steps address
37
+ agents by id and those agents must resolve in the same runtime:
38
+
39
+ ```ruby
40
+ newsroom = Insika.system do
41
+ provider :deepseek
42
+
43
+ agent("writer") { model "deepseek-v4-flash"; instructions "Write ONE paragraph." }
44
+ agent("editor") { model "deepseek-v4-flash"; instructions "Rewrite as ONE sentence." }
45
+
46
+ workflow "publish",
47
+ description: "Draft a paragraph, then tighten it.",
48
+ input: { type: "object", properties: { topic: { type: "string" } }, required: ["topic"] },
49
+ output: { type: "object", properties: { headline: { type: "string" } } } do |input, ctx|
50
+ draft = ctx.ask("writer", "Topic: #{input['topic']}")
51
+ { "headline" => ctx.ask("editor", draft) }
52
+ end
53
+ end
54
+
55
+ newsroom.run("publish", input: { "topic" => "Ruby fibers" })
56
+ # => { "headline" => "…" }
57
+ ```
58
+
59
+ What the block receives:
60
+
61
+ | | What it does |
62
+ |---|---|
63
+ | `input` | the validated input Hash (string keys) |
64
+ | `ctx.ask(agent, message)` | one turn against one agent of the system → its text |
65
+ | `ctx.gather(*blocks, max: 8)` | runs blocks **concurrently**, returns values **in order** |
66
+ | `ctx.context` / `ctx.tools` | the turn's ContextPackage and its resolved tool instances — escape hatches, rarely needed |
67
+
68
+ The workflow's return value **is** its output: whatever Hash (or value) you
69
+ return is what `run` returns, what `:workflow_completed` carries, and what the
70
+ `output` schema validates.
71
+
72
+ ## What you get for free
73
+
74
+ - **A durable run.** The run id *is* a Task — checkpointed and recoverable, so a
75
+ crash mid-run does not lose the record. `run_id` comes back from every trigger.
76
+ - **Schemas at the edges.** `input:` is validated **synchronously**: a bad input
77
+ is refused with **no run created** (a `422` over HTTP). `output:` is validated
78
+ when the workflow returns; a violation fails the run at the `workflow_schema`
79
+ stage instead of handing a malformed result downstream. Both take a JSON Schema
80
+ Hash or any dry-schema-compatible validator.
81
+ - **Events.** `:workflow_started` and `:workflow_completed` (with the typed
82
+ output) on the turn's event stream, alongside every step's own turn events.
83
+ - **Every step is visible.** Each `ctx.ask` is its own turn — its own Task, its
84
+ own trace in the Studio. A workflow is not a black box.
85
+
86
+ ## Over HTTP
87
+
88
+ Workflows are exposed when the deployment injects the registry (a `serve`d system
89
+ with at least one workflow does it automatically):
90
+
91
+ ```bash
92
+ GET /v1/workflows # discovery: names, descriptions, I/O schemas
93
+ POST /v1/workflows/publish # 202 { run_id, task_id } — fire and observe
94
+ POST /v1/workflows/publish?stream=true # SSE of the run, ending at workflow_completed
95
+ ```
96
+
97
+ ```jsonc
98
+ // POST body
99
+ { "agent": "writer", // the profile the run executes under
100
+ "input": { "topic": "Ruby fibers" },
101
+ "session_id": null } // optional
102
+ ```
103
+
104
+ Observe an async run with `GET /v1/tasks/:run_id` or
105
+ `GET /v1/events?task_id=:run_id`. A bad input is a `422`; an unknown workflow or
106
+ agent is a `404`.
107
+
108
+ ## The five patterns
109
+
110
+ Each has a runnable script in
111
+ [`examples/agentic-workflows/`](https://github.com/guizaols/insika/tree/main/examples/agentic-workflows/).
112
+
113
+ ### Sequential (prompt chaining)
114
+
115
+ Each step's output feeds the next. The order is code, so it is identical on
116
+ every run.
117
+
118
+ ```ruby
119
+ draft = ctx.ask("writer", input["topic"])
120
+ ctx.ask("editor", draft)
121
+ ```
122
+
123
+ ### Routing
124
+
125
+ One cheap classification turn, then a specialist with a short, focused prompt —
126
+ instead of one agent carrying every instruction it might ever need.
127
+
128
+ ```ruby
129
+ label = ctx.ask("router", input["message"]).downcase # constrain the label set
130
+ lane = label.include?("billing") ? "billing" : "technical"
131
+ ctx.ask(lane, input["message"])
132
+ ```
133
+
134
+ **Do the branching in Ruby, with a fallback.** A closed label set plus an `else`
135
+ means an unexpected answer degrades instead of picking a random branch.
136
+
137
+ ### Parallel (fan-out / fan-in)
138
+
139
+ An LLM turn is almost all *waiting* on the provider, so independent turns overlap
140
+ on the fiber scheduler: wall-clock is the **slowest** branch, not the sum.
141
+
142
+ ```ruby
143
+ security, performance = ctx.gather(-> { ctx.ask("security", code) },
144
+ -> { ctx.ask("performance", code) })
145
+ ctx.ask("lead", "security: #{security}\nperformance: #{performance}") # fan-in
146
+ ```
147
+
148
+ `gather` returns values in declaration order and caps concurrency at `max:`
149
+ (default 8) — the cap matters, because N simultaneous calls hit provider rate
150
+ limits and per-agent token ceilings.
151
+
152
+ ### Evaluator-optimizer
153
+
154
+ Produce, judge, revise. The judge's reason becomes the revision instruction.
155
+
156
+ ```ruby
157
+ loop do
158
+ verdict = ctx.ask("critic", "Brief: #{brief}\nTagline: #{tagline}")
159
+ break if verdict.include?("PASS") || attempts >= MAX_ATTEMPTS # the cap is CODE
160
+ tagline = ctx.ask("writer", "Fix this: #{verdict}")
161
+ attempts += 1
162
+ end
163
+ ```
164
+
165
+ **Two things must be code, not prompt:** the attempt cap (a loop that asks the
166
+ model when to stop can run forever) and the honest report of whether it actually
167
+ passed. Have the judge answer in a shape you can branch on (`VERDICT: PASS|FAIL`).
168
+
169
+ ### Orchestrator-workers
170
+
171
+ The model chooses the workers. This is delegation, not a workflow — see
172
+ [Agents](AGENTS.md#delegation-subagents).
173
+
174
+ ```ruby
175
+ agent "lead" do
176
+ instructions "You have no expertise yourself — always delegate, then report."
177
+ subagents "security", "performance"
178
+ end
179
+ ```
180
+
181
+ ## Gotchas
182
+
183
+ - **`ctx.ask` is stateless.** A step is a unit of work, not a conversation: no
184
+ session is threaded, so the agent sees only the message you pass. Pass what it
185
+ needs.
186
+ - **Nothing forces the model to delegate.** If a parent answers alone, the task
187
+ list shows only the parent — and it is the *prompt* that needs work: say
188
+ plainly that it has no expertise of its own. (The engine helps by naming the
189
+ spawnable agent ids in the tool contract, but it cannot make the model use them.)
190
+ - **The `agent:` of a run is the profile it executes under** — its policy,
191
+ limits, and guardrails. It is not "the agent that does the work"; the steps
192
+ pick those themselves.
193
+ - **A workflow turn assembles no chat.** Stage 5 is skipped: there is no
194
+ model call for the *workflow itself*, only for the turns its steps start.
195
+ One consequence worth naming: `queue_mode: "steer"`
196
+ ([Agents](AGENTS.md#steer--the-message-arrives-while-the-turn-is-already-running))
197
+ cannot append into a running workflow — there is no chat to append to, and no tool
198
+ batch of the engine's to use as a boundary. A message that arrives while a workflow
199
+ run is in flight becomes the next turn on the session, which is what `followup`
200
+ does. Steering is refused at the door, not silently dropped.
201
+ - **A running workflow reaches no boundary until it returns.** The engine's stage
202
+ boundaries sit *around* the workflow call, not inside it, so `queue_mode:
203
+ "interrupt"` (and a plain `cancel_task`) is observed only once the workflow is done:
204
+ the steps run, and then the run is abandoned without publishing or persisting. If you
205
+ need a run to stop early, that decision belongs inside the workflow — it is ordinary
206
+ Ruby, so `return` on the condition you care about.
207
+ - **Braces bite in Ruby.** `agent "x" { … }` is a syntax error (`{` binds to the
208
+ argument); write `agent("x") { … }` or `agent "x" do … end`.
209
+
210
+ ## Troubleshooting
211
+
212
+ | Symptom | Cause |
213
+ |---|---|
214
+ | `GET /v1/workflows` 404s | no workflow declared, so the routes are not exposed (parity) |
215
+ | `422` on trigger | the input violates `input:` — the message names the offending field |
216
+ | `404` on trigger | unknown workflow name, or the `agent:` does not exist |
217
+ | run fails at `workflow_schema` | your return value violates `output:` |
218
+ | a step "does nothing" | the agent id is not in this system — `ask` raises with the list of valid ids |
219
+
220
+ ## See also
221
+
222
+ - [Agents](AGENTS.md) — profiles, and delegation via `subagents`.
223
+ - [Tools](TOOLS.md) — what a single agent can call inside one turn.
224
+ - [Architecture](ARCHITECTURE.md) — the turn pipeline a run goes through, and checkpointing.
225
+ - [`examples/agentic-workflows/`](https://github.com/guizaols/insika/tree/main/examples/agentic-workflows/) — the five patterns, runnable.
data/docs/build.md ADDED
@@ -0,0 +1,14 @@
1
+ ---
2
+ title: Build an agent
3
+ nav_order: 3
4
+ has_children: true
5
+ permalink: /build/
6
+ ---
7
+
8
+ # Build an agent
9
+
10
+ An agent is data: a profile, its tools, its skills, and what fills its prompt. Start
11
+ with [Agents](AGENTS.md), then add capability one doc at a time — and when config is not
12
+ enough, [Plugins](PLUGINS.md) covers the two ways to extend the engine itself. Every one of these has a
13
+ runnable counterpart under
14
+ [`examples/`](https://github.com/guizaols/insika/tree/main/examples/).
data/docs/index.md ADDED
@@ -0,0 +1,68 @@
1
+ ---
2
+ title: Home
3
+ nav_order: 1
4
+ permalink: /
5
+ ---
6
+
7
+ # Insika
8
+ {: .fs-9 }
9
+
10
+ Your agent is the idea. Insika is what holds it up in production.
11
+ {: .fs-6 .fw-300 }
12
+
13
+ [Build your first agent](RUNNING-LOCAL.md){: .btn .btn-primary .fs-5 .mb-4 .mb-md-0 .mr-2 }
14
+ [View on GitHub](https://github.com/guizaols/insika){: .btn .fs-5 .mb-4 .mb-md-0 }
15
+
16
+ ---
17
+
18
+ *Insika* is Zulu for the pillar that carries a structure — the part nobody admires and
19
+ everything rests on. A turn that survives a crash, tools that cannot wander off, limits
20
+ that hold under load, and an API your clients already speak. Build the agent; the
21
+ scaffolding is already here.
22
+
23
+ Concretely: a Ruby runtime for **LLM agents in production** — a durable, resumable turn
24
+ pipeline behind an **OpenAI-Responses-compatible** HTTP API (`POST /v1/responses`), with
25
+ tools, skills, cross-session memory, per-agent policy, content-safety guardrails, and a
26
+ web control UI. Point an existing Responses client at it and serve many agents from one
27
+ deployment.
28
+
29
+ ## Your first agent
30
+
31
+ Ruby `>= 3.3` and a provider key (the demo uses DeepSeek). The whole program:
32
+
33
+ ```ruby
34
+ require "insika"
35
+
36
+ assistant = Insika.agent("assistant") do
37
+ model "deepseek-v4-flash"
38
+ provider :deepseek
39
+ instructions "You are Bia, a concise and friendly assistant. Answer briefly."
40
+ end
41
+
42
+ puts assistant.reply("hi, what can you do?") # one turn, in-process
43
+ ```
44
+
45
+ Swap `reply` for `serve` and the same agent is a server — the control UI at `/studio`
46
+ plus the drop-in API, on `:9292`. → [Running locally](RUNNING-LOCAL.md)
47
+
48
+ ## Or let your coding agent read the docs
49
+
50
+ A running instance serves this same documentation as raw markdown, plus a
51
+ skill-structured prompt that walks a coding agent through building your first agent:
52
+
53
+ ```
54
+ Read http://localhost:9292/start.md then help me build my first agent
55
+ ```
56
+
57
+ Alongside it, `GET /models.json` (configured providers, model ids, defaults — no
58
+ secrets), `GET /docs` and `GET /docs/<name>.md`. Public and on by default when you
59
+ `serve`; opt-in in production (`INSIKA_ONBOARDING=1`).
60
+
61
+ ## Where to go next
62
+
63
+ - **[Understand the idea](understand.md)** — why a runtime rather than a DIY loop, and how a turn actually runs.
64
+ - **[Build an agent](build.md)** — agents, tools, skills, context, and the local loop.
65
+ - **[Ship it](ship.md)** — security, confined execution, deployment.
66
+ - **[Operate & prove it](operate.md)** — observability, the benchmark, load testing, evals, refinement.
67
+
68
+ Pre-release: APIs may still change and nothing is tagged yet. Licensed MIT.
@@ -0,0 +1,126 @@
1
+ # Insika — build your first agent
2
+
3
+ > **You are a coding agent** (Claude Code, Cursor, an IDE assistant, …) reading this
4
+ > on behalf of a developer who just pointed you at a running Insika. Treat this file
5
+ > as a **skill**: follow the steps in order, apply the RULES literally, and stop at the
6
+ > self-check. Do **not** improvise beyond it.
7
+
8
+ Insika is a Ruby runtime for LLM agents in production. Your job here is the smallest
9
+ possible one: get the developer a **first working agent**, defined in Ruby with the
10
+ public DSL (`Insika.agent { … }`), talking to a real model — nothing more.
11
+
12
+ The machine-readable list of models this engine already knows about is at
13
+ **{{MODELS_URL}}** — fetch it before you write any `model`/`provider` line. The full
14
+ docs are mirrored as raw markdown under **{{DOCS_URL}}**.
15
+
16
+ ---
17
+
18
+ ## Step 0 — Gather context (do this first, silently)
19
+
20
+ RULES — verify, do not assume:
21
+
22
+ - **Ruby ≥ 3.3.** Run `ruby -v`. If lower, stop and tell the developer; do not try to
23
+ upgrade Ruby for them.
24
+ - **The gem/library must be loadable.** In a project that already depends on Insika,
25
+ `require "insika"` works. Otherwise add it to the `Gemfile` (or `bundle add`) — do
26
+ not vendor or copy source files.
27
+ - **A provider key must come from the developer, via the environment.** Look for one
28
+ already exported (e.g. `DEEPSEEK_API_KEY`, `OPENAI_API_KEY`). If none is set, **ask
29
+ the developer for the provider and confirm the env var is exported.** See the hard
30
+ constraint on keys below.
31
+ - **Fetch {{MODELS_URL}}.** It tells you which providers and model ids this engine is
32
+ configured for, the platform default, and the valid `thinking` levels. Use those
33
+ exact ids.
34
+
35
+ ## Step 1 — Decide (RULES, not taste)
36
+
37
+ | Question | RULE |
38
+ |---|---|
39
+ | How many agents? | **Exactly one.** A first agent is a single `Insika.agent`. Resist adding a second. |
40
+ | Which model/provider? | Use an id from **{{MODELS_URL}}**. If it lists a `default`, use that. Never guess a model id. |
41
+ | Where does it live? | One new Ruby file (e.g. `my_agent.rb`), or extend `examples/quickstart.rb` if present. One file. |
42
+ | Tools / skills / memory? | **None yet.** Ship a plain conversational agent first; add capability only after it replies. |
43
+ | Reply or serve? | Start with `reply` (one in-process turn). Only switch to `serve` once `reply` works. |
44
+
45
+ ## Step 2 — Build (exact shape)
46
+
47
+ Write **one** file. This is the whole program:
48
+
49
+ ```ruby
50
+ require "insika"
51
+
52
+ assistant = Insika.agent("assistant") do
53
+ provider :deepseek # ← the provider slug from {{MODELS_URL}}
54
+ model "deepseek-v4-flash" # ← a model id from {{MODELS_URL}}
55
+ instructions "You are a concise, friendly assistant. Answer briefly."
56
+ end
57
+
58
+ puts assistant.reply(ARGV.join(" ").empty? ? "hi, what can you do?" : ARGV.join(" "))
59
+ ```
60
+
61
+ Notes that are RULES, not options:
62
+
63
+ - The block is **config that generates data** — `assistant.to_pack` is a plain
64
+ provisioning pack. Do not reach past the DSL into internal classes; everything a
65
+ first agent needs is a DSL method (`model`, `provider`, `instructions`, `tools`,
66
+ `skill`, `memory`, `temperature`, …).
67
+ - The provider key is read from the environment by name (`<PROVIDER>_API_KEY`). Do
68
+ **not** write it into the file, the DSL, or a committed config.
69
+
70
+ ## Step 3 — Run it
71
+
72
+ ```bash
73
+ bundle install
74
+ DEEPSEEK_API_KEY=sk-... ruby my_agent.rb "hi, what can you do?"
75
+ ```
76
+
77
+ Once that prints a reply, turning the same agent into a server is one line — swap
78
+ `reply` for `serve`:
79
+
80
+ ```ruby
81
+ assistant.serve # control UI at /studio + drop-in POST /v1/responses on :9292
82
+ ```
83
+
84
+ Over the drop-in API the `model` field is the **agent id**:
85
+
86
+ ```bash
87
+ curl -N http://localhost:9292/v1/responses \
88
+ -H "Authorization: Bearer local-demo" \
89
+ -H "Content-Type: application/json" \
90
+ -d '{"model":"assistant","user":"chat-1","stream":true,"input":"hi"}'
91
+ ```
92
+
93
+ ## Step 4 — Self-check before you report done
94
+
95
+ - [ ] `ruby -v` is ≥ 3.3.
96
+ - [ ] The `model`/`provider` you wrote appear in **{{MODELS_URL}}**.
97
+ - [ ] The provider key is exported in the environment, **not** written into any file.
98
+ - [ ] Running the file printed a real model reply (not an auth/model error).
99
+ - [ ] Exactly one agent, one file. No tools, skills, workflows, or extra agents added
100
+ "to test things".
101
+
102
+ If any box is unchecked, fix that one thing — do not add scope to work around it.
103
+
104
+ ---
105
+
106
+ ## Hard constraints (known failure modes — never do these)
107
+
108
+ - **Never invent, guess, or hard-code an API key.** If no key is available, ask. A
109
+ placeholder like `sk-xxxx` is not a fix — it produces a confusing auth failure.
110
+ - **Never guess a model id.** Use only ids from {{MODELS_URL}}. A wrong id fails at
111
+ the provider, not in the engine, and wastes the developer's time.
112
+ - **Do not create a workflow, tool, or second agent just to test the first one.** The
113
+ test is `reply` / one `curl`. Extra machinery is scope you were not asked for.
114
+ - **Do not bypass the DSL / config-over-code.** If something seems to need a private
115
+ class, it is almost certainly a DSL method you have not used yet — check the docs at
116
+ {{DOCS_URL}} before reaching deeper.
117
+ - **Keep secrets in the environment.** No keys in source, in the pack, or in commits.
118
+
119
+ ## Where to go next
120
+
121
+ - **{{MODELS_URL}}** — live list of configured models, the default, and `thinking` levels.
122
+ - **{{DOCS_URL}}** — the docs index (README, running locally, deploy, guardrails, …), raw markdown.
123
+ - Add capability only after the first reply works: `tools`, `skill`, `memory`,
124
+ `temperature`, `data_tool` — each is a DSL method documented in the README.
125
+
126
+ This is `rails new` reimplemented as a prompt — and you are the generator.
data/docs/operate.md ADDED
@@ -0,0 +1,12 @@
1
+ ---
2
+ title: Operate & prove it
3
+ nav_order: 5
4
+ has_children: true
5
+ permalink: /operate/
6
+ ---
7
+
8
+ # Operate & prove it
9
+
10
+ Turns as traces and metrics, the engine's measured overhead, how to load-test it
11
+ yourself, the cases that grade an agent, and reading a live agent's own traffic back as
12
+ a report of what broke.
data/docs/ship.md ADDED
@@ -0,0 +1,10 @@
1
+ ---
2
+ title: Ship it
3
+ nav_order: 4
4
+ has_children: true
5
+ permalink: /ship/
6
+ ---
7
+
8
+ # Ship it
9
+
10
+ What stands between your agent and the open internet, and how to put it on a server.
@@ -0,0 +1,10 @@
1
+ ---
2
+ title: Understand the idea
3
+ nav_order: 2
4
+ has_children: true
5
+ permalink: /understand/
6
+ ---
7
+
8
+ # Understand the idea
9
+
10
+ What Insika is for, what it replaces, and how a turn actually runs end to end.