insika 0.0.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +361 -0
- data/LICENSE +21 -0
- data/README.md +136 -2
- data/bin/insika +366 -0
- data/docs/AGENTS.md +618 -0
- data/docs/ARCHITECTURE.md +333 -0
- data/docs/BENCHMARK.md +114 -0
- data/docs/CHANNELS.md +453 -0
- data/docs/CONTEXT.md +117 -0
- data/docs/DEPLOY.md +354 -0
- data/docs/EMBEDDING.md +198 -0
- data/docs/EVALS.md +273 -0
- data/docs/LOADTEST.md +232 -0
- data/docs/OBSERVABILITY.md +374 -0
- data/docs/PLUGINS.md +211 -0
- data/docs/REFINEMENT.md +477 -0
- data/docs/RELEASING.md +70 -0
- data/docs/RUNNING-LOCAL.md +153 -0
- data/docs/SANDBOX.md +114 -0
- data/docs/SECURITY.md +375 -0
- data/docs/SKILLS.md +284 -0
- data/docs/TOOLS.md +302 -0
- data/docs/WHY.md +137 -0
- data/docs/WORKFLOWS.md +225 -0
- data/docs/build.md +14 -0
- data/docs/index.md +68 -0
- data/docs/onboarding/start.md +126 -0
- data/docs/operate.md +12 -0
- data/docs/ship.md +10 -0
- data/docs/understand.md +10 -0
- data/lib/insika/agent_file_store.rb +125 -0
- data/lib/insika/agent_profile.rb +255 -0
- data/lib/insika/alert_dispatcher.rb +139 -0
- data/lib/insika/allowlist.rb +28 -0
- data/lib/insika/baseline_store.rb +74 -0
- data/lib/insika/budget_ledger.rb +135 -0
- data/lib/insika/capability/resolved_tool.rb +34 -0
- data/lib/insika/capability_registry.rb +112 -0
- data/lib/insika/channel_delivery.rb +153 -0
- data/lib/insika/channel_registry.rb +30 -0
- data/lib/insika/channels/relay.rb +178 -0
- data/lib/insika/channels/web/widget.js +283 -0
- data/lib/insika/channels/web.rb +211 -0
- data/lib/insika/channels/webhook.rb +58 -0
- data/lib/insika/chat_builder.rb +303 -0
- data/lib/insika/checkpoint.rb +13 -0
- data/lib/insika/checkpoint_store.rb +153 -0
- data/lib/insika/circuit_state.rb +114 -0
- data/lib/insika/coercion.rb +58 -0
- data/lib/insika/command.rb +32 -0
- data/lib/insika/command_bus.rb +39 -0
- data/lib/insika/commands/agent_payload.rb +43 -0
- data/lib/insika/commands/approve_action.rb +46 -0
- data/lib/insika/commands/cancel_task.rb +33 -0
- data/lib/insika/commands/create_agent.rb +54 -0
- data/lib/insika/commands/create_session.rb +67 -0
- data/lib/insika/commands/delete_agent.rb +33 -0
- data/lib/insika/commands/delete_agent_file.rb +50 -0
- data/lib/insika/commands/delete_data_tool.rb +33 -0
- data/lib/insika/commands/delete_llm_provider.rb +36 -0
- data/lib/insika/commands/delete_mcp.rb +30 -0
- data/lib/insika/commands/delete_skill.rb +43 -0
- data/lib/insika/commands/delete_system_file.rb +29 -0
- data/lib/insika/commands/gate_refinement.rb +245 -0
- data/lib/insika/commands/import_mcp_tools.rb +48 -0
- data/lib/insika/commands/import_tools.rb +81 -0
- data/lib/insika/commands/issue_tenant_token.rb +41 -0
- data/lib/insika/commands/memory_add_note.rb +32 -0
- data/lib/insika/commands/memory_forget_fact.rb +32 -0
- data/lib/insika/commands/memory_put_fact.rb +35 -0
- data/lib/insika/commands/pause_task.rb +29 -0
- data/lib/insika/commands/resolve_refinement.rb +126 -0
- data/lib/insika/commands/restore_agent_file.rb +36 -0
- data/lib/insika/commands/restore_data_tool.rb +34 -0
- data/lib/insika/commands/restore_system_file.rb +31 -0
- data/lib/insika/commands/resume_task.rb +85 -0
- data/lib/insika/commands/revoke_token.rb +39 -0
- data/lib/insika/commands/rotate_tenant_token.rb +43 -0
- data/lib/insika/commands/run_refinement.rb +133 -0
- data/lib/insika/commands/send_message.rb +150 -0
- data/lib/insika/commands/set_agent_tools.rb +39 -0
- data/lib/insika/commands/set_skill_agents.rb +112 -0
- data/lib/insika/commands/trigger_workflow.rb +80 -0
- data/lib/insika/commands/update_agent.rb +49 -0
- data/lib/insika/commands/update_settings.rb +33 -0
- data/lib/insika/commands/upsert_llm_provider.rb +34 -0
- data/lib/insika/commands/upsert_mcp.rb +32 -0
- data/lib/insika/commands/write_agent_file.rb +57 -0
- data/lib/insika/commands/write_data_tool.rb +43 -0
- data/lib/insika/commands/write_golden.rb +58 -0
- data/lib/insika/commands/write_skill.rb +60 -0
- data/lib/insika/commands/write_system_file.rb +31 -0
- data/lib/insika/config_store.rb +89 -0
- data/lib/insika/context/builder.rb +166 -0
- data/lib/insika/context/catalog_provider.rb +23 -0
- data/lib/insika/context/fragment.rb +43 -0
- data/lib/insika/context/priority.rb +30 -0
- data/lib/insika/context/provider.rb +19 -0
- data/lib/insika/context/providers/memory.rb +60 -0
- data/lib/insika/context/providers/prompt.rb +105 -0
- data/lib/insika/context/providers/request.rb +32 -0
- data/lib/insika/context/providers/session.rb +123 -0
- data/lib/insika/context/providers/skill.rb +24 -0
- data/lib/insika/context/providers/skill_trigger.rb +128 -0
- data/lib/insika/context/providers/tool_search.rb +20 -0
- data/lib/insika/context_trace_store.rb +92 -0
- data/lib/insika/delegation_store.rb +153 -0
- data/lib/insika/doctor.rb +539 -0
- data/lib/insika/dsl/definition.rb +55 -0
- data/lib/insika/dsl/runtime.rb +382 -0
- data/lib/insika/dsl/server_boot.rb +98 -0
- data/lib/insika/dsl/system.rb +93 -0
- data/lib/insika/dsl/workflow_adapter.rb +59 -0
- data/lib/insika/dsl.rb +364 -0
- data/lib/insika/edge_limiter.rb +268 -0
- data/lib/insika/egress_guard.rb +75 -0
- data/lib/insika/env_schema.rb +249 -0
- data/lib/insika/errors.rb +201 -0
- data/lib/insika/evals/assertions.rb +247 -0
- data/lib/insika/evals/baseline.rb +69 -0
- data/lib/insika/evals/golden.rb +172 -0
- data/lib/insika/evals/judge.rb +225 -0
- data/lib/insika/evals/pairwise.rb +178 -0
- data/lib/insika/evals/report.rb +115 -0
- data/lib/insika/evals/runner.rb +141 -0
- data/lib/insika/evals/transport.rb +178 -0
- data/lib/insika/event.rb +18 -0
- data/lib/insika/event_stream.rb +132 -0
- data/lib/insika/executor.rb +1995 -0
- data/lib/insika/frontmatter.rb +42 -0
- data/lib/insika/golden_store.rb +145 -0
- data/lib/insika/hooks.rb +48 -0
- data/lib/insika/http_client.rb +63 -0
- data/lib/insika/inbound_log.rb +84 -0
- data/lib/insika/llm_configurator.rb +99 -0
- data/lib/insika/llm_provider_store.rb +83 -0
- data/lib/insika/loop_detector.rb +143 -0
- data/lib/insika/mcp_http_client.rb +67 -0
- data/lib/insika/mcp_store.rb +115 -0
- data/lib/insika/mcp_tool_ingestor.rb +143 -0
- data/lib/insika/memory_store.rb +93 -0
- data/lib/insika/message_origin.rb +76 -0
- data/lib/insika/middleware.rb +36 -0
- data/lib/insika/model_policy.rb +52 -0
- data/lib/insika/model_resolver.rb +176 -0
- data/lib/insika/model_selection.rb +115 -0
- data/lib/insika/onboarding.rb +208 -0
- data/lib/insika/outbox_store.rb +166 -0
- data/lib/insika/overlay_tool_registry.rb +102 -0
- data/lib/insika/pack.rb +102 -0
- data/lib/insika/pack_importer.rb +123 -0
- data/lib/insika/pending_action_store.rb +120 -0
- data/lib/insika/plugin/loader.rb +356 -0
- data/lib/insika/plugin.rb +35 -0
- data/lib/insika/policy/engine.rb +83 -0
- data/lib/insika/policy/policy.rb +120 -0
- data/lib/insika/policy_registry.rb +23 -0
- data/lib/insika/profile_source.rb +143 -0
- data/lib/insika/prompt_catalog.rb +61 -0
- data/lib/insika/provider_error_classifier.rb +160 -0
- data/lib/insika/queue_policy.rb +167 -0
- data/lib/insika/recovery.rb +168 -0
- data/lib/insika/refinement/candidate.rb +159 -0
- data/lib/insika/refinement/evidence_collector.rb +371 -0
- data/lib/insika/refinement/gate.rb +234 -0
- data/lib/insika/refinement/panel.rb +222 -0
- data/lib/insika/refinement/proposer.rb +262 -0
- data/lib/insika/refinement_store.rb +295 -0
- data/lib/insika/registry.rb +59 -0
- data/lib/insika/reliability.rb +185 -0
- data/lib/insika/safety/config.rb +109 -0
- data/lib/insika/safety/detectors.rb +176 -0
- data/lib/insika/safety/factory.rb +102 -0
- data/lib/insika/safety/input_guardrail.rb +102 -0
- data/lib/insika/safety/moderator.rb +94 -0
- data/lib/insika/safety/output_filter.rb +79 -0
- data/lib/insika/safety/output_validator.rb +101 -0
- data/lib/insika/safety/safe_responses.rb +47 -0
- data/lib/insika/sandbox/boundary.rb +93 -0
- data/lib/insika/sandbox/docker.rb +74 -0
- data/lib/insika/sandbox/local.rb +33 -0
- data/lib/insika/sandbox/runner.rb +80 -0
- data/lib/insika/sandbox.rb +85 -0
- data/lib/insika/schema_guard.rb +147 -0
- data/lib/insika/secret_masking.rb +34 -0
- data/lib/insika/server/a2a/agent_card.rb +27 -0
- data/lib/insika/server/a2a/app.rb +112 -0
- data/lib/insika/server/a2a/client.rb +101 -0
- data/lib/insika/server/a2a/errors.rb +32 -0
- data/lib/insika/server/a2a/http.rb +42 -0
- data/lib/insika/server/a2a/message.rb +27 -0
- data/lib/insika/server/a2a/protocol.rb +45 -0
- data/lib/insika/server/a2a/remotes.rb +25 -0
- data/lib/insika/server/a2a/task_projection.rb +40 -0
- data/lib/insika/server/app.rb +1022 -0
- data/lib/insika/server/boot.rb +119 -0
- data/lib/insika/server/rack_app.rb +118 -0
- data/lib/insika/server/responses.rb +165 -0
- data/lib/insika/server/sse_body.rb +96 -0
- data/lib/insika/server/tenant_auth.rb +61 -0
- data/lib/insika/session_actor.rb +162 -0
- data/lib/insika/session_store.rb +143 -0
- data/lib/insika/settings_store.rb +154 -0
- data/lib/insika/shutdown.rb +125 -0
- data/lib/insika/skill_catalog.rb +220 -0
- data/lib/insika/skill_store.rb +127 -0
- data/lib/insika/steer_injector.rb +110 -0
- data/lib/insika/store.rb +52 -0
- data/lib/insika/stores/memory.rb +123 -0
- data/lib/insika/stores/sqlite.rb +183 -0
- data/lib/insika/studio/app.rb +1693 -0
- data/lib/insika/studio/assets/dist/application.css +1 -0
- data/lib/insika/studio/assets/dist/application.js +70 -0
- data/lib/insika/studio/forms.rb +335 -0
- data/lib/insika/studio/nav_icons.rb +31 -0
- data/lib/insika/studio/views/_message.erb +44 -0
- data/lib/insika/studio/views/agent_detail.erb +285 -0
- data/lib/insika/studio/views/agents.erb +63 -0
- data/lib/insika/studio/views/approvals.erb +41 -0
- data/lib/insika/studio/views/chats.erb +34 -0
- data/lib/insika/studio/views/evals.erb +83 -0
- data/lib/insika/studio/views/home.erb +72 -0
- data/lib/insika/studio/views/layout.erb +94 -0
- data/lib/insika/studio/views/login.erb +17 -0
- data/lib/insika/studio/views/mcp.erb +91 -0
- data/lib/insika/studio/views/not_found.erb +5 -0
- data/lib/insika/studio/views/playground.erb +47 -0
- data/lib/insika/studio/views/refinement.erb +234 -0
- data/lib/insika/studio/views/session.erb +137 -0
- data/lib/insika/studio/views/settings.erb +168 -0
- data/lib/insika/studio/views/skills.erb +141 -0
- data/lib/insika/studio/views/system_files.erb +65 -0
- data/lib/insika/studio/views/task.erb +105 -0
- data/lib/insika/studio/views/tasks.erb +33 -0
- data/lib/insika/studio/views/tool_edit.erb +107 -0
- data/lib/insika/studio/views/tools.erb +89 -0
- data/lib/insika/subagent_graph.rb +96 -0
- data/lib/insika/system_file_store.rb +96 -0
- data/lib/insika/task_actor.rb +128 -0
- data/lib/insika/task_store.rb +250 -0
- data/lib/insika/telemetry/pricing.rb +104 -0
- data/lib/insika/telemetry/recorder.rb +228 -0
- data/lib/insika/telemetry.rb +127 -0
- data/lib/insika/testing/store_contract.rb +270 -0
- data/lib/insika/tick.rb +122 -0
- data/lib/insika/token_estimator.rb +16 -0
- data/lib/insika/token_store.rb +168 -0
- data/lib/insika/tool_assembly.rb +140 -0
- data/lib/insika/tool_catalog.rb +89 -0
- data/lib/insika/tool_definition.rb +518 -0
- data/lib/insika/tool_envelope.rb +140 -0
- data/lib/insika/tool_manifest.rb +218 -0
- data/lib/insika/tool_output_compressor.rb +100 -0
- data/lib/insika/tool_registry.rb +21 -0
- data/lib/insika/tool_store.rb +135 -0
- data/lib/insika/tool_trace_store.rb +92 -0
- data/lib/insika/tools/a2a_remote.rb +48 -0
- data/lib/insika/tools/agent_enum.rb +68 -0
- data/lib/insika/tools/concurrency.rb +54 -0
- data/lib/insika/tools/data_defined_tool.rb +219 -0
- data/lib/insika/tools/load_skill.rb +99 -0
- data/lib/insika/tools/remember.rb +53 -0
- data/lib/insika/tools/stuck_signal.rb +44 -0
- data/lib/insika/tools/subagent.rb +75 -0
- data/lib/insika/tools/subagents.rb +77 -0
- data/lib/insika/tools/tool_search.rb +94 -0
- data/lib/insika/turn_output.rb +139 -0
- data/lib/insika/turn_state.rb +162 -0
- data/lib/insika/turn_timing.rb +56 -0
- data/lib/insika/usage_ledger.rb +47 -0
- data/lib/insika/version.rb +3 -1
- data/lib/insika/wiring/graph.rb +249 -0
- data/lib/insika/workflow.rb +185 -0
- data/lib/insika/workflow_registry.rb +33 -0
- data/lib/insika.rb +220 -4
- metadata +412 -8
data/docs/WHY.md
ADDED
|
@@ -0,0 +1,137 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Why Insika
|
|
3
|
+
parent: Understand the idea
|
|
4
|
+
nav_order: 1
|
|
5
|
+
permalink: /why/
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Why Insika
|
|
9
|
+
|
|
10
|
+
**Ruby is ready for production AI.** The narrative that you must reach for Python
|
|
11
|
+
to ship serious LLM applications is out of date. [RubyLLM](https://rubyllm.com)
|
|
12
|
+
gave the language a first-class, provider-agnostic LLM client — chat, tools,
|
|
13
|
+
streaming, embeddings, moderation — a genuinely valuable foundation that closed
|
|
14
|
+
the "Ruby can't do AI" gap. Insika is the next layer up: it takes those
|
|
15
|
+
primitives and makes an agent *dependable in production*.
|
|
16
|
+
|
|
17
|
+
Because **an agent without a harness is just a chat loop.** The model does the
|
|
18
|
+
reasoning, but everything that makes an agent survive contact with real traffic —
|
|
19
|
+
durable state, tools, guardrails, evals, operations — lives in the engine around
|
|
20
|
+
it. Those harnesses exist, mature, for other ecosystems. Ruby teams have mostly
|
|
21
|
+
been told to glue libraries together. They shouldn't have to.
|
|
22
|
+
|
|
23
|
+
Insika is an **agent runtime for Ruby**: the turn pipeline, the operational
|
|
24
|
+
surface, and the safety layer, in one deployable piece — behind an
|
|
25
|
+
OpenAI-Responses-compatible API. It stands on RubyLLM for the provider layer and
|
|
26
|
+
adds everything between a chat call and a production agent.
|
|
27
|
+
|
|
28
|
+
And it does it *ergonomically*. A complete, running agent is a few lines of Ruby —
|
|
29
|
+
the same speed-to-first-agent story Python teams tell, without leaving Ruby:
|
|
30
|
+
|
|
31
|
+
```ruby
|
|
32
|
+
require "insika"
|
|
33
|
+
|
|
34
|
+
agent = Insika.agent("assistant") do
|
|
35
|
+
model "deepseek-v4-flash"
|
|
36
|
+
provider :deepseek
|
|
37
|
+
instructions "You are a concise, friendly assistant."
|
|
38
|
+
end
|
|
39
|
+
|
|
40
|
+
puts agent.reply("hi, what can you do?") # one turn, in-process
|
|
41
|
+
agent.serve # ...or a full server: /studio + /v1/responses
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Every capability below is reachable the same way — see [`examples/`](https://github.com/guizaols/insika/tree/main/examples/),
|
|
45
|
+
one small runnable project per capability.
|
|
46
|
+
|
|
47
|
+
## Four ways to run an agent
|
|
48
|
+
|
|
49
|
+
There are, roughly, four ways to put an LLM agent into production. They're not
|
|
50
|
+
wrong — they're different amounts of "build it yourself":
|
|
51
|
+
|
|
52
|
+
- **Raw SDK loop.** Call the provider SDK directly and hand-roll the
|
|
53
|
+
reason→act→observe loop. Minimal to start; everything operational is on you.
|
|
54
|
+
- **Assemble a framework.** Compose library pieces — a tool abstraction here, a
|
|
55
|
+
memory store there, your own persistence and moderation. Flexible, but you own
|
|
56
|
+
the integration and the gaps between the parts.
|
|
57
|
+
- **Hosted agent gateway.** A managed service runs the agent behind an API. You
|
|
58
|
+
get operations for free, but your agents, prompts, and conversation data live in
|
|
59
|
+
someone else's system.
|
|
60
|
+
- **An agent runtime (this project).** The loop, durability, tools, safety, and an
|
|
61
|
+
operations UI as one thing you deploy in your own infrastructure.
|
|
62
|
+
|
|
63
|
+
How they compare on what production actually demands:
|
|
64
|
+
|
|
65
|
+
| | Raw SDK loop | Assemble a framework | Hosted gateway | **Insika (runtime)** |
|
|
66
|
+
|---|:---:|:---:|:---:|:---:|
|
|
67
|
+
| Get a first agent talking | ✅ | ⚠️ | ✅ | ✅ |
|
|
68
|
+
| Add/change a tool without a redeploy | ❌ | ❌ | ⚠️ | ✅ |
|
|
69
|
+
| Durable, resumable turns (crash mid-conversation) | ❌ | ⚠️ | ✅ | ✅ |
|
|
70
|
+
| Content-safety guardrails built in | ❌ | ⚠️ | ⚠️ | ✅ |
|
|
71
|
+
| Evals / regression gating as a primitive | ❌ | ⚠️ | ⚠️ | ✅ |
|
|
72
|
+
| Operations UI included | ❌ | ❌ | ✅ | ✅ |
|
|
73
|
+
| Runs in your infrastructure (you own the data) | ✅ | ✅ | ❌ | ✅ |
|
|
74
|
+
| One container, no orchestrator to start | ✅ | ⚠️ | — | ✅ |
|
|
75
|
+
|
|
76
|
+
✅ built in · ⚠️ possible, but you build/integrate it · ❌ not addressed · — n/a
|
|
77
|
+
|
|
78
|
+
## What we do differently
|
|
79
|
+
|
|
80
|
+
- **Tools are data, not code.** Define a tool with a JSON Schema manifest and a
|
|
81
|
+
declarative binding, and the running agent picks it up — no rebuild, no redeploy.
|
|
82
|
+
Import whole toolsets from an MCP server or a manifest at runtime. In most setups
|
|
83
|
+
a new tool is a code change; here it's a row.
|
|
84
|
+
- **Durability is the default.** Every turn checkpoints; sessions and tasks recover
|
|
85
|
+
across restarts; continuous replication is one env var away, with a documented
|
|
86
|
+
restore drill. If the box dies mid-conversation, the conversation doesn't.
|
|
87
|
+
- **Safety is built in, not bolted on.** Input guardrails (prompt-injection and
|
|
88
|
+
abuse handling that fails gracefully *without* burning a model turn), content
|
|
89
|
+
moderation, PII/secret redaction in the output stream, and a post-turn validator
|
|
90
|
+
— configured, not hand-rolled, and opt-in per agent.
|
|
91
|
+
- **Evals are a primitive.** Golden conversations, LLM-as-judge, and baseline
|
|
92
|
+
gating run against the *same* API your users hit — so a prompt or model change
|
|
93
|
+
that regresses behavior fails loudly before it ships.
|
|
94
|
+
- **An operations UI is included.** Studio lets an operator manage agents, tools,
|
|
95
|
+
approvals, and traces — pause a task, approve a sensitive action, inspect a
|
|
96
|
+
tool-call trace — without writing code. Headless when you want it, operable when
|
|
97
|
+
you need it.
|
|
98
|
+
- **One container, one file.** SQLite in WAL mode with streaming replication. No
|
|
99
|
+
queue cluster, no orchestrator, no managed database required to start — and it
|
|
100
|
+
runs a real production workload today.
|
|
101
|
+
- **Fibers, not a thread per request.** An LLM turn is almost all *waiting* on the
|
|
102
|
+
provider. Insika runs on the [Async](https://github.com/socketry/async) fiber
|
|
103
|
+
scheduler and serves under [Falcon](https://github.com/socketry/falcon), so one
|
|
104
|
+
process handles thousands of concurrent turns on a handful of connections instead
|
|
105
|
+
of pinning a heavyweight thread per call. This is exactly the model
|
|
106
|
+
[RubyLLM's async guide](https://rubyllm.com/async/) recommends — *"Falcon is a
|
|
107
|
+
Ruby application server built on fibers; with Falcon, async just works"* — and it
|
|
108
|
+
composes for free, because RubyLLM's HTTP cooperates with Ruby's fiber scheduler.
|
|
109
|
+
It's a big part of why one box goes so far.
|
|
110
|
+
- **The engine gets out of the way.** On top of that concurrency, all the machinery
|
|
111
|
+
above adds **well under a millisecond of overhead per turn**; a single process
|
|
112
|
+
sustains thousands of turns per second of pure engine work. A neutral,
|
|
113
|
+
provider-free benchmark in the repo reproduces it with one command — see
|
|
114
|
+
[BENCHMARK.md](BENCHMARK.md). The rest of a turn's latency is the model, not us.
|
|
115
|
+
*(That number is engine overhead only — end-to-end latency is provider-bound and
|
|
116
|
+
not claimed here.)*
|
|
117
|
+
|
|
118
|
+
## When a lighter tool is the right call
|
|
119
|
+
|
|
120
|
+
Insika is for when an agent has to run **unattended, durably, and safely, in
|
|
121
|
+
production**. If that's not you yet, reach for less:
|
|
122
|
+
|
|
123
|
+
- If you only need LLM calls with tools in Ruby and nothing operational around
|
|
124
|
+
them, **[RubyLLM](https://rubyllm.com) alone is excellent** — and it's exactly
|
|
125
|
+
what Insika builds on. Reach for Insika when "nothing operational around them"
|
|
126
|
+
stops being true: when the agent has to run unattended, recover from crashes,
|
|
127
|
+
enforce safety, and be operable by someone who isn't you.
|
|
128
|
+
- If your team is committed to another language ecosystem, use a runtime native to
|
|
129
|
+
it — the ideas here travel, the code doesn't.
|
|
130
|
+
- If you want a minimal terminal coding assistant rather than a product platform,
|
|
131
|
+
a single-purpose CLI agent will be simpler.
|
|
132
|
+
|
|
133
|
+
## See also
|
|
134
|
+
|
|
135
|
+
- [README](https://github.com/guizaols/insika#readme) — quickstart and the drop-in `/v1/responses` API.
|
|
136
|
+
- [BENCHMARK.md](BENCHMARK.md) — the neutral, reproducible engine benchmark.
|
|
137
|
+
- [OBSERVABILITY.md](OBSERVABILITY.md) — OpenTelemetry tracing (opt-in).
|
data/docs/WORKFLOWS.md
ADDED
|
@@ -0,0 +1,225 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Workflows
|
|
3
|
+
parent: Build an agent
|
|
4
|
+
nav_order: 5
|
|
5
|
+
permalink: /workflows/
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Workflows
|
|
9
|
+
|
|
10
|
+
A single agent's tool-loop is one shape of work: the model decides, calls a tool,
|
|
11
|
+
looks at the result, decides again. Plenty of real work has a shape you already
|
|
12
|
+
know — draft then edit, classify then answer, three reviewers then a summary,
|
|
13
|
+
try until it passes. For those, the choice of "what happens next" belongs in
|
|
14
|
+
**Ruby**, not in a prompt.
|
|
15
|
+
|
|
16
|
+
That is a **workflow**: deterministic orchestration around agent turns.
|
|
17
|
+
|
|
18
|
+
## The one decision that matters
|
|
19
|
+
|
|
20
|
+
**Who chooses the next step — your code, or the model?**
|
|
21
|
+
|
|
22
|
+
| | **Workflow** | **Delegation (`subagents`)** |
|
|
23
|
+
|---|---|---|
|
|
24
|
+
| Chooses the next step | your Ruby | the model |
|
|
25
|
+
| Cost of deciding | zero | a model call |
|
|
26
|
+
| Repeatable | yes — same path every run | no |
|
|
27
|
+
| Reach for it when | the shape is known in advance | the work depends on what the user said |
|
|
28
|
+
| Surface | `workflow` in a system + `POST /v1/workflows/:name` | `subagents` on a profile + `spawn_subagent(s)` |
|
|
29
|
+
|
|
30
|
+
They compose: a workflow step can `ask` an agent that itself delegates. Start
|
|
31
|
+
with a workflow and hand the choice to the model only where the choice is
|
|
32
|
+
genuinely open — see [Agents](AGENTS.md#delegation-subagents) for delegation.
|
|
33
|
+
|
|
34
|
+
## Declaring one
|
|
35
|
+
|
|
36
|
+
Workflows live in a **system** (`Insika.system`), because their steps address
|
|
37
|
+
agents by id and those agents must resolve in the same runtime:
|
|
38
|
+
|
|
39
|
+
```ruby
|
|
40
|
+
newsroom = Insika.system do
|
|
41
|
+
provider :deepseek
|
|
42
|
+
|
|
43
|
+
agent("writer") { model "deepseek-v4-flash"; instructions "Write ONE paragraph." }
|
|
44
|
+
agent("editor") { model "deepseek-v4-flash"; instructions "Rewrite as ONE sentence." }
|
|
45
|
+
|
|
46
|
+
workflow "publish",
|
|
47
|
+
description: "Draft a paragraph, then tighten it.",
|
|
48
|
+
input: { type: "object", properties: { topic: { type: "string" } }, required: ["topic"] },
|
|
49
|
+
output: { type: "object", properties: { headline: { type: "string" } } } do |input, ctx|
|
|
50
|
+
draft = ctx.ask("writer", "Topic: #{input['topic']}")
|
|
51
|
+
{ "headline" => ctx.ask("editor", draft) }
|
|
52
|
+
end
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
newsroom.run("publish", input: { "topic" => "Ruby fibers" })
|
|
56
|
+
# => { "headline" => "…" }
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
What the block receives:
|
|
60
|
+
|
|
61
|
+
| | What it does |
|
|
62
|
+
|---|---|
|
|
63
|
+
| `input` | the validated input Hash (string keys) |
|
|
64
|
+
| `ctx.ask(agent, message)` | one turn against one agent of the system → its text |
|
|
65
|
+
| `ctx.gather(*blocks, max: 8)` | runs blocks **concurrently**, returns values **in order** |
|
|
66
|
+
| `ctx.context` / `ctx.tools` | the turn's ContextPackage and its resolved tool instances — escape hatches, rarely needed |
|
|
67
|
+
|
|
68
|
+
The workflow's return value **is** its output: whatever Hash (or value) you
|
|
69
|
+
return is what `run` returns, what `:workflow_completed` carries, and what the
|
|
70
|
+
`output` schema validates.
|
|
71
|
+
|
|
72
|
+
## What you get for free
|
|
73
|
+
|
|
74
|
+
- **A durable run.** The run id *is* a Task — checkpointed and recoverable, so a
|
|
75
|
+
crash mid-run does not lose the record. `run_id` comes back from every trigger.
|
|
76
|
+
- **Schemas at the edges.** `input:` is validated **synchronously**: a bad input
|
|
77
|
+
is refused with **no run created** (a `422` over HTTP). `output:` is validated
|
|
78
|
+
when the workflow returns; a violation fails the run at the `workflow_schema`
|
|
79
|
+
stage instead of handing a malformed result downstream. Both take a JSON Schema
|
|
80
|
+
Hash or any dry-schema-compatible validator.
|
|
81
|
+
- **Events.** `:workflow_started` and `:workflow_completed` (with the typed
|
|
82
|
+
output) on the turn's event stream, alongside every step's own turn events.
|
|
83
|
+
- **Every step is visible.** Each `ctx.ask` is its own turn — its own Task, its
|
|
84
|
+
own trace in the Studio. A workflow is not a black box.
|
|
85
|
+
|
|
86
|
+
## Over HTTP
|
|
87
|
+
|
|
88
|
+
Workflows are exposed when the deployment injects the registry (a `serve`d system
|
|
89
|
+
with at least one workflow does it automatically):
|
|
90
|
+
|
|
91
|
+
```bash
|
|
92
|
+
GET /v1/workflows # discovery: names, descriptions, I/O schemas
|
|
93
|
+
POST /v1/workflows/publish # 202 { run_id, task_id } — fire and observe
|
|
94
|
+
POST /v1/workflows/publish?stream=true # SSE of the run, ending at workflow_completed
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
```jsonc
|
|
98
|
+
// POST body
|
|
99
|
+
{ "agent": "writer", // the profile the run executes under
|
|
100
|
+
"input": { "topic": "Ruby fibers" },
|
|
101
|
+
"session_id": null } // optional
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
Observe an async run with `GET /v1/tasks/:run_id` or
|
|
105
|
+
`GET /v1/events?task_id=:run_id`. A bad input is a `422`; an unknown workflow or
|
|
106
|
+
agent is a `404`.
|
|
107
|
+
|
|
108
|
+
## The five patterns
|
|
109
|
+
|
|
110
|
+
Each has a runnable script in
|
|
111
|
+
[`examples/agentic-workflows/`](https://github.com/guizaols/insika/tree/main/examples/agentic-workflows/).
|
|
112
|
+
|
|
113
|
+
### Sequential (prompt chaining)
|
|
114
|
+
|
|
115
|
+
Each step's output feeds the next. The order is code, so it is identical on
|
|
116
|
+
every run.
|
|
117
|
+
|
|
118
|
+
```ruby
|
|
119
|
+
draft = ctx.ask("writer", input["topic"])
|
|
120
|
+
ctx.ask("editor", draft)
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
### Routing
|
|
124
|
+
|
|
125
|
+
One cheap classification turn, then a specialist with a short, focused prompt —
|
|
126
|
+
instead of one agent carrying every instruction it might ever need.
|
|
127
|
+
|
|
128
|
+
```ruby
|
|
129
|
+
label = ctx.ask("router", input["message"]).downcase # constrain the label set
|
|
130
|
+
lane = label.include?("billing") ? "billing" : "technical"
|
|
131
|
+
ctx.ask(lane, input["message"])
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
**Do the branching in Ruby, with a fallback.** A closed label set plus an `else`
|
|
135
|
+
means an unexpected answer degrades instead of picking a random branch.
|
|
136
|
+
|
|
137
|
+
### Parallel (fan-out / fan-in)
|
|
138
|
+
|
|
139
|
+
An LLM turn is almost all *waiting* on the provider, so independent turns overlap
|
|
140
|
+
on the fiber scheduler: wall-clock is the **slowest** branch, not the sum.
|
|
141
|
+
|
|
142
|
+
```ruby
|
|
143
|
+
security, performance = ctx.gather(-> { ctx.ask("security", code) },
|
|
144
|
+
-> { ctx.ask("performance", code) })
|
|
145
|
+
ctx.ask("lead", "security: #{security}\nperformance: #{performance}") # fan-in
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
`gather` returns values in declaration order and caps concurrency at `max:`
|
|
149
|
+
(default 8) — the cap matters, because N simultaneous calls hit provider rate
|
|
150
|
+
limits and per-agent token ceilings.
|
|
151
|
+
|
|
152
|
+
### Evaluator-optimizer
|
|
153
|
+
|
|
154
|
+
Produce, judge, revise. The judge's reason becomes the revision instruction.
|
|
155
|
+
|
|
156
|
+
```ruby
|
|
157
|
+
loop do
|
|
158
|
+
verdict = ctx.ask("critic", "Brief: #{brief}\nTagline: #{tagline}")
|
|
159
|
+
break if verdict.include?("PASS") || attempts >= MAX_ATTEMPTS # the cap is CODE
|
|
160
|
+
tagline = ctx.ask("writer", "Fix this: #{verdict}")
|
|
161
|
+
attempts += 1
|
|
162
|
+
end
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
**Two things must be code, not prompt:** the attempt cap (a loop that asks the
|
|
166
|
+
model when to stop can run forever) and the honest report of whether it actually
|
|
167
|
+
passed. Have the judge answer in a shape you can branch on (`VERDICT: PASS|FAIL`).
|
|
168
|
+
|
|
169
|
+
### Orchestrator-workers
|
|
170
|
+
|
|
171
|
+
The model chooses the workers. This is delegation, not a workflow — see
|
|
172
|
+
[Agents](AGENTS.md#delegation-subagents).
|
|
173
|
+
|
|
174
|
+
```ruby
|
|
175
|
+
agent "lead" do
|
|
176
|
+
instructions "You have no expertise yourself — always delegate, then report."
|
|
177
|
+
subagents "security", "performance"
|
|
178
|
+
end
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
## Gotchas
|
|
182
|
+
|
|
183
|
+
- **`ctx.ask` is stateless.** A step is a unit of work, not a conversation: no
|
|
184
|
+
session is threaded, so the agent sees only the message you pass. Pass what it
|
|
185
|
+
needs.
|
|
186
|
+
- **Nothing forces the model to delegate.** If a parent answers alone, the task
|
|
187
|
+
list shows only the parent — and it is the *prompt* that needs work: say
|
|
188
|
+
plainly that it has no expertise of its own. (The engine helps by naming the
|
|
189
|
+
spawnable agent ids in the tool contract, but it cannot make the model use them.)
|
|
190
|
+
- **The `agent:` of a run is the profile it executes under** — its policy,
|
|
191
|
+
limits, and guardrails. It is not "the agent that does the work"; the steps
|
|
192
|
+
pick those themselves.
|
|
193
|
+
- **A workflow turn assembles no chat.** Stage 5 is skipped: there is no
|
|
194
|
+
model call for the *workflow itself*, only for the turns its steps start.
|
|
195
|
+
One consequence worth naming: `queue_mode: "steer"`
|
|
196
|
+
([Agents](AGENTS.md#steer--the-message-arrives-while-the-turn-is-already-running))
|
|
197
|
+
cannot append into a running workflow — there is no chat to append to, and no tool
|
|
198
|
+
batch of the engine's to use as a boundary. A message that arrives while a workflow
|
|
199
|
+
run is in flight becomes the next turn on the session, which is what `followup`
|
|
200
|
+
does. Steering is refused at the door, not silently dropped.
|
|
201
|
+
- **A running workflow reaches no boundary until it returns.** The engine's stage
|
|
202
|
+
boundaries sit *around* the workflow call, not inside it, so `queue_mode:
|
|
203
|
+
"interrupt"` (and a plain `cancel_task`) is observed only once the workflow is done:
|
|
204
|
+
the steps run, and then the run is abandoned without publishing or persisting. If you
|
|
205
|
+
need a run to stop early, that decision belongs inside the workflow — it is ordinary
|
|
206
|
+
Ruby, so `return` on the condition you care about.
|
|
207
|
+
- **Braces bite in Ruby.** `agent "x" { … }` is a syntax error (`{` binds to the
|
|
208
|
+
argument); write `agent("x") { … }` or `agent "x" do … end`.
|
|
209
|
+
|
|
210
|
+
## Troubleshooting
|
|
211
|
+
|
|
212
|
+
| Symptom | Cause |
|
|
213
|
+
|---|---|
|
|
214
|
+
| `GET /v1/workflows` 404s | no workflow declared, so the routes are not exposed (parity) |
|
|
215
|
+
| `422` on trigger | the input violates `input:` — the message names the offending field |
|
|
216
|
+
| `404` on trigger | unknown workflow name, or the `agent:` does not exist |
|
|
217
|
+
| run fails at `workflow_schema` | your return value violates `output:` |
|
|
218
|
+
| a step "does nothing" | the agent id is not in this system — `ask` raises with the list of valid ids |
|
|
219
|
+
|
|
220
|
+
## See also
|
|
221
|
+
|
|
222
|
+
- [Agents](AGENTS.md) — profiles, and delegation via `subagents`.
|
|
223
|
+
- [Tools](TOOLS.md) — what a single agent can call inside one turn.
|
|
224
|
+
- [Architecture](ARCHITECTURE.md) — the turn pipeline a run goes through, and checkpointing.
|
|
225
|
+
- [`examples/agentic-workflows/`](https://github.com/guizaols/insika/tree/main/examples/agentic-workflows/) — the five patterns, runnable.
|
data/docs/build.md
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Build an agent
|
|
3
|
+
nav_order: 3
|
|
4
|
+
has_children: true
|
|
5
|
+
permalink: /build/
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Build an agent
|
|
9
|
+
|
|
10
|
+
An agent is data: a profile, its tools, its skills, and what fills its prompt. Start
|
|
11
|
+
with [Agents](AGENTS.md), then add capability one doc at a time — and when config is not
|
|
12
|
+
enough, [Plugins](PLUGINS.md) covers the two ways to extend the engine itself. Every one of these has a
|
|
13
|
+
runnable counterpart under
|
|
14
|
+
[`examples/`](https://github.com/guizaols/insika/tree/main/examples/).
|
data/docs/index.md
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Home
|
|
3
|
+
nav_order: 1
|
|
4
|
+
permalink: /
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Insika
|
|
8
|
+
{: .fs-9 }
|
|
9
|
+
|
|
10
|
+
Your agent is the idea. Insika is what holds it up in production.
|
|
11
|
+
{: .fs-6 .fw-300 }
|
|
12
|
+
|
|
13
|
+
[Build your first agent](RUNNING-LOCAL.md){: .btn .btn-primary .fs-5 .mb-4 .mb-md-0 .mr-2 }
|
|
14
|
+
[View on GitHub](https://github.com/guizaols/insika){: .btn .fs-5 .mb-4 .mb-md-0 }
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
*Insika* is Zulu for the pillar that carries a structure — the part nobody admires and
|
|
19
|
+
everything rests on. A turn that survives a crash, tools that cannot wander off, limits
|
|
20
|
+
that hold under load, and an API your clients already speak. Build the agent; the
|
|
21
|
+
scaffolding is already here.
|
|
22
|
+
|
|
23
|
+
Concretely: a Ruby runtime for **LLM agents in production** — a durable, resumable turn
|
|
24
|
+
pipeline behind an **OpenAI-Responses-compatible** HTTP API (`POST /v1/responses`), with
|
|
25
|
+
tools, skills, cross-session memory, per-agent policy, content-safety guardrails, and a
|
|
26
|
+
web control UI. Point an existing Responses client at it and serve many agents from one
|
|
27
|
+
deployment.
|
|
28
|
+
|
|
29
|
+
## Your first agent
|
|
30
|
+
|
|
31
|
+
Ruby `>= 3.3` and a provider key (the demo uses DeepSeek). The whole program:
|
|
32
|
+
|
|
33
|
+
```ruby
|
|
34
|
+
require "insika"
|
|
35
|
+
|
|
36
|
+
assistant = Insika.agent("assistant") do
|
|
37
|
+
model "deepseek-v4-flash"
|
|
38
|
+
provider :deepseek
|
|
39
|
+
instructions "You are Bia, a concise and friendly assistant. Answer briefly."
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
puts assistant.reply("hi, what can you do?") # one turn, in-process
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Swap `reply` for `serve` and the same agent is a server — the control UI at `/studio`
|
|
46
|
+
plus the drop-in API, on `:9292`. → [Running locally](RUNNING-LOCAL.md)
|
|
47
|
+
|
|
48
|
+
## Or let your coding agent read the docs
|
|
49
|
+
|
|
50
|
+
A running instance serves this same documentation as raw markdown, plus a
|
|
51
|
+
skill-structured prompt that walks a coding agent through building your first agent:
|
|
52
|
+
|
|
53
|
+
```
|
|
54
|
+
Read http://localhost:9292/start.md then help me build my first agent
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Alongside it, `GET /models.json` (configured providers, model ids, defaults — no
|
|
58
|
+
secrets), `GET /docs` and `GET /docs/<name>.md`. Public and on by default when you
|
|
59
|
+
`serve`; opt-in in production (`INSIKA_ONBOARDING=1`).
|
|
60
|
+
|
|
61
|
+
## Where to go next
|
|
62
|
+
|
|
63
|
+
- **[Understand the idea](understand.md)** — why a runtime rather than a DIY loop, and how a turn actually runs.
|
|
64
|
+
- **[Build an agent](build.md)** — agents, tools, skills, context, and the local loop.
|
|
65
|
+
- **[Ship it](ship.md)** — security, confined execution, deployment.
|
|
66
|
+
- **[Operate & prove it](operate.md)** — observability, the benchmark, load testing, evals, refinement.
|
|
67
|
+
|
|
68
|
+
Pre-release: APIs may still change and nothing is tagged yet. Licensed MIT.
|
|
@@ -0,0 +1,126 @@
|
|
|
1
|
+
# Insika — build your first agent
|
|
2
|
+
|
|
3
|
+
> **You are a coding agent** (Claude Code, Cursor, an IDE assistant, …) reading this
|
|
4
|
+
> on behalf of a developer who just pointed you at a running Insika. Treat this file
|
|
5
|
+
> as a **skill**: follow the steps in order, apply the RULES literally, and stop at the
|
|
6
|
+
> self-check. Do **not** improvise beyond it.
|
|
7
|
+
|
|
8
|
+
Insika is a Ruby runtime for LLM agents in production. Your job here is the smallest
|
|
9
|
+
possible one: get the developer a **first working agent**, defined in Ruby with the
|
|
10
|
+
public DSL (`Insika.agent { … }`), talking to a real model — nothing more.
|
|
11
|
+
|
|
12
|
+
The machine-readable list of models this engine already knows about is at
|
|
13
|
+
**{{MODELS_URL}}** — fetch it before you write any `model`/`provider` line. The full
|
|
14
|
+
docs are mirrored as raw markdown under **{{DOCS_URL}}**.
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## Step 0 — Gather context (do this first, silently)
|
|
19
|
+
|
|
20
|
+
RULES — verify, do not assume:
|
|
21
|
+
|
|
22
|
+
- **Ruby ≥ 3.3.** Run `ruby -v`. If lower, stop and tell the developer; do not try to
|
|
23
|
+
upgrade Ruby for them.
|
|
24
|
+
- **The gem/library must be loadable.** In a project that already depends on Insika,
|
|
25
|
+
`require "insika"` works. Otherwise add it to the `Gemfile` (or `bundle add`) — do
|
|
26
|
+
not vendor or copy source files.
|
|
27
|
+
- **A provider key must come from the developer, via the environment.** Look for one
|
|
28
|
+
already exported (e.g. `DEEPSEEK_API_KEY`, `OPENAI_API_KEY`). If none is set, **ask
|
|
29
|
+
the developer for the provider and confirm the env var is exported.** See the hard
|
|
30
|
+
constraint on keys below.
|
|
31
|
+
- **Fetch {{MODELS_URL}}.** It tells you which providers and model ids this engine is
|
|
32
|
+
configured for, the platform default, and the valid `thinking` levels. Use those
|
|
33
|
+
exact ids.
|
|
34
|
+
|
|
35
|
+
## Step 1 — Decide (RULES, not taste)
|
|
36
|
+
|
|
37
|
+
| Question | RULE |
|
|
38
|
+
|---|---|
|
|
39
|
+
| How many agents? | **Exactly one.** A first agent is a single `Insika.agent`. Resist adding a second. |
|
|
40
|
+
| Which model/provider? | Use an id from **{{MODELS_URL}}**. If it lists a `default`, use that. Never guess a model id. |
|
|
41
|
+
| Where does it live? | One new Ruby file (e.g. `my_agent.rb`), or extend `examples/quickstart.rb` if present. One file. |
|
|
42
|
+
| Tools / skills / memory? | **None yet.** Ship a plain conversational agent first; add capability only after it replies. |
|
|
43
|
+
| Reply or serve? | Start with `reply` (one in-process turn). Only switch to `serve` once `reply` works. |
|
|
44
|
+
|
|
45
|
+
## Step 2 — Build (exact shape)
|
|
46
|
+
|
|
47
|
+
Write **one** file. This is the whole program:
|
|
48
|
+
|
|
49
|
+
```ruby
|
|
50
|
+
require "insika"
|
|
51
|
+
|
|
52
|
+
assistant = Insika.agent("assistant") do
|
|
53
|
+
provider :deepseek # ← the provider slug from {{MODELS_URL}}
|
|
54
|
+
model "deepseek-v4-flash" # ← a model id from {{MODELS_URL}}
|
|
55
|
+
instructions "You are a concise, friendly assistant. Answer briefly."
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
puts assistant.reply(ARGV.join(" ").empty? ? "hi, what can you do?" : ARGV.join(" "))
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Notes that are RULES, not options:
|
|
62
|
+
|
|
63
|
+
- The block is **config that generates data** — `assistant.to_pack` is a plain
|
|
64
|
+
provisioning pack. Do not reach past the DSL into internal classes; everything a
|
|
65
|
+
first agent needs is a DSL method (`model`, `provider`, `instructions`, `tools`,
|
|
66
|
+
`skill`, `memory`, `temperature`, …).
|
|
67
|
+
- The provider key is read from the environment by name (`<PROVIDER>_API_KEY`). Do
|
|
68
|
+
**not** write it into the file, the DSL, or a committed config.
|
|
69
|
+
|
|
70
|
+
## Step 3 — Run it
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
bundle install
|
|
74
|
+
DEEPSEEK_API_KEY=sk-... ruby my_agent.rb "hi, what can you do?"
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Once that prints a reply, turning the same agent into a server is one line — swap
|
|
78
|
+
`reply` for `serve`:
|
|
79
|
+
|
|
80
|
+
```ruby
|
|
81
|
+
assistant.serve # control UI at /studio + drop-in POST /v1/responses on :9292
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
Over the drop-in API the `model` field is the **agent id**:
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
curl -N http://localhost:9292/v1/responses \
|
|
88
|
+
-H "Authorization: Bearer local-demo" \
|
|
89
|
+
-H "Content-Type: application/json" \
|
|
90
|
+
-d '{"model":"assistant","user":"chat-1","stream":true,"input":"hi"}'
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
## Step 4 — Self-check before you report done
|
|
94
|
+
|
|
95
|
+
- [ ] `ruby -v` is ≥ 3.3.
|
|
96
|
+
- [ ] The `model`/`provider` you wrote appear in **{{MODELS_URL}}**.
|
|
97
|
+
- [ ] The provider key is exported in the environment, **not** written into any file.
|
|
98
|
+
- [ ] Running the file printed a real model reply (not an auth/model error).
|
|
99
|
+
- [ ] Exactly one agent, one file. No tools, skills, workflows, or extra agents added
|
|
100
|
+
"to test things".
|
|
101
|
+
|
|
102
|
+
If any box is unchecked, fix that one thing — do not add scope to work around it.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## Hard constraints (known failure modes — never do these)
|
|
107
|
+
|
|
108
|
+
- **Never invent, guess, or hard-code an API key.** If no key is available, ask. A
|
|
109
|
+
placeholder like `sk-xxxx` is not a fix — it produces a confusing auth failure.
|
|
110
|
+
- **Never guess a model id.** Use only ids from {{MODELS_URL}}. A wrong id fails at
|
|
111
|
+
the provider, not in the engine, and wastes the developer's time.
|
|
112
|
+
- **Do not create a workflow, tool, or second agent just to test the first one.** The
|
|
113
|
+
test is `reply` / one `curl`. Extra machinery is scope you were not asked for.
|
|
114
|
+
- **Do not bypass the DSL / config-over-code.** If something seems to need a private
|
|
115
|
+
class, it is almost certainly a DSL method you have not used yet — check the docs at
|
|
116
|
+
{{DOCS_URL}} before reaching deeper.
|
|
117
|
+
- **Keep secrets in the environment.** No keys in source, in the pack, or in commits.
|
|
118
|
+
|
|
119
|
+
## Where to go next
|
|
120
|
+
|
|
121
|
+
- **{{MODELS_URL}}** — live list of configured models, the default, and `thinking` levels.
|
|
122
|
+
- **{{DOCS_URL}}** — the docs index (README, running locally, deploy, guardrails, …), raw markdown.
|
|
123
|
+
- Add capability only after the first reply works: `tools`, `skill`, `memory`,
|
|
124
|
+
`temperature`, `data_tool` — each is a DSL method documented in the README.
|
|
125
|
+
|
|
126
|
+
This is `rails new` reimplemented as a prompt — and you are the generator.
|
data/docs/operate.md
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Operate & prove it
|
|
3
|
+
nav_order: 5
|
|
4
|
+
has_children: true
|
|
5
|
+
permalink: /operate/
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Operate & prove it
|
|
9
|
+
|
|
10
|
+
Turns as traces and metrics, the engine's measured overhead, how to load-test it
|
|
11
|
+
yourself, the cases that grade an agent, and reading a live agent's own traffic back as
|
|
12
|
+
a report of what broke.
|
data/docs/ship.md
ADDED