silas 0.5.0 → 0.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +145 -0
  3. data/DEPLOY.md +111 -0
  4. data/README.md +87 -256
  5. data/app/helpers/silas/inbox/trace_helper.rb +29 -6
  6. data/app/models/concerns/silas/inbox/broadcastable.rb +12 -0
  7. data/app/views/layouts/silas/inbox.html.erb +89 -30
  8. data/app/views/silas/inbox/invocations/_approval_card.html.erb +11 -2
  9. data/app/views/silas/inbox/invocations/_invocation.html.erb +17 -5
  10. data/app/views/silas/inbox/sessions/_cost.html.erb +3 -1
  11. data/app/views/silas/inbox/sessions/_row.html.erb +14 -0
  12. data/app/views/silas/inbox/sessions/index.html.erb +16 -15
  13. data/app/views/silas/inbox/sessions/show.html.erb +10 -1
  14. data/app/views/silas/inbox/turns/_header.html.erb +31 -25
  15. data/app/views/silas/inbox/turns/_turn.html.erb +5 -3
  16. data/docs/agents.md +81 -0
  17. data/docs/budgets.md +67 -0
  18. data/docs/cancellation.md +41 -0
  19. data/docs/channels.md +290 -0
  20. data/docs/configuration.md +106 -0
  21. data/docs/connections.md +58 -0
  22. data/docs/conventions.md +161 -0
  23. data/docs/evals.md +95 -0
  24. data/docs/guarantees.md +76 -0
  25. data/docs/inbox-and-api.md +91 -0
  26. data/docs/memory.md +35 -0
  27. data/docs/sandbox.md +44 -0
  28. data/docs/tools.md +77 -0
  29. data/docs/tutorial.md +272 -0
  30. data/docs/vs-eve.md +93 -0
  31. data/docs/why-silas.md +87 -0
  32. data/lib/generators/silas/install/install_generator.rb +9 -0
  33. data/lib/generators/silas/install/templates/claude_skill.md +136 -0
  34. data/lib/generators/silas/install/templates/ruby_llm.rb +4 -1
  35. data/lib/silas/connection.rb +16 -0
  36. data/lib/silas/eval/dsl.rb +7 -2
  37. data/lib/silas/version.rb +1 -1
  38. metadata +22 -1
data/docs/evals.md ADDED
@@ -0,0 +1,95 @@
1
+ # Evals
2
+
3
+ Agent evals answer the question tests can't: *given these model decisions, does
4
+ the durable machinery do the right thing?* A scenario scripts the **model's
5
+ decisions** step by step; everything else is real — the real registry resolves
6
+ your real tools, the real Ledger enforces effect modes and approvals, and the
7
+ assertions read the genuine durable transcript, not a mock.
8
+
9
+ They run keyless and deterministically, which makes them a **deploy gate**:
10
+
11
+ ```sh
12
+ bin/rails silas:eval # loads <eval_dir>/**/*_eval.rb; exits 1 on failure
13
+ ```
14
+
15
+ The installer wires this into `bin/ci` (or tells you to add it to yours), and
16
+ the application template generates a working example in
17
+ `test/agent_evals/refund_desk_eval.rb`.
18
+
19
+ ## A scenario
20
+
21
+ ```ruby
22
+ Silas::Eval.scenario "over the gate: a £64 refund holds at the signal" do
23
+ input "The walnut monitor stand (order R-1002) arrived cracked."
24
+
25
+ on_step 0, call: { name: "lookup_order", arguments: { number: "R-1002" } }
26
+ on_step 1, call: { name: "issue_refund",
27
+ arguments: { number: "R-1002", amount_pence: 6400, reason: "arrived cracked" } }
28
+
29
+ expect do
30
+ assert_parked tool: "issue_refund" # the turn holds at zero compute
31
+ assert_no_tool_called "issue_refund" # and the refund row does NOT exist
32
+ end
33
+ end
34
+ ```
35
+
36
+ `on_step index, ...` is the script: what the "model" returns on that model
37
+ call. `text:` is an assistant text block, `call:`/`calls:` are tool calls. A
38
+ step index with no entry ends the scripted conversation.
39
+
40
+ ## The DSL
41
+
42
+ | Call | Meaning |
43
+ |---|---|
44
+ | `input "…"` | The turn's input. Required. |
45
+ | `on_step i, text:, call:, calls:, data:` | The scripted model output for step `i`. `data:` scripts a **structured answer** (the `final_answer` schema case) — it becomes the same block the real adapter persists, so `assert_answer_data` reads it back. |
46
+ | `approve tool: "name"` | When the turn parks on this tool, approve it (recorded as approved by `"eval"` — the same `approve!` the inbox uses) and resume. This is how you assert exactly-once **across** a hold. |
47
+ | `stub_tool "name", effect_mode:, approval: { \|**args\| … }` | Replace a real tool with a stub for a side-effect-free scenario. Unstubbed tools stay real. |
48
+ | `mode :real` | Drive a real model instead of the script. **Skipped automatically when `ANTHROPIC_API_KEY` is unset**, so the gate stays green offline. |
49
+ | `max_steps n` | Override the per-turn step cap for this scenario. |
50
+ | `metadata h` / tags on `scenario "…", tags: [...]` | Stamped on the session / used for filtering. |
51
+ | `expect { … }` | The assertion block. Required. Runs after the turn reaches its resting state. |
52
+
53
+ ## Assertions
54
+
55
+ All assertions read only the durable rows, and they **collect** failures — one
56
+ run reports every miss, not just the first.
57
+
58
+ | Assertion | Checks |
59
+ |---|---|
60
+ | `assert_tool_called "name", times: 1` | The tool actually executed (`times:` for exactly-N — the exactly-once assertion). |
61
+ | `assert_no_tool_called "name"` | No execution happened (e.g. because the turn parked first). |
62
+ | `assert_tool_arg "name", :key, value` (or a block predicate) | An argument the model passed. |
63
+ | `assert_parked tool: "name"` | The turn is waiting, held on that tool's approval (or in-doubt). |
64
+ | `assert_approved tool: "name"` | The invocation's approval state is `approved`. |
65
+ | `assert_turn_completed` / `assert_turn_failed(reason: "…")` | Terminal state. |
66
+ | `assert_final_matches(/…/)` | The final answer text. |
67
+ | `assert_answer_data(key: :verdict, value: "approve")` | The structured `final_answer` payload (whole-Hash and predicate forms too). |
68
+ | `assert_no_hallucinated_price(allowed: [])` | Every money amount in the final answer traces to a number the agent actually saw (tool results or the input), allowing pence↔pounds scaling. |
69
+ | `assert_rubric "…" ` | LLM-graded check against `config.eval_grader` — **skips offline** rather than failing the gate (set `SILAS_EVAL_STRICT=1` to make skips fail). |
70
+
71
+ ## How a scenario runs (so you can trust it)
72
+
73
+ The driver starts a real session (`Silas.agent.start`), runs the real
74
+ `AgentLoopJob` inline on the `:test` queue adapter, and — for each `approve
75
+ tool:` — performs the same `approve!` an operator would, then resumes the
76
+ loop. Scripted mode swaps only the inference adapter; `mode :real` uses your
77
+ configured one.
78
+
79
+ Two consequences worth knowing:
80
+
81
+ - **Rows persist.** Scenarios commit ordinary durable rows to the database of
82
+ the environment you run them in (they need real seeds for tools that read
83
+ your tables — the template's CI runs `db:seed silas:eval`). Run against a
84
+ scratch database if residue bothers you.
85
+ - **Config is restored per run**, but scenarios execute sequentially in one
86
+ process — don't have two scenarios fight over global state you set yourself.
87
+
88
+ ## Configuration
89
+
90
+ ```ruby
91
+ Silas.configure do |c|
92
+ c.eval_dir = "test/agent_evals" # where *_eval.rb files live
93
+ # c.eval_grader = ->(prompt) { … } # custom LLM grader for assert_rubric
94
+ end
95
+ ```
@@ -0,0 +1,76 @@
1
+ # Guarantees
2
+
3
+ Durability in Silas is a contract, not a slide. Everything on this page is
4
+ verified by the in-repo chaos harness (`chaos_host/bin/chaos`), which
5
+ `kill -9`s live agents hundreds of times per release — the current gate is
6
+ 100/100 completions per mode across a 295-run matrix, zero duplicate effects,
7
+ byte-identical replay, on SQLite and Postgres. Results live in
8
+ `chaos_host/results/` and every run is reproducible.
9
+
10
+ ## Turns survive hard process death
11
+
12
+ Worker `kill -9`, whole-tree `kill -9`, SIGTERM deploys — a turn resumes from
13
+ its last completed step. Each step checkpoints through Active Job
14
+ Continuations; the resume replays completed work **from rows**, never by
15
+ re-calling the model or re-running settled tools.
16
+
17
+ ## `transactional!` tools execute exactly once
18
+
19
+ The tool's database writes and the ledger's record of "this ran" commit or
20
+ roll back **in one transaction**:
21
+
22
+ - **Crash before commit** → both roll back; resume runs the tool from a clean
23
+ slate. No orphan effect.
24
+ - **Crash after commit** → the ledger row says `completed`; resume skips the
25
+ step. No second effect.
26
+
27
+ No idempotency key is required, because the dedup lives in your database — the
28
+ same transaction as the effect. This is the guarantee that requires the ledger
29
+ to live *inside* your app: a runtime outside your database can retry and
30
+ reconcile, but it can't make its record and your effect one atomic commit.
31
+
32
+ ## Ambiguity parks instead of guessing
33
+
34
+ The default effect mode is `at_most_once!`: when a crash makes an execution
35
+ ambiguous ("did the email send?"), the invocation parks **in doubt** for a
36
+ human verdict rather than re-firing blind. `idempotent!` is the explicit
37
+ opt-in to automatic re-runs. Never double-pay; sometimes ask.
38
+
39
+ ## Approvals park at zero compute
40
+
41
+ A turn awaiting a human exits the worker entirely — no held thread, no polling
42
+ loop. Approving enqueues a fresh job that replays from rows. Parks expire
43
+ (default 7 days, `config.approval_ttl`) rather than ghosting forever.
44
+
45
+ ## Errors can't strand a turn
46
+
47
+ Transient model errors (rate limits, overloads, timeouts) back off and retry
48
+ from the checkpoint. Exhausted retries and permanent rejections (bad key, bad
49
+ request) expire pending approvals and fail the turn loudly. And a turn can
50
+ never sit in `running` forever: the recurring `DeadJobRescuerJob` retries jobs
51
+ failed by dead-worker reaping and fails turns stranded by a loop job that died
52
+ outside the retry list. The rescuer is part of the contract — keep it in
53
+ `config/recurring.yml`, and monitor worker liveness (the rescuer can requeue
54
+ work; it cannot conjure a consumer).
55
+
56
+ ## Deploys can't corrupt a run
57
+
58
+ Instructions are snapshotted per turn. Tools, skills, connections, and the
59
+ final-answer schema are model-visible state, captured in a definitions digest —
60
+ a deploy that changes them while a turn is parked fails that turn loudly on
61
+ resume (`NondeterminismError`) instead of quietly resuming into a different
62
+ agent. Settle parked turns before shipping agent changes.
63
+
64
+ ## Compaction can't corrupt a replay
65
+
66
+ Long conversations summarise past `config.compact_at` — but a summary is a
67
+ **persisted, exactly-once effect** (claimed compare-and-swap, written once),
68
+ never a rebuild-time computation. The message array a resumed turn sees is
69
+ byte-identical to the one the crashed turn saw.
70
+
71
+ ## The one rule you owe the contract
72
+
73
+ Run agents on Solid Queue (or `:inline` for scripts) — never ActiveJob's
74
+ in-process `:async` adapter, which runs a re-enqueued continuation
75
+ concurrently with the original and double-executes steps. Silas raises on it
76
+ in production; `bin/rails silas:doctor` flags it everywhere.
@@ -0,0 +1,91 @@
1
+ # Inbox & API
2
+
3
+ Everything an operator needs ships in the gem — mount the engine (the
4
+ installer does this) and it's there. No dashboard to build, no second console
5
+ to deploy.
6
+
7
+ ## The inbox — `/silas/inbox`
8
+
9
+ A session rail grouped **Held / Working / Filed**, web chat (start a session
10
+ or reply from the browser — same durable loop), a live step-trace that
11
+ streams tokens over Turbo as the agent runs, approval and question cards
12
+ hoisted to the top of the session, a full audit trail (every tool call's
13
+ arguments and result, who cleared what and why), cancel, raise-budget, and
14
+ per-session token/cost accounting priced from RubyLLM's model registry.
15
+
16
+ <img src="https://raw.githubusercontent.com/danielstpaul/silas/main/docs/img/silas-inbox-held.png" width="740"
17
+ alt="A session held at the signal: the amber approval card for issue_refund hoisted above the trace, with the turn marked HELD and a held-at-the-signal stub in place">
18
+
19
+ A £64 refund holding at the signal: the card carries the exact arguments, the
20
+ turn costs nothing while it waits, and the trace below keeps a one-line stub
21
+ where the invocation held.
22
+
23
+ It's **deny-by-default** — invisible until you wire auth:
24
+
25
+ ```ruby
26
+ Silas.configure do |c|
27
+ # The lambda DENIES by rendering (or head-ing) and PASSES by not rendering.
28
+ c.inbox_auth = ->(controller) { controller.head :not_found unless controller.current_user&.admin? }
29
+ # c.inbox_public_read = true # read-only demo mode; writes stay gated
30
+ # c.model_prices["your-fine-tune"] = { in: 3000, out: 15_000 } # price overrides
31
+ end
32
+ ```
33
+
34
+ `config.inbox_actor` names the identity recorded on approvals (defaults to
35
+ `current_user&.email || "inbox"`). Live streaming activates automatically when
36
+ the host app has `turbo-rails` (every default Rails app does); without it the
37
+ trace falls back to a polling refresh. The gem takes no turbo dependency.
38
+
39
+ ## The JSON API — `/silas/api/v1`
40
+
41
+ The same surface over HTTP, also deny-by-default (`config.api_auth`, same
42
+ contract as the inbox; `config.api_actor` names the API identity):
43
+
44
+ ```sh
45
+ curl -X POST .../silas/api/v1/sessions -d "input=Refund order 42, £12.50"
46
+ curl .../silas/api/v1/sessions/1?trace=1 # turns + steps + tool calls
47
+ curl .../silas/api/v1/sessions/1/approvals # what's held
48
+ curl -X POST .../silas/api/v1/approvals/7/approve # the same approve! as the inbox
49
+ curl -X POST .../silas/api/v1/approvals/9/answer -d "answer=Use the June invoice"
50
+ curl -X POST .../silas/api/v1/sessions/1/turns -d "input=Now email them" # 409 if busy
51
+ curl -X POST .../silas/api/v1/turns/9/cancel
52
+ curl -N .../silas/api/v1/sessions/1/stream # server-sent events
53
+ ```
54
+
55
+ The stream is SSE at **row granularity** — turn, completed-step, and
56
+ invocation changes, at-least-once with `Last-Event-ID` resume (ids are
57
+ epoch-ms watermarks). `?poll=1` returns the backlog and closes, curl-friendly;
58
+ streams close themselves after `config.api_stream_max_duration` and clients
59
+ reconnect. Per-token streaming is deliberately the browser/Turbo feature:
60
+ deltas live in the worker process, and the gem requires no cross-process bus.
61
+
62
+ ## Streaming — decoration over durable rows
63
+
64
+ Turns stream everywhere it helps: `silas:chat` prints tokens as they arrive,
65
+ and the inbox renders them live (coalesced to ~10Hz). Deltas are never
66
+ persisted, never fed back to the model, and absent from replays — a replayed
67
+ step renders from its row — so streaming adds zero risk to the durability
68
+ contract. Custom sinks subscribe to the `"delta.silas"` notification
69
+ (`{ session_id:, turn_id:, step_id:, step_index:, text: }`, where `text` is
70
+ the accumulated string so far; filter by ids — notifications are
71
+ process-global).
72
+
73
+ ## Structured answers
74
+
75
+ Give the turn's final answer a schema in `agent.yml`:
76
+
77
+ ```yaml
78
+ final_answer:
79
+ type: object
80
+ properties:
81
+ verdict: { type: string }
82
+ amount_pence: { type: integer }
83
+ required: [verdict]
84
+ ```
85
+
86
+ `Turn#answer_data` returns the parsed Hash (`answer_text` stays for prose
87
+ agents); the API carries it as `answer_data`, and evals assert on it with
88
+ `assert_answer_data(key: :verdict, value: "approve")`. Rendered through
89
+ RubyLLM's `with_schema`, so each provider's native structured-output mode is
90
+ used. The schema is model-visible state — changing it mid-turn fails the turn
91
+ loudly rather than resuming into a different contract.
data/docs/memory.md ADDED
@@ -0,0 +1,35 @@
1
+ # Memory
2
+
3
+ Silas memory is **graph-shaped, not a graph database**: facts stored as
4
+ `subject · attribute · content` triples with provenance and supersession —
5
+ "author:jane · report_format: prefers CSV", and a new value retires the old
6
+ one instead of piling up beside it.
7
+
8
+ ## Approval-gated by default
9
+
10
+ The first time the model calls `remember`, nothing persists — the memory
11
+ **parks as a card in your inbox** and you approve or decline it
12
+ (`config.memory_approval = :always`, the default). An agent that can silently
13
+ write its own long-term memory is an agent that can silently drift;
14
+ the gate keeps a person in that loop. `:never` auto-approves once you trust
15
+ an agent's judgment, and `config.memory = false` removes the tools entirely.
16
+
17
+ ## Recall
18
+
19
+ A handful of the most recent relevant memories (default 8,
20
+ `config.memory_injection_limit`) are injected into each turn automatically;
21
+ the `recall` tool digs deeper on demand. Memories are private per agent, or
22
+ `shared: true` for the whole staff.
23
+
24
+ ## What belongs here — and what doesn't
25
+
26
+ Memory is for the fuzzy residue with no natural home: preferences, standing
27
+ context, things a colleague would jot in a notebook. Your **domain data does
28
+ not belong here** — it belongs in your own tables, which your tools already
29
+ read. If it has a schema, it's a model; if it's a remark, it's a memory.
30
+
31
+ ## The fine print
32
+
33
+ Memory tools are model-visible state, so toggling `config.memory` (like any
34
+ builtin) changes the definitions digest — settle parked turns before flipping
35
+ it, or they fail loudly on resume ([guarantees](guarantees.md)).
data/docs/sandbox.md ADDED
@@ -0,0 +1,44 @@
1
+ # Sandbox
2
+
3
+ Code execution is **off by default** (`config.sandbox = :none`). Configure a
4
+ sandbox and the `run_code` tool is advertised to the model automatically —
5
+ always `at_most_once!`, because an exec is an external effect.
6
+
7
+ ## Built-in: `:docker`
8
+
9
+ A hardened container seam — resource-capped, network-off by default, with
10
+ knobs for image, memory, CPUs, pids, and timeout
11
+ ([configuration](configuration.md)). It's honest-but-interim: a container is
12
+ weaker than a microVM, and it runs on the same host as your ledger and
13
+ `RAILS_MASTER_KEY`. Treat it as suitable for code **you** wrote, not code the
14
+ model wrote.
15
+
16
+ ## Real isolation: hermetic
17
+
18
+ For untrusted or model-generated code, the companion gem
19
+ [hermetic](https://github.com/danielstpaul/hermetic) drops straight in:
20
+
21
+ ```ruby
22
+ # Gemfile: gem "hermetic" (zero runtime deps)
23
+ Silas.configure do |c|
24
+ c.sandbox = Hermetic.gvisor(image: "python:3.12-slim")
25
+ # or .docker / .firecracker(kernel:, rootfs:) / .hosted(:e2b, api_key:)
26
+ end
27
+ ```
28
+
29
+ Two properties carry through the seam:
30
+
31
+ - **The trust axis is visible.** Every hermetic backend exposes `trust`
32
+ (`:vendor` / `:remote` / `:vm` / `:host`) and `off_host?`, so you can refuse
33
+ to run untrusted code on the box that holds your secrets — and pair any
34
+ local backend with `executor:` to push execution to a dedicated sandbox
35
+ host.
36
+ - **The ledger guard is auto-armed.** Configuring a hermetic backend loads its
37
+ Silas shim, so a sandbox exec attempted inside a ledger transaction fails
38
+ loudly (sandbox-backed tools must be `at_most_once!`, never
39
+ `transactional!`).
40
+
41
+ ## Bring your own
42
+
43
+ `config.sandbox` accepts any object responding to `#run` — the seam is the
44
+ contract, the backends are interchangeable.
data/docs/tools.md ADDED
@@ -0,0 +1,77 @@
1
+ # Tools & approvals
2
+
3
+ A tool is one file in `app/agent/tools/`. The filename is the tool's identity;
4
+ the keyword signature of `#call` is the schema the model sees. No registry, no
5
+ wiring — add a file, restart, and the model can use it.
6
+
7
+ ```ruby
8
+ # app/agent/tools/issue_refund.rb -> the "issue_refund" tool
9
+ class Agent::Tools::IssueRefund < Silas::Tool
10
+ description "Refund part or all of an order."
11
+ param :amount_pence, :integer, desc: "Amount in pence (1800 = £18.00)"
12
+ approval ->(session:, input:) { input[:amount_pence] > 2_500 ? :user_approval : :approved }
13
+ transactional!
14
+
15
+ def call(number:, amount_pence:, reason:)
16
+ order = Order.find_by!(number: number)
17
+ refund = order.refunds.create!(amount_pence:, reason:)
18
+ { refunded_pence: refund.amount_pence, order: order.number }
19
+ end
20
+ end
21
+ ```
22
+
23
+ The rules:
24
+
25
+ - **Keywords only** — the keyword signature *is* the schema; `param` refines
26
+ types and descriptions.
27
+ - **Return a Hash** (anything else is wrapped as `{"value" => ...}`). Raising
28
+ records a failed invocation the model sees — don't rescue-and-swallow.
29
+ - `session` is available inside `call` (the `Silas::Session` row), so tools
30
+ can read channel metadata or scope queries.
31
+
32
+ ## Effect modes — decide by where the side effect lives
33
+
34
+ | The tool… | Declare | What you get |
35
+ |---|---|---|
36
+ | writes this app's database | `transactional!` | effect + ledger commit atomically → **exactly-once**, even through `kill -9` |
37
+ | calls anything external (email, HTTP, Slack…) | `at_most_once!` (the default) | a crash mid-call leaves it **in doubt** → parks for a human verdict, never re-fires blind |
38
+ | only reads, safe to repeat | `idempotent!` | crash replays may re-run it freely |
39
+
40
+ Never mark an external call `transactional!` — the ledger cannot roll back a
41
+ sent email. Model money as rows in your own database wherever you can; that's
42
+ what upgrades the guarantee to exactly-once
43
+ ([guarantees](guarantees.md)).
44
+
45
+ ## Approval — who holds the lever
46
+
47
+ ```ruby
48
+ approval :never # default — runs without asking
49
+ approval :always # every call parks for a human
50
+ approval :once # one approval per identical (tool, arguments) pair per session
51
+ approval ->(session:, input:) { ... } # decide from the arguments
52
+ ```
53
+
54
+ A lambda returns `:user_approval`, `:approved`, `:not_applicable`, or
55
+ `{denied: "reason"}` (the denial rides back to the model as the tool result).
56
+ Parking costs nothing — the turn **holds at the signal** at zero compute until
57
+ someone clears it from the inbox, Slack, a signed email link, or the JSON API,
58
+ all calling the same `approve!`/`decline!`. Parks expire after
59
+ `config.approval_ttl` (default 7 days).
60
+
61
+ The built-in **`ask_question`** tool is the reverse direction: the agent parks
62
+ to ask the operator something — information, not permission — and their typed
63
+ answer resumes the turn as the tool result.
64
+
65
+ ## Skills — playbooks, loaded on demand
66
+
67
+ A skill is a markdown file in `app/agent/skills/` with a `description:`
68
+ frontmatter line. The description is always visible to the model as a routing
69
+ hint; the body loads only when the model asks for it (`load_skill`), keeping
70
+ the always-on prompt small. Put procedures in skills; put identity in
71
+ `instructions.md`.
72
+
73
+ ## Remote tools
74
+
75
+ `app/agent/connections/*.yml` plugs a remote MCP server's tools in under the
76
+ same ledger, namespaced `<connection>__<tool>` — see
77
+ [connections](connections.md).