nexo_ai 0.8.0 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +145 -0
- data/docs/concurrency.md +28 -0
- data/docs/loops.md +27 -0
- data/docs/mcp.md +26 -0
- data/docs/sandboxes.md +175 -39
- data/docs/skills.md +82 -3
- data/docs/workflows.md +47 -1
- data/lib/nexo/agent.rb +199 -3
- data/lib/nexo/configuration.rb +8 -0
- data/lib/nexo/loops/ruby_llm.rb +30 -2
- data/lib/nexo/sandbox.rb +106 -0
- data/lib/nexo/sandboxes/container.rb +52 -6
- data/lib/nexo/sandboxes/remote.rb +22 -1
- data/lib/nexo/skills.rb +78 -3
- data/lib/nexo/version.rb +1 -1
- data/lib/nexo/workflow.rb +125 -10
- data/lib/nexo.rb +6 -0
- metadata +2 -2
data/docs/workflows.md
CHANGED
|
@@ -115,7 +115,7 @@ the count staged.
|
|
|
115
115
|
|
|
116
116
|
`artifact(name, content:)` records a **named deliverable** on the run — a digest,
|
|
117
117
|
a report, an improved file, a generated script. The body is written to the
|
|
118
|
-
sandbox at
|
|
118
|
+
sandbox at `artifacts/<name>` (so later steps can read it) and recorded on the
|
|
119
119
|
run. `run.artifacts` reads it back as an **ordered array** of string-keyed hashes
|
|
120
120
|
(`{"name" =>, "content" =>, "at" =>}`) in both stores:
|
|
121
121
|
|
|
@@ -150,6 +150,52 @@ artifact("digest.md", from: "app/templates/digest.md.erb",
|
|
|
150
150
|
|
|
151
151
|
See [`examples/artifact_from_template.rb`](../examples/artifact_from_template.rb) for
|
|
152
152
|
the full offline flow (`ruby -Ilib examples/artifact_from_template.rb`).
|
|
153
|
+
### Agent output — `produces`
|
|
154
|
+
|
|
155
|
+
`from:` renders ERB and is **only** for templates you wrote. An agent's output is model
|
|
156
|
+
output, so it is copied **verbatim** instead. An agent declares what it produces, and may
|
|
157
|
+
produce as many artifacts as it likes:
|
|
158
|
+
|
|
159
|
+
```ruby
|
|
160
|
+
class Publisher < Nexo::Agent
|
|
161
|
+
skills :dashboard_designer
|
|
162
|
+
produces "dashboard.html", "digest.json", "out/*.csv"
|
|
163
|
+
end
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
`run_agent` copies those out of the sandbox and records them on the run the moment the
|
|
167
|
+
agent finishes — **including when it suspended for approval or raised**, which are exactly
|
|
168
|
+
the paths where results would otherwise be lost. Under the hood that is
|
|
169
|
+
`artifact(name, path:)`, the verbatim third mode, which you can also call directly.
|
|
170
|
+
|
|
171
|
+
Why this matters: a workflow builds **one** sandbox that every `run_agent` borrows, so a
|
|
172
|
+
hand-off between stages is just "agent B reads a path agent A wrote". On `:local` that is a
|
|
173
|
+
real directory and survives anything. On `:docker`/`:apple` it does not — `Container#close`
|
|
174
|
+
is `rm -f`, and a run's sandbox is released on **every** terminal path, `suspended`
|
|
175
|
+
included. So pausing for a human approval destroyed everything produced before the pause,
|
|
176
|
+
while the identical code on `:local` kept it. Declared artifacts survive the teardown.
|
|
177
|
+
|
|
178
|
+
Declared, never inferred: sweeping the sandbox would also collect staged skill scripts,
|
|
179
|
+
templates and scratch files, and naming outputs is the only honest way for an agent to say
|
|
180
|
+
it produced nothing. A declared artifact that was never written is skipped, not fatal.
|
|
181
|
+
|
|
182
|
+
Bytes that are not valid UTF-8 are Base64-wrapped so they survive a JSON column. Read a
|
|
183
|
+
body back with `Nexo::Workflow.artifact_body(art)` rather than `art["content"]`, which is
|
|
184
|
+
Base64 text for binary.
|
|
185
|
+
|
|
186
|
+
### Handing output to the next stage — `restore_artifacts`
|
|
187
|
+
|
|
188
|
+
```ruby
|
|
189
|
+
run_agent("extract") # Extractor produces gmail.json
|
|
190
|
+
restore_artifacts # put recorded artifacts back in the sandbox
|
|
191
|
+
run_agent("synthesize") # Synthesizer reads gmail.json
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
`Skills.materialize` gets a skill's files *into* a sandbox; this is the other direction and
|
|
195
|
+
then back in again, so a pipeline behaves the same whether the tier is persistent or
|
|
196
|
+
ephemeral — including across a suspend and resume, where the sandbox that held the files no
|
|
197
|
+
longer exists. Pass `only:` to restore a subset.
|
|
198
|
+
|
|
153
199
|
|
|
154
200
|
The `artifacts` column ships with fresh installs. Apps installed before this
|
|
155
201
|
release add it with:
|
data/lib/nexo/agent.rb
CHANGED
|
@@ -22,7 +22,7 @@ module Nexo
|
|
|
22
22
|
# configuration instead of silently resetting to defaults.
|
|
23
23
|
CONFIG_IVARS = %i[
|
|
24
24
|
@model @assume_model_exists @provider @sandbox @permissions @instructions
|
|
25
|
-
@skills @mcp @mcp_allow @fetch_allow @search_backend
|
|
25
|
+
@skills @mcp @mcp_allow @fetch_allow @search_backend @requires @produces
|
|
26
26
|
].freeze
|
|
27
27
|
|
|
28
28
|
class << self
|
|
@@ -102,6 +102,59 @@ module Nexo
|
|
|
102
102
|
names.empty? ? (@skills || []) : (@skills = ((@skills || []) + names).uniq)
|
|
103
103
|
end
|
|
104
104
|
|
|
105
|
+
# What this agent needs from whatever sandbox it runs in — checked once,
|
|
106
|
+
# before the first turn, against Sandbox#environment:
|
|
107
|
+
#
|
|
108
|
+
# class Publisher < Nexo::Agent
|
|
109
|
+
# skills :dashboard_designer
|
|
110
|
+
# requires commands: {"ruby" => ">= 3.1"}, locale: :utf8
|
|
111
|
+
# end
|
|
112
|
+
#
|
|
113
|
+
# Declared HERE, in Nexo's own vocabulary, rather than in the skill file:
|
|
114
|
+
# whoever wires an agent to a sandbox is the only person who can *fix* a
|
|
115
|
+
# gap, so the declaration and the fix live in the same place. A skill states
|
|
116
|
+
# its needs in prose via +compatibility:+, which is the spec's field for it
|
|
117
|
+
# and is aimed at a human or a model.
|
|
118
|
+
#
|
|
119
|
+
# +commands:+ maps a command that must be on +PATH+ to a version constraint
|
|
120
|
+
# — a Gem::Requirement string (+">= 3.1"+) or +"*"+ for "any version". A
|
|
121
|
+
# command whose version cannot be read (busybox +sh+ prints none) satisfies
|
|
122
|
+
# any constraint by being present: an unreadable version is not evidence of
|
|
123
|
+
# a wrong one. +locale:+ takes +:utf8+ (any UTF-8 locale — the useful case)
|
|
124
|
+
# or an exact String.
|
|
125
|
+
#
|
|
126
|
+
# Deliberately coarse, and never packages: gems, wheels and npm modules
|
|
127
|
+
# belong to the image, not here (see Sandbox#environment).
|
|
128
|
+
def requires(commands: nil, locale: nil)
|
|
129
|
+
return @requires if commands.nil? && locale.nil?
|
|
130
|
+
|
|
131
|
+
@requires = {commands: commands || {}, locale: locale}
|
|
132
|
+
end
|
|
133
|
+
|
|
134
|
+
# The artifacts this agent produces — sandbox paths, or globs, that a
|
|
135
|
+
# workflow copies out of the sandbox and records on the run as soon as the
|
|
136
|
+
# agent finishes:
|
|
137
|
+
#
|
|
138
|
+
# class Publisher < Nexo::Agent
|
|
139
|
+
# produces "dashboard.html", "digest.json", "out/*.csv"
|
|
140
|
+
# end
|
|
141
|
+
#
|
|
142
|
+
# Accumulating and deduped, like +skills+: an agent may produce many
|
|
143
|
+
# artifacts, and several +produces+ lines add up rather than replacing.
|
|
144
|
+
#
|
|
145
|
+
# This is the agent's OUTPUT, not a template. Workflow#artifact's +from:+
|
|
146
|
+
# mode renders ERB and is only ever for files you wrote; what an agent
|
|
147
|
+
# produces is model output and is copied verbatim (Workflow#artifact +path:+).
|
|
148
|
+
#
|
|
149
|
+
# Declared rather than inferred because a sweep of the sandbox would also
|
|
150
|
+
# collect staged skill scripts, templates and scratch files — and because
|
|
151
|
+
# naming outputs is the only honest way for an agent to say it produced
|
|
152
|
+
# nothing. A declared artifact that is absent when the agent finishes is
|
|
153
|
+
# skipped, not fatal.
|
|
154
|
+
def produces(*names)
|
|
155
|
+
names.empty? ? (@produces || []) : (@produces = ((@produces || []) + names).uniq)
|
|
156
|
+
end
|
|
157
|
+
|
|
105
158
|
# Declares an MCP server for this agent (Spec 6). Accumulating: multiple
|
|
106
159
|
# +mcp+ lines are collected. With no args (+name+ nil and +opts+ empty) it
|
|
107
160
|
# reads the list (default +[]+); otherwise it appends the friendly,
|
|
@@ -246,6 +299,7 @@ module Nexo
|
|
|
246
299
|
apply_mcp(c)
|
|
247
300
|
apply_fetch(c)
|
|
248
301
|
apply_search(c)
|
|
302
|
+
apply_tool_concurrency(c)
|
|
249
303
|
c
|
|
250
304
|
end
|
|
251
305
|
|
|
@@ -254,9 +308,87 @@ module Nexo
|
|
|
254
308
|
# so swapping +loop:+ swaps the engine without touching this class. The
|
|
255
309
|
# optional +&on_event+ block receives +(type, payload)+ progress events.
|
|
256
310
|
def prompt(text, max_turns: 25, &on_event)
|
|
311
|
+
verify_environment!
|
|
257
312
|
@loop.run(agent: self, prompt: text, max_turns: max_turns, &on_event)
|
|
258
313
|
end
|
|
259
314
|
|
|
315
|
+
# Checks the agent's declared ::requires against its sandbox, once, before the
|
|
316
|
+
# first turn. A no-op — and zero cost, no probe at all — when nothing is
|
|
317
|
+
# declared, which is the default.
|
|
318
|
+
#
|
|
319
|
+
# Fails BEFORE the model is called, because the alternative is what this
|
|
320
|
+
# exists to prevent: the agent spends turns deciding to run a script, runs it,
|
|
321
|
+
# and gets +sh: ruby: not found+ or an Encoding::InvalidByteSequenceError three
|
|
322
|
+
# frames into a JSON parse. One legible sentence naming what is missing and
|
|
323
|
+
# where beats a stack trace after the fact.
|
|
324
|
+
#
|
|
325
|
+
# Raises Nexo::EnvironmentError listing every unmet requirement at once, so a
|
|
326
|
+
# misprovisioned image is fixed in one pass rather than one round trip per
|
|
327
|
+
# missing command.
|
|
328
|
+
def verify_environment!
|
|
329
|
+
return if @environment_verified
|
|
330
|
+
|
|
331
|
+
req = self.class.requires
|
|
332
|
+
@environment_verified = true
|
|
333
|
+
return if req.nil?
|
|
334
|
+
|
|
335
|
+
missing = unmet_requirements(req, @sandbox.environment)
|
|
336
|
+
return if missing.empty?
|
|
337
|
+
|
|
338
|
+
raise Nexo::EnvironmentError,
|
|
339
|
+
"#{self.class} cannot run here: #{missing.join("; ")}. " \
|
|
340
|
+
"Provision the sandbox, or drop the `requires` declaration."
|
|
341
|
+
end
|
|
342
|
+
|
|
343
|
+
private
|
|
344
|
+
|
|
345
|
+
# The declared requirements this environment does not meet, as human-readable
|
|
346
|
+
# phrases. Reports the probe's own failure first when it could not run at all:
|
|
347
|
+
# "no ruby" and "I never got to look" are different problems with different
|
|
348
|
+
# fixes, and saying the first when the second is true sends the reader to the
|
|
349
|
+
# wrong place.
|
|
350
|
+
def unmet_requirements(req, env)
|
|
351
|
+
return ["could not inspect the sandbox (#{env[:error]})"] if env[:error]
|
|
352
|
+
|
|
353
|
+
missing = req[:commands].filter_map do |name, constraint|
|
|
354
|
+
found = env[:commands][name.to_s]
|
|
355
|
+
next "no #{name} on PATH" if found.nil?
|
|
356
|
+
next unless version_short?(found[:version], constraint)
|
|
357
|
+
|
|
358
|
+
"#{name} #{found[:version]} does not satisfy #{constraint}"
|
|
359
|
+
end
|
|
360
|
+
missing << locale_complaint(req[:locale], env[:locale]) if locale_unmet?(req[:locale], env[:locale])
|
|
361
|
+
missing
|
|
362
|
+
end
|
|
363
|
+
|
|
364
|
+
# Whether a found command's version fails +constraint+. A missing version, an
|
|
365
|
+
# unparseable constraint, or +"*"+ all pass: presence is the requirement, and
|
|
366
|
+
# an unreadable version is not evidence of a wrong one.
|
|
367
|
+
def version_short?(version, constraint)
|
|
368
|
+
return false if version.nil? || constraint.nil? || constraint.to_s.strip == "*"
|
|
369
|
+
|
|
370
|
+
!Gem::Requirement.new(constraint.to_s).satisfied_by?(Gem::Version.new(version))
|
|
371
|
+
rescue ArgumentError
|
|
372
|
+
false
|
|
373
|
+
end
|
|
374
|
+
|
|
375
|
+
# +:utf8+ asks for any UTF-8 locale (the only case that comes up in practice);
|
|
376
|
+
# a String asks for that exact locale. An unset locale never satisfies either
|
|
377
|
+
# — that is the container default, and the bug this catches.
|
|
378
|
+
def locale_unmet?(required, actual)
|
|
379
|
+
return false if required.nil?
|
|
380
|
+
return true if actual.nil?
|
|
381
|
+
|
|
382
|
+
(required == :utf8) ? !actual.match?(/utf-?8/i) : actual != required.to_s
|
|
383
|
+
end
|
|
384
|
+
|
|
385
|
+
def locale_complaint(required, actual)
|
|
386
|
+
want = (required == :utf8) ? "a UTF-8 locale" : "locale #{required}"
|
|
387
|
+
actual.nil? ? "no locale set (needs #{want})" : "locale #{actual} is not #{want}"
|
|
388
|
+
end
|
|
389
|
+
|
|
390
|
+
public
|
|
391
|
+
|
|
260
392
|
# The agent's Nexo permission mode mapped onto an opt-in backend's own
|
|
261
393
|
# permission vocabulary (see PERMISSION_MODE_MAP). Consumed by
|
|
262
394
|
# Loops::AgentSDK; the default Loops::RubyLLM does its gating inside the
|
|
@@ -336,7 +468,7 @@ module Nexo
|
|
|
336
468
|
texts = []
|
|
337
469
|
texts << @instructions if @instructions
|
|
338
470
|
texts << @sandbox.instructions if @sandbox.instructions
|
|
339
|
-
self.class.skills.each { |name| texts << Skills.find(name)
|
|
471
|
+
self.class.skills.each { |name| texts << skill_instructions(Skills.find(name)) }
|
|
340
472
|
|
|
341
473
|
texts.each_with_index do |text, i|
|
|
342
474
|
chat = chat.with_instructions(text, append: i.positive?)
|
|
@@ -344,6 +476,32 @@ module Nexo
|
|
|
344
476
|
chat
|
|
345
477
|
end
|
|
346
478
|
|
|
479
|
+
# One skill's contribution to the system prompt: its body, plus its
|
|
480
|
+
# +compatibility:+ frontmatter when set.
|
|
481
|
+
#
|
|
482
|
+
# +compatibility:+ is the Agent Skills spec's own field for what a skill needs
|
|
483
|
+
# in order to run ("requires a Ruby interpreter and a UTF-8 locale"), and it is
|
|
484
|
+
# free text by design — the spec deliberately does not make it machine-checkable.
|
|
485
|
+
# Nexo parsed it and then dropped it, so the one place an author can state a
|
|
486
|
+
# requirement never reached the model. That matters now that a skill can ship a
|
|
487
|
+
# script Nexo stages into a sandbox (Skills.materialize): the environment is no
|
|
488
|
+
# longer the author's machine, and the model is the one deciding whether to run
|
|
489
|
+
# the thing.
|
|
490
|
+
#
|
|
491
|
+
# Labelled rather than concatenated, so the model can tell a requirement from an
|
|
492
|
+
# instruction. Absent or blank +compatibility:+ contributes nothing, leaving the
|
|
493
|
+
# prompt byte-for-byte as before for every skill that does not set it. +license:+
|
|
494
|
+
# and +allowed-tools:+ are deliberately NOT surfaced: the first is prompt noise,
|
|
495
|
+
# and the second would be a second source of truth about what an agent may do,
|
|
496
|
+
# competing with Nexo::Permissions, which is the real gate.
|
|
497
|
+
def skill_instructions(skill)
|
|
498
|
+
body = skill.content.to_s
|
|
499
|
+
compat = skill.compatibility.to_s.strip if skill.respond_to?(:compatibility)
|
|
500
|
+
return body if compat.nil? || compat.empty?
|
|
501
|
+
|
|
502
|
+
"#{body}\n\nCompatibility: #{compat}".strip
|
|
503
|
+
end
|
|
504
|
+
|
|
347
505
|
# Lazily connects the declared MCP servers and attaches their tools, each
|
|
348
506
|
# wrapped in a MCP::GatedTool so every invocation is authorized through this
|
|
349
507
|
# agent's Permissions first. Attached after the four sandbox tools and the
|
|
@@ -360,11 +518,49 @@ module Nexo
|
|
|
360
518
|
|
|
361
519
|
@mcp_clients ||= self.class.mcp.map { |cfg| Nexo::MCP.build(**cfg) }
|
|
362
520
|
gated = @mcp_clients.flat_map(&:tools).map do |tool|
|
|
363
|
-
Nexo::MCP::GatedTool.new(tool: tool, permissions: @permissions)
|
|
521
|
+
wrap_mcp_tool(Nexo::MCP::GatedTool.new(tool: tool, permissions: @permissions))
|
|
364
522
|
end
|
|
365
523
|
chat.with_tools(*gated) unless gated.empty?
|
|
366
524
|
end
|
|
367
525
|
|
|
526
|
+
# Applies Nexo.config.tool_concurrency to the chat, LAST — every +with_tools+ call
|
|
527
|
+
# resets the chat's concurrency to its current value, so setting it before the
|
|
528
|
+
# tools are attached would be undone. +nil+ leaves RubyLLM's own setting alone.
|
|
529
|
+
def apply_tool_concurrency(chat)
|
|
530
|
+
mode = Nexo.config.tool_concurrency
|
|
531
|
+
return if mode.nil?
|
|
532
|
+
return unless chat.respond_to?(:with_tools)
|
|
533
|
+
|
|
534
|
+
chat.with_tools(concurrency: mode)
|
|
535
|
+
end
|
|
536
|
+
|
|
537
|
+
# Hook: called once per MCP tool, AFTER it is wrapped in a permission-gating
|
|
538
|
+
# MCP::GatedTool and before it is attached. Returns the tool to attach; the
|
|
539
|
+
# default returns it unchanged.
|
|
540
|
+
#
|
|
541
|
+
# Override to decorate every MCP tool an agent gets — capping an oversized reply,
|
|
542
|
+
# recording timings, redacting a field. A third-party MCP server's response shape
|
|
543
|
+
# is not yours to change, and its tools are attached by the harness rather than by
|
|
544
|
+
# you, so without this seam the only way in was to override the private
|
|
545
|
+
# +#apply_mcp+.
|
|
546
|
+
#
|
|
547
|
+
# class MailAgent < Nexo::Agent
|
|
548
|
+
# mcp :mail, transport: :stdio, command: "apple-mail-mcp"
|
|
549
|
+
#
|
|
550
|
+
# def wrap_mcp_tool(tool)
|
|
551
|
+
# CappedTool.new(tool: tool, max_chars: 4_000)
|
|
552
|
+
# end
|
|
553
|
+
# end
|
|
554
|
+
#
|
|
555
|
+
# A wrapper must keep the duck type the chat relies on — +#name+, +#description+,
|
|
556
|
+
# +#params_schema+ and +#call+. GatedTool delegates the rest through
|
|
557
|
+
# +method_missing+, so wrappers compose. Gating happens underneath, so a wrapper
|
|
558
|
+
# only ever sees an already-authorized call: wrapping cannot widen what the agent
|
|
559
|
+
# may do.
|
|
560
|
+
def wrap_mcp_tool(tool)
|
|
561
|
+
tool
|
|
562
|
+
end
|
|
563
|
+
|
|
368
564
|
# Attaches a single Nexo::Tools::Fetch scoped to the agent's +fetch_allow+
|
|
369
565
|
# hosts (Spec 9). Returns early when no host is declared, so an agent that never
|
|
370
566
|
# calls +fetch_allow+ gets no fetch tool. Attached right after +apply_mcp+ so
|
data/lib/nexo/configuration.rb
CHANGED
|
@@ -25,6 +25,14 @@ module Nexo
|
|
|
25
25
|
# which always runs its own reactor when called.
|
|
26
26
|
attr_accessor :concurrency
|
|
27
27
|
|
|
28
|
+
# How RubyLLM executes MULTIPLE tool calls returned in a single assistant turn:
|
|
29
|
+
# +:fibers+ (via +async+, the same driver Nexo.concurrent uses), +:threads+, or
|
|
30
|
+
# +false+ for one at a time. +nil+ (the default) leaves RubyLLM's own setting
|
|
31
|
+
# alone, so existing behaviour is unchanged until you opt in.
|
|
32
|
+
#
|
|
33
|
+
# Distinct from #concurrency, which governs how Nexo itself runs blocking work.
|
|
34
|
+
attr_accessor :tool_concurrency
|
|
35
|
+
|
|
28
36
|
# Max in-flight tasks for Nexo.concurrent fan-out (Spec 5), default +8+.
|
|
29
37
|
# The single most important knob for staying under provider rate limits.
|
|
30
38
|
attr_accessor :max_in_flight
|
data/lib/nexo/loops/ruby_llm.rb
CHANGED
|
@@ -22,10 +22,19 @@ module Nexo
|
|
|
22
22
|
# +chat:+ lets a Nexo::Session inject the hydrated, continuing chat so the
|
|
23
23
|
# loop runs over the persisted thread; left nil it builds the agent's own
|
|
24
24
|
# fresh chat exactly as before — the default (no-session) path is unchanged.
|
|
25
|
+
# +max_turns+ is a BUDGET, not a hard stop. ruby_llm runs the whole tool loop
|
|
26
|
+
# inside #ask and its callbacks cannot halt it, so this loop cannot cut a run
|
|
27
|
+
# short. What it can do is COUNT the tool calls and say so: exceeding the budget
|
|
28
|
+
# emits a +:turn_limit_exceeded+ event (once) carrying the count and the limit.
|
|
29
|
+
#
|
|
30
|
+
# Previously the parameter was accepted and never read, which made it look like
|
|
31
|
+
# a safety bound it has never been. Treat it as telemetry: if you need a hard
|
|
32
|
+
# ceiling, bound the work inside your tools, or use Loops::AgentSDK whose engine
|
|
33
|
+
# enforces its own max_turns natively.
|
|
25
34
|
def run(agent:, prompt:, max_turns: 25, chat: nil, &on_event)
|
|
26
35
|
chat ||= agent.chat
|
|
27
36
|
|
|
28
|
-
wire_observability(chat, &on_event)
|
|
37
|
+
wire_observability(chat, max_turns: max_turns, &on_event)
|
|
29
38
|
|
|
30
39
|
response = chat.ask(prompt)
|
|
31
40
|
on_event&.call(:done, response)
|
|
@@ -47,20 +56,39 @@ module Nexo
|
|
|
47
56
|
# prompt observes with its own block rather than the first one. A fresh chat
|
|
48
57
|
# (the default per-prompt path) simply wires once — byte-for-byte the prior
|
|
49
58
|
# behavior.
|
|
50
|
-
def wire_observability(chat, &on_event)
|
|
59
|
+
def wire_observability(chat, max_turns: nil, &on_event)
|
|
51
60
|
return unless chat.respond_to?(:before_tool_call) && chat.respond_to?(:after_tool_result)
|
|
52
61
|
|
|
53
62
|
chat.instance_variable_set(:@nexo_on_event, on_event)
|
|
63
|
+
chat.instance_variable_set(:@nexo_max_turns, max_turns)
|
|
64
|
+
# The count belongs to the PROMPT, not the chat: a continuing Session runs the
|
|
65
|
+
# loop repeatedly over one chat, and each prompt gets its own budget.
|
|
66
|
+
chat.instance_variable_set(:@nexo_turns, 0)
|
|
67
|
+
chat.instance_variable_set(:@nexo_turns_reported, false)
|
|
54
68
|
return if chat.instance_variable_get(:@nexo_observed)
|
|
55
69
|
|
|
56
70
|
chat.instance_variable_set(:@nexo_observed, true)
|
|
57
71
|
chat.before_tool_call do |tc|
|
|
72
|
+
count_turn(chat)
|
|
58
73
|
chat.instance_variable_get(:@nexo_on_event)&.call(:tool_call, tc)
|
|
59
74
|
end
|
|
60
75
|
chat.after_tool_result do |r|
|
|
61
76
|
chat.instance_variable_get(:@nexo_on_event)&.call(:tool_result, r)
|
|
62
77
|
end
|
|
63
78
|
end
|
|
79
|
+
|
|
80
|
+
# Counts one tool call and reports the first time the budget is passed. Reported
|
|
81
|
+
# once per prompt, not per call, so a long run does not drown its own event log.
|
|
82
|
+
def count_turn(chat)
|
|
83
|
+
limit = chat.instance_variable_get(:@nexo_max_turns)
|
|
84
|
+
turns = chat.instance_variable_get(:@nexo_turns).to_i + 1
|
|
85
|
+
chat.instance_variable_set(:@nexo_turns, turns)
|
|
86
|
+
return unless limit && turns > limit
|
|
87
|
+
return if chat.instance_variable_get(:@nexo_turns_reported)
|
|
88
|
+
|
|
89
|
+
chat.instance_variable_set(:@nexo_turns_reported, true)
|
|
90
|
+
chat.instance_variable_get(:@nexo_on_event)&.call(:turn_limit_exceeded, {turns: turns, max_turns: limit})
|
|
91
|
+
end
|
|
64
92
|
end
|
|
65
93
|
end
|
|
66
94
|
end
|
data/lib/nexo/sandbox.rb
CHANGED
|
@@ -60,5 +60,111 @@ module Nexo
|
|
|
60
60
|
def mtime(path)
|
|
61
61
|
nil
|
|
62
62
|
end
|
|
63
|
+
|
|
64
|
+
# Commands the default #environment probe looks for. Deliberately short —
|
|
65
|
+
# each entry costs a +command -v+ plus one +--version+ when found — and
|
|
66
|
+
# extensible per call for anything else a skill's script might need.
|
|
67
|
+
PROBE_COMMANDS = %w[ruby python3 node sh].freeze
|
|
68
|
+
|
|
69
|
+
# The shape #environment always answers with, built fresh on every call.
|
|
70
|
+
# Deliberately a method and not a frozen constant: the +:commands+ Hash is
|
|
71
|
+
# mutated while parsing, and +CONST.dup+ is shallow — sharing one inner Hash
|
|
72
|
+
# let every probe accumulate into the constant, so a sandbox with no shell
|
|
73
|
+
# answered with the previous sandbox's findings. +:error+ is nil on a
|
|
74
|
+
# successful probe and carries the reason when one could not be run.
|
|
75
|
+
def self.empty_environment
|
|
76
|
+
{commands: {}, locale: nil, error: nil}
|
|
77
|
+
end
|
|
78
|
+
|
|
79
|
+
# One POSIX +sh+ script, no interpreter required on the far side, so the probe
|
|
80
|
+
# works on a busybox image. +command -v+ locates each command and a single
|
|
81
|
+
# +--version+ reports it; the format placeholder is filled by #environment.
|
|
82
|
+
PROBE_SCRIPT = <<~SH
|
|
83
|
+
printf 'locale=%%s\\n' "${LC_ALL:-${LC_CTYPE:-${LANG:-}}}"
|
|
84
|
+
for c in %<commands>s; do
|
|
85
|
+
p=$(command -v "$c" 2>/dev/null) || continue
|
|
86
|
+
v=$("$c" --version 2>/dev/null | head -1)
|
|
87
|
+
printf 'cmd=%%s\\t%%s\\t%%s\\n' "$c" "$p" "$v"
|
|
88
|
+
done
|
|
89
|
+
SH
|
|
90
|
+
|
|
91
|
+
# What this execution environment actually provides, as data:
|
|
92
|
+
#
|
|
93
|
+
# sandbox.environment
|
|
94
|
+
# # => { commands: { "ruby" => { path: "/usr/local/bin/ruby", version: "4.0.0" } },
|
|
95
|
+
# # locale: "C.UTF-8" }
|
|
96
|
+
#
|
|
97
|
+
# #instructions describes the environment *for the model*; this is the same
|
|
98
|
+
# question answered *for code*, so a caller can check before staging a skill's
|
|
99
|
+
# script rather than discovering the answer as a stack trace several turns in.
|
|
100
|
+
# A container typically has no locale at all, under which Ruby's default
|
|
101
|
+
# external encoding is US-ASCII and a bare File.read on a UTF-8 file raises —
|
|
102
|
+
# and an image can carry a full Ruby toolchain and still report no locale, so
|
|
103
|
+
# the two are reported as independent axes.
|
|
104
|
+
#
|
|
105
|
+
# Deliberately coarse: commands on +PATH+ and the locale, never packages. Gems,
|
|
106
|
+
# wheels and npm modules belong to whoever builds the image, and modelling them
|
|
107
|
+
# here would be a cross-language dependency resolver competing with the manifest
|
|
108
|
+
# every ecosystem already has.
|
|
109
|
+
#
|
|
110
|
+
# Costs one #shell round trip and is memoized for the sandbox's lifetime (0.14s
|
|
111
|
+
# on +:local+, 0.25–0.48s on a container, measured). A sandbox with no shell
|
|
112
|
+
# A sandbox with no shell (Virtual) reports empty. A probe that fails for any
|
|
113
|
+
# other reason — the container would not start, the client is unreachable —
|
|
114
|
+
# also reports empty, but carries the reason under +:error+: this is
|
|
115
|
+
# diagnostics and must never be the reason a run dies, yet "I probed and found
|
|
116
|
+
# nothing" and "I could not probe" are different answers and a caller building
|
|
117
|
+
# an error message needs to tell them apart.
|
|
118
|
+
def environment(commands: PROBE_COMMANDS)
|
|
119
|
+
@environment ||= {}
|
|
120
|
+
@environment[commands] ||= probe_environment(commands)
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
private
|
|
124
|
+
|
|
125
|
+
# Runs the probe and parses its output. Rescues everything: see #environment.
|
|
126
|
+
def probe_environment(commands)
|
|
127
|
+
return Sandbox.empty_environment.merge(error: "sandbox has no shell") unless supports?(:shell)
|
|
128
|
+
|
|
129
|
+
out = shell(format(PROBE_SCRIPT, commands: Array(commands).join(" ")))
|
|
130
|
+
parse_environment(out[:stdout].to_s)
|
|
131
|
+
rescue StandardError, NotImplementedError => e
|
|
132
|
+
Sandbox.empty_environment.merge(error: "#{e.class}: #{salient_line(e.message)}")
|
|
133
|
+
end
|
|
134
|
+
|
|
135
|
+
# The most useful single line of a multi-line runtime failure. Neither the
|
|
136
|
+
# first line nor a blind truncation works: a container runtime prints a
|
|
137
|
+
# progress banner first ("[1/6] Fetching image", "container start failed")
|
|
138
|
+
# and puts the actual cause last ("The volume is read only"). Prefer the last
|
|
139
|
+
# line that announces an error, else the last non-empty line.
|
|
140
|
+
def salient_line(message)
|
|
141
|
+
lines = message.to_s.lines.map(&:strip).reject(&:empty?)
|
|
142
|
+
# Written as a plain scan on purpose. The idiomatic spellings deadlock the
|
|
143
|
+
# linter — Style/ReverseFind rewrites `reverse.find` to Enumerable#rfind,
|
|
144
|
+
# which does not exist on the Ruby 3.3 floor this gem supports, and
|
|
145
|
+
# Performance/Detect rewrites `select.last` straight back to `reverse.find`.
|
|
146
|
+
line = nil
|
|
147
|
+
lines.each { |l| line = l if l.match?(/error|fail|cause/i) }
|
|
148
|
+
line ||= lines.last
|
|
149
|
+
line.to_s[0, 300]
|
|
150
|
+
end
|
|
151
|
+
|
|
152
|
+
# Turns the probe's line protocol into the #environment Hash. An empty locale
|
|
153
|
+
# line means "unset", which is the interesting case, so it maps to +nil+ rather
|
|
154
|
+
# than an empty String.
|
|
155
|
+
def parse_environment(stdout)
|
|
156
|
+
env = Sandbox.empty_environment
|
|
157
|
+
stdout.each_line do |line|
|
|
158
|
+
case line.chomp
|
|
159
|
+
when /\Alocale=(.*)\z/
|
|
160
|
+
env[:locale] = $1.empty? ? nil : $1
|
|
161
|
+
when /\Acmd=([^\t]*)\t([^\t]*)\t(.*)\z/
|
|
162
|
+
# A command that prints no version (busybox sh) still counts as present;
|
|
163
|
+
# an unreadable version is not evidence of a wrong one.
|
|
164
|
+
env[:commands][$1] = {path: $2, version: $3[/(\d+(?:\.\d+)+)/]}
|
|
165
|
+
end
|
|
166
|
+
end
|
|
167
|
+
env
|
|
168
|
+
end
|
|
63
169
|
end
|
|
64
170
|
end
|
|
@@ -39,6 +39,23 @@ module Nexo
|
|
|
39
39
|
# supported local runtimes in v1.
|
|
40
40
|
RUNTIMES = {docker: "docker", apple: "container"}.freeze
|
|
41
41
|
|
|
42
|
+
# What each runtime's CLI can actually express. Verified live against
|
|
43
|
+
# +container+ 1.2.2 on macOS and Docker 29.4: Apple's CLI rejects
|
|
44
|
+
# +--security-opt+ and +--pids-limit+ with "Unknown option", which aborts
|
|
45
|
+
# +container run+ before the sandbox ever starts.
|
|
46
|
+
#
|
|
47
|
+
# +writable_tmpfs+ is not about the flag being accepted — Apple accepts
|
|
48
|
+
# +--tmpfs+ — but about what it does: combined with +--read-only+ the mount is
|
|
49
|
+
# NOT writable there, so a read-only rootfs leaves no usable scratch.
|
|
50
|
+
#
|
|
51
|
+
# Anything a runtime cannot honor is reported through #hardening_gaps rather
|
|
52
|
+
# than dropped quietly: running with weaker isolation, or with a workspace you
|
|
53
|
+
# cannot write to, must be visible.
|
|
54
|
+
CAPABILITIES = {
|
|
55
|
+
docker: {security_opt: true, pids_limit: true, writable_tmpfs: true},
|
|
56
|
+
apple: {security_opt: false, pids_limit: false, writable_tmpfs: false}
|
|
57
|
+
}.freeze
|
|
58
|
+
|
|
42
59
|
# The container working directory (default +/workspace+, a container path),
|
|
43
60
|
# the selected +runtime+ (+:docker+ / +:apple+), and the required +image+.
|
|
44
61
|
attr_reader :cwd, :runtime, :image
|
|
@@ -81,10 +98,20 @@ module Nexo
|
|
|
81
98
|
out[:stdout]
|
|
82
99
|
end
|
|
83
100
|
|
|
84
|
-
# Writes +content+ to +path+ (guarded) inside the container
|
|
85
|
-
# on stdin, never interpolated into the argv,
|
|
101
|
+
# Writes +content+ to +path+ (guarded) inside the container, creating parent
|
|
102
|
+
# directories first. Content travels on stdin, never interpolated into the argv,
|
|
103
|
+
# so arbitrary bytes are safe.
|
|
104
|
+
#
|
|
105
|
+
# Matches Local#write on both counts, which it previously did not: without the
|
|
106
|
+
# +mkdir+ a nested path failed with "Directory nonexistent", and because the
|
|
107
|
+
# exit status was discarded the caller was told the write had succeeded. Raises
|
|
108
|
+
# +IOError+ on failure, mirroring how #read raises +Errno::ENOENT+.
|
|
86
109
|
def write(path, content)
|
|
87
|
-
|
|
110
|
+
full = guard_path(path)
|
|
111
|
+
out = exec_stdin!(content, "sh", "-c", 'mkdir -p "$(dirname "$0")" && cat > "$0"', full)
|
|
112
|
+
raise IOError, "container write failed: #{path}: #{out[:stderr].strip}" unless out[:status].zero?
|
|
113
|
+
|
|
114
|
+
full
|
|
88
115
|
end
|
|
89
116
|
|
|
90
117
|
# Returns the container paths matching the glob +pattern+ (guarded). Empty
|
|
@@ -120,6 +147,20 @@ module Nexo
|
|
|
120
147
|
@cid = nil
|
|
121
148
|
end
|
|
122
149
|
|
|
150
|
+
# Hardening the caller asked for that this runtime cannot express, as short
|
|
151
|
+
# human-readable strings. Empty on +:docker+. Callers that require a guarantee
|
|
152
|
+
# should check this rather than assume every runtime honors every knob.
|
|
153
|
+
def hardening_gaps
|
|
154
|
+
gaps = []
|
|
155
|
+
gaps << "--security-opt no-new-privileges is not supported by #{@runtime}" unless capable?(:security_opt)
|
|
156
|
+
gaps << "--pids-limit is not supported by #{@runtime}" if @pids_limit && !capable?(:pids_limit)
|
|
157
|
+
if @readonly_rootfs && !capable?(:writable_tmpfs)
|
|
158
|
+
gaps << "#{@runtime} --tmpfs is not writable under --read-only; " \
|
|
159
|
+
"set readonly_rootfs: false or add a :rw bind for a writable workspace"
|
|
160
|
+
end
|
|
161
|
+
gaps
|
|
162
|
+
end
|
|
163
|
+
|
|
123
164
|
# A short, human-readable description of the execution environment for the
|
|
124
165
|
# agent to inject later (consumed by the Refinements spec; until then the
|
|
125
166
|
# method simply exists and is correct).
|
|
@@ -160,11 +201,11 @@ module Nexo
|
|
|
160
201
|
argv = [@bin, "run", "-d", "--name", @name,
|
|
161
202
|
"--label", "nexo.sandbox.id=#{@name}",
|
|
162
203
|
"--network", @network.to_s,
|
|
163
|
-
"--cap-drop", "ALL"
|
|
164
|
-
|
|
204
|
+
"--cap-drop", "ALL"]
|
|
205
|
+
argv += ["--security-opt", "no-new-privileges"] if capable?(:security_opt)
|
|
165
206
|
@cap_add.each { |cap| argv += ["--cap-add", cap.to_s] }
|
|
166
207
|
argv += ["--read-only", "--tmpfs", "#{@cwd}:rw"] if @readonly_rootfs
|
|
167
|
-
argv += ["--pids-limit", @pids_limit.to_s]
|
|
208
|
+
argv += ["--pids-limit", @pids_limit.to_s] if @pids_limit && capable?(:pids_limit)
|
|
168
209
|
argv += ["--memory", @memory.to_s] unless @memory.nil?
|
|
169
210
|
argv += ["--cpus", @cpus.to_s] unless @cpus.nil?
|
|
170
211
|
argv += ["--user", @user.to_s] unless @user.nil?
|
|
@@ -177,6 +218,11 @@ module Nexo
|
|
|
177
218
|
argv
|
|
178
219
|
end
|
|
179
220
|
|
|
221
|
+
# Whether this runtime's CLI can express +knob+.
|
|
222
|
+
def capable?(knob)
|
|
223
|
+
CAPABILITIES.fetch(@runtime, CAPABILITIES[:docker]).fetch(knob, true)
|
|
224
|
+
end
|
|
225
|
+
|
|
180
226
|
# Builds one +-v host:ctr:mode+ spec. Binds default to +:ro+; a Hash form
|
|
181
227
|
# +{ to:, mode: :rw }+ makes a single bind writable.
|
|
182
228
|
def bind_spec(host, dst)
|
|
@@ -25,8 +25,29 @@ module Nexo
|
|
|
25
25
|
# default sandbox stays +:virtual+.
|
|
26
26
|
class Remote < Sandbox
|
|
27
27
|
# Stores any object responding to +read+/+write+/+exec+/+close+.
|
|
28
|
-
|
|
28
|
+
#
|
|
29
|
+
# +instructions:+ describes the remote environment for the agent's system
|
|
30
|
+
# prompt — working directory, available tooling, what is writable. Local and
|
|
31
|
+
# Container derive that themselves; a remote sandbox cannot, because only the
|
|
32
|
+
# shim knows where it points. Supplying it is strongly recommended: the tier
|
|
33
|
+
# most likely to surprise a weak tool-caller is the one it knows least about.
|
|
34
|
+
def initialize(client:, instructions: nil)
|
|
29
35
|
@client = client
|
|
36
|
+
@instructions = instructions
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
# A short, plain-text description of the execution environment. Falls back to an
|
|
40
|
+
# honest generic statement rather than +nil+, so an agent is never left assuming
|
|
41
|
+
# it runs on the host.
|
|
42
|
+
#
|
|
43
|
+
# NOTE: unlike Local and Container, +Remote+ performs NO path confinement — every
|
|
44
|
+
# path is passed to the client untouched. Confining the agent to a working
|
|
45
|
+
# directory is the client's responsibility, not Nexo's.
|
|
46
|
+
def instructions
|
|
47
|
+
@instructions ||
|
|
48
|
+
"You run inside a remote sandbox managed by an external provider. The " \
|
|
49
|
+
"working directory, available tooling, and writable paths are defined by " \
|
|
50
|
+
"that provider."
|
|
30
51
|
end
|
|
31
52
|
|
|
32
53
|
# Reads +path+ via the client.
|