nexo_ai 0.8.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/docs/workflows.md CHANGED
@@ -115,7 +115,7 @@ the count staged.
115
115
 
116
116
  `artifact(name, content:)` records a **named deliverable** on the run — a digest,
117
117
  a report, an improved file, a generated script. The body is written to the
118
- sandbox at `/artifacts/<name>` (so later steps can read it) and recorded on the
118
+ sandbox at `artifacts/<name>` (so later steps can read it) and recorded on the
119
119
  run. `run.artifacts` reads it back as an **ordered array** of string-keyed hashes
120
120
  (`{"name" =>, "content" =>, "at" =>}`) in both stores:
121
121
 
@@ -150,6 +150,52 @@ artifact("digest.md", from: "app/templates/digest.md.erb",
150
150
 
151
151
  See [`examples/artifact_from_template.rb`](../examples/artifact_from_template.rb) for
152
152
  the full offline flow (`ruby -Ilib examples/artifact_from_template.rb`).
153
+ ### Agent output — `produces`
154
+
155
+ `from:` renders ERB and is **only** for templates you wrote. An agent's output is model
156
+ output, so it is copied **verbatim** instead. An agent declares what it produces, and may
157
+ produce as many artifacts as it likes:
158
+
159
+ ```ruby
160
+ class Publisher < Nexo::Agent
161
+ skills :dashboard_designer
162
+ produces "dashboard.html", "digest.json", "out/*.csv"
163
+ end
164
+ ```
165
+
166
+ `run_agent` copies those out of the sandbox and records them on the run the moment the
167
+ agent finishes — **including when it suspended for approval or raised**, which are exactly
168
+ the paths where results would otherwise be lost. Under the hood that is
169
+ `artifact(name, path:)`, the verbatim third mode, which you can also call directly.
170
+
171
+ Why this matters: a workflow builds **one** sandbox that every `run_agent` borrows, so a
172
+ hand-off between stages is just "agent B reads a path agent A wrote". On `:local` that is a
173
+ real directory and survives anything. On `:docker`/`:apple` it does not — `Container#close`
174
+ is `rm -f`, and a run's sandbox is released on **every** terminal path, `suspended`
175
+ included. So pausing for a human approval destroyed everything produced before the pause,
176
+ while the identical code on `:local` kept it. Declared artifacts survive the teardown.
177
+
178
+ Declared, never inferred: sweeping the sandbox would also collect staged skill scripts,
179
+ templates and scratch files, and naming outputs is the only honest way for an agent to say
180
+ it produced nothing. A declared artifact that was never written is skipped, not fatal.
181
+
182
+ Bytes that are not valid UTF-8 are Base64-wrapped so they survive a JSON column. Read a
183
+ body back with `Nexo::Workflow.artifact_body(art)` rather than `art["content"]`, which is
184
+ Base64 text for binary.
185
+
186
+ ### Handing output to the next stage — `restore_artifacts`
187
+
188
+ ```ruby
189
+ run_agent("extract") # Extractor produces gmail.json
190
+ restore_artifacts # put recorded artifacts back in the sandbox
191
+ run_agent("synthesize") # Synthesizer reads gmail.json
192
+ ```
193
+
194
+ `Skills.materialize` gets a skill's files *into* a sandbox; this is the other direction and
195
+ then back in again, so a pipeline behaves the same whether the tier is persistent or
196
+ ephemeral — including across a suspend and resume, where the sandbox that held the files no
197
+ longer exists. Pass `only:` to restore a subset.
198
+
153
199
 
154
200
  The `artifacts` column ships with fresh installs. Apps installed before this
155
201
  release add it with:
data/lib/nexo/agent.rb CHANGED
@@ -22,7 +22,7 @@ module Nexo
22
22
  # configuration instead of silently resetting to defaults.
23
23
  CONFIG_IVARS = %i[
24
24
  @model @assume_model_exists @provider @sandbox @permissions @instructions
25
- @skills @mcp @mcp_allow @fetch_allow @search_backend
25
+ @skills @mcp @mcp_allow @fetch_allow @search_backend @requires @produces
26
26
  ].freeze
27
27
 
28
28
  class << self
@@ -102,6 +102,59 @@ module Nexo
102
102
  names.empty? ? (@skills || []) : (@skills = ((@skills || []) + names).uniq)
103
103
  end
104
104
 
105
+ # What this agent needs from whatever sandbox it runs in — checked once,
106
+ # before the first turn, against Sandbox#environment:
107
+ #
108
+ # class Publisher < Nexo::Agent
109
+ # skills :dashboard_designer
110
+ # requires commands: {"ruby" => ">= 3.1"}, locale: :utf8
111
+ # end
112
+ #
113
+ # Declared HERE, in Nexo's own vocabulary, rather than in the skill file:
114
+ # whoever wires an agent to a sandbox is the only person who can *fix* a
115
+ # gap, so the declaration and the fix live in the same place. A skill states
116
+ # its needs in prose via +compatibility:+, which is the spec's field for it
117
+ # and is aimed at a human or a model.
118
+ #
119
+ # +commands:+ maps a command that must be on +PATH+ to a version constraint
120
+ # — a Gem::Requirement string (+">= 3.1"+) or +"*"+ for "any version". A
121
+ # command whose version cannot be read (busybox +sh+ prints none) satisfies
122
+ # any constraint by being present: an unreadable version is not evidence of
123
+ # a wrong one. +locale:+ takes +:utf8+ (any UTF-8 locale — the useful case)
124
+ # or an exact String.
125
+ #
126
+ # Deliberately coarse, and never packages: gems, wheels and npm modules
127
+ # belong to the image, not here (see Sandbox#environment).
128
+ def requires(commands: nil, locale: nil)
129
+ return @requires if commands.nil? && locale.nil?
130
+
131
+ @requires = {commands: commands || {}, locale: locale}
132
+ end
133
+
134
+ # The artifacts this agent produces — sandbox paths, or globs, that a
135
+ # workflow copies out of the sandbox and records on the run as soon as the
136
+ # agent finishes:
137
+ #
138
+ # class Publisher < Nexo::Agent
139
+ # produces "dashboard.html", "digest.json", "out/*.csv"
140
+ # end
141
+ #
142
+ # Accumulating and deduped, like +skills+: an agent may produce many
143
+ # artifacts, and several +produces+ lines add up rather than replacing.
144
+ #
145
+ # This is the agent's OUTPUT, not a template. Workflow#artifact's +from:+
146
+ # mode renders ERB and is only ever for files you wrote; what an agent
147
+ # produces is model output and is copied verbatim (Workflow#artifact +path:+).
148
+ #
149
+ # Declared rather than inferred because a sweep of the sandbox would also
150
+ # collect staged skill scripts, templates and scratch files — and because
151
+ # naming outputs is the only honest way for an agent to say it produced
152
+ # nothing. A declared artifact that is absent when the agent finishes is
153
+ # skipped, not fatal.
154
+ def produces(*names)
155
+ names.empty? ? (@produces || []) : (@produces = ((@produces || []) + names).uniq)
156
+ end
157
+
105
158
  # Declares an MCP server for this agent (Spec 6). Accumulating: multiple
106
159
  # +mcp+ lines are collected. With no args (+name+ nil and +opts+ empty) it
107
160
  # reads the list (default +[]+); otherwise it appends the friendly,
@@ -246,6 +299,7 @@ module Nexo
246
299
  apply_mcp(c)
247
300
  apply_fetch(c)
248
301
  apply_search(c)
302
+ apply_tool_concurrency(c)
249
303
  c
250
304
  end
251
305
 
@@ -254,9 +308,87 @@ module Nexo
254
308
  # so swapping +loop:+ swaps the engine without touching this class. The
255
309
  # optional +&on_event+ block receives +(type, payload)+ progress events.
256
310
  def prompt(text, max_turns: 25, &on_event)
311
+ verify_environment!
257
312
  @loop.run(agent: self, prompt: text, max_turns: max_turns, &on_event)
258
313
  end
259
314
 
315
+ # Checks the agent's declared ::requires against its sandbox, once, before the
316
+ # first turn. A no-op — and zero cost, no probe at all — when nothing is
317
+ # declared, which is the default.
318
+ #
319
+ # Fails BEFORE the model is called, because the alternative is what this
320
+ # exists to prevent: the agent spends turns deciding to run a script, runs it,
321
+ # and gets +sh: ruby: not found+ or an Encoding::InvalidByteSequenceError three
322
+ # frames into a JSON parse. One legible sentence naming what is missing and
323
+ # where beats a stack trace after the fact.
324
+ #
325
+ # Raises Nexo::EnvironmentError listing every unmet requirement at once, so a
326
+ # misprovisioned image is fixed in one pass rather than one round trip per
327
+ # missing command.
328
+ def verify_environment!
329
+ return if @environment_verified
330
+
331
+ req = self.class.requires
332
+ @environment_verified = true
333
+ return if req.nil?
334
+
335
+ missing = unmet_requirements(req, @sandbox.environment)
336
+ return if missing.empty?
337
+
338
+ raise Nexo::EnvironmentError,
339
+ "#{self.class} cannot run here: #{missing.join("; ")}. " \
340
+ "Provision the sandbox, or drop the `requires` declaration."
341
+ end
342
+
343
+ private
344
+
345
+ # The declared requirements this environment does not meet, as human-readable
346
+ # phrases. Reports the probe's own failure first when it could not run at all:
347
+ # "no ruby" and "I never got to look" are different problems with different
348
+ # fixes, and saying the first when the second is true sends the reader to the
349
+ # wrong place.
350
+ def unmet_requirements(req, env)
351
+ return ["could not inspect the sandbox (#{env[:error]})"] if env[:error]
352
+
353
+ missing = req[:commands].filter_map do |name, constraint|
354
+ found = env[:commands][name.to_s]
355
+ next "no #{name} on PATH" if found.nil?
356
+ next unless version_short?(found[:version], constraint)
357
+
358
+ "#{name} #{found[:version]} does not satisfy #{constraint}"
359
+ end
360
+ missing << locale_complaint(req[:locale], env[:locale]) if locale_unmet?(req[:locale], env[:locale])
361
+ missing
362
+ end
363
+
364
+ # Whether a found command's version fails +constraint+. A missing version, an
365
+ # unparseable constraint, or +"*"+ all pass: presence is the requirement, and
366
+ # an unreadable version is not evidence of a wrong one.
367
+ def version_short?(version, constraint)
368
+ return false if version.nil? || constraint.nil? || constraint.to_s.strip == "*"
369
+
370
+ !Gem::Requirement.new(constraint.to_s).satisfied_by?(Gem::Version.new(version))
371
+ rescue ArgumentError
372
+ false
373
+ end
374
+
375
+ # +:utf8+ asks for any UTF-8 locale (the only case that comes up in practice);
376
+ # a String asks for that exact locale. An unset locale never satisfies either
377
+ # — that is the container default, and the bug this catches.
378
+ def locale_unmet?(required, actual)
379
+ return false if required.nil?
380
+ return true if actual.nil?
381
+
382
+ (required == :utf8) ? !actual.match?(/utf-?8/i) : actual != required.to_s
383
+ end
384
+
385
+ def locale_complaint(required, actual)
386
+ want = (required == :utf8) ? "a UTF-8 locale" : "locale #{required}"
387
+ actual.nil? ? "no locale set (needs #{want})" : "locale #{actual} is not #{want}"
388
+ end
389
+
390
+ public
391
+
260
392
  # The agent's Nexo permission mode mapped onto an opt-in backend's own
261
393
  # permission vocabulary (see PERMISSION_MODE_MAP). Consumed by
262
394
  # Loops::AgentSDK; the default Loops::RubyLLM does its gating inside the
@@ -336,7 +468,7 @@ module Nexo
336
468
  texts = []
337
469
  texts << @instructions if @instructions
338
470
  texts << @sandbox.instructions if @sandbox.instructions
339
- self.class.skills.each { |name| texts << Skills.find(name).content }
471
+ self.class.skills.each { |name| texts << skill_instructions(Skills.find(name)) }
340
472
 
341
473
  texts.each_with_index do |text, i|
342
474
  chat = chat.with_instructions(text, append: i.positive?)
@@ -344,6 +476,32 @@ module Nexo
344
476
  chat
345
477
  end
346
478
 
479
+ # One skill's contribution to the system prompt: its body, plus its
480
+ # +compatibility:+ frontmatter when set.
481
+ #
482
+ # +compatibility:+ is the Agent Skills spec's own field for what a skill needs
483
+ # in order to run ("requires a Ruby interpreter and a UTF-8 locale"), and it is
484
+ # free text by design — the spec deliberately does not make it machine-checkable.
485
+ # Nexo parsed it and then dropped it, so the one place an author can state a
486
+ # requirement never reached the model. That matters now that a skill can ship a
487
+ # script Nexo stages into a sandbox (Skills.materialize): the environment is no
488
+ # longer the author's machine, and the model is the one deciding whether to run
489
+ # the thing.
490
+ #
491
+ # Labelled rather than concatenated, so the model can tell a requirement from an
492
+ # instruction. Absent or blank +compatibility:+ contributes nothing, leaving the
493
+ # prompt byte-for-byte as before for every skill that does not set it. +license:+
494
+ # and +allowed-tools:+ are deliberately NOT surfaced: the first is prompt noise,
495
+ # and the second would be a second source of truth about what an agent may do,
496
+ # competing with Nexo::Permissions, which is the real gate.
497
+ def skill_instructions(skill)
498
+ body = skill.content.to_s
499
+ compat = skill.compatibility.to_s.strip if skill.respond_to?(:compatibility)
500
+ return body if compat.nil? || compat.empty?
501
+
502
+ "#{body}\n\nCompatibility: #{compat}".strip
503
+ end
504
+
347
505
  # Lazily connects the declared MCP servers and attaches their tools, each
348
506
  # wrapped in a MCP::GatedTool so every invocation is authorized through this
349
507
  # agent's Permissions first. Attached after the four sandbox tools and the
@@ -360,11 +518,49 @@ module Nexo
360
518
 
361
519
  @mcp_clients ||= self.class.mcp.map { |cfg| Nexo::MCP.build(**cfg) }
362
520
  gated = @mcp_clients.flat_map(&:tools).map do |tool|
363
- Nexo::MCP::GatedTool.new(tool: tool, permissions: @permissions)
521
+ wrap_mcp_tool(Nexo::MCP::GatedTool.new(tool: tool, permissions: @permissions))
364
522
  end
365
523
  chat.with_tools(*gated) unless gated.empty?
366
524
  end
367
525
 
526
+ # Applies Nexo.config.tool_concurrency to the chat, LAST — every +with_tools+ call
527
+ # resets the chat's concurrency to its current value, so setting it before the
528
+ # tools are attached would be undone. +nil+ leaves RubyLLM's own setting alone.
529
+ def apply_tool_concurrency(chat)
530
+ mode = Nexo.config.tool_concurrency
531
+ return if mode.nil?
532
+ return unless chat.respond_to?(:with_tools)
533
+
534
+ chat.with_tools(concurrency: mode)
535
+ end
536
+
537
+ # Hook: called once per MCP tool, AFTER it is wrapped in a permission-gating
538
+ # MCP::GatedTool and before it is attached. Returns the tool to attach; the
539
+ # default returns it unchanged.
540
+ #
541
+ # Override to decorate every MCP tool an agent gets — capping an oversized reply,
542
+ # recording timings, redacting a field. A third-party MCP server's response shape
543
+ # is not yours to change, and its tools are attached by the harness rather than by
544
+ # you, so without this seam the only way in was to override the private
545
+ # +#apply_mcp+.
546
+ #
547
+ # class MailAgent < Nexo::Agent
548
+ # mcp :mail, transport: :stdio, command: "apple-mail-mcp"
549
+ #
550
+ # def wrap_mcp_tool(tool)
551
+ # CappedTool.new(tool: tool, max_chars: 4_000)
552
+ # end
553
+ # end
554
+ #
555
+ # A wrapper must keep the duck type the chat relies on — +#name+, +#description+,
556
+ # +#params_schema+ and +#call+. GatedTool delegates the rest through
557
+ # +method_missing+, so wrappers compose. Gating happens underneath, so a wrapper
558
+ # only ever sees an already-authorized call: wrapping cannot widen what the agent
559
+ # may do.
560
+ def wrap_mcp_tool(tool)
561
+ tool
562
+ end
563
+
368
564
  # Attaches a single Nexo::Tools::Fetch scoped to the agent's +fetch_allow+
369
565
  # hosts (Spec 9). Returns early when no host is declared, so an agent that never
370
566
  # calls +fetch_allow+ gets no fetch tool. Attached right after +apply_mcp+ so
@@ -25,6 +25,14 @@ module Nexo
25
25
  # which always runs its own reactor when called.
26
26
  attr_accessor :concurrency
27
27
 
28
+ # How RubyLLM executes MULTIPLE tool calls returned in a single assistant turn:
29
+ # +:fibers+ (via +async+, the same driver Nexo.concurrent uses), +:threads+, or
30
+ # +false+ for one at a time. +nil+ (the default) leaves RubyLLM's own setting
31
+ # alone, so existing behaviour is unchanged until you opt in.
32
+ #
33
+ # Distinct from #concurrency, which governs how Nexo itself runs blocking work.
34
+ attr_accessor :tool_concurrency
35
+
28
36
  # Max in-flight tasks for Nexo.concurrent fan-out (Spec 5), default +8+.
29
37
  # The single most important knob for staying under provider rate limits.
30
38
  attr_accessor :max_in_flight
@@ -22,10 +22,19 @@ module Nexo
22
22
  # +chat:+ lets a Nexo::Session inject the hydrated, continuing chat so the
23
23
  # loop runs over the persisted thread; left nil it builds the agent's own
24
24
  # fresh chat exactly as before — the default (no-session) path is unchanged.
25
+ # +max_turns+ is a BUDGET, not a hard stop. ruby_llm runs the whole tool loop
26
+ # inside #ask and its callbacks cannot halt it, so this loop cannot cut a run
27
+ # short. What it can do is COUNT the tool calls and say so: exceeding the budget
28
+ # emits a +:turn_limit_exceeded+ event (once) carrying the count and the limit.
29
+ #
30
+ # Previously the parameter was accepted and never read, which made it look like
31
+ # a safety bound it has never been. Treat it as telemetry: if you need a hard
32
+ # ceiling, bound the work inside your tools, or use Loops::AgentSDK whose engine
33
+ # enforces its own max_turns natively.
25
34
  def run(agent:, prompt:, max_turns: 25, chat: nil, &on_event)
26
35
  chat ||= agent.chat
27
36
 
28
- wire_observability(chat, &on_event)
37
+ wire_observability(chat, max_turns: max_turns, &on_event)
29
38
 
30
39
  response = chat.ask(prompt)
31
40
  on_event&.call(:done, response)
@@ -47,20 +56,39 @@ module Nexo
47
56
  # prompt observes with its own block rather than the first one. A fresh chat
48
57
  # (the default per-prompt path) simply wires once — byte-for-byte the prior
49
58
  # behavior.
50
- def wire_observability(chat, &on_event)
59
+ def wire_observability(chat, max_turns: nil, &on_event)
51
60
  return unless chat.respond_to?(:before_tool_call) && chat.respond_to?(:after_tool_result)
52
61
 
53
62
  chat.instance_variable_set(:@nexo_on_event, on_event)
63
+ chat.instance_variable_set(:@nexo_max_turns, max_turns)
64
+ # The count belongs to the PROMPT, not the chat: a continuing Session runs the
65
+ # loop repeatedly over one chat, and each prompt gets its own budget.
66
+ chat.instance_variable_set(:@nexo_turns, 0)
67
+ chat.instance_variable_set(:@nexo_turns_reported, false)
54
68
  return if chat.instance_variable_get(:@nexo_observed)
55
69
 
56
70
  chat.instance_variable_set(:@nexo_observed, true)
57
71
  chat.before_tool_call do |tc|
72
+ count_turn(chat)
58
73
  chat.instance_variable_get(:@nexo_on_event)&.call(:tool_call, tc)
59
74
  end
60
75
  chat.after_tool_result do |r|
61
76
  chat.instance_variable_get(:@nexo_on_event)&.call(:tool_result, r)
62
77
  end
63
78
  end
79
+
80
+ # Counts one tool call and reports the first time the budget is passed. Reported
81
+ # once per prompt, not per call, so a long run does not drown its own event log.
82
+ def count_turn(chat)
83
+ limit = chat.instance_variable_get(:@nexo_max_turns)
84
+ turns = chat.instance_variable_get(:@nexo_turns).to_i + 1
85
+ chat.instance_variable_set(:@nexo_turns, turns)
86
+ return unless limit && turns > limit
87
+ return if chat.instance_variable_get(:@nexo_turns_reported)
88
+
89
+ chat.instance_variable_set(:@nexo_turns_reported, true)
90
+ chat.instance_variable_get(:@nexo_on_event)&.call(:turn_limit_exceeded, {turns: turns, max_turns: limit})
91
+ end
64
92
  end
65
93
  end
66
94
  end
data/lib/nexo/sandbox.rb CHANGED
@@ -60,5 +60,111 @@ module Nexo
60
60
  def mtime(path)
61
61
  nil
62
62
  end
63
+
64
+ # Commands the default #environment probe looks for. Deliberately short —
65
+ # each entry costs a +command -v+ plus one +--version+ when found — and
66
+ # extensible per call for anything else a skill's script might need.
67
+ PROBE_COMMANDS = %w[ruby python3 node sh].freeze
68
+
69
+ # The shape #environment always answers with, built fresh on every call.
70
+ # Deliberately a method and not a frozen constant: the +:commands+ Hash is
71
+ # mutated while parsing, and +CONST.dup+ is shallow — sharing one inner Hash
72
+ # let every probe accumulate into the constant, so a sandbox with no shell
73
+ # answered with the previous sandbox's findings. +:error+ is nil on a
74
+ # successful probe and carries the reason when one could not be run.
75
+ def self.empty_environment
76
+ {commands: {}, locale: nil, error: nil}
77
+ end
78
+
79
+ # One POSIX +sh+ script, no interpreter required on the far side, so the probe
80
+ # works on a busybox image. +command -v+ locates each command and a single
81
+ # +--version+ reports it; the format placeholder is filled by #environment.
82
+ PROBE_SCRIPT = <<~SH
83
+ printf 'locale=%%s\\n' "${LC_ALL:-${LC_CTYPE:-${LANG:-}}}"
84
+ for c in %<commands>s; do
85
+ p=$(command -v "$c" 2>/dev/null) || continue
86
+ v=$("$c" --version 2>/dev/null | head -1)
87
+ printf 'cmd=%%s\\t%%s\\t%%s\\n' "$c" "$p" "$v"
88
+ done
89
+ SH
90
+
91
+ # What this execution environment actually provides, as data:
92
+ #
93
+ # sandbox.environment
94
+ # # => { commands: { "ruby" => { path: "/usr/local/bin/ruby", version: "4.0.0" } },
95
+ # # locale: "C.UTF-8" }
96
+ #
97
+ # #instructions describes the environment *for the model*; this is the same
98
+ # question answered *for code*, so a caller can check before staging a skill's
99
+ # script rather than discovering the answer as a stack trace several turns in.
100
+ # A container typically has no locale at all, under which Ruby's default
101
+ # external encoding is US-ASCII and a bare File.read on a UTF-8 file raises —
102
+ # and an image can carry a full Ruby toolchain and still report no locale, so
103
+ # the two are reported as independent axes.
104
+ #
105
+ # Deliberately coarse: commands on +PATH+ and the locale, never packages. Gems,
106
+ # wheels and npm modules belong to whoever builds the image, and modelling them
107
+ # here would be a cross-language dependency resolver competing with the manifest
108
+ # every ecosystem already has.
109
+ #
110
+ # Costs one #shell round trip and is memoized for the sandbox's lifetime (0.14s
111
+ # on +:local+, 0.25–0.48s on a container, measured). A sandbox with no shell
112
+ # A sandbox with no shell (Virtual) reports empty. A probe that fails for any
113
+ # other reason — the container would not start, the client is unreachable —
114
+ # also reports empty, but carries the reason under +:error+: this is
115
+ # diagnostics and must never be the reason a run dies, yet "I probed and found
116
+ # nothing" and "I could not probe" are different answers and a caller building
117
+ # an error message needs to tell them apart.
118
+ def environment(commands: PROBE_COMMANDS)
119
+ @environment ||= {}
120
+ @environment[commands] ||= probe_environment(commands)
121
+ end
122
+
123
+ private
124
+
125
+ # Runs the probe and parses its output. Rescues everything: see #environment.
126
+ def probe_environment(commands)
127
+ return Sandbox.empty_environment.merge(error: "sandbox has no shell") unless supports?(:shell)
128
+
129
+ out = shell(format(PROBE_SCRIPT, commands: Array(commands).join(" ")))
130
+ parse_environment(out[:stdout].to_s)
131
+ rescue StandardError, NotImplementedError => e
132
+ Sandbox.empty_environment.merge(error: "#{e.class}: #{salient_line(e.message)}")
133
+ end
134
+
135
+ # The most useful single line of a multi-line runtime failure. Neither the
136
+ # first line nor a blind truncation works: a container runtime prints a
137
+ # progress banner first ("[1/6] Fetching image", "container start failed")
138
+ # and puts the actual cause last ("The volume is read only"). Prefer the last
139
+ # line that announces an error, else the last non-empty line.
140
+ def salient_line(message)
141
+ lines = message.to_s.lines.map(&:strip).reject(&:empty?)
142
+ # Written as a plain scan on purpose. The idiomatic spellings deadlock the
143
+ # linter — Style/ReverseFind rewrites `reverse.find` to Enumerable#rfind,
144
+ # which does not exist on the Ruby 3.3 floor this gem supports, and
145
+ # Performance/Detect rewrites `select.last` straight back to `reverse.find`.
146
+ line = nil
147
+ lines.each { |l| line = l if l.match?(/error|fail|cause/i) }
148
+ line ||= lines.last
149
+ line.to_s[0, 300]
150
+ end
151
+
152
+ # Turns the probe's line protocol into the #environment Hash. An empty locale
153
+ # line means "unset", which is the interesting case, so it maps to +nil+ rather
154
+ # than an empty String.
155
+ def parse_environment(stdout)
156
+ env = Sandbox.empty_environment
157
+ stdout.each_line do |line|
158
+ case line.chomp
159
+ when /\Alocale=(.*)\z/
160
+ env[:locale] = $1.empty? ? nil : $1
161
+ when /\Acmd=([^\t]*)\t([^\t]*)\t(.*)\z/
162
+ # A command that prints no version (busybox sh) still counts as present;
163
+ # an unreadable version is not evidence of a wrong one.
164
+ env[:commands][$1] = {path: $2, version: $3[/(\d+(?:\.\d+)+)/]}
165
+ end
166
+ end
167
+ env
168
+ end
63
169
  end
64
170
  end
@@ -39,6 +39,23 @@ module Nexo
39
39
  # supported local runtimes in v1.
40
40
  RUNTIMES = {docker: "docker", apple: "container"}.freeze
41
41
 
42
+ # What each runtime's CLI can actually express. Verified live against
43
+ # +container+ 1.2.2 on macOS and Docker 29.4: Apple's CLI rejects
44
+ # +--security-opt+ and +--pids-limit+ with "Unknown option", which aborts
45
+ # +container run+ before the sandbox ever starts.
46
+ #
47
+ # +writable_tmpfs+ is not about the flag being accepted — Apple accepts
48
+ # +--tmpfs+ — but about what it does: combined with +--read-only+ the mount is
49
+ # NOT writable there, so a read-only rootfs leaves no usable scratch.
50
+ #
51
+ # Anything a runtime cannot honor is reported through #hardening_gaps rather
52
+ # than dropped quietly: running with weaker isolation, or with a workspace you
53
+ # cannot write to, must be visible.
54
+ CAPABILITIES = {
55
+ docker: {security_opt: true, pids_limit: true, writable_tmpfs: true},
56
+ apple: {security_opt: false, pids_limit: false, writable_tmpfs: false}
57
+ }.freeze
58
+
42
59
  # The container working directory (default +/workspace+, a container path),
43
60
  # the selected +runtime+ (+:docker+ / +:apple+), and the required +image+.
44
61
  attr_reader :cwd, :runtime, :image
@@ -81,10 +98,20 @@ module Nexo
81
98
  out[:stdout]
82
99
  end
83
100
 
84
- # Writes +content+ to +path+ (guarded) inside the container. Content travels
85
- # on stdin, never interpolated into the argv, so arbitrary bytes are safe.
101
+ # Writes +content+ to +path+ (guarded) inside the container, creating parent
102
+ # directories first. Content travels on stdin, never interpolated into the argv,
103
+ # so arbitrary bytes are safe.
104
+ #
105
+ # Matches Local#write on both counts, which it previously did not: without the
106
+ # +mkdir+ a nested path failed with "Directory nonexistent", and because the
107
+ # exit status was discarded the caller was told the write had succeeded. Raises
108
+ # +IOError+ on failure, mirroring how #read raises +Errno::ENOENT+.
86
109
  def write(path, content)
87
- exec_stdin!(content, "sh", "-c", 'cat > "$0"', guard_path(path))
110
+ full = guard_path(path)
111
+ out = exec_stdin!(content, "sh", "-c", 'mkdir -p "$(dirname "$0")" && cat > "$0"', full)
112
+ raise IOError, "container write failed: #{path}: #{out[:stderr].strip}" unless out[:status].zero?
113
+
114
+ full
88
115
  end
89
116
 
90
117
  # Returns the container paths matching the glob +pattern+ (guarded). Empty
@@ -120,6 +147,20 @@ module Nexo
120
147
  @cid = nil
121
148
  end
122
149
 
150
+ # Hardening the caller asked for that this runtime cannot express, as short
151
+ # human-readable strings. Empty on +:docker+. Callers that require a guarantee
152
+ # should check this rather than assume every runtime honors every knob.
153
+ def hardening_gaps
154
+ gaps = []
155
+ gaps << "--security-opt no-new-privileges is not supported by #{@runtime}" unless capable?(:security_opt)
156
+ gaps << "--pids-limit is not supported by #{@runtime}" if @pids_limit && !capable?(:pids_limit)
157
+ if @readonly_rootfs && !capable?(:writable_tmpfs)
158
+ gaps << "#{@runtime} --tmpfs is not writable under --read-only; " \
159
+ "set readonly_rootfs: false or add a :rw bind for a writable workspace"
160
+ end
161
+ gaps
162
+ end
163
+
123
164
  # A short, human-readable description of the execution environment for the
124
165
  # agent to inject later (consumed by the Refinements spec; until then the
125
166
  # method simply exists and is correct).
@@ -160,11 +201,11 @@ module Nexo
160
201
  argv = [@bin, "run", "-d", "--name", @name,
161
202
  "--label", "nexo.sandbox.id=#{@name}",
162
203
  "--network", @network.to_s,
163
- "--cap-drop", "ALL",
164
- "--security-opt", "no-new-privileges"]
204
+ "--cap-drop", "ALL"]
205
+ argv += ["--security-opt", "no-new-privileges"] if capable?(:security_opt)
165
206
  @cap_add.each { |cap| argv += ["--cap-add", cap.to_s] }
166
207
  argv += ["--read-only", "--tmpfs", "#{@cwd}:rw"] if @readonly_rootfs
167
- argv += ["--pids-limit", @pids_limit.to_s] unless @pids_limit.nil?
208
+ argv += ["--pids-limit", @pids_limit.to_s] if @pids_limit && capable?(:pids_limit)
168
209
  argv += ["--memory", @memory.to_s] unless @memory.nil?
169
210
  argv += ["--cpus", @cpus.to_s] unless @cpus.nil?
170
211
  argv += ["--user", @user.to_s] unless @user.nil?
@@ -177,6 +218,11 @@ module Nexo
177
218
  argv
178
219
  end
179
220
 
221
+ # Whether this runtime's CLI can express +knob+.
222
+ def capable?(knob)
223
+ CAPABILITIES.fetch(@runtime, CAPABILITIES[:docker]).fetch(knob, true)
224
+ end
225
+
180
226
  # Builds one +-v host:ctr:mode+ spec. Binds default to +:ro+; a Hash form
181
227
  # +{ to:, mode: :rw }+ makes a single bind writable.
182
228
  def bind_spec(host, dst)
@@ -25,8 +25,29 @@ module Nexo
25
25
  # default sandbox stays +:virtual+.
26
26
  class Remote < Sandbox
27
27
  # Stores any object responding to +read+/+write+/+exec+/+close+.
28
- def initialize(client:)
28
+ #
29
+ # +instructions:+ describes the remote environment for the agent's system
30
+ # prompt — working directory, available tooling, what is writable. Local and
31
+ # Container derive that themselves; a remote sandbox cannot, because only the
32
+ # shim knows where it points. Supplying it is strongly recommended: the tier
33
+ # most likely to surprise a weak tool-caller is the one it knows least about.
34
+ def initialize(client:, instructions: nil)
29
35
  @client = client
36
+ @instructions = instructions
37
+ end
38
+
39
+ # A short, plain-text description of the execution environment. Falls back to an
40
+ # honest generic statement rather than +nil+, so an agent is never left assuming
41
+ # it runs on the host.
42
+ #
43
+ # NOTE: unlike Local and Container, +Remote+ performs NO path confinement — every
44
+ # path is passed to the client untouched. Confining the agent to a working
45
+ # directory is the client's responsibility, not Nexo's.
46
+ def instructions
47
+ @instructions ||
48
+ "You run inside a remote sandbox managed by an external provider. The " \
49
+ "working directory, available tooling, and writable paths are defined by " \
50
+ "that provider."
30
51
  end
31
52
 
32
53
  # Reads +path+ via the client.