maf-sandbox-codeact 0.2.2__tar.gz → 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: maf-sandbox-codeact
3
- Version: 0.2.2
3
+ Version: 0.3.0
4
4
  Summary: CodeAct as a Microsoft Agent Framework tool — the model writes a short Python program and it runs inside a sandbox — written against maf-sandbox so it runs on any sandbox backend.
5
5
  Keywords: codeact,code-execution,sandbox,agent-framework,microsoft-agent-framework
6
6
  Author: SOKOLAI BV
@@ -15,7 +15,7 @@ Classifier: Programming Language :: Python :: 3.12
15
15
  Classifier: Programming Language :: Python :: 3.13
16
16
  Classifier: Programming Language :: Python :: 3.14
17
17
  Classifier: Topic :: Software Development :: Interpreters
18
- Requires-Dist: maf-sandbox>=0.8.0,<0.11
18
+ Requires-Dist: maf-sandbox>=0.11.0,<0.13
19
19
  Requires-Dist: agent-framework-core>=1.13.0,<2
20
20
  Requires-Python: >=3.12, <3.15
21
21
  Project-URL: Homepage, https://www.sokol.ai
@@ -60,20 +60,20 @@ Pass `router=None` — or a router with no backend — and you get `[]` back: an
60
60
 
61
61
  One tool, `execute_code`. The program is written to a directory of its own and run as the argv `["python3", ".../program.py"]`, and the result is its stdout, its stderr when it wrote any, and its exit code when that was not zero. There is no REPL echo, so a program that computes without printing returns a sentence saying so.
62
62
 
63
- **Every call gets a fresh directory, and that is load-bearing rather than hygiene.** `acquire` is get-or-create, so the same sandbox serves every call in a conversation. Without a per-call directory a file deleted from the workspace between rounds would still be there for the next program to read as current, and last round's output file would be collected as this round's — a stale answer presented as a live one, in a kind whose whole job is transforming files.
63
+ **Every call gets a fresh directory, and that is load-bearing rather than hygiene.** `acquire` is get-or-create, so the same sandbox serves every call in a conversation. Without a per-call directory a file deleted from the file store between rounds would still be there for the next program to read as current, and last round's output file would be collected as this round's — a stale answer presented as a live one, in a kind whose whole job is transforming files.
64
64
 
65
65
  Two further channels exist and neither is on by default. Wire neither and this is the stdout-only kind it has always been.
66
66
 
67
67
  ### Files in
68
68
 
69
- Pass a `workspace_store` and the tool grows a `files` parameter:
69
+ Pass a `file_store` and the tool grows a `files` parameter:
70
70
 
71
71
  ```python
72
72
  tools = make_codeact_tools(router, "data-analyst", context,
73
- workspace_store=store, image=...)
73
+ file_store=store, image=...)
74
74
  ```
75
75
 
76
- Each named file is read from the store and written into the program's working directory under its own name, so `data/sales.csv` is what the program opens. **The caller's listing is the authority**: only a name present in `WorkspaceContext.list_files` is ever shared, so a name the model invented — or read out of a file it was given — has nowhere to go. A name outside the listing comes back as a refusal naming the near misses; a name that traverses comes back as a refusal that echoes nothing.
76
+ Each named file is read from the store and written into the program's working directory under its own name, so `data/sales.csv` is what the program opens. **The caller's listing is the authority**: only a name present in `CallerContext.list_files` is ever shared, so a name the model invented — or read out of a file it was given — has nowhere to go. A name outside the listing comes back as a refusal naming the near misses; a name that traverses comes back as a refusal that echoes nothing.
77
77
 
78
78
  ### Files out
79
79
 
@@ -103,11 +103,17 @@ Either way the kind requires `FILES_OUT` and **never** `FILES_LIST`: it collects
103
103
 
104
104
  **Nothing is dispatchable from inside.** There is no host-tool registry in this version, and that emptiness is the security story rather than a missing feature: the program cannot open a socket and cannot call a host function, so it initiates nothing. The output sink does not change that — the kind calls it host-side, after the program has exited, and nothing inside the sandbox can reach it. A host wanting a hard stop denies `FILES_OUT`.
105
105
 
106
- **A workspace store is ingress, and it is the host's own.** "Nothing can get in" describes what the *program* can initiate, not what the host puts there. With `workspace_store` wired, caller-selected files are written into the sandbox before the program runs — deliberately, and constrained to the caller's listing, so the model cannot widen the set. What that content *is* remains the host's to know: a workspace file may itself carry text from somewhere untrusted, and a program that parses it is running on input the sandbox did not vet. Wire no store and this paragraph does not apply.
106
+ **A file store is ingress, and it is the host's own.** "Nothing can get in" describes what the *program* can initiate, not what the host puts there. With `file_store` wired, caller-selected files are written into the sandbox before the program runs — deliberately, and constrained to the caller's listing, so the model cannot widen the set. What that content *is* remains the host's to know: a file in the store may itself carry text from somewhere untrusted, and a program that parses it is running on input the sandbox did not vet. Wire no store and this paragraph does not apply.
107
107
 
108
108
  **The tool declares no `source_integrity`.** The library's default is `"trusted"`, which is right for a workload whose result is a compiler's own diagnostics and wrong for this one: what comes back is whatever a model-written `print(...)` chose to emit. Undeclared, MAF's information-flow tracker applies its untrusted default and the result taints the conversation — the fail-safe direction, and the honest one.
109
109
 
110
- **Isolation is the host's call, and a store changes what that call is about.** This kind does not raise `SandboxSpec.min_isolation`, so the router's floor governs — `MICROVM` unless the host opted down. A kind that ran code influenced by untrusted external content would pin the floor itself, and this one cannot know whether it is one: with no store, the program's only input is source the model wrote, and opting down to `CONTAINER` weighs model-written code against a shared kernel. **With a store, the program also reads whatever those files contain**, so the floor should be chosen against the provenance of the workspace, not against this kind's defaults. Only the host knows that.
110
+ **Isolation is the host's call, and a store changes what that call is about.** This kind does not raise `SandboxSpec.min_isolation`, so the router's floor governs — `MICROVM` unless the host opted down. A kind that ran code influenced by untrusted external content would pin the floor itself, and this one cannot know whether it is one: with no store, the program's only input is source the model wrote, and opting down to `CONTAINER` weighs model-written code against a shared kernel. **With a store, the program also reads whatever those files contain**, so the floor should be chosen against the provenance of the file store, not against this kind's defaults. Only the host knows that.
111
+
112
+ ## Upgrading to 0.3
113
+
114
+ `0.3.0` follows `maf-sandbox` 0.11, which retired the word `workspace` from the vocabulary. It requires that release.
115
+
116
+ **`make_codeact_tools` takes `file_store` where it took `workspace_store`.** This one is **keyword-only**, so unlike the Bicep kind's there is no positional call that survives untouched: every host wiring a store has an edit to make. A host that wires none is unaffected, since the parameter defaults to `None`.
111
117
 
112
118
  ## What this version is not
113
119
 
@@ -35,20 +35,20 @@ Pass `router=None` — or a router with no backend — and you get `[]` back: an
35
35
 
36
36
  One tool, `execute_code`. The program is written to a directory of its own and run as the argv `["python3", ".../program.py"]`, and the result is its stdout, its stderr when it wrote any, and its exit code when that was not zero. There is no REPL echo, so a program that computes without printing returns a sentence saying so.
37
37
 
38
- **Every call gets a fresh directory, and that is load-bearing rather than hygiene.** `acquire` is get-or-create, so the same sandbox serves every call in a conversation. Without a per-call directory a file deleted from the workspace between rounds would still be there for the next program to read as current, and last round's output file would be collected as this round's — a stale answer presented as a live one, in a kind whose whole job is transforming files.
38
+ **Every call gets a fresh directory, and that is load-bearing rather than hygiene.** `acquire` is get-or-create, so the same sandbox serves every call in a conversation. Without a per-call directory a file deleted from the file store between rounds would still be there for the next program to read as current, and last round's output file would be collected as this round's — a stale answer presented as a live one, in a kind whose whole job is transforming files.
39
39
 
40
40
  Two further channels exist and neither is on by default. Wire neither and this is the stdout-only kind it has always been.
41
41
 
42
42
  ### Files in
43
43
 
44
- Pass a `workspace_store` and the tool grows a `files` parameter:
44
+ Pass a `file_store` and the tool grows a `files` parameter:
45
45
 
46
46
  ```python
47
47
  tools = make_codeact_tools(router, "data-analyst", context,
48
- workspace_store=store, image=...)
48
+ file_store=store, image=...)
49
49
  ```
50
50
 
51
- Each named file is read from the store and written into the program's working directory under its own name, so `data/sales.csv` is what the program opens. **The caller's listing is the authority**: only a name present in `WorkspaceContext.list_files` is ever shared, so a name the model invented — or read out of a file it was given — has nowhere to go. A name outside the listing comes back as a refusal naming the near misses; a name that traverses comes back as a refusal that echoes nothing.
51
+ Each named file is read from the store and written into the program's working directory under its own name, so `data/sales.csv` is what the program opens. **The caller's listing is the authority**: only a name present in `CallerContext.list_files` is ever shared, so a name the model invented — or read out of a file it was given — has nowhere to go. A name outside the listing comes back as a refusal naming the near misses; a name that traverses comes back as a refusal that echoes nothing.
52
52
 
53
53
  ### Files out
54
54
 
@@ -78,11 +78,17 @@ Either way the kind requires `FILES_OUT` and **never** `FILES_LIST`: it collects
78
78
 
79
79
  **Nothing is dispatchable from inside.** There is no host-tool registry in this version, and that emptiness is the security story rather than a missing feature: the program cannot open a socket and cannot call a host function, so it initiates nothing. The output sink does not change that — the kind calls it host-side, after the program has exited, and nothing inside the sandbox can reach it. A host wanting a hard stop denies `FILES_OUT`.
80
80
 
81
- **A workspace store is ingress, and it is the host's own.** "Nothing can get in" describes what the *program* can initiate, not what the host puts there. With `workspace_store` wired, caller-selected files are written into the sandbox before the program runs — deliberately, and constrained to the caller's listing, so the model cannot widen the set. What that content *is* remains the host's to know: a workspace file may itself carry text from somewhere untrusted, and a program that parses it is running on input the sandbox did not vet. Wire no store and this paragraph does not apply.
81
+ **A file store is ingress, and it is the host's own.** "Nothing can get in" describes what the *program* can initiate, not what the host puts there. With `file_store` wired, caller-selected files are written into the sandbox before the program runs — deliberately, and constrained to the caller's listing, so the model cannot widen the set. What that content *is* remains the host's to know: a file in the store may itself carry text from somewhere untrusted, and a program that parses it is running on input the sandbox did not vet. Wire no store and this paragraph does not apply.
82
82
 
83
83
  **The tool declares no `source_integrity`.** The library's default is `"trusted"`, which is right for a workload whose result is a compiler's own diagnostics and wrong for this one: what comes back is whatever a model-written `print(...)` chose to emit. Undeclared, MAF's information-flow tracker applies its untrusted default and the result taints the conversation — the fail-safe direction, and the honest one.
84
84
 
85
- **Isolation is the host's call, and a store changes what that call is about.** This kind does not raise `SandboxSpec.min_isolation`, so the router's floor governs — `MICROVM` unless the host opted down. A kind that ran code influenced by untrusted external content would pin the floor itself, and this one cannot know whether it is one: with no store, the program's only input is source the model wrote, and opting down to `CONTAINER` weighs model-written code against a shared kernel. **With a store, the program also reads whatever those files contain**, so the floor should be chosen against the provenance of the workspace, not against this kind's defaults. Only the host knows that.
85
+ **Isolation is the host's call, and a store changes what that call is about.** This kind does not raise `SandboxSpec.min_isolation`, so the router's floor governs — `MICROVM` unless the host opted down. A kind that ran code influenced by untrusted external content would pin the floor itself, and this one cannot know whether it is one: with no store, the program's only input is source the model wrote, and opting down to `CONTAINER` weighs model-written code against a shared kernel. **With a store, the program also reads whatever those files contain**, so the floor should be chosen against the provenance of the file store, not against this kind's defaults. Only the host knows that.
86
+
87
+ ## Upgrading to 0.3
88
+
89
+ `0.3.0` follows `maf-sandbox` 0.11, which retired the word `workspace` from the vocabulary. It requires that release.
90
+
91
+ **`make_codeact_tools` takes `file_store` where it took `workspace_store`.** This one is **keyword-only**, so unlike the Bicep kind's there is no positional call that survives untouched: every host wiring a store has an edit to make. A host that wires none is unaffected, since the parameter defaults to `None`.
86
92
 
87
93
  ## What this version is not
88
94
 
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "maf-sandbox-codeact"
3
- version = "0.2.2"
3
+ version = "0.3.0"
4
4
  description = "CodeAct as a Microsoft Agent Framework tool — the model writes a short Python program and it runs inside a sandbox — written against maf-sandbox so it runs on any sandbox backend."
5
5
  readme = "README.md"
6
6
  requires-python = ">=3.12,<3.15"
@@ -24,7 +24,7 @@ classifiers = [
24
24
  "Topic :: Software Development :: Interpreters",
25
25
  ]
26
26
  dependencies = [
27
- "maf-sandbox>=0.8.0,<0.11",
27
+ "maf-sandbox>=0.11.0,<0.13",
28
28
  "agent-framework-core>=1.13.0,<2",
29
29
  ]
30
30
 
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "maf-sandbox-codeact"
3
- version = "0.2.2"
3
+ version = "0.3.0"
4
4
  description = "CodeAct as a Microsoft Agent Framework tool — the model writes a short Python program and it runs inside a sandbox — written against maf-sandbox so it runs on any sandbox backend."
5
5
  readme = "README.md"
6
6
  requires-python = ">=3.12,<3.15"
@@ -34,7 +34,7 @@ dependencies = [
34
34
  # published wheel resolves a maf-sandbox base missing those imports. The number itself is
35
35
  # below, and `scripts/set_dependents_range.py` moves it — so this says which release the
36
36
  # floor is for, never which release that is.
37
- "maf-sandbox>=0.8.0,<0.11",
37
+ "maf-sandbox>=0.11.0,<0.13",
38
38
  # The `@tool` decorator. MAF is a genuine dependency of the tool itself, not of the host
39
39
  # application.
40
40
  "agent-framework-core>=1.13.0,<2",
@@ -9,7 +9,7 @@ pull surface, so the same tool runs unchanged against ACA Sandboxes, a Docker co
9
9
  in-process fake.
10
10
 
11
11
  Three channels, and the host chooses which of them exist. Stdout is always there. A
12
- **workspace store** adds a ``files`` parameter, so a program can transform files that already
12
+ **file store** adds a ``files`` parameter, so a program can transform files that already
13
13
  exist rather than only data the model wrote into its own source. An **output sink** plus a
14
14
  :class:`CodeactOutputs` mode adds a way for files the program produces to reach host state.
15
15
  Wire neither and this is the stdout-only kind it has always been, with nothing dispatchable
@@ -28,6 +28,7 @@ from uuid import uuid4
28
28
 
29
29
  from maf_sandbox import (
30
30
  DEFAULT_TRANSFER_LIMITS,
31
+ CallerContext,
31
32
  Capability,
32
33
  DeclaredOutput,
33
34
  ExecResult,
@@ -38,7 +39,6 @@ from maf_sandbox import (
38
39
  SandboxRouter,
39
40
  SandboxSpec,
40
41
  TransferLimits,
41
- WorkspaceContext,
42
42
  collect_outputs,
43
43
  error_detail,
44
44
  validate_artifact_name,
@@ -159,9 +159,9 @@ def codeact_sandbox_spec(
159
159
  def make_codeact_tools(
160
160
  router: SandboxRouter | None,
161
161
  agent_dir: str,
162
- context: WorkspaceContext,
162
+ context: CallerContext,
163
163
  *,
164
- workspace_store: "AgentFileStore | None" = None,
164
+ file_store: "AgentFileStore | None" = None,
165
165
  output_sink: OutputSink | None = None,
166
166
  outputs: CodeactOutputs = CodeactOutputs.NONE,
167
167
  outbound_max_confidentiality: str | None = None,
@@ -174,15 +174,15 @@ def make_codeact_tools(
174
174
  """Return the ``[execute_code]`` tool list, or ``[]`` when no sandbox is available.
175
175
 
176
176
  The tool's *signature* follows the channels the host wired: ``files`` appears only with a
177
- ``workspace_store``, and ``outputs`` only under :data:`CodeactOutputs.DECLARED`. A model is
177
+ ``file_store``, and ``outputs`` only under :data:`CodeactOutputs.DECLARED`. A model is
178
178
  never shown a parameter this deployment cannot honour.
179
179
 
180
180
  Args:
181
181
  router: The sandbox router, or ``None`` when sandboxing is not configured.
182
182
  agent_dir: The agent's directory name. Baked into the sandbox key at factory time
183
183
  rather than taken from the model at call time.
184
- context: How to read the caller's scope and thread, and how to enumerate the workspace.
185
- workspace_store: The agent's workspace store. Given one, the tool takes a ``files``
184
+ context: How to read the caller's scope and thread, and how to enumerate the file store.
185
+ file_store: The agent's file store. Given one, the tool takes a ``files``
186
186
  parameter and shares those files into the sandbox; the caller's listing is the
187
187
  authority on which names exist, exactly as it is for the Bicep kind.
188
188
  output_sink: Where produced files land. Required by any mode but
@@ -255,7 +255,7 @@ def make_codeact_tools(
255
255
  image, image_id, outputs=outputs, files_in=files_in, files_out=files_out
256
256
  )
257
257
  return sandboxed_tool(
258
- lambda session: _execute_code_tool(session, workspace_store, outputs, exec_timeout_seconds),
258
+ lambda session: _execute_code_tool(session, file_store, outputs, exec_timeout_seconds),
259
259
  router=router,
260
260
  context=context,
261
261
  agent_dir=agent_dir,
@@ -322,8 +322,8 @@ _DESCRIPTION_MANIFEST = f"""**To produce files, write them into the working dire
322
322
  _DESCRIPTION_ARG_CODE = """code: The Python source to run. The standard library, plus
323
323
  whatever the sandbox image ships."""
324
324
 
325
- _DESCRIPTION_ARG_FILES = """files: Workspace-relative paths to share into the sandbox, or
326
- omit for none. Only files in your workspace listing can be shared."""
325
+ _DESCRIPTION_ARG_FILES = """files: Store-relative paths to share into the sandbox, or
326
+ omit for none. Only files in your file store listing can be shared."""
327
327
 
328
328
  _DESCRIPTION_ARG_OUTPUTS = """outputs: The file names your program will write into its
329
329
  working directory, or omit if it writes none."""
@@ -369,7 +369,7 @@ def _execute_code_tool(
369
369
  """Build the ``execute_code`` body for one attached tool.
370
370
 
371
371
  Four signatures over one implementation, because MAF derives the tool's schema from the
372
- function's parameters: a host that wired no workspace store must not be shown ``files``.
372
+ function's parameters: a host that wired no file store must not be shown ``files``.
373
373
  """
374
374
 
375
375
  async def run(code: str, files: list[str] | None, declared: list[str] | None) -> str:
@@ -455,10 +455,10 @@ async def _execute(
455
455
  if over_cap is not None:
456
456
  return over_cap
457
457
  if store is not None:
458
- resolved = await _resolve_workspace_files(session, store, files, reserved=reserved)
458
+ resolved = await _resolve_listed_files(session, store, files, reserved=reserved)
459
459
  if isinstance(resolved, str):
460
460
  return resolved
461
- read = await _read_workspace_files(store, resolved, tally)
461
+ read = await _read_listed_files(store, resolved, tally)
462
462
  if isinstance(read, str):
463
463
  return read
464
464
  shared = read
@@ -513,7 +513,7 @@ async def _execute(
513
513
  # --- Files in ------------------------------------------------------------------------------
514
514
 
515
515
 
516
- async def _resolve_workspace_files(
516
+ async def _resolve_listed_files(
517
517
  session: SandboxToolSession,
518
518
  store: "AgentFileStore",
519
519
  files: list[str],
@@ -553,7 +553,7 @@ async def _resolve_workspace_files(
553
553
  return f"Error: {name!r} was listed twice."
554
554
  if name not in known:
555
555
  logger.warning(
556
- "execute_code: %r is not in this tool's workspace listing (%d file(s) visible) "
556
+ "execute_code: %r is not in this tool's file store listing (%d file(s) visible) "
557
557
  "— the store wired here may be narrower than the agent's",
558
558
  name,
559
559
  len(listing),
@@ -566,7 +566,7 @@ async def _resolve_workspace_files(
566
566
  return resolved
567
567
 
568
568
 
569
- #: Capped so a large workspace cannot flood the model's context.
569
+ #: Capped so a large file store cannot flood the model's context.
570
570
  _LISTING_HINT_MAX = 20
571
571
 
572
572
 
@@ -582,7 +582,7 @@ def _listing_hint(name: str, listing: list[str]) -> str:
582
582
  return f"Files visible here: {', '.join(shown)}{more}."
583
583
 
584
584
 
585
- async def _read_workspace_files(
585
+ async def _read_listed_files(
586
586
  store: "AgentFileStore", names: list[str], tally: "_InboundTally"
587
587
  ) -> list[tuple[str, str]] | str:
588
588
  """Read every requested file into memory, or answer with the refusal.
@@ -601,14 +601,14 @@ async def _read_workspace_files(
601
601
  try:
602
602
  content = await store.read(name)
603
603
  except Exception as exc: # noqa: BLE001
604
- logger.warning("execute_code: could not read %r from workspace: %s", name, exc)
605
- return f"Error: could not read {name!r} from workspace"
604
+ logger.warning("execute_code: could not read %r from the file store: %s", name, exc)
605
+ return f"Error: could not read {name!r} from the file store"
606
606
  if content is None:
607
607
  # A store read can miss without raising (the file was listed, then removed). Writing
608
608
  # `None` through would put the string "None" into the sandbox for the program to
609
609
  # parse.
610
610
  logger.warning("execute_code: %r is listed but has no content", name)
611
- return f"Error: {name!r} is listed in the workspace but has no content"
611
+ return f"Error: {name!r} is listed in the file store but has no content"
612
612
  over_cap = tally.add(name, content)
613
613
  if over_cap is not None:
614
614
  return over_cap
@@ -677,7 +677,7 @@ class _InboundTally:
677
677
 
678
678
 
679
679
  async def _write_shared(sandbox: "Sandbox", name: str, guest_path: str, content: str) -> str | None:
680
- """Put one already-read workspace file into the run's directory, or answer with the refusal."""
680
+ """Put one already-read file store file into the run's directory, or answer with the refusal."""
681
681
  try:
682
682
  await sandbox.write_file(guest_path, content)
683
683
  except Exception as exc: # noqa: BLE001