@argszero/cordis-plugin-sandbox-grant-advisor 0.1.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,31 +1,59 @@
1
1
  # @argszero/cordis-plugin-sandbox-grant-advisor
2
2
 
3
- Turns a Windows sandbox **ACL provisioning failure with no path forward** into a
4
- diagnosis the model — and the user reading the transcript — can act on.
3
+ Turns a sandbox environment failure that has **no path forward** into a
4
+ diagnosis the model — and the user reading the transcript — can act on. Two
5
+ signatures, one mechanism:
5
6
 
6
7
  ```
7
- SetNamedSecurityInfoW failed (Win32 5): grantWrite(D:\ws)
8
+ SetNamedSecurityInfoW failed (Win32 5): grantWrite(D:\ws) # Windows workspace ACL
9
+ PTY shell exited during startup # persistent shell × confining mode
8
10
  ```
9
11
 
10
- Three reports describe this exact line: [discussion #7538], [discussion #7622],
11
- [discussion #7646]. In each one every sandboxed command fails the same way,
12
- before it runs, and the error names neither the missing right nor a remedy.
13
-
14
12
  **This plugin is the stopgap for "the error does not name the outstanding
15
- condition".** It does not repair anything: no ACL is written, no privilege is
16
- requested, nothing is elevated.
17
-
18
- ## The failure it recognizes
13
+ condition".** It repairs nothing: no ACL is written, no privilege is requested,
14
+ nothing is elevated, no preset is installed and no mode is changed.
15
+
16
+ ## The two failures it recognizes
17
+
18
+ Both are recognized on the public **`tools/post-execute`** waterfall
19
+ (`@deepseek-ai/dsh-tools`) — the one seam that has all three of: the failure text
20
+ (providers propagate their error unchanged and the tool pipeline settles it as an
21
+ `isError` result), an agent identity to attribute it to (`exec.agent`), and a
22
+ channel that speaks to the model in the same step (`PostToolDecision`'s
23
+ `additionalContexts`, which the agent loop turns into a durable user-role message
24
+ — `packages/core/agent-loop/src/tool-calls.ts`).
25
+
26
+ That seam, not `ctx.sandbox.confine`: `confine(argv, policy, signal)` sees the
27
+ confinement failure too, but its signature carries no agent, so a wrapper could
28
+ detect the condition and never deliver a word about it to the session that is
29
+ stuck.
30
+
31
+ ### 1. Workspace provisioning — the Windows ACL failure (`acl-provisioning`)
32
+
33
+ Four reports describe this exact line: [discussion #7538], [discussion #7622],
34
+ [discussion #7646], [discussion #7720]. In each one every sandboxed command fails
35
+ the same way, before it runs, and the error names neither the missing right nor
36
+ a remedy.
37
+
38
+ `#7720` is worth reading for where the failure lands: the grant is materialized
39
+ at sandbox **initialization**, so this is not one refused operation but *every*
40
+ shell tool at once — the reporter could not run `netstat` or even `icacls` to
41
+ diagnose the error they were staring at (on `0.1.5-rc.3` the same directory
42
+ worked, because confinement was skipped silently rather than failing closed).
43
+ They also report the two remedies that look right and are not
44
+ (`takeown /F <dir> /R /D Y`, `icacls <dir> /reset /T /C`), which is why the
45
+ advisory names them with the reason each fails instead of leaving the reader to
46
+ discover it.
19
47
 
20
48
  The Windows backend provisions a workspace by writing the directory's DACL and
21
49
  its mandatory-integrity label in **one** `SetNamedSecurityInfoW` call
22
- (`packages/sandbox/sandbox-windows-acl/src/acl.ts`:
50
+ (`packages/sandbox/sandbox-windows-acl/src/acl.ts`):
23
51
 
24
52
  ```ts
25
53
  if (applyResult !== abi.ERROR_SUCCESS) throwWin32(api, 'SetNamedSecurityInfoW', applyResult, `${label}(${path})`)
26
54
  ```
27
55
 
28
- ). Two consequences follow from that one line:
56
+ Two consequences follow from that one line:
29
57
 
30
58
  1. **The label lives in the SACL, and its half is what gets refused.** The
31
59
  owner's implicit rights cover only `READ_CONTROL` and `WRITE_DAC`, so the
@@ -45,53 +73,134 @@ cached when it throws** — so the same failure repeats per command (850 calls
45
73
  across 39 sessions in #7622; 52,588 output tokens with no output in #7538),
46
74
  which is why the loop cannot separate it from ordinary command noise.
47
75
 
48
- ## What it does
49
-
50
- One listener on the public **`tools/post-execute`** waterfall
51
- (`@deepseek-ai/dsh-tools`). That seam — not `ctx.sandbox.confine` — because it is
52
- the only one that has all three of: the failure text (providers propagate their
53
- error unchanged, and the tool pipeline settles it as an `isError` result), an
54
- agent identity to attribute it to (`exec.agent`), and a channel that speaks to
55
- the model in the same step (`PostToolDecision`'s `additionalContexts`, which the
56
- agent loop turns into a durable user-role message —
57
- `packages/core/agent-loop/src/tool-calls.ts`).
58
-
59
- 1. **One durable advisory per agent.** On the first recognized failure, the
60
- result is enriched with a user-role notice that names the missing right, the
61
- discriminator, and the unelevated fix:
62
-
63
- ```
64
- Sandbox provisioning failed — no sandboxed command can run in this workspace until its ACL applies.
65
-
66
- What was reported:
67
- SetNamedSecurityInfoW failed (Win32 5): grantWrite(D:\ws)
68
-
69
- Why it is refused while the directory looks writable: that call is a MERGED write ...
70
- ... the label half additionally needs WRITE_OWNER on the directory. ...
71
-
72
- Confirm the cause (unelevated) — `icacls` is a normal user command:
73
- icacls "D:\ws"
74
-
75
- Fix it (unelevated, one line) and then run the command again:
76
- PowerShell: icacls "D:\ws" /grant "$env:USERNAME:(OI)(CI)F"
77
- cmd: icacls "D:\ws" /grant "%USERNAME%:(OI)(CI)F"
78
- ```
79
-
80
- The notice carries its own producer-owned `source.kind`
81
- (`sandbox-grant-advisor`) — not the retired `plugin` wrapper, which the
82
- current session format refuses — and a bounded one-line `summary` for the
83
- transcript row. The host log gets one matching `warn` line, so the fact
84
- survives outside the transcript too.
85
- 2. **An optional, bounded fail-fast half** (`enforceAfter`, default **0** =
86
- off). It refuses a call **before dispatch** only when both hold: the
87
- environment has failed provisioning at least `enforceAfter` times, **and**
88
- this exact call (tool + canonical arguments) is one this plugin watched fail.
89
- The budget is `maxDenials` (default 2), after which the call proceeds again.
90
- The budget is per **episode of brokenness**: a call that finally succeeds stops
91
- being a denial target and re-arms it, so an environment that breaks twice can
92
- be refused twice — while a session can always make progress by spending the
76
+ On the first recognized failure, the result is enriched with:
77
+
78
+ ```
79
+ Sandbox provisioning failed — no sandboxed command can run in this workspace until its ACL applies.
80
+
81
+ What was reported:
82
+ SetNamedSecurityInfoW failed (Win32 5): grantWrite(D:\ws)
83
+
84
+ Why it is refused while the directory looks writable: that call is a MERGED write ...
85
+ ... the label half additionally needs WRITE_OWNER on the directory. ...
86
+
87
+ Confirm the cause (unelevated) — `icacls` is a normal user command:
88
+ icacls "D:\ws"
89
+
90
+ Fix it (unelevated, one line) and then run the command again:
91
+ PowerShell: icacls "D:\ws" /grant "$env:USERNAME:(OI)(CI)F"
92
+ cmd: icacls "D:\ws" /grant "%USERNAME%:(OI)(CI)F"
93
+
94
+ What will NOT fix it — both look like the right move, and both were tried and reported:
95
+ takeown /F "D:\ws" /R /D Y
96
+ makes you the owner, but ownership's implicit rights are READ_CONTROL and WRITE_DAC only.
97
+ The owner does not implicitly hold WRITE_OWNER, which is the right this call needs.
98
+ icacls "D:\ws" /reset /T /C
99
+ restores inheritance, and inheritance is what supplied the Modify-only ACE above.
100
+ ```
101
+
102
+ ### 2. Persistent shell startup (`pty-startup`)
103
+
104
+ [Discussion #7638] reports the second shape: with the **`minimal` preset** on
105
+ Windows, and under a **confining** sandbox mode (`workspace-write` / `read-only`,
106
+ not `danger-full-access`), **every** shell call dies instantly with
107
+
108
+ ```
109
+ PTY shell exited during startup
110
+ ```
111
+
112
+ The terminal backend spawns the shell through the sandbox
113
+ (`packages/terminal/terminal-bash/src/index.ts`; the throw is in
114
+ `src/session.ts` and `src/index.ts`, both on the same `waitReason ===
115
+ 'session_exit'` branch), and there the pseudo-console cannot be created at all,
116
+ so the child exits before its first prompt. Retrying never helps; the message
117
+ points at no cause.
118
+
119
+ The reporter's own three-arm control makes the sandbox mode the discriminator:
120
+ minimal × confining fails, minimal × `danger-full-access` succeeds, `standard`
121
+ (one-shot shell) × confining succeeds. That is why the advisory is only ever
122
+ built with the **resolved** mode the failing call actually ran under — from
123
+ `ctx.sandboxPolicy.resolve({ session })`, the same resolver the terminal layer
124
+ calls before spawning, with the same session.
125
+
126
+ The advisory that follows is addressed to **two different readers**:
127
+
128
+ ```
129
+ Persistent shell failed to start — command execution is unavailable in this session, and retrying cannot fix it.
130
+
131
+ What was reported:
132
+ PTY shell exited during startup
133
+
134
+ The `bash` tool is a PERSISTENT PTY session (a shell that stays alive between calls), and this session's sandbox mode is
135
+ `workspace-write` — not `danger-full-access`. A confining mode spawns the shell through the sandbox, and there the
136
+ terminal backend cannot create the pseudo-console at all, so the child exits before its first prompt. ...
137
+
138
+ Do NOT retry, and do not look for a command that fixes it: every attempt will fail identically, and there is no
139
+ shell to run a command in. Use your file read/write tools instead, and hand the choice below to the user.
140
+
141
+ What unblocks the session — the user's decision, not the model's:
142
+ 1. switch the agent preset to `standard`, whose shell tool is a one-shot subprocess (no PTY) and works
143
+ under the sandbox; or
144
+ 2. override the `preset-minimal` row in your profile patch — `$DSH_HOME/profiles/<profile>/cordis.patch.yml`, or
145
+ `$DSH_HOME/cordis.patch.yml` for every profile — replacing its `persistent-shell` group with
146
+ `@deepseek-ai/dsh-tool-pwsh` (a one-shot subprocess, no PTY); the patch layer is yours, so an upgrade
147
+ will not overwrite it; or
148
+ 3. run the session with `danger-full-access`, which drops the very confinement the sandbox exists to give.
149
+ Prefer 1 or 2.
150
+ ```
151
+
152
+ The model's instruction is to **stop** — not to run a command (there is no shell
153
+ to run it in) and not to call a fallback shell tool (`minimal` mounts exactly
154
+ **one** platform-selected persistent shell and **no** one-shot shell, by design:
155
+ `.agents/notes/implemented/simplification/2026-09-03-minimal-profiles-persistent-shell-only.md`).
156
+ Naming a tool the failing composition does not mount would be a wrong remedy,
157
+ which is the main risk this family's text is written to avoid.
158
+
159
+ The remedy is a **patch layer**, not a directory. The pre-declarative
160
+ `$DSH_HOME/.agent-presets/<id>/` preset folder is a plausible-looking trap: it
161
+ still reads as the natural place to put a preset, and nothing in the harness
162
+ reads it any more (`@deepseek-ai/dsh-agent-preset-registry`: the registry
163
+ "neither scans directories nor accepts preset paths"). Preset changes are
164
+ `@deepseek-ai/dsh-agent-preset` rows — an `insert` for a new one, a patch keyed
165
+ by row id (`preset-minimal`) for a change to a shipped one. A test arm asserts
166
+ the advisory never names the dead directory.
167
+
168
+ ## What it does with a recognized failure
169
+
170
+ 1. **One durable advisory per agent, per family.** An agent that hits both
171
+ families is told about **both**, once each. The notice carries its own
172
+ producer-owned `source.kind` (`sandbox-grant-advisor`) — not the retired
173
+ `plugin` wrapper, which the current session format refuses — and a bounded
174
+ one-line `summary` for the transcript row. The host log gets one matching
175
+ `warn` line, so the fact survives outside the transcript too.
176
+ 2. **A disclosure when it withholds.** The PTY advisory is only sent when the
177
+ resolved mode actually confines. If the mode is `danger-full-access`, or
178
+ cannot be resolved at all (no `sandboxPolicy` service mounted, no agent
179
+ session, a resolver that throws), the failure is left exactly as it was
180
+ **and the host log says so once**. Silence alone would make "the sandbox is
181
+ not the cause" and "this plugin could not tell" indistinguishable from the
182
+ outside. Withholding is never a guess: an unresolvable mode is *not* an
183
+ invitation to fall back to the deployment default.
184
+ 3. **An optional, bounded fail-fast half** (`enforceAfter`, default **0** =
185
+ off) — **ACL family only**. It refuses a call **before dispatch**
186
+ (`tools/pre-execute`) only when both hold: the environment has failed
187
+ provisioning at least `enforceAfter` times, **and** this exact call (tool +
188
+ canonical arguments) is one this plugin watched fail. The budget is
189
+ `maxDenials` (default 2), after which the call proceeds again. The budget is
190
+ per **episode of brokenness**: a call that finally succeeds stops being a
191
+ denial target and re-arms it, so an environment that breaks twice can be
192
+ refused twice — while a session can always make progress by spending the
93
193
  budget it has.
94
194
 
195
+ **Why the blocking half does not extend to the PTY family** (it is
196
+ ACL-only by construction, in the parameter type): the ACL remedy is a command
197
+ the user can run *while the session continues*, so refusing further identical
198
+ calls cannot make the session unfinishable — spending the budget always lets
199
+ the call through, and a repaired environment is discovered by exactly that.
200
+ The PTY remedy is a preset swap, which happens **between** sessions;
201
+ refusing calls there could only pad a session that is already unable to do
202
+ the thing being refused.
203
+
95
204
  ## Install
96
205
 
97
206
  ```sh
@@ -124,6 +233,11 @@ Mount it by adding the patch to your profile, or apply the shipped
124
233
  enforceAfter: 3
125
234
  ```
126
235
 
236
+ `include` / `exclude` narrow **both** families: an untracked call is
237
+ transparent to the plugin entirely, so a watched-out shell call produces no PTY
238
+ advisory (and no withholding note either — the configuration said "not our
239
+ story", which is different from "we could not tell").
240
+
127
241
  ## What it deliberately refuses to explain
128
242
 
129
243
  Recognition is narrow, because a classifier that names the wrong cause is worse
@@ -135,42 +249,59 @@ than one that stays silent.
135
249
  problem, and neither are the `LocalFree`, `LockFileEx`,
136
250
  `SetConsoleCtrlHandler` or `SetEnvironmentVariableW` failures thrown by the
137
251
  same package.
252
+ - **The PTY family is matched on a whole line, not a substring.** The producer's
253
+ message *is* the sentence (`PTY shell exited during startup`) with no detail
254
+ field at all, so any longer line that merely contains it is something
255
+ **quoting** it — a transcript, a log a failing command printed, a pasted issue
256
+ body — and the harness is not the producer.
257
+ - **The sibling throw is not classified.** `PTY shell did not reach readiness
258
+ before startup timeout` means the shell started and then did not reach a
259
+ prompt: a different cause space (a slow or blocked shell) with a different
260
+ remedy.
138
261
  - **The Win32 code is kept, not flattened.** `ERROR_ACCESS_DENIED` (5) is the
139
262
  case the documented prerequisite explains; another code gets a different
140
263
  paragraph that says so instead of borrowing the same sentence.
141
264
  - **A successful command whose *output* contains the line is not a failure.**
142
265
  The gate is the result's error state, not the presence of the text — reading a
143
266
  log file that quotes the error must not trigger advice.
144
- - **Only one advisory per agent.** The environment is explained once; repeating
145
- it per failed command would be noise competing with the failure itself.
267
+ - **Only one advisory per agent, per family.** The environment is explained
268
+ once; repeating it per failed command would be noise competing with the
269
+ failure itself.
146
270
 
147
271
  ## Honest boundaries
148
272
 
149
- - **The Windows path itself cannot be witnessed on macOS**, where this plugin was
150
- built. What the test suite proves is the decision layer — classification, the
151
- once-per-agent rule, the fail-fast budget and its self-feeding guard, and the
152
- wiring to a real cordis `Context` and the real `ToolRuntime` — driven by
153
- fixtures that throw the producer's exact error shape (`Win32Error`,
154
- `packages/subprocess/win32-process/src/errors.ts`). It does **not** prove that
155
- `icacls ... :(OI)(CI)F` fixes a given machine; that is the user's one-line
156
- experiment, and the advisory says so.
273
+ - **The Windows path itself cannot be witnessed on macOS**, where this plugin
274
+ was built. What the test suite proves is the decision layer — classification
275
+ of both families, the once-per-agent-per-family rule, the sandbox-mode gate
276
+ and its fail-closed behaviour, the fail-fast budget and its self-feeding
277
+ guard, and the wiring to a real cordis `Context` and the real `ToolRuntime` —
278
+ driven by fixtures that throw the producers' exact error shapes
279
+ (`Win32Error`, `packages/subprocess/win32-process/src/errors.ts`; the
280
+ terminal throws, `packages/terminal/terminal-bash/src/{index,session}.ts`). It
281
+ does **not** prove that `icacls ... :(OI)(CI)F` fixes a given machine, nor that
282
+ a given Windows host reproduces the PTY startup failure; those are the user's
283
+ one-line experiment and the reporter's own control, and both advisories say
284
+ where they stop.
157
285
  - **It repairs nothing and elevates nothing.** If the directory really is
158
- Full-control for the caller, the remaining hypothesis is `SeSecurityPrivilege`
159
- — i.e. the backend's documented prerequisite would be wrong. That is an
160
- upstream question; the advisory states the discriminator rather than assuming
161
- the answer.
162
- - **Delivery to the model is the agent loop's.** `additionalContexts` are ferried
163
- on the settled result here and appended as durable user-role events by
286
+ Full-control for the caller, the remaining ACL hypothesis is
287
+ `SeSecurityPrivilege` — i.e. the backend's documented prerequisite would be
288
+ wrong. That is an upstream question; the advisory states the discriminator
289
+ rather than assuming the answer.
290
+ - **Delivery to the model is the agent loop's.** `additionalContexts` are
291
+ ferried on the settled result here and appended as durable user-role events by
164
292
  `agent-loop`; a direct `ctx.tools.execute()` caller with no agent gets no
165
293
  advisory (and no agent to explain anything to).
166
294
  - **It complements `@argszero/cordis-plugin-repeat-guard-escalation`, it does not
167
295
  replace it.** That guard keys on **call identity** (identical arguments
168
296
  retried); this one keys on the **environment signature**, which is how several
169
297
  *different* commands share one cause. Mounting both is sensible.
170
- - **The real fix is upstream.** `grantWrite` already computes
171
- `hasExactGrant` / `hasExactDeny` / `hasExactLabel` and discards which one was
172
- false, so the diagnostic that turns a 52-minute detour into one line belongs at
173
- that site — next to the preflight the grant's lazy materialization wants.
298
+ - **The real fix is upstream, in both families.** For the ACL failure,
299
+ `grantWrite` already computes `hasExactGrant` / `hasExactDeny` /
300
+ `hasExactLabel` and discards which one was false, so the diagnostic that turns
301
+ a 52-minute detour into one line belongs at that site. For the PTY failure,
302
+ the startup path should either report "this sandbox mode is incompatible with
303
+ the PTY backend" or fall back to a one-shot shell. This plugin is the stopgap
304
+ for both.
174
305
 
175
306
  ## Compatibility
176
307
 
@@ -188,6 +319,11 @@ Harness peers — all three carry the **same** range, quoted in full on purpose
188
319
  that reason.
189
320
  - `@deepseek-ai/cordis@^4.0.2`.
190
321
 
322
+ `@deepseek-ai/dsh-sandbox-policy` is **not** a peer: the PTY family's mode
323
+ lookup is a guarded, structural one (`ctx.get('sandboxPolicy')`) precisely so a
324
+ composition that does not mount the service degrades to silence instead of
325
+ failing to load. See `src/mode.ts`.
326
+
191
327
  Probed at the newest build of every line the range admits — `0.1.2-rc.1`,
192
328
  `0.1.3-alpha.2`, `0.1.5-rc.3`, `0.1.6-alpha.2`, `0.1.7-rc.1` (the build the third
193
329
  report ran) — with `npm run test:probe-lines`, which derives those builds from
@@ -200,6 +336,7 @@ left claimed.
200
336
  ```sh
201
337
  npm install
202
338
  npm test # tsc, then the suite (real cordis + real ToolRuntime)
339
+ npm run test:inject # defect injection: mutate the source, rebuild, require the suite to go red
203
340
  npm run test:probe-lines # install the newest build of each admitted line and run the suite against it
204
341
  npm run test:probe-lines -- 0.1.7-rc.1 # one line only
205
342
  ```
@@ -207,6 +344,17 @@ npm run test:probe-lines -- 0.1.7-rc.1 # one line only
207
344
  The suite is mostly control arms: a guard that explains the wrong failure, or
208
345
  refuses a call that would have worked, is worse than one that stays silent.
209
346
 
347
+ `test:inject` exists because an arm nobody has seen fail proves nothing. It
348
+ mutates the decision layer one defect at a time — the mode gate removed, the
349
+ PTY message matched as a substring, the preset remedy pointed back at the dead
350
+ legacy directory, the two families collapsed into one bookkeeping slot, the
351
+ advisory delivered per call instead of per agent — and requires that specific
352
+ arms fail. It reports `SILENT ARMS: none` when every arm bites, restores the
353
+ source in a `finally`, and prints `EQUIVALENT` (with the reason) for a mutation
354
+ the current runtime cannot distinguish rather than counting it as a pass.
355
+
210
356
  [discussion #7538]: https://github.com/deepseek-ai/deepseek-harness/discussions/7538
211
357
  [discussion #7622]: https://github.com/deepseek-ai/deepseek-harness/discussions/7622
212
358
  [discussion #7646]: https://github.com/deepseek-ai/deepseek-harness/discussions/7646
359
+ [discussion #7720]: https://github.com/deepseek-ai/deepseek-harness/discussions/7720
360
+ [discussion #7638]: https://github.com/deepseek-ai/deepseek-harness/discussions/7638
package/cordis.patch.yml CHANGED
@@ -45,6 +45,24 @@
45
45
  # is off by default on purpose: it may only refuse a call it has watched fail in
46
46
  # this environment, and it is bounded by `maxDenials`, because a plugin that can
47
47
  # stop command execution must never be the reason a session cannot finish.
48
+ #
49
+ # A SECOND family is recognized on the same seam (#7638): with the `minimal`
50
+ # preset and a confining sandbox mode, every shell call dies instantly with
51
+ #
52
+ # PTY shell exited during startup
53
+ #
54
+ # The terminal backend cannot create the pseudo-console inside the sandbox, so
55
+ # the child exits before its first prompt; retrying never helps and `minimal`
56
+ # mounts no fallback shell tool. That advisory states the resolved mode, tells
57
+ # the model to STOP rather than retry, and hands the user a preset choice — it
58
+ # never names a command to run (there is no shell to run it in) and never names
59
+ # a shell tool the failing composition does not mount. It is sent only when the
60
+ # mode the call actually ran under confines; otherwise the failure is left
61
+ # untouched and the host log says so once (silence alone would read as "the
62
+ # sandbox is not the cause", which this plugin cannot claim). The blocking half
63
+ # above deliberately does NOT cover this family: its remedy is a patch-layer
64
+ # change the user makes between sessions, not a command that repairs the
65
+ # running one.
48
66
 
49
67
  - insert:
50
68
  - id: sandbox-grant-advisor