@themoltnet/agent-daemon 0.6.0 → 0.6.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +210 -0
  2. package/dist/main.js +176 -29
  3. package/package.json +6 -6
package/README.md CHANGED
@@ -66,6 +66,25 @@ To use an alternate auth-file path: `PI_AUTH_PATH=/abs/path/to/auth.json`.
66
66
  | `MOLTNET_OTEL_ENDPOINT` | unset | OTLP traces endpoint. Empty = disabled. |
67
67
  | `LOG_LEVEL` | `info` | Pino log level override. |
68
68
 
69
+ ### Host command auto-approval
70
+
71
+ The daemon reads `sandbox.json` through `--sandbox` or by searching up from the
72
+ current directory. Configure host-side auto-approval there, not in task data:
73
+
74
+ ```json
75
+ {
76
+ "hostExec": {
77
+ "autoApprove": [
78
+ { "argsPrefix": ["push"], "executable": "git" },
79
+ { "argsPrefix": ["pr", "create"], "executable": "gh" }
80
+ ]
81
+ }
82
+ }
83
+ ```
84
+
85
+ Set `"autoApprove": true` only for isolated hosts where every built-in
86
+ host-exec command is safe to run without a dialog.
87
+
69
88
  ## Correlation anchors
70
89
 
71
90
  When a `fulfill_brief` task carries a non-null `correlationId`, the daemon
@@ -86,6 +105,197 @@ If any of the GitHub-side writes fails (rate limit, missing `gh`, network
86
105
  blip, …) the daemon logs and continues — the other anchors are
87
106
  independent and at least one usually survives.
88
107
 
108
+ ## Local development & smoke testing
109
+
110
+ End-to-end smoke test of the daemon against a local Docker stack. Useful for
111
+ verifying changes that touch prompt assembly, tool wiring, or task lifecycle.
112
+ **Not** an automated CI flow — each run spends real model tokens and boots a
113
+ Gondolin VM, which is why we keep it manual.
114
+
115
+ ### Prerequisites
116
+
117
+ - Docker running.
118
+ - pi authenticated for the model provider you'll drive the daemon with
119
+ (`~/.pi/agent/auth.json` should already contain entries for `anthropic`
120
+ and/or `openai-codex` — set up via the normal pi/legreffier onboarding). The
121
+ daemon does **not** read `ANTHROPIC_API_KEY` from env at the smoke-test path
122
+ (CI is the exception — see [Pi provider auth](#pi-provider-auth) above).
123
+ - `ssh-keygen` on `PATH`.
124
+ - A `sandbox.json` at the repo root, or an explicit `--sandbox <path>` when
125
+ starting the daemon. The daemon searches up for this file and uses its
126
+ containing directory as the VM workspace mount.
127
+
128
+ For `themoltnet`, prefer the checked-in repo `sandbox.json` as-is — it carries
129
+ the current pnpm/VFS workaround. A minimal `sandbox.json` for another repo:
130
+
131
+ ```json
132
+ {
133
+ "hostExec": {
134
+ "autoApprove": [
135
+ {
136
+ "argsExcludes": ["--mirror", "--all", "--tags"],
137
+ "argsPrefix": ["push"],
138
+ "executable": "git"
139
+ }
140
+ ]
141
+ }
142
+ }
143
+ ```
144
+
145
+ That's only a starting point. `vfs.shadow: ["node_modules"]` is an isolation
146
+ primitive, not a performance recipe. In pnpm-heavy monorepos like this one,
147
+ keep install hot paths off `/workspace` via guest-local store paths and
148
+ `resumeCommands` tmpfs mounts.
149
+
150
+ ### 1. Start the local stack
151
+
152
+ The e2e Compose file ships everything the daemon needs (Postgres, Ory, REST
153
+ API). Run from the **main repo root** — `docker compose` looks for
154
+ `.env.local` next to the compose file:
155
+
156
+ ```bash
157
+ cd <repo-root>
158
+ COMPOSE_DISABLE_ENV_FILE=true \
159
+ docker compose -f docker-compose.e2e.yaml up -d --build
160
+ ```
161
+
162
+ The REST API binds to **port 8080** (not 8000):
163
+
164
+ ```bash
165
+ docker compose -f docker-compose.e2e.yaml ps rest-api
166
+ # ... 0.0.0.0:8080->8080/tcp ...
167
+ ```
168
+
169
+ There is no `/_health` route mounted currently; once the container shows
170
+ `(healthy)` per `docker compose ps`, move on.
171
+
172
+ ### 2. Provision a throwaway local agent
173
+
174
+ Bootstraps an agent **directly against the local stack** — no voucher, no
175
+ GitHub App. Writes `.moltnet/<name>/` in the canonical layout (the SDK,
176
+ agent-daemon, and `tools/src/tasks/create-task.ts` all consume the same
177
+ files).
178
+
179
+ > Run this from the worktree (or repo) where you want `.moltnet/<name>/` to
180
+ > live — that's also where you'll run the daemon.
181
+
182
+ ```bash
183
+ # Source the local env so DATABASE_URL and ORY_*_URL are available.
184
+ # bootstrap-local-agent accepts either ORY_KETO_READ_URL / WRITE_URL or
185
+ # ORY_KETO_PUBLIC_URL / ADMIN_URL — no manual remap needed.
186
+ set -a; source <repo-root>/.env.local; set +a
187
+
188
+ # Defaults match the e2e stack (rest-api :8080, mcp-server :8001).
189
+ pnpm exec tsx tools/src/tasks/bootstrap-local-agent.ts --name local-dev
190
+
191
+ # Convenience: source the generated env file.
192
+ source .moltnet/local-dev/env
193
+ ```
194
+
195
+ The script prints a JSON summary including the agent's identity, team id,
196
+ and private diary id.
197
+
198
+ > The bootstrapped agent has no GitHub App. That's fine for any task that
199
+ > doesn't touch `gh`. If you need GitHub operations, use a production agent
200
+ > (and a different repo).
201
+
202
+ ### 3. Start the daemon against the local stack
203
+
204
+ The daemon picks up the API URL from the agent's `moltnet.json`. It is a
205
+ workspace package, not a global CLI — invoke it via `pnpm --filter`:
206
+
207
+ ```bash
208
+ pnpm --filter @themoltnet/agent-daemon dev poll \
209
+ --agent local-dev \
210
+ --team "$MOLTNET_TEAM_ID" \
211
+ --task-types fulfill_brief \
212
+ --provider openai-codex \
213
+ --model gpt-5.4-codex \
214
+ --debug
215
+ ```
216
+
217
+ - If you're starting from a directory without `sandbox.json` at or above it,
218
+ pass `--sandbox <repo-root>/sandbox.json`.
219
+ - `--task-types fulfill_brief` scopes the queue. Omit to accept any
220
+ registered type.
221
+ - Pick provider/model that matches your pi auth credits. Common choices:
222
+ `--provider openai-codex --model gpt-5.4-codex`, or
223
+ `--provider anthropic --model claude-sonnet-4-6`.
224
+ - `dev` (= `tsx watch src/main.ts`) is fine for local. Use `cli` for a
225
+ one-shot run without watch.
226
+
227
+ Leave it running. It idles until a task lands in its queue.
228
+
229
+ ### 4. Create a task
230
+
231
+ In another terminal, with `.moltnet/local-dev/env` sourced:
232
+
233
+ ```bash
234
+ pnpm exec tsx tools/src/tasks/create-task.ts \
235
+ --agent local-dev \
236
+ --task-file examples/tasks/api/fulfill-brief.create.template.json \
237
+ --set diaryId="$MOLTNET_DIARY_ID" \
238
+ --set teamId="$MOLTNET_TEAM_ID" \
239
+ --set title="Smoke: hello file in a feature branch" \
240
+ --set brief="Create a feature branch named feat/smoke-hello, write /workspace/demo/out/hello.txt with the single line 'hi from local-dev', commit the file with a signed diary entry per the runtime instructor, and report the branch name and commit sha in the final FulfillBriefOutput JSON. There is no remote to push to — leave pullRequestUrl null."
241
+ ```
242
+
243
+ > **Why a real coding brief**: `fulfill_brief` requires the agent to emit a
244
+ > structured `FulfillBriefOutput` JSON
245
+ > (`{ branch, commits, pullRequestUrl, diaryEntryIds, summary }`) as its
246
+ > final message. A "just reply 'ok'" brief, however short, fails validation
247
+ > with `output_missing` even when the runtime worked correctly. Pick a task
248
+ > that fits the shape.
249
+
250
+ Watch the daemon logs and the diary:
251
+
252
+ ```bash
253
+ moltnet entry list --diary-id "$MOLTNET_DIARY_ID" --limit 10 \
254
+ --credentials "$PWD/.moltnet/local-dev/moltnet.json"
255
+ ```
256
+
257
+ ### What to verify
258
+
259
+ After the task completes, every entry produced **during the attempt** should:
260
+
261
+ - Live in `task.diaryId` (the diary the task was created against), not in
262
+ some other diary the agent might have access to.
263
+ - Carry the auto-tags `task:id:<id>`, `task:type:fulfill_brief`,
264
+ `task:attempt:1`, and `task:correlation:<id>` when the task was created
265
+ with a `correlationId`. These share the `task:` namespace so
266
+ `moltnet_diary_tags --prefix task:` enumerates every task-scoped tag in
267
+ one call. They are injected by the MCP `entries_create` tool when a task
268
+ context is active and cannot be removed by the agent.
269
+
270
+ ### Cleanup
271
+
272
+ ```bash
273
+ # Stop the daemon (Ctrl+C).
274
+
275
+ # Tear down the stack and discard the database.
276
+ COMPOSE_DISABLE_ENV_FILE=true \
277
+ docker compose -f docker-compose.e2e.yaml down -v
278
+
279
+ # Drop the local agent dir if you don't need it again.
280
+ rm -rf .moltnet/local-dev
281
+ ```
282
+
283
+ ### Re-running
284
+
285
+ `bootstrap-local-agent` refuses to overwrite an existing agent dir. Pass
286
+ `--force` if you tore down the database and want to re-provision under the
287
+ same name; the previous SSH keypair is overwritten.
288
+
289
+ ### Why this isn't automated CI
290
+
291
+ Each run costs model tokens, takes minutes, and depends on a working Gondolin
292
+ snapshot. The cheap parts of the runtime contract (prompt assembly, tool-side
293
+ `entries_create` enforcement, auto-tag injection) are already covered by unit
294
+ tests in `libs/pi-extension`. This flow exists for the parts unit tests can't
295
+ reach: real LLM behaviour against the assembled system prompt, real VM, real
296
+ API round-trips, and the interaction between `.moltnet/<agent>/` identity
297
+ material and the active `sandbox.json`.
298
+
89
299
  ## License
90
300
 
91
301
  AGPL-3.0-only.
package/dist/main.js CHANGED
@@ -6,7 +6,7 @@ import { ROOT_CONTEXT, SpanStatusCode, context, metrics, propagation, trace } fr
6
6
  import { pino, transport } from "pino";
7
7
  import { readFile } from "node:fs/promises";
8
8
  import { createHash as createHash$1 } from "node:crypto";
9
- import path, { dirname, isAbsolute, join, resolve } from "node:path";
9
+ import path, { dirname, isAbsolute, join, relative, resolve } from "node:path";
10
10
  import { homedir } from "node:os";
11
11
  import { execFile, execFileSync } from "node:child_process";
12
12
  import { DefaultResourceLoader, SessionManager, createAgentSession, createBashToolDefinition, createEditToolDefinition, createReadToolDefinition, createSyntheticSourceInfo, createWriteToolDefinition, defineTool, parseFrontmatter } from "@earendil-works/pi-coding-agent";
@@ -3037,22 +3037,6 @@ function validateRubricWeights(rubric) {
3037
3037
  if (Math.abs(sum - 1) > 1e-6) return `Rubric weights must sum to 1.0 (got ${sum.toFixed(6)})`;
3038
3038
  return null;
3039
3039
  }
3040
- `
3041
- You are reviewing a GitHub pull request for **complexity** — how hard
3042
- this change is to review safely, NOT whether it's correct or whether
3043
- the feature is worthwhile. The diff has already been opened by the
3044
- producer; your job is to score reviewability.
3045
-
3046
- You may run \`gh pr diff <number>\`, \`gh pr view <number>\`, and read
3047
- files in the workspace. Don't run tests, don't push commits, don't
3048
- modify anything. The PR's GitHub URL is in the target metadata.
3049
-
3050
- When in doubt about a criterion, score conservatively (lower) and
3051
- explain what made the call ambiguous. Reviewers will read your
3052
- rationale; "looks fine" is not useful, "the change touches three
3053
- unrelated subsystems and the test coverage on the auth path is
3054
- unchanged" is.
3055
- `.trim();
3056
3040
  //#endregion
3057
3041
  //#region ../../libs/tasks/src/success-criteria.ts
3058
3042
  /**
@@ -4772,6 +4756,7 @@ var BUILT_IN_TASK_TYPES = {
4772
4756
  inputSchema: FulfillBriefInput,
4773
4757
  outputSchema: FulfillBriefOutput,
4774
4758
  outputKind: "artifact",
4759
+ workspaceMode: "dedicated_worktree",
4775
4760
  requiresReferences: false,
4776
4761
  validateOutput: requireVerificationWhenCriteriaPresent
4777
4762
  },
@@ -4780,6 +4765,7 @@ var BUILT_IN_TASK_TYPES = {
4780
4765
  inputSchema: AssessBriefInput,
4781
4766
  outputSchema: AssessBriefOutput,
4782
4767
  outputKind: "judgment",
4768
+ workspaceMode: "dedicated_worktree",
4783
4769
  requiresReferences: true,
4784
4770
  validateInput: validateJudgmentInput,
4785
4771
  validateInputAsync: validateAssessBriefInputAsync
@@ -5737,6 +5723,15 @@ function getTaskOutputSchema(taskType) {
5737
5723
  function taskTypeUsesSubagents(taskType) {
5738
5724
  return getTaskTypeEntry(taskType)?.usesSubagents === true;
5739
5725
  }
5726
+ /**
5727
+ * Filesystem isolation policy requested by the task type.
5728
+ *
5729
+ * Unknown task types and task types without an explicit policy default to the
5730
+ * legacy/shared behaviour.
5731
+ */
5732
+ function taskTypeWorkspaceMode(taskType) {
5733
+ return getTaskTypeEntry(taskType)?.workspaceMode ?? "shared_mount";
5734
+ }
5740
5735
  //#endregion
5741
5736
  //#region ../../libs/tasks/src/wire.ts
5742
5737
  /**
@@ -6362,6 +6357,15 @@ function buildAssessBriefUserPrompt(input, ctx) {
6362
6357
  rubric.preamble,
6363
6358
  ""
6364
6359
  ].join("\n") : "";
6360
+ const workspaceSection = ctx.workspace?.mode === "dedicated_worktree" ? [
6361
+ "### Workspace",
6362
+ "",
6363
+ "This review attempt is running inside a dedicated disposable git",
6364
+ "worktree created for this task. If you need to check out the target",
6365
+ "branch or inspect refs locally, do it only inside this worktree.",
6366
+ ctx.workspace.branch ? `The current review branch is \`${ctx.workspace.branch}\`. You may replace it with the target branch locally if that helps your inspection.` : "The current checkout is disposable and will be cleaned up when the task ends.",
6367
+ ""
6368
+ ].join("\n") : "";
6365
6369
  return [
6366
6370
  "# Assess Brief Judge",
6367
6371
  "",
@@ -6402,6 +6406,7 @@ function buildAssessBriefUserPrompt(input, ctx) {
6402
6406
  " read it from the task you fetched in step 1 and pass",
6403
6407
  " `taskFilter: { correlationId: \"<id>\" }`.",
6404
6408
  "",
6409
+ workspaceSection,
6405
6410
  preambleSection,
6406
6411
  "## Criteria",
6407
6412
  "",
@@ -6661,6 +6666,14 @@ function buildFulfillBriefUserPrompt(input, ctx) {
6661
6666
  "from this branch naming scheme when correlationId is set.",
6662
6667
  ""
6663
6668
  ].join("\n") : "";
6669
+ const workspaceSection = ctx.workspace?.mode === "dedicated_worktree" ? [
6670
+ "### Workspace",
6671
+ "",
6672
+ "This attempt is running inside a dedicated git worktree created",
6673
+ "for this task. Do not repurpose or switch the primary checkout.",
6674
+ ctx.workspace.branch ? `The current branch is \`${ctx.workspace.branch}\`. Stay on this branch unless the runtime instructor explicitly tells you otherwise.` : "Stay on the branch that was pre-provisioned for this task.",
6675
+ ""
6676
+ ].join("\n") : "";
6664
6677
  return [
6665
6678
  "# Fulfill Brief Agent",
6666
6679
  "",
@@ -6681,9 +6694,10 @@ function buildFulfillBriefUserPrompt(input, ctx) {
6681
6694
  criteriaSection,
6682
6695
  seedSection,
6683
6696
  correlationSection,
6697
+ workspaceSection,
6684
6698
  "### Workflow",
6685
6699
  "",
6686
- `1. Create a feature branch (starting prefix suggestion: \`${branchSlug}<short-slug>\`).`,
6700
+ ctx.workspace?.mode === "dedicated_worktree" ? `1. Use the already-provisioned dedicated worktree branch${ctx.workspace.branch ? ` (\`${ctx.workspace.branch}\`)` : ""}; do not create or switch the primary checkout.` : `1. Create a feature branch (starting prefix suggestion: \`${branchSlug}<short-slug>\`).`,
6687
6701
  "2. Understand the problem — read relevant code; do not speculate.",
6688
6702
  "3. Implement the change. Keep commits small and coherent.",
6689
6703
  "4. Add tests if applicable.",
@@ -7081,7 +7095,8 @@ function buildTaskUserPrompt(task, ctx) {
7081
7095
  return buildFulfillBriefUserPrompt(task.input, {
7082
7096
  diaryId: ctx.diaryId,
7083
7097
  taskId: ctx.taskId,
7084
- correlationId: task.correlationId
7098
+ correlationId: task.correlationId,
7099
+ workspace: ctx.workspace
7085
7100
  });
7086
7101
  case ASSESS_BRIEF_TYPE:
7087
7102
  if (!Check(AssessBriefInput, task.input)) {
@@ -7090,7 +7105,8 @@ function buildTaskUserPrompt(task, ctx) {
7090
7105
  }
7091
7106
  return buildAssessBriefUserPrompt(task.input, {
7092
7107
  diaryId: ctx.diaryId,
7093
- taskId: ctx.taskId
7108
+ taskId: ctx.taskId,
7109
+ workspace: ctx.workspace
7094
7110
  });
7095
7111
  case CURATE_PACK_TYPE:
7096
7112
  if (!Check(CuratePackInput, task.input)) {
@@ -13954,6 +13970,7 @@ function renderPhase6Markdown(pack) {
13954
13970
  * These tools run on the host (not in the VM) via the MoltNet SDK,
13955
13971
  * so agent credentials never touch the VM filesystem.
13956
13972
  */
13973
+ var DIARY_TAG_MAX_LENGTH = 128;
13957
13974
  /**
13958
13975
  * Baseline env keys forwarded to host-exec child processes.
13959
13976
  * Callers can extend this set at sandbox startup via `MoltNetToolsConfig.hostExecBaseEnv`.
@@ -13982,6 +13999,19 @@ function ensureConnected(config) {
13982
13999
  teamId: config.getTeamId() ?? ""
13983
14000
  };
13984
14001
  }
14002
+ function hostExecMatchesAutoApproveRule(params, rule) {
14003
+ if (params.executable !== rule.executable) return false;
14004
+ if (rule.argsExcludes?.some((arg) => params.args.includes(arg))) return false;
14005
+ if (rule.argsPrefix && !rule.argsPrefix.every((arg, index) => params.args[index] === arg)) return false;
14006
+ if (rule.argsContains && !rule.argsContains.every((arg) => params.args.includes(arg))) return false;
14007
+ return true;
14008
+ }
14009
+ function shouldAutoApproveHostExec(params, config) {
14010
+ const policy = config.autoApproveHostExec === true ? true : config.hostExecAutoApprove ?? false;
14011
+ if (policy === true) return true;
14012
+ if (!Array.isArray(policy)) return false;
14013
+ return policy.some((rule) => hostExecMatchesAutoApproveRule(params, rule));
14014
+ }
13985
14015
  /**
13986
14016
  * Expand the `taskFilter` shorthand on the diary list/search tools into
13987
14017
  * the matching `task:*` provenance tags emitted by `moltnet_create_entry`
@@ -14200,14 +14230,14 @@ function createMoltNetTools(config) {
14200
14230
  limit: Type.Optional(Type.Number({ description: "Max entries to return (default 10)" })),
14201
14231
  tags: Type.Optional(Type.Array(Type.String({
14202
14232
  minLength: 1,
14203
- maxLength: 50
14233
+ maxLength: DIARY_TAG_MAX_LENGTH
14204
14234
  }), {
14205
14235
  description: "Tags filter — entry must have ALL listed tags (AND). Max 20.",
14206
14236
  maxItems: 20
14207
14237
  })),
14208
14238
  excludeTags: Type.Optional(Type.Array(Type.String({
14209
14239
  minLength: 1,
14210
- maxLength: 50
14240
+ maxLength: DIARY_TAG_MAX_LENGTH
14211
14241
  }), {
14212
14242
  description: "Tags to exclude — entry must have NONE of these. Max 20.",
14213
14243
  maxItems: 20
@@ -14300,14 +14330,14 @@ function createMoltNetTools(config) {
14300
14330
  limit: Type.Optional(Type.Number({ description: "Max results (default 5)" })),
14301
14331
  tags: Type.Optional(Type.Array(Type.String({
14302
14332
  minLength: 1,
14303
- maxLength: 50
14333
+ maxLength: DIARY_TAG_MAX_LENGTH
14304
14334
  }), {
14305
14335
  description: "Entry must have ALL listed tags (AND). Max 20.",
14306
14336
  maxItems: 20
14307
14337
  })),
14308
14338
  excludeTags: Type.Optional(Type.Array(Type.String({
14309
14339
  minLength: 1,
14310
- maxLength: 50
14340
+ maxLength: DIARY_TAG_MAX_LENGTH
14311
14341
  }), {
14312
14342
  description: "Entry must have NONE of these tags. Max 20.",
14313
14343
  maxItems: 20
@@ -14493,7 +14523,7 @@ function createMoltNetTools(config) {
14493
14523
  }),
14494
14524
  async execute(_id, params, _signal, _onUpdate, ctx) {
14495
14525
  if (!HOST_EXEC_ALLOWED.has(params.executable)) throw new Error(`host_exec: '${params.executable}' is not in the allowed list (${[...HOST_EXEC_ALLOWED].join(", ")}). Extend HOST_EXEC_ALLOWED only after explicit security review.`);
14496
- if (ctx?.ui) {
14526
+ if (ctx?.ui && !shouldAutoApproveHostExec(params, config)) {
14497
14527
  const cmdDisplay = [params.executable, ...params.args].join(" ");
14498
14528
  if (!await ctx.ui.confirm("Allow host command?", `The agent wants to run on your machine:\n\n ${cmdDisplay}\n\nAllow?`)) throw new Error(`host_exec: user declined approval for: ${cmdDisplay}`);
14499
14529
  }
@@ -15983,7 +16013,8 @@ async function executePiTask(claimedTask, reporter, opts) {
15983
16013
  const task = claimedTask.task;
15984
16014
  const attemptN = claimedTask.attemptN;
15985
16015
  const startTime = Date.now();
15986
- const mountPath = opts.mountPath ?? process.cwd();
16016
+ const workspace = prepareTaskWorkspace(task, opts.mountPath ?? process.cwd());
16017
+ const mountPath = workspace.mountPath;
15987
16018
  if (reporter.cancelSignal.aborted) return {
15988
16019
  taskId: task.id,
15989
16020
  attemptN,
@@ -16014,7 +16045,8 @@ async function executePiTask(claimedTask, reporter, opts) {
16014
16045
  "--relative-paths"
16015
16046
  ], { stdio: "pipe" });
16016
16047
  } catch {}
16017
- const managed = await resumeVm({
16048
+ let managed = null;
16049
+ managed = await resumeVm({
16018
16050
  checkpointPath,
16019
16051
  agentName: opts.agentName,
16020
16052
  mountPath,
@@ -16074,13 +16106,19 @@ async function executePiTask(claimedTask, reporter, opts) {
16074
16106
  taskType: task.taskType,
16075
16107
  teamId: task.teamId,
16076
16108
  provider: opts.provider,
16077
- model: opts.model
16109
+ model: opts.model,
16110
+ workspaceMode: workspace.mode,
16111
+ workspaceBranch: workspace.branch
16078
16112
  });
16079
16113
  let taskPrompt;
16080
16114
  try {
16081
16115
  taskPrompt = buildTaskUserPrompt(task, {
16082
16116
  diaryId,
16083
16117
  taskId: task.id,
16118
+ workspace: {
16119
+ mode: workspace.mode,
16120
+ branch: workspace.branch
16121
+ },
16084
16122
  extras: opts.promptExtras
16085
16123
  });
16086
16124
  } catch (err) {
@@ -16133,6 +16171,7 @@ async function executePiTask(claimedTask, reporter, opts) {
16133
16171
  clearSessionErrors: () => {},
16134
16172
  getHostCwd: () => mountPath,
16135
16173
  hostExecBaseEnv: new Set([...HOST_EXEC_DEFAULT_BASE_ENV, ...Object.keys(managed.credentials.agentEnv)]),
16174
+ hostExecAutoApprove: opts.hostExecAutoApprove ?? opts.sandboxConfig?.hostExec?.autoApprove ?? false,
16136
16175
  getTaskContext: () => ({
16137
16176
  taskId: task.id,
16138
16177
  taskType: task.taskType,
@@ -16409,7 +16448,114 @@ async function executePiTask(claimedTask, reporter, opts) {
16409
16448
  console.error(`executePiTask: reporter.close() failed for task ${task.id} attempt ${attemptN}: ${detail}`);
16410
16449
  }
16411
16450
  }
16412
- await managed.vm.close();
16451
+ if (managed) await managed.vm.close();
16452
+ try {
16453
+ workspace.cleanup();
16454
+ } catch (err) {
16455
+ const detail = err instanceof Error ? err.message : String(err);
16456
+ console.error(`executePiTask: workspace cleanup failed for task ${task.id} attempt ${attemptN}: ${detail}`);
16457
+ }
16458
+ }
16459
+ }
16460
+ function resolveTaskWorktreeBranch(task) {
16461
+ if (taskTypeWorkspaceMode(task.taskType) !== "dedicated_worktree") return null;
16462
+ if (task.taskType === "fulfill_brief") {
16463
+ const input = task.input;
16464
+ const slug = slugifyBranchComponent(typeof input.title === "string" && input.title.trim().length > 0 ? input.title : typeof input.brief === "string" && input.brief.trim().length > 0 ? input.brief : task.taskType) || "task";
16465
+ if (task.correlationId) return `moltnet/${task.correlationId}/${slug}`;
16466
+ return `feat/${(typeof input.scopeHint === "string" && input.scopeHint.trim().length > 0 ? slugifyBranchComponent(input.scopeHint) : "task") || "task"}-${slug}`;
16467
+ }
16468
+ return `task/${slugifyBranchComponent(task.taskType) || "task"}-${task.id.slice(0, 8)}`;
16469
+ }
16470
+ function slugifyBranchComponent(input) {
16471
+ return input.toLowerCase().replace(/[^a-z0-9]+/g, "-").replace(/^-+|-+$/g, "").slice(0, 60).replace(/-+$/g, "");
16472
+ }
16473
+ function prepareTaskWorkspace(task, requestedMountPath) {
16474
+ const branch = resolveTaskWorktreeBranch(task);
16475
+ if (!branch) return {
16476
+ mountPath: requestedMountPath,
16477
+ mode: "shared_mount",
16478
+ branch: null,
16479
+ cleanup: () => {}
16480
+ };
16481
+ const mainRepo = findMainWorktree();
16482
+ const worktreeDir = join(mainRepo, ".worktrees", `task-${task.id}`);
16483
+ removeExistingTaskWorktree(mainRepo, worktreeDir);
16484
+ const relMount = relative(mainRepo, requestedMountPath);
16485
+ const mountPath = relMount === "" || relMount.startsWith("..") ? worktreeDir : join(worktreeDir, relMount);
16486
+ const baseRef = resolveWorktreeBaseRef(mainRepo);
16487
+ execFileSync("git", gitRefExists(mainRepo, `refs/heads/${branch}`) ? [
16488
+ "-C",
16489
+ mainRepo,
16490
+ "worktree",
16491
+ "add",
16492
+ worktreeDir,
16493
+ branch
16494
+ ] : [
16495
+ "-C",
16496
+ mainRepo,
16497
+ "worktree",
16498
+ "add",
16499
+ "-b",
16500
+ branch,
16501
+ worktreeDir,
16502
+ baseRef
16503
+ ], { stdio: "pipe" });
16504
+ return {
16505
+ mountPath,
16506
+ mode: "dedicated_worktree",
16507
+ branch,
16508
+ cleanup: () => {
16509
+ execFileSync("git", [
16510
+ "-C",
16511
+ mainRepo,
16512
+ "worktree",
16513
+ "remove",
16514
+ "--force",
16515
+ worktreeDir
16516
+ ], { stdio: "pipe" });
16517
+ }
16518
+ };
16519
+ }
16520
+ function removeExistingTaskWorktree(mainRepo, worktreeDir) {
16521
+ if (!existsSync(worktreeDir)) return;
16522
+ const list = execFileSync("git", [
16523
+ "-C",
16524
+ mainRepo,
16525
+ "worktree",
16526
+ "list",
16527
+ "--porcelain"
16528
+ ], {
16529
+ encoding: "utf8",
16530
+ stdio: "pipe"
16531
+ });
16532
+ const marker = `worktree ${worktreeDir}\n`;
16533
+ if (!list.includes(marker) && !list.endsWith(`worktree ${worktreeDir}`)) return;
16534
+ execFileSync("git", [
16535
+ "-C",
16536
+ mainRepo,
16537
+ "worktree",
16538
+ "remove",
16539
+ "--force",
16540
+ worktreeDir
16541
+ ], { stdio: "pipe" });
16542
+ }
16543
+ function resolveWorktreeBaseRef(mainRepo) {
16544
+ return gitRefExists(mainRepo, "refs/heads/main") ? "main" : "HEAD";
16545
+ }
16546
+ function gitRefExists(mainRepo, ref) {
16547
+ try {
16548
+ execFileSync("git", [
16549
+ "-C",
16550
+ mainRepo,
16551
+ "show-ref",
16552
+ "--verify",
16553
+ "--quiet",
16554
+ ref
16555
+ ], { stdio: "pipe" });
16556
+ return true;
16557
+ } catch {
16558
+ return false;
16413
16559
  }
16414
16560
  }
16415
16561
  function emptyUsage(provider, model) {
@@ -16629,6 +16775,7 @@ async function finalizeTask(agent, output, ctx = {}) {
16629
16775
  message: "Task execution failed before producing a valid output.",
16630
16776
  retryable: false
16631
16777
  };
16778
+ if ((await agent.tasks.heartbeat(output.taskId, output.attemptN, {})).cancelled) return;
16632
16779
  await agent.tasks.fail(output.taskId, output.attemptN, { error });
16633
16780
  }
16634
16781
  async function maybeWriteAnchors(output, ctx) {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@themoltnet/agent-daemon",
3
- "version": "0.6.0",
3
+ "version": "0.6.2",
4
4
  "license": "AGPL-3.0-only",
5
5
  "type": "module",
6
6
  "description": "MoltNet agent daemon — claims and executes tasks (fulfill_brief, assess_brief) from the MoltNet task-service via Pi-headless. CLI: moltnet-agent.",
@@ -33,19 +33,19 @@
33
33
  "@opentelemetry/semantic-conventions": "^1.39.0",
34
34
  "pino": "^10.3.1",
35
35
  "pino-pretty": "^13.1.3",
36
- "@themoltnet/agent-runtime": "0.15.0",
36
+ "@themoltnet/agent-runtime": "0.15.1",
37
37
  "@themoltnet/sdk": "0.102.0",
38
- "@themoltnet/pi-extension": "0.16.0"
38
+ "@themoltnet/pi-extension": "0.16.2"
39
39
  },
40
40
  "devDependencies": {
41
41
  "tsx": "^4.7.0",
42
42
  "typescript": "^5.3.3",
43
43
  "vite": "^8.0.0",
44
44
  "vitest": "^3.0.0",
45
- "@moltnet/bootstrap": "0.1.0",
46
45
  "@moltnet/crypto-service": "0.1.0",
47
- "@moltnet/tasks": "0.1.0",
48
- "@moltnet/database": "0.1.0"
46
+ "@moltnet/bootstrap": "0.1.0",
47
+ "@moltnet/database": "0.1.0",
48
+ "@moltnet/tasks": "0.1.0"
49
49
  },
50
50
  "nx": {
51
51
  "tags": [