@livx.cc/agentx 0.99.47 → 0.99.49

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -92,6 +92,7 @@ Beyond file tools, the runtime ships the higher-altitude pieces too — each an
92
92
  - **DuplexAgent** (`src/duplex.ts`) — voice-optimized three-tier engine (reflex/act/think): a fast reflex agent streams instant replies and self-selects escalation — `Act` for standard tool work (Sonnet-class), `Think` for deep reasoning (Opus-class, configurable, default on). Results are pushed back and re-voiced by the reflex (turn mutex, coalesced completions, `TaskStatus`/`CancelTask`). See [`mind/10`](./mind/10-duplex.md).
93
93
  - **Scheduler** (`src/scheduler.ts` + `cli/osScheduler.ts`) — one-off (`{at}`), interval (`{everyMs}`), cron (`{cron}`) via `ScheduleTask`/`ScheduleList`/`ScheduleCancel`/`Wakeup`. In-session jobs fire while the session is alive (persisted, re-armed on `--resume`); far one-offs (or `backend:'os'`) register with the OS scheduler (launchd / crontab / at) and **survive quitting** — the fired job headless-resumes the session (`agentx -p … --resume <id> --yes`). The `PushNotification` tool (osascript / notify-send) alerts the user out-of-band; `Read` on a `.pdf` returns extracted text (poppler's pdftotext, disk mode). **`RemoteTrigger`** invokes another agentx session on this machine: a session open in a live terminal receives the prompt as an injected turn (per-session unix socket, same-user only); otherwise it's resumed headless and the final answer comes back. See [`mind/12`](./mind/12-scheduler.md).
94
94
  - **Budget kill-switches** — always-on per-run guards (`maxTokens`/`timeoutMs`/`maxRepeats`/`maxToolCalls`/`signal` → `finishReason` `budget`/`timeout`/`loop`/`max_tool_calls`/`aborted`) protect the API spend against runaway loops. The *enforceable* billing cap is server-side in the web key-proxy: a VFS-backed budget config (`/.agent/budget.json`, USD-metered, hot-reloaded, $100/wk default) a browser client can't bypass. See [`web/`](./web) and [`mind/06`](./mind/06-agent-features.md).
95
+ - **Jobs & liveness (never kill a working tool)** — every long-lived tool call is a job in `agent.jobs` (`JobRegistry`: id, kind `shell|mcp|host|native|bash|task`, status, `startedAt`, `lastActivityAt`, output tail, result). The model follows jobs with `JobStatus` / `JobOutput` (tail or `offset`) / `JobWait` (bounded — use instead of sleeping) / `JobKill`, and a job that finishes after its caller stopped waiting is **reported back** (re-opening a stopped turn). Liveness comes from output chunks, MCP `notifications/progress`, delegated tool events, and (for adopted shell jobs) process-tree CPU time via `ps` (group + descendants that left it). A job is killed only when it is **stuck** — silent past its kind's idle window with no CPU progress (`shell`: 10 min; unknown CPU never counts as idle) — for `mcp` only after the server has sent progress and then stopped, since most servers never report progress — or past its hard cap (`mcp`: 30 min; delegated host/native tools: the transport tool deadline). A delegated **`native`** tool (cursor's own shell/mcp/task) additionally forfeits its hold after `stallMs`-scale silence (`AgentOptions.nativeHoldMs`, 5 min): unlike a `host` job — whose pending dispatch promise is independent proof it is alive — a native job is only the provider's own claim, and it suppresses the very stall detection that would catch that provider going quiet. Further `running` events `touch()` it, so a native tool that keeps reporting progress holds indefinitely — and the model sees `killed: stuck: idle 600s`; the run continues, and every kill is logged with its evidence. Explicit `Shell({background:true})` jobs are never idle-reaped (a quiet server is healthy). The model **stall watchdog** (`stallMs`, 60s of stream silence) only fires when no job is holding the stream; a detached job does not mask a silent model, and a step whose tools already ran is never replayed. A hold is never silent — every suppressed window logs and notifies `waiting on <kind> "<label>" (<age>, quiet <n>s)`, so a wedged tool is visible in seconds instead of stalling the run unseen. Knobs: `new Agent({ jobs: new JobRegistry({ idleMs, hardCapMs, foregroundMs, sweepMs, killGraceMs }) })`; share the same registry with `new ShellJobRegistry({ jobs })` (the CLI and `fullAgentOptions` do).
95
96
 
96
97
  ## The `agentx` CLI
97
98
 
@@ -118,7 +119,7 @@ agentx --resume <id> "…" # resume a specific session
118
119
  - **Raise, don't guess** (`AskUserQuestion` + `finishReason: 'needs_input'`) — when a decision is genuinely the user's, the agent asks. Interactively that's an arrow-select prompt. **Headless** (`-p`, piped, no TTY) there's nobody to prompt, so instead of blocking on a prompt no one can see — or, worse, "proceeding with best judgment" — the run **stops and hands the question to its caller**: exit code `2`, a yellow `? needs input` footer, and `{finishReason:'needs_input', question:{…}, sessionId}` in `--output-format json`. The transcript is preserved, so the caller answers and continues (`--resume <id> -p "<answer>"`) instead of re-running. A delegated agent silently deciding something that was never its to decide is the failure this closes.
119
120
  - **Tab-completion** — `Tab` completes `/<command>` names and `@<path>` file/dir references (descends subdirs, dotfiles hidden unless typed) straight from the working tree.
120
121
  - **Duplex mode** — `agentx --duplex` runs the full standard REPL (slash commands, sessions, postures, rewind, MCP) with the three-tier engine driving turns: a fast voice model (`--voice-model`, default `groq/openai/gpt-oss-120b`) answers every line instantly and delegates real work to background workers built with the same wiring as a normal run (fs mode, permissions, MCP); worker activity shows as dim chrome and results are re-voiced when ready. Switch any tier live with `/model` (opens a reflex/act/think picker), or the `/voice-model` · `/think-model` shortcuts. `/tasks` lists background tasks, inspects a task's live output tail, and cancels a running one from a picker (Esc mid-turn cancels the foreground turn; Esc again at the idle prompt cancels running workers).
121
- - **MCP servers** — declare `mcpServers: { name: { command, args } | { url } }` in config and they're auto-mounted at startup (in parallel, with an optional `mountTimeoutMs` deadline so one slow/dead server never blocks the rest): the client does the JSON-RPC handshake (stdio or HTTP) + `tools/list`, and the discovered tools appear as `mcp__<name>__<tool>` in `/tools` (inspect with `/mcp`). A bad server is logged and skipped, never blocking the agent. For large tool sets, **deferred mode** (`makeMcpToolSearch` / `mountMcpDeferred`) exposes just two bounded tools (`ToolSearch` + `McpCall`) instead of N defs — dodging the provider tool-cap and improving selection accuracy; the CLI applies this automatically past 12 mounted tools (a 42-tool server was costing ~80k tok/turn in schema alone), and permission rules written against the real `mcp__<name>__<tool>` names still match through `McpCall`. **`mountMcpCatalog`** goes further: a cached, hash-keyed catalog + lazy connect means a turn that uses no MCP tool opens **zero** connections, and one that uses a tool connects exactly that server — latency scales with tools-used, not servers-configured. A down server is **negative-cached** (`failureCooldownMs`) so it never re-floors a later turn at the deadline. For zero turn-path latency even on a cold process, call **`warmMcpCatalog`** at boot + on a timer (off-turn discovery) and mount with **`{ discover: 'cache-only' }`** — the turn then never synchronously connects: it serves the warmed catalog and discovers any miss in the background.
122
+ - **MCP servers** — declare `mcpServers: { name: { command, args } | { url } }` in config and they're auto-mounted at startup (in parallel, with an optional `mountTimeoutMs` deadline so one slow/dead server never blocks the rest): the client does the JSON-RPC handshake (stdio or HTTP) + `tools/list`, and the discovered tools appear as `mcp__<name>__<tool>` in `/tools` (inspect with `/mcp`). A bad server is logged and skipped, never blocking the agent. For large tool sets, **deferred mode** (`makeMcpToolSearch` / `mountMcpDeferred`) exposes just two bounded tools (`ToolSearch` + `McpCall`) instead of N defs — dodging the provider tool-cap and improving selection accuracy; the CLI applies this automatically past 12 mounted tools (a 42-tool server was costing ~80k tok/turn in schema alone), and permission rules written against the real `mcp__<name>__<tool>` names still match through `McpCall`. **`mountMcpCatalog`** goes further: a cached, hash-keyed catalog + lazy connect means a turn that uses no MCP tool opens **zero** connections, and one that uses a tool connects exactly that server — latency scales with tools-used, not servers-configured. A down server is **negative-cached** (`failureCooldownMs`) so it never re-floors a later turn at the deadline. For zero turn-path latency even on a cold process, call **`warmMcpCatalog`** at boot + on a timer (off-turn discovery) and mount with **`{ discover: 'cache-only' }`** — the turn then never synchronously connects: it serves the warmed catalog and discovers any miss in the background. **Slow tools are jobs, not timeouts**: `timeoutMs` (default 30s) bounds protocol requests (`initialize`, `tools/list`) only; a `tools/call` still running after the soft foreground timeout (`foregroundMs.mcp`, 30s) returns `[still running] … background job job-N` while the request keeps going, and its result is delivered when it lands. The client sends a `progressToken` and reads progress over stdio and incrementally over HTTP/SSE; a call is cancelled (`notifications/cancelled`) only by `JobKill`, the stuck-job reaper, or its hard cap `callTimeoutMs` (per server, default 30 min).
122
123
 
123
124
  ## 🧬 It improves itself
124
125
 
@@ -1,5 +1,5 @@
1
1
  import { IFilesystem } from '@livx.cc/wcli/core';
2
- import { M as Message, H as HostBridge, A as AgentTool, C as ChatLike, g as MessageContent, U as UserQuestion } from './tools-BqL8Lk4J.js';
2
+ import { M as Message, H as HostBridge, A as AgentTool, C as ChatLike, k as JobRegistry, o as MessageContent, U as UserQuestion } from './tools-HsbxgqjF.js';
3
3
 
4
4
  /**
5
5
  * Hooks — deterministic interception points around tool execution, run by the
@@ -244,6 +244,15 @@ declare class AgentOptions {
244
244
  * arrives for this long. The between-steps `timeoutMs` can't preempt an in-flight `chat()`, so a
245
245
  * provider that goes silent mid-stream would otherwise park the loop forever. 0 = off. */
246
246
  stallMs: number;
247
+ /** IDLE bound on a NATIVE delegated tool's hold over the stall watchdog (cursor's own shell/mcp/task:
248
+ * `running` → silence → `completed`). Such a hold is the provider's own unverified CLAIM that it is
249
+ * busy, and it SUSPENDS the very stall detection that would catch that provider going silent — so its
250
+ * only backstop was `hardCapMs = toolHoldMs`, the TRANSPORT deadline (≥30m, sized for host-tool
251
+ * round-trips a native call never makes). A wedged native call therefore bought 30m of total silence,
252
+ * ×2 attempts = the 61-minute zero-output run (blank 2026-09-18). Idle, not absolute: every further
253
+ * `running` event `touch()`es the job, so a native tool that keeps reporting progress holds
254
+ * indefinitely — what expires is silence, not work. 0 = off (hard cap only). */
255
+ nativeHoldMs: number;
247
256
  /** Stop if the identical tool-call batch (name+args) repeats this many times in a row. 0 = off. */
248
257
  maxRepeats: number;
249
258
  /** Cumulative cap on tool calls dispatched across the run. 0 = unbounded. */
@@ -367,6 +376,12 @@ declare class AgentOptions {
367
376
  /** Enable `bash({background:true})` + JobOutput/JobStatus/JobKill — sandbox background jobs (overlay-isolated,
368
377
  * committed on completion, drained at turn end). Useful when the VFS backend is slow (remote) or for sub-agents. */
369
378
  backgroundJobs: boolean;
379
+ /** The job registry every long-lived tool call lives in (MCP calls past their soft timeout, delegated host/native
380
+ * tools, and — when the host shares it with its `ShellJobRegistry` — shell jobs). Its liveness is what the stall
381
+ * watchdog and the stuck-job reaper decide from. Default: a fresh registry per Agent. Knobs (idle windows, hard
382
+ * caps, MCP soft timeout) live on `JobRegistryOptions`. Completion notices reach the model via `inject` unless
383
+ * the host set its own `onExit`. */
384
+ jobs?: JobRegistry;
370
385
  /** Plan mode: block mutating tools until the agent calls `ExitPlanMode` (host-approved). */
371
386
  planMode: boolean;
372
387
  /** Permission policy gating each tool call (allow / ask / deny). */
@@ -421,17 +436,18 @@ declare class Agent {
421
436
  /** Per-run ground truth: every tool name actually INVOKED this run (native dispatch + delegated
422
437
  * runtime activity). The ledger that makes a fabricated tool claim provable. Reset per run. */
423
438
  private invokedTools;
424
- /** Delegated-runtime HOST tools currently executing (cursor toolExecutor). While > 0 the provider is
425
- * not silent — it is waiting on US — so the idle-stall watchdog must not fire (see armStallWatchdog). */
426
- private hostToolsInFlight;
427
- /** The hold is bounded by the transport tool deadline: a host tool that never settles (the helper already gave
428
- * up on it) must not disable stall protection for the rest of the request — or this instance. */
429
- private hostToolHoldUntil;
430
- /** Delegated-runtime NATIVE tools (cursor's own shell/mcp/task) seen `running` but not yet `completed`/`error`,
431
- * id → hold-until. The runtime streams nothing while one runs, so it holds the watchdog like a host tool —
432
- * bounded by the same tool deadline (`toolHoldMs`) so a lost `completed` event can't disable it. */
433
- private nativeToolsInFlight;
439
+ /** Every long-lived tool call is a job here (see AgentOptions.jobs). Delegated-runtime HOST tools (cursor
440
+ * toolExecutor) and NATIVE tools (cursor's own shell/mcp/task, `running` → `completed`) are tracked as
441
+ * `holds` jobs: while one runs the provider is not silent — it is waiting on a tool — so the idle-stall
442
+ * watchdog must not fire. Each is bounded by the transport tool deadline (`hardCapMs = toolHoldMs`), so a
443
+ * tool that never settles (or a lost `completed` event) is reaped and cannot disable stall protection.
444
+ * A NATIVE job carries a second, much tighter bound (`idleMs = nativeHoldMs`): the transport deadline is
445
+ * sized for a host-tool round-trip it never makes, and silence is the only thing it can actually prove. */
446
+ readonly jobs: JobRegistry;
447
+ /** Native delegated tool activity id → its job (per step attempt). */
448
+ private nativeJobs;
434
449
  private toolHoldMs;
450
+ private nativeHoldMs;
435
451
  /** A tool (host or native) already executed in the current step attempt: a stall retry replays the step
436
452
  * from the pre-step transcript, which would re-run it — so such a stall is not retried. */
437
453
  private toolRanThisAttempt;
@@ -541,7 +557,8 @@ declare class Agent {
541
557
  * `stallMs = 0` disables the timer but still links parent-abort → child so cancellation propagates.
542
558
  */
543
559
  private armStallWatchdog;
544
- /** A host or native delegated tool is running within its deadline — the provider is quiet, not stalled. */
560
+ /** A host or native delegated tool is running and not stuck — the provider is quiet, not stalled. The reap runs
561
+ * FIRST, so a holder past its deadline is released (logged with evidence) instead of holding forever. */
545
562
  private toolHoldActive;
546
563
  /** A clean-stop turn that carries no assistant text AND no tool calls (native or delegated) — a
547
564
  * non-result the model produced by going silent. Callers retry it rather than report "done". */
package/dist/cli.d.ts CHANGED
@@ -1,7 +1,7 @@
1
1
  #!/usr/bin/env bun
2
- import { H as Hooks, i as RunResult, R as ReasoningEffort, A as Agent } from './Agent-BWbQ4Sqp.js';
2
+ import { H as Hooks, i as RunResult, R as ReasoningEffort, A as Agent } from './Agent-CXkPfYX2.js';
3
3
  import { IFilesystem } from '@livx.cc/wcli/core';
4
- import { M as Message, U as UserQuestion, H as HostBridge, c as ContentPart, g as MessageContent } from './tools-BqL8Lk4J.js';
4
+ import { M as Message, U as UserQuestion, H as HostBridge, c as ContentPart, o as MessageContent } from './tools-HsbxgqjF.js';
5
5
 
6
6
  /**
7
7
  * On-disk session store for the CLI: each conversation is one JSON file at