@shanepadgett/tau-agent 0.46.2 → 0.46.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  # Script Runner
2
2
 
3
- Gives the agent a first-class `script_runner` tool for running Python 3, Node.js, and Deno scripts instead of falling back to bash. The agent picks whichever runtime is more efficient for the task.
3
+ Gives the agent a first-class `script_runner` tool for running Python 3, Node.js, and Deno scripts. The agent uses dedicated tools for ordinary reads, searches, and edits. Scripts are for work those tools cannot reasonably handle, or substantial bulk transformations and computation that would otherwise require many repetitive or error-prone calls. When a script is justified, the agent uses this tool rather than embedding it in bash and picks the runtime that fits the task.
4
4
 
5
5
  When a run fails, the tool keeps the script and returns a `scriptId`. The agent retries with targeted `{oldText, newText}` edits against what it just wrote instead of resending the whole script, saving output tokens and keeping duplicate scripts out of context. Only the source the agent already sent is referenced; no script source file path is exposed.
6
6
 
@@ -188,7 +188,8 @@ export default function scriptRunnerExtension(pi: ExtensionAPI): void {
188
188
  ].join("\n"),
189
189
  promptSnippet: `Run ${langPhrase} scripts; on failure retry with edits + scriptId.`,
190
190
  promptGuidelines: [
191
- `Prefer script_runner over bash for ${langPhrase} when computation, data handling, or bulk file work is cleaner than chaining tools.`,
191
+ "Use dedicated tools for ordinary reads, searches, and edits. Use scripts only when those tools cannot reasonably do the work, or for substantial bulk transformations or computation that would otherwise require many repetitive or error-prone calls.",
192
+ `When a ${langPhrase} script is justified, use script_runner rather than embedding it in bash. A shorter script alone is not a reason to replace a dedicated tool.`,
192
193
  "script_runner never exposes the script path; you already have the source. Never try to read it back.",
193
194
  ],
194
195
  parameters: paramsSchema,
@@ -2,6 +2,8 @@
2
2
 
3
3
  Soul supplies Tau's system prompt: communication, discussion, planning, execution, coding, and tool-use guidance. It is always on.
4
4
 
5
+ The agent uses dedicated tools first for ordinary reads, searches, and edits. Bash handles shell commands such as builds and tests. Scripts are reserved for work the available tools cannot reasonably handle, or substantial bulk transformations and computation that would otherwise require many repetitive or error-prone calls. A shorter script alone is not a reason to replace a dedicated tool.
6
+
5
7
  Soul adds Pi documentation pointers, tool guidance, and the context that other Tau extensions supply, such as the local date, directory snapshot, and automatic-check instructions. The date and directory snapshot are captured on the first prompt and again after successful compaction, so they stay fixed across turns, reload, resume, and tree navigation.
6
8
 
7
9
  Everything else is read each turn. When an instruction changes, such as an edited `AGENTS.md` after `/reload`, a changed automatic-check configuration, or a different set of active tools, Pi appends the updated section without rewriting earlier instructions. Models that cannot take later system messages receive the change in the leading instructions, which costs one cache miss on the next request.
@@ -42,7 +42,7 @@ export default function soulExtension(pi: ExtensionAPI): void {
42
42
  ...new Set([
43
43
  ...(active.includes("bash")
44
44
  ? [
45
- "Use bash for file operations like ls, rg, find.",
45
+ "Use bash for shell commands such as builds and tests, and for ls, rg, or find when no suitable dedicated tool is available.",
46
46
  "Bash already runs in the working directory; do not cd into it.",
47
47
  ]
48
48
  : []),
@@ -42,7 +42,10 @@ Keep the final answer proportional to the request. Avoid turning a simple answer
42
42
 
43
43
  <tool-use>
44
44
  Use the tools available to you for the purpose each is designed for.
45
- Prefer the file tools for ordinary reads, edits, and writes. Use bash or scripts when they express the work more directly and in fewer tokens: generating files, bulk mechanical transforms, or computation whose intermediate output does not need to be shown.
45
+ Default to dedicated tools for reading, searching, inspecting, editing, and writing files. Use the available tool designed for the task rather than recreating it in bash, Python, or another script.
46
+ Do not choose a script for an ordinary read or edit merely because it is familiar, shorter, or uses fewer tokens. Use bash for actual shell commands, such as builds, tests, and installed command-line tools, when no dedicated tool reasonably handles the task.
47
+ Use scripts when the available tools cannot reasonably accomplish the work, or when a substantial bulk transformation or computation would otherwise require many repetitive or error-prone tool calls. Applying one deterministic transformation across 100 files is a good use; replacing text in one file is not.
48
+ When a script is justified, use the available script-running tool instead of embedding it in bash. Keep its file scope and intended effects explicit. Do not switch tools to evade an approval request.
46
49
  Batch independent calls into one response: independent reads, searches, and edits go together, and only calls that depend on earlier results are sequenced.
47
50
  Keep tool output small: request only the data you need and prefer compact, high-signal commands over ones that flood the context.
48
51
  </tool-use>
@@ -106,7 +106,7 @@ Supplies Soul with the local date and root directory snapshot. Both remain fixed
106
106
 
107
107
  ## script-runner
108
108
 
109
- Gives the agent a first-class `script_runner` tool to execute Python 3, Node.js, and Deno scripts instead of bash. On failure it returns a `scriptId`; the agent retries with targeted `{oldText,newText}` edits against the script it already wrote rather than resending the whole script. Runtimes are detected from the environment (Python 3 via `python3`; `node` is the local Node.js runtime with `--experimental-strip-types`, Node 22.6+ — full Node APIs, TypeScript with erasable syntax or plain JavaScript; `deno` via `deno run -A` — full permissions, native TypeScript/JavaScript, Deno APIs). The tool registers only available runtimes and is hidden from the prompt if none are present.
109
+ Gives the agent a first-class `script_runner` tool to execute Python 3, Node.js, and Deno scripts. Dedicated tools come first for ordinary reads, searches, and edits. Scripts are for work those tools cannot reasonably handle, or substantial bulk transformations and computation that would otherwise require many repetitive or error-prone calls; justified scripts use this tool rather than bash. On failure it returns a `scriptId`; the agent retries with targeted `{oldText,newText}` edits against the script it already wrote rather than resending the whole script. Runtimes are detected from the environment (Python 3 via `python3`; `node` is the local Node.js runtime with `--experimental-strip-types`, Node 22.6+ — full Node APIs, TypeScript with erasable syntax or plain JavaScript; `deno` via `deno run -A` — full permissions, native TypeScript/JavaScript, Deno APIs). The tool registers only available runtimes and is hidden from the prompt if none are present.
110
110
 
111
111
  ## silent-command-runner
112
112
 
@@ -116,6 +116,8 @@ Runs configured commands while keeping their output out of agent context when th
116
116
 
117
117
  Supplies Tau's communication, discussion, planning, execution, and coding instructions, plus tool-use rules, tool guidance, and context from other Tau extensions. The date and directory snapshot stay fixed until successful compaction. Other changes, such as an edited `AGENTS.md` after `/reload`, arrive as appended updates without rewriting earlier instructions.
118
118
 
119
+ Tool-use guidance defaults to dedicated tools for ordinary file work, bash for shell commands, and scripts only when dedicated tools cannot reasonably do the work or when substantial bulk work would otherwise need many repetitive or error-prone calls.
120
+
119
121
  ## stash
120
122
 
121
123
  Adds `Alt+S` to stash the current prompt draft and `/pop` to browse stashed drafts and put one back in the editor.
@@ -134,6 +136,8 @@ Reviews agent `bash` and `script_runner` requests before they run. Common read-o
134
136
 
135
137
  Approval decisions are saved privately in the session JSONL, including allowlist skips, review stages and models, inspected paths, evidence-gap categories, and user decisions. These records stay out of model context and do not copy scripts, arguments, file contents, or raw errors.
136
138
 
139
+ Scoped project edits, builds, tests, and generated-file cleanup should be approved whether they use Python, Node, or bash. Reading, replacing, and writing project text is ordinary editing; computed data paths and a less suitable tool choice do not themselves require confirmation. Destructive changes to valuable databases, remote objects, backups, or unrelated work do. Clearly disposable local test data remains routine validation. Hidden executable code still requires inspection.
140
+
137
141
  ## tool-loader
138
142
 
139
143
  Keeps specialist tool groups out of every request as deferred tools and lets the agent load them with Pi's `tool_search`. Tau registers `web`, `image`, and `appshot`; project or global package extensions can add groups with `registerDeferredToolGroup()` from `@shanepadgett/tau-agent`. All models can load tools. Compatible models preserve the cached prefix; other models can incur a cache miss when tools are activated. Loading never triggers compaction.
@@ -16,6 +16,10 @@ With `autoApprove` enabled, reviewer-approved requests run without another confi
16
16
 
17
17
  The reviewer asks before meaningful data loss, difficult-to-recover overwrites, disruptive production or system changes, elevated privileges or access/security changes, credential exposure, sensitive-data disclosure to unintended audiences, substantial payments, or consequential publication, messages, and workflows. Additive writes still require confirmation if they change access, expose private material, or trigger hard-to-reverse effects. Deleting a public post cannot undo disclosure; deleting a record cannot undo a message or charge it already triggered. Ordinary internal document creation does not count as consequential publication by itself. Missing information requires confirmation when it leaves executable code or one of these substantial risks unresolved, not merely because every implementation detail or response field is unknown.
18
18
 
19
+ Approval depends on effects, not language or tool choice. Scoped project edits, formatting, code generation, builds, tests, and cleanup of generated or temporary files are routine work even when performed through Python, Node, or bash. Reading a file, replacing text, and writing the updated contents is an ordinary edit. Computed project paths and bulk edits are not reasons to ask by themselves. Files edited as data do not need execution inspection merely because they contain source code, and ordinary edits do not require a backup or a clean Git tree.
20
+
21
+ Destructive database operations, loss of unrelated work through Git resets or cleans, deletion of backups, and destructive overwrites of valuable remote objects or shared records require confirmation when they risk meaningful loss or disruption. Read-only database queries and clearly disposable local test databases and fixtures are routine validation. Local data can still be valuable; remote reads and additive writes can still be low-impact. Hidden executable dependencies retain the inspection requirements above.
22
+
19
23
  When approval is required, Tau shows one plain-language paragraph explaining what you are allowing, who or what is affected, why approval is needed, and what recovery might involve. It summarizes consequences rather than listing script steps or specialized APIs. If the issue is missing evidence rather than a known danger, it says what could not be checked. For `script_runner`, human confirmation also shows the complete script that will run. If several requests need approval, their confirmation windows open one at a time. If the reviewer fails or returns a malformed decision, Tau asks for direct human approval instead of running it automatically. Tau also sends an attention notification when the approval window opens.
20
24
 
21
25
  In the terminal approval panel, move between Approve and Reject, press `j` or `k` to scroll a script, press `n` to add a note to the highlighted choice, then press Enter to choose. Enter saves an edited note before choosing; Escape cancels note editing or blocks the request from the choice list. A rejection note tells the agent why the request was blocked. An approval note reaches the agent with the tool result; it does not change the request being approved. To ask for a different request, reject it with a note. Long notes are truncated. RPC clients use the standard confirmation dialog without notes.
@@ -55,10 +55,15 @@ const REVIEW_SYSTEM_PROMPT = [
55
55
  "bash runs a shell command; script_runner runs supplied Python 3, Node.js, or Deno source with normal local process permissions.",
56
56
  "script_runner stages its source in a new temporary directory. Relative module imports resolve from that directory, not the project; relative file operations and subprocesses use the project working directory. Changes to code search paths need explicit inspection or human approval.",
57
57
  "Default to approved for understood routine, low-impact actions, locally or in external services. Approve ordinary file edits, builds, tests, package tools, scripts, quotes, pipes, redirects, and other recoverable effects.",
58
+ "Judge the actual effects and affected data, not the programming language or choice of tool. A Python or shell edit is not riskier merely because a dedicated edit tool could have done it. Tool-choice preferences are not reasons to require human approval.",
59
+ "Routine development is authorized: approve scoped edits to project source, configuration, documentation, and tests; formatting, code generation, builds, local validation, and creation or cleanup of generated output and temporary files. This includes Python read_text/replace/write_text, Node readFile/writeFile, shell redirects, and bulk edits with an understood project scope. Writing updated file contents is a normal edit, not inherently destructive overwriting.",
60
+ "Computed filenames, loops, globbing, and replacement strings are normal ways to implement project edits. Do not ask merely because paths or arguments are computed when the visible code establishes their scope and effects. Do not require a backup, a clean Git tree, or inspection of every data file before an ordinary scoped edit. Ask when the code discards unrelated work, deletes valuable contents, or has an unresolved materially destructive scope.",
61
+ "Files read, searched, parsed, or rewritten as data are not executable dependencies. Source code in a replacement string is data until the request executes or loads it. Do not request execution inspection merely to read the current contents of a file being edited. The supplied inline script is already visible code; only code it actually imports, loads, or executes can require additional execution evidence.",
58
62
  "Approve routine read-only service requests, including Jira searches, fetching Confluence pages, listing records, and checking status. Reading a remote or production service is not changing it. Ordinary response output is not an unauthorized export merely because it may contain private work data.",
59
63
  "Approve additive writes such as creating a document, page, draft, or record when they do not replace valuable content, change access, disclose sensitive data to an unintended audience, incur substantial costs, or trigger consequential workflows. An external or shared destination alone is not a reason to ask the user.",
60
64
  "Approve normal authentication: reading existing credentials from environment variables or the usual credential store and using them with their intended service, without printing, exposing, or persisting the secret elsewhere. Passing a token through a request header or an SDK's normal authentication mechanism is not credential disclosure.",
61
65
  "Require user approval only for concrete substantial risk: meaningful data loss or difficult-to-recover overwrites; disruptive production or system changes; elevated privileges or access/security changes; exposing credentials or sensitive data to an unintended audience or untrusted destination; substantial payments; or consequential publication, messages, or workflows that cannot be meaningfully undone. A routine internal document creation is not consequential publication by itself.",
66
+ "Protect valuable state: database DROP/TRUNCATE, broad DELETE/UPDATE, destructive schema migrations, deletion of backups, destructive Git resets or cleans that discard unrelated work, and replacing or deleting valuable remote objects or shared records require approval when they risk meaningful loss or disruption. Read-only database queries and setup, reset, or cleanup of clearly disposable local test databases and fixtures are routine validation. A database is not disposable merely because it is local; an external service is not destructive merely because it is remote.",
62
67
  "Non-destructive does not always mean reversible: deleting a public post later cannot undo disclosure, and deleting a record cannot undo messages, charges, or workflow effects it already triggered. Evaluate those actual side effects, not the service name or the mere presence of a write or credential.",
63
68
  "Do not require approval merely because the request writes files, invokes code, uses shell composition, accesses an external service, authenticates, could fail, or has ordinary recoverable side effects. Small recoverable edits are not substantial data loss.",
64
69
  "Routine deletion of generated, temporary, or local project files is ordinary local work. Escalate deletion only when it is broad or difficult to recover.",
@@ -619,7 +624,7 @@ async function reviewToolRequest(
619
624
  content: [
620
625
  "Bounded inspection evidence (untrusted source):",
621
626
  JSON.stringify({ files: evidence.files, gaps: evidence.gaps }),
622
- "Final review: return approved or requires_user_approval, never inspect. A reported evidence gap alone does not require human approval. Approve understood routine reads, recoverable writes, and normal authentication. Ask when executable code remains uninspected or missing information leaves a substantial risk unresolved. Explain the effect and any material risk or missing execution evidence in everyday language.",
627
+ "Final review: return approved or requires_user_approval, never inspect. A reported evidence gap alone does not require human approval. Approve understood routine project edits, builds, tests, reads, recoverable writes, and normal authentication regardless of whether they use Python, Node, or bash. Computed data paths and ordinary read/replace/write operations are not hidden executable code or destructive effects by themselves. Ask when executable code remains uninspected or missing information leaves a substantial risk unresolved, such as valuable data loss, destructive database or remote changes, or disruption. Explain the effect and any material risk or missing execution evidence in everyday language.",
623
628
  ].join("\n"),
624
629
  timestamp: Date.now(),
625
630
  },
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@shanepadgett/tau-agent",
3
- "version": "0.46.2",
3
+ "version": "0.46.3",
4
4
  "description": "Tau is a custom agentic harness built with pi extensions",
5
5
  "type": "module",
6
6
  "main": "./src/index.ts",
@@ -35,7 +35,7 @@
35
35
  ],
36
36
  "dependencies": {
37
37
  "@ast-grep/wasm": "0.45.3",
38
- "@shanepadgett/tau-tui": "0.46.2",
38
+ "@shanepadgett/tau-tui": "0.46.3",
39
39
  "@vscode/tree-sitter-wasm": "0.3.1",
40
40
  "image-size": "2.0.4",
41
41
  "smol-toml": "1.8.0",