zames_pro 2.55.0 → 2.59.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,220 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented in this file.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
+ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
+
8
+ ## [Unreleased]
9
+
10
+ ## [2.59.0]
11
+
12
+ ### Added
13
+
14
+ - `/improve [id]` — a backlog-driven self-improvement loop. It picks the
15
+ next open item from `BACKLOG.md` (highest priority, or an explicit id),
16
+ runs the standard task loop on it (implement, typecheck/lint/test, mark
17
+ the item done, add a CHANGELOG entry) and leaves the changes uncommitted
18
+ for review. This is the reproducible hand-off path for a fresh agent.
19
+
20
+ ## [2.58.0]
21
+
22
+ ### Added
23
+
24
+ - `--output-format json|jsonl` for one-shot runs: transcript events are
25
+ streamed as JSON lines on stdout (human text goes to stderr), so zames
26
+ pipes into `jq` and CI pipelines. `Transcript` gained an `onLine` mirror.
27
+ - Loop detection: four identical tool calls in a row trigger a nudge to
28
+ change approach instead of repeating the same call forever.
29
+ - MCP configs pinned to `@latest`/`@next` now warn (in `/mcp` and at
30
+ startup) so the tool set does not silently change between runs.
31
+ - `/help` is grouped into sections (session / workspace / git / agent /
32
+ context / files) instead of one long list.
33
+ - `Grep` accepts a comma-separated `include` list.
34
+ - System prompt: a "Choosing the right tool" section (Edit vs MultiEdit vs
35
+ ApplyPatch, Grep vs Glob vs LS, Git tools vs Bash).
36
+
37
+ ### Changed
38
+
39
+ - `coverage-gate` enforces floors for 8 more modules and prints the
40
+ lowest-covered modules in the CI log.
41
+ - README: Why zames?, an FAQ and a machine-readable-output section.
42
+ - Dependabot config for npm and GitHub Actions.
43
+
44
+ ## [2.57.1]
45
+
46
+ ### Changed
47
+
48
+ - CI workflows use actions/checkout@v5 and actions/setup-node@v5 (the v4
49
+ actions run on the deprecated Node 20 runtime and showed up as an
50
+ annotation on every run).
51
+
52
+ ### Added
53
+
54
+ - README badges (npm version/downloads, CI status, license, Node, PRs) and a
55
+ Features section, so the GitHub landing page and the npm page show what the
56
+ project is at a glance.
57
+ - Community files: issue forms (bug report, feature request), a PR template
58
+ and CODE_OF_CONDUCT.md.
59
+ - package.json: a fuller description and more keywords (npm search), and
60
+ CHANGELOG.md is now included in the published files.
61
+
62
+ ## [2.57.0]
63
+
64
+ ### Added
65
+
66
+ - Plan mode (read-only): start with --plan or toggle with /plan [on|off].
67
+ In this mode mutating tools (Write/Edit/MultiEdit/ApplyPatch/Bash, GitAdd/
68
+ GitCommit/GitPush) are removed from the tool set entirely, so the agent can
69
+ investigate without touching the tree.
70
+ - src/fsutil.ts - one shared atomic writer (writeFileAtomic/writeJsonAtomic).
71
+
72
+ ### Changed
73
+
74
+ - config, sessions and undo now share the single atomic writer (the
75
+ temp-file+rename logic used to be copy-pasted in three places).
76
+
77
+ ### Fixed
78
+
79
+ - BACKLOG.md is now git-ignored and untracked: it is the agent's own
80
+ improvement-notes file, rewritten constantly, and used to be published with
81
+ the package.
82
+ - GitPush validates the branch name before building the shell command, so a
83
+ model-supplied branch cannot smuggle shell metacharacters.
84
+ - SECURITY.md no longer promises a /permissions command / alwaysConfirm list
85
+ that was removed in 2.55.0; AGENTS.md cleaned of the same dead refs.
86
+
87
+ ## [2.56.0]
88
+
89
+ ### Added
90
+
91
+ - `/context` — show what is loaded into the prompt (AGENTS.md, MEMORY.md,
92
+ skills, custom commands) plus the system-prompt size.
93
+ - `/retry` — resend the last task into the same chat (handy after a truncated
94
+ or empty answer).
95
+ - `/rename <title>` — set the current session title (shown in `/sessions`).
96
+ - `/diffstat` — a one-line `git diff --stat` change summary.
97
+ - `/copy` — copy the last assistant answer to the OS clipboard (wl-copy /
98
+ xclip / xsel / pbcopy / clip).
99
+ - `GitShow` and `GitBranchList` git tools (read-only).
100
+ - The status line now shows the elapsed time of the current phase ("45s",
101
+ "2m 05s") in BOTH the LineEditor and the non-TTY spinner, and the spinner
102
+ shows the same "tasks: 2/5" badge the editor does.
103
+ - A one-time hint when the context reaches ~80% while auto-compact is off,
104
+ pointing at `/compact`.
105
+ - Tests: dead-i18n-key guard, agent-facing-no-Cyrillic guard, CLI flag
106
+ consistency, binary content-type handling for `WebFetch`.
107
+
108
+ ### Changed
109
+
110
+ - `WebFetch` returns a short human note for binary content types (PDF, images,
111
+ archives, ...) instead of dumping raw bytes into the model context.
112
+ - `npm run build` now wipes `dist/` first (`scripts/clean-dist.mjs`), so a
113
+ removed source file no longer leaves an orphan `.js` in the package.
114
+
115
+ ### Fixed
116
+
117
+ - `/undo` failure reasons are localized: `src/undo.ts` returned hardcoded
118
+ Russian strings that showed up in an English UI. It now returns machine codes
119
+ ('empty' / 'disabled') localized by the caller.
120
+ - Removed the dead `--project` flag (it only swallowed the following argument)
121
+ and the no-op `--calibrate` flag together with their help/i18n entries.
122
+ - Removed the orphaned `confirm.hint` / `confirm.ask_label` i18n keys left by
123
+ the confirm-subsystem removal.
124
+ - README Node requirement aligned with `engines` / `.nvmrc` (>= 20).
125
+ - `GitBranchList` quotes its `--format` argument so the POSIX shell does not
126
+ choke on the parentheses/brackets in the format string.
127
+
128
+ ## [2.55.0]
129
+
130
+ ### Added
131
+
132
+ - `!command` in the prompt runs a shell command directly, bypassing the model
133
+ (like Claude Code's bash mode). The same sandbox guard as the `Bash` tool
134
+ applies, so a direct command cannot leave the project.
135
+ - `/remember <text>` appends a durable note to the project `MEMORY.md`, so a
136
+ fact worth keeping across sessions is written to disk instead of living only
137
+ in the chat.
138
+ - `/skills`, `/memory`, `/remember`, `/init` and `/mcp` are now listed in the
139
+ `/help` output (they were only in the «/» completion list before).
140
+
141
+ ### Changed
142
+
143
+ - The shell runner moved to `src/shell.ts`, shared by the `Bash` tool and the
144
+ new `!command` escape (one sandbox guard, one abort wiring).
145
+ - `/self-review` now copies root configs (`package.json`, `eslint.config.js`,
146
+ `tsconfig*.json`, `.prettierrc`) and `scripts/**` into a read-only `_context/`
147
+ folder of the snapshot, so the reviewer can read them. The editable set
148
+ (diff/apply) stays `.ts`-only.
149
+ - `package.json` gained `packageManager: npm@11.19.0`.
150
+
151
+ ### Fixed
152
+
153
+ - The status line no longer prefixes a running-tool indicator with the
154
+ previous answer's "done" phase (e.g. `✓ done: running Bash`): `toolCall()`
155
+ clears the send-phase state before starting the tool animation.
156
+
157
+ ### Removed
158
+
159
+ - The never-wired tool-confirmation subsystem: `src/confirm.ts`
160
+ (`ConfirmManager`/`formatDiffPreview`) was fully implemented and tested but
161
+ never called by the runtime, so `confirmation.write/edit/bash`,
162
+ `alwaysConfirm` and the `/permissions` command did nothing. Removed the
163
+ module, its tests, the config keys, the `/permissions` command and the
164
+ related i18n strings.
165
+
166
+ ## [2.54.0]
167
+
168
+ ### Added
169
+
170
+ - New unit tests for previously uncovered modules: `diff.ts` (unified/color
171
+ diff), `spinner.ts` (ellipsis/dots/UI surface), `confirm.ts`
172
+ (confirmation logic + alwaysConfirm fallback) and `markdown.ts`
173
+ (rendering of headings, lists, code, tables).
174
+
175
+ - `WebFetch` now blocks private/loopback/link-local addresses (SSRF guard,
176
+ including cloud metadata `169.254.169.254`) and retries transient network
177
+ errors and 5xx responses with exponential backoff.
178
+ - CI prints per-file test coverage (`node --experimental-test-coverage`) so
179
+ untested modules are visible in the log (informational, does not fail).
180
+
181
+ ### Changed
182
+
183
+ - The `pre-push` git hook no longer runs the full test suite (it was fragile
184
+ and timed out on tag pushes); tests stay in `pre-commit` and CI.
185
+
186
+ ## [2.53.1]
187
+
188
+ ### Fixed
189
+
190
+ - `require('fs')` in ESM modules (`src/attachments.ts`, `src/self-review.ts`)
191
+ threw `ReferenceError` at runtime: WSL clipboard detection always returned
192
+ `false` and `/self-review` found no sources when run from the built `dist/`.
193
+ - Config, session and undo index writes are now atomic (temp file + rename),
194
+ so a crash or Ctrl+C mid-write can no longer corrupt `~/.zames/config.json`
195
+ or a saved session.
196
+
197
+ ### Changed
198
+
199
+ - Removed the unused `marked` dependency (Markdown is rendered by `markdansi`).
200
+ - CI installs dependencies with `npm ci` for reproducible builds.
201
+ - The i18n test now validates the whole `CATALOG` instead of a small sample.
202
+ - Added `.nvmrc` (Node 24).
203
+
204
+ ## [2.53.0]
205
+
206
+ ### Added
207
+
208
+ - Session banner warns when no saved DeepSeek session exists.
209
+ - `--no-color` flag (explicit `NO_COLOR`).
210
+
211
+ [Unreleased]: https://github.com/Viqto0r/zames_pro/compare/v2.59.0...HEAD
212
+ [2.59.0]: https://github.com/Viqto0r/zames_pro/compare/v2.58.0...v2.59.0
213
+ [2.58.0]: https://github.com/Viqto0r/zames_pro/compare/v2.57.1...v2.58.0
214
+ [2.57.1]: https://github.com/Viqto0r/zames_pro/compare/v2.57.0...v2.57.1
215
+ [2.57.0]: https://github.com/Viqto0r/zames_pro/compare/v2.56.0...v2.57.0
216
+ [2.56.0]: https://github.com/Viqto0r/zames_pro/compare/v2.55.0...v2.56.0
217
+ [2.55.0]: https://github.com/Viqto0r/zames_pro/compare/v2.54.0...v2.55.0
218
+ [2.54.0]: https://github.com/Viqto0r/zames_pro/compare/v2.53.1...v2.54.0
219
+ [2.53.1]: https://github.com/Viqto0r/zames_pro/compare/v2.53.0...v2.53.1
220
+ [2.53.0]: https://github.com/Viqto0r/zames_pro/releases/tag/v2.53.0
package/README.md CHANGED
@@ -1,14 +1,53 @@
1
1
  # zames_pro
2
2
 
3
+ [![npm version](https://img.shields.io/npm/v/zames_pro.svg)](https://www.npmjs.com/package/zames_pro)
4
+ [![npm downloads](https://img.shields.io/npm/dm/zames_pro.svg)](https://www.npmjs.com/package/zames_pro)
5
+ [![tests](https://github.com/Viqto0r/zames_pro/actions/workflows/test.yml/badge.svg)](https://github.com/Viqto0r/zames_pro/actions/workflows/test.yml)
6
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
7
+ [![Node.js](https://img.shields.io/badge/node-%3E%3D20-brightgreen.svg)](package.json)
8
+ [![PRs welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg)](CONTRIBUTING.md)
9
+
3
10
  ![zames logo](https://raw.githubusercontent.com/Viqto0r/zames_pro/master/logo.jpg)
4
11
 
5
12
  A terminal coding agent that works on top of [chat.deepseek.com](https://chat.deepseek.com/) through Playwright.
6
13
  In spirit it is similar to Claude Code / Codex CLI: it starts in the current
7
14
  directory, reads and edits files, runs commands, and commits to git.
8
15
 
16
+ > No API key required — it drives the DeepSeek web chat like a regular user
17
+ > through a real (headless) browser.
18
+
19
+ ## Features
20
+
21
+ - **Tools like Claude Code / Codex** — `Read`, `Write`, `Edit`, `Bash`,
22
+ `Glob`, `Grep`, plus `MultiEdit`, `ApplyPatch`, `LS`, `TodoWrite`, git and web
23
+ tools. Every edit is backed by `/undo`.
24
+ - **Runs while you keep typing** — queue messages during a task; they run right
25
+ after it, in the same chat (like typing during generation on the web).
26
+ - **Plan mode** — `/plan` (or `--plan`) drops all mutating tools, so the agent
27
+ can investigate the code without touching the tree.
28
+ - **Project context** — reads `AGENTS.md`, `MEMORY.md`, skills (`SKILL.md`) and
29
+ custom commands from the repo and `~/.zames`, the same idea as Codex / Claude
30
+ Code.
31
+ - **MCP support** — plug in external tool servers (e.g. `@playwright/mcp`).
32
+ - **Self-review** — `/self-review` snapshots `src/` so the agent can review and
33
+ fix itself in a sandbox (`/self-fix`, `/self-apply`).
34
+ - **Scheduling** — `/loop`, `/cron` and `/jobs` repeat tasks on a timer.
35
+ - **Bilingual UI** — Russian / English (`/config lang`).
36
+
37
+ ## Why zames?
38
+
39
+ - **No API key, no per-token bill** — it uses your own DeepSeek chat account,
40
+ not the paid API. Good for long, tool-heavy tasks.
41
+ - **Same workflow as Claude Code / Codex** — tools, `AGENTS.md`, skills, MCP
42
+ and slash commands, so it feels familiar from day one.
43
+ - **Runs unattended** — headless by default, resumable sessions, and `/loop`
44
+ plus `/cron` for scheduled work.
45
+ - **Local-first** — the browser profile, credentials and logs never leave your
46
+ machine, and the agent is sandboxed to the project directory.
47
+
9
48
  ## Requirements
10
49
 
11
- - Node.js >= 18
50
+ - Node.js >= 20 (CI and development use Node 24; see `.nvmrc`)
12
51
  - A DeepSeek account. On first launch zames asks for your DeepSeek
13
52
  login/password in the terminal (and stores them in `~/.zames/config.json`
14
53
  after a successful sign-in, so a later logout is handled automatically
@@ -46,6 +85,13 @@ version). You do not need `--headed` just to log in.
46
85
  Credentials and toggles can also be edited from `/config`
47
86
  (`browser.auth.username`, `browser.auth.password`, `browser.auth.saveSession`).
48
87
 
88
+ ## Links
89
+
90
+ - npm: <https://www.npmjs.com/package/zames_pro>
91
+ - Changelog: [`CHANGELOG.md`](CHANGELOG.md)
92
+ - Contributing: [`CONTRIBUTING.md`](CONTRIBUTING.md)
93
+ - Security policy: [`SECURITY.md`](SECURITY.md)
94
+
49
95
  ## Installation
50
96
 
51
97
  ```bash
@@ -127,11 +173,30 @@ zames --resend-prompt
127
173
  zames --dir <path>
128
174
  zames --headless
129
175
  zames --headed
176
+ zames --plan
177
+ zames --output-format jsonl
130
178
  zames --debug
131
179
  zames --version
132
180
  zames --help
133
181
  ```
134
182
 
183
+ ### Machine-readable output (scripts / CI)
184
+
185
+ For one-shot runs (`--task`) the agent can emit its event stream as JSON lines
186
+ on stdout, keeping all human text on stderr:
187
+
188
+ ```bash
189
+ zames --task "summarize the diff" --output-format jsonl
190
+ ```
191
+
192
+ Each line is one event (`tool_call`, `tool_result`, `assistant_final`, ...), so
193
+ it pipes straight into `jq`:
194
+
195
+ ```bash
196
+ zames --task "..." --output-format jsonl \
197
+ | jq -r 'select(.type=="assistant_final") | .message'
198
+ ```
199
+
135
200
  ## Tools
136
201
 
137
202
  The agent has the same style of tools as Claude Code / Codex CLI:
@@ -142,7 +207,8 @@ The agent has the same style of tools as Claude Code / Codex CLI:
142
207
  - **Extra tools** — `LS` (list a directory), `MultiEdit` (several edits to one
143
208
  file applied atomically), `TodoWrite` (session task checklist), `ApplyPatch`
144
209
  (multi-file patch in Codex's V4A format: `*** Begin Patch` … `*** End Patch`).
145
- - **Git** — `GitStatus`, `GitDiff`, `GitLog`, `GitAdd`, `GitCommit`, `GitPush`.
210
+ - **Git** — `GitStatus`, `GitDiff`, `GitLog`, `GitShow`, `GitBranchList`,
211
+ `GitAdd`, `GitCommit`, `GitPush`.
146
212
  - **Web** — `WebFetch`, `WebSearch`.
147
213
  - **Service** — `respond` (final answer to the operator, ends the task).
148
214
 
@@ -158,6 +224,13 @@ config commands (/new, /chats, /resume, /cd, /status, /config,
158
224
  Codex CLI:
159
225
 
160
226
  - /diff [--staged] — show the working-tree git diff (--staged for the index).
227
+ - /diffstat — a one-line change summary (`git diff --stat`).
228
+ - /context — show what is loaded into the prompt (AGENTS.md, MEMORY.md,
229
+ skills, custom commands) and the system-prompt size.
230
+ - /retry — resend the last task into the same chat (handy after a truncated
231
+ or empty answer).
232
+ - /rename <title> — set the current session title (shown in /sessions).
233
+ - /copy — copy the last assistant answer to the OS clipboard.
161
234
  - /cost (alias /usage) — session stats: tasks, tool calls, duration, and the
162
235
  context size in tokens (DeepSeek's `accumulated_token_usage`).
163
236
  - /export [file] — write the session transcript to a Markdown file
@@ -174,6 +247,14 @@ Codex CLI:
174
247
  instead of the whole list.
175
248
  - /review [focus] [--staged] — ask the agent to review uncommitted changes
176
249
  and report findings (no code changes).
250
+ - /improve [id] — self-improvement loop: take the next open item from
251
+ `BACKLOG.md` (or a specific id, e.g. `/improve B3`), implement it, run the
252
+ typecheck/lint/tests, mark it done and add a CHANGELOG entry. Nothing is
253
+ committed — the changes stay in the working tree for review.
254
+ - /plan [on|off] — plan (read-only) mode. While it is on, the mutating tools
255
+ (Write/Edit/MultiEdit/ApplyPatch/Bash, GitAdd/GitCommit/GitPush) are removed
256
+ from the tool set, so the agent can investigate without touching the tree.
257
+ Start in it with `--plan`.
177
258
  - /compact — ask DeepSeek to compress the current chat into a handover
178
259
  summary, then open a NEW chat, resend the system prompt and post the summary
179
260
  as the carried-over context. Use it when the context gets long.
@@ -305,6 +386,34 @@ locale lives in `ui.locale` in the config file.
305
386
 
306
387
  Agent data is stored in `~/.zames`: browser profile, logs, undo history, self-review snapshots.
307
388
 
389
+ ## FAQ
390
+
391
+ **Is this an official DeepSeek product?**
392
+ No. zames drives the public chat.deepseek.com web UI through a real browser,
393
+ like a regular user. It is not affiliated with DeepSeek.
394
+
395
+ **Do I need an API key?**
396
+ No. You sign in with your own DeepSeek account once; the session is stored in
397
+ `~/.zames/profile` and reused.
398
+
399
+ **Which model does it use?**
400
+ Whichever the DeepSeek web chat uses (DeepSeek-V3, or the reasoning model with
401
+ "Deep thinking" on). zames never calls the API directly.
402
+
403
+ **Does it work headless / on a server?**
404
+ Yes — headless is the default, and `--output-format jsonl` makes it scriptable.
405
+ `--headed` is only needed for a manual sign-in or selector debugging.
406
+
407
+ **Is automating the web UI allowed?**
408
+ That depends on DeepSeek's terms; automating a website may violate them. Use at
409
+ your own risk and keep the send throttle (15s by default) so you do not hammer
410
+ the rate limit.
411
+
412
+ **How is it different from Claude Code / Codex?**
413
+ Same shape (tools, `AGENTS.md`, skills, MCP, slash commands) but it runs on your
414
+ DeepSeek account instead of an API, as a browser automation rather than a
415
+ first-party API client.
416
+
308
417
  ## License
309
418
 
310
419
  MIT
@@ -146,6 +146,12 @@ export async function runAgentLoop({ browser, tools, task, workdir, maxIteration
146
146
  'Look at the uploaded images in the chat; file copies are in <project>/tmp.)';
147
147
  }
148
148
  transcript?.log('task', { task });
149
+ // Each guard keeps its OWN cap IN ADDITION to the shared budget below.
150
+ // Reason: the shared budget is replenished after every successful tool, so
151
+ // over a long chain of tools a single misbehaving guard could re-ask more
152
+ // than intended; the per-guard cap keeps any one failure mode bounded. The
153
+ // named counters are all reset together after a tool runs (see the reset
154
+ // block after the tool loop).
149
155
  let malformedRetries = 0;
150
156
  const MAX_MALFORMED_RETRIES = 3;
151
157
  let stallRetries = 0;
@@ -190,6 +196,16 @@ export async function runAgentLoop({ browser, tools, task, workdir, maxIteration
190
196
  // Retry budget for a browser.ask() TIMEOUT (not for content): the send did
191
197
  // not come back in time. This is separate from unparsedRetries because a
192
198
  // timeout is an infrastructure failure, not a model protocol violation.
199
+ // Loop detection: the model sometimes repeats the SAME tool call with the
200
+ // SAME arguments (e.g. Read the same file) turn after turn. There was no
201
+ // explicit detector for that — the stale-answer guard only catches a
202
+ // repeated ANSWER, not a repeated call. Track the signature of the last
203
+ // batch of calls; after MAX_REPEAT_CALLS identical batches in a row, stop
204
+ // running it and nudge the model to change approach.
205
+ let lastCallSig = '';
206
+ let repeatCallCount = 0;
207
+ const MAX_REPEAT_CALLS = 4;
208
+ let repeatWarned = false;
193
209
  let afterToolRetries = 0;
194
210
  // toolsRanInTask: how many tools actually ran in THIS task. The protocol
195
211
  // guard uses it to tell "started, then slipped into chat mode" (a reasoning
@@ -197,6 +213,10 @@ export async function runAgentLoop({ browser, tools, task, workdir, maxIteration
197
213
  // short answer. We key on the STRUCTURE (work already started), not words.
198
214
  let toolsRanInTask = 0;
199
215
  const MAX_AFTER_TOOL_RETRIES = Math.max(0, Math.floor(maxAfterToolRetries));
216
+ // One-time hint when the context nears full while auto-compact is OFF:
217
+ // the operator is told about /compact instead of silently degrading. Set
218
+ // once per task so the terminal is not spammed.
219
+ let contextHintShown = false;
200
220
  // Token count at which the last auto-compact fired. Re-arm only after
201
221
  // the (fresh) chat grows past this plus a margin, so a chat that starts
202
222
  // above the threshold does not compact on every tool call.
@@ -639,6 +659,38 @@ export async function runAgentLoop({ browser, tools, task, workdir, maxIteration
639
659
  // message must NOT be delivered before the tools' results are known.
640
660
  // The model will get the tool results and can call respond again.
641
661
  const callsToRun = realCalls.length > 0 ? realCalls : calls;
662
+ // Stable signature of this batch (sorted keys) for loop detection.
663
+ const callSig = callsToRun
664
+ .map((c) => c.tool +
665
+ ':' +
666
+ JSON.stringify(c.args, Object.keys(c.args || {}).sort()))
667
+ .join('|');
668
+ if (callSig && callSig === lastCallSig)
669
+ repeatCallCount++;
670
+ else {
671
+ lastCallSig = callSig;
672
+ repeatCallCount = 1;
673
+ repeatWarned = false;
674
+ }
675
+ if (repeatCallCount >= MAX_REPEAT_CALLS) {
676
+ transcript?.log('tool_loop_detected', {
677
+ sig: callSig.slice(0, 300),
678
+ count: repeatCallCount,
679
+ });
680
+ if (!repeatWarned) {
681
+ repeatWarned = true;
682
+ safeWarning(translate(locale)('msg.tool_loop'));
683
+ }
684
+ // Reset so a single nudge does not carry into the next different call.
685
+ repeatCallCount = 0;
686
+ lastCallSig = '';
687
+ message =
688
+ 'You have called the SAME tool with the SAME arguments several times ' +
689
+ 'in a row. Stop repeating it. Either try a different tool or approach, ' +
690
+ 'or, if the task is done, call respond with the final message. Do NOT ' +
691
+ 'repeat the identical call.';
692
+ continue;
693
+ }
642
694
  for (const call of callsToRun) {
643
695
  const tool = tools.find((t) => t.name === call.tool);
644
696
  if (!tool) {
@@ -784,6 +836,21 @@ export async function runAgentLoop({ browser, tools, task, workdir, maxIteration
784
836
  }
785
837
  }
786
838
  }
839
+ else if (getTokenUsage && !contextHintShown) {
840
+ // Auto-compact is OFF (the default). The context can still fill up and
841
+ // silently degrade the answers, so tell the operator ONCE that /compact
842
+ // exists. Same 80% threshold the status line uses to turn red.
843
+ const tokens = getTokenUsage();
844
+ if (tokens !== null && Number.isFinite(tokens) && tokens > 0) {
845
+ const pct = (tokens / contextLimit) * 100;
846
+ if (pct >= 80) {
847
+ contextHintShown = true;
848
+ safeWarning(translate(locale)('context.near_full', {
849
+ pct: String(Math.round(pct)),
850
+ }));
851
+ }
852
+ }
853
+ }
787
854
  }
788
855
  return 'Iteration limit reached.';
789
856
  }
@@ -367,6 +367,38 @@ function tryCommand(bin, args) {
367
367
  return null;
368
368
  }
369
369
  }
370
+ // Write text to the OS clipboard. Mirrors readClipboardImageDetailed: every
371
+ // platform has its own tool, and a failure is returned (not thrown) so the
372
+ // caller can tell the operator which tool was missing.
373
+ export function writeClipboardText(text) {
374
+ const data = Buffer.from(String(text ?? ''), 'utf-8');
375
+ const win = process.platform === 'win32';
376
+ const mac = process.platform === 'darwin';
377
+ const candidates = win
378
+ ? [{ bin: 'clip', args: [] }]
379
+ : mac
380
+ ? [{ bin: 'pbcopy', args: [] }]
381
+ : [
382
+ { bin: 'wl-copy', args: [] },
383
+ { bin: 'xclip', args: ['-selection', 'clipboard'] },
384
+ { bin: 'xsel', args: ['--clipboard', '--input'] },
385
+ ];
386
+ for (const c of candidates) {
387
+ try {
388
+ execFileSync(c.bin, c.args, {
389
+ input: data,
390
+ stdio: ['pipe', 'ignore', 'ignore'],
391
+ timeout: 8000,
392
+ windowsHide: true,
393
+ });
394
+ return { ok: true, via: c.bin };
395
+ }
396
+ catch {
397
+ // try the next tool
398
+ }
399
+ }
400
+ return { ok: false, via: candidates.map((c) => c.bin).join('/') };
401
+ }
370
402
  export async function findFileByName(workdir, name) {
371
403
  const candidate = path.resolve(workdir, name);
372
404
  if (await fs.stat(candidate).catch(() => null))
package/dist/commands.js CHANGED
@@ -25,6 +25,58 @@ export function formatDiff(diffText, opts = {}, t = translate('en')) {
25
25
  export function diffGitArgs(staged = false) {
26
26
  return staged ? 'git diff --staged' : 'git diff';
27
27
  }
28
+ // ---------- /diffstat ----------
29
+ /** Render `git diff --stat` output (a compact change summary) for display. */
30
+ export function formatDiffStat(statText, opts = {}, t = translate('en')) {
31
+ const maxLines = opts.maxLines ?? 60;
32
+ const trimmed = String(statText ?? '')
33
+ .split(CR + NL)
34
+ .join(NL)
35
+ .trimEnd();
36
+ if (!trimmed || trimmed === '(command produced no output)') {
37
+ return t('diff.no_changes');
38
+ }
39
+ const lines = trimmed.split(NL);
40
+ if (lines.length <= maxLines)
41
+ return trimmed;
42
+ const head = lines.slice(0, maxLines).join(NL);
43
+ return (head + NL + t('diff.more_lines', { n: String(lines.length - maxLines) }));
44
+ }
45
+ // ---------- /context ----------
46
+ /** One renderable line per loaded context source (file + char count). */
47
+ export function formatContextSources(agents, memory, skills, commands, opts = {}, t = translate('en')) {
48
+ const lines = [];
49
+ lines.push(t('context.title'));
50
+ if (typeof opts.systemPromptChars === 'number') {
51
+ lines.push(t('context.system_prompt', { n: String(opts.systemPromptChars) }));
52
+ }
53
+ const fileLine = (p, content) => ' ' + p + ' (' + content.length + ' chars)';
54
+ lines.push(t('context.agents'));
55
+ if (agents.length)
56
+ for (const a of agents)
57
+ lines.push(fileLine(a.path, a.content));
58
+ else
59
+ lines.push(' ' + t('context.none'));
60
+ lines.push(t('context.memory'));
61
+ if (memory.length)
62
+ for (const m of memory)
63
+ lines.push(fileLine(m.path, m.content));
64
+ else
65
+ lines.push(' ' + t('context.none'));
66
+ lines.push(t('context.skills'));
67
+ if (skills.length)
68
+ for (const s of skills)
69
+ lines.push(' ' + s.name + (s.path ? ' ' + s.path : ''));
70
+ else
71
+ lines.push(' ' + t('context.none'));
72
+ lines.push(t('context.commands'));
73
+ if (commands.length)
74
+ for (const c of commands)
75
+ lines.push(' /' + c.name);
76
+ else
77
+ lines.push(' ' + t('context.none'));
78
+ return lines.join(NL);
79
+ }
28
80
  export function summarizeTranscript(entries) {
29
81
  const stats = {
30
82
  turns: 0,
@@ -655,6 +707,79 @@ export function mergeMessages(msgs) {
655
707
  const text = header + NL + NL + parts.join(NL + NL + '---' + NL + NL);
656
708
  return { text, attachments };
657
709
  }
710
+ // Parse BACKLOG.md into actionable items. Recognizes headings of the form
711
+ // `### A1. ...`, `### D7. ...` and marks an item done when its heading carries
712
+ // a `[x]` marker. Pure; unit-tested. We deliberately do NOT parse the whole
713
+ // markdown structure — only the headings the file contract promises.
714
+ export function parseBacklogItems(text) {
715
+ const out = [];
716
+ let priority = '';
717
+ for (const raw of String(text ?? '').split(NL)) {
718
+ const line = raw.trimEnd();
719
+ const p = /^##\s+(P[0-3])\b/.exec(line);
720
+ if (p) {
721
+ priority = p[1];
722
+ continue;
723
+ }
724
+ const m = /^###\s+([A-Z]\d+)\.\s+(.*)$/.exec(line);
725
+ if (!m)
726
+ continue;
727
+ const id = m[1];
728
+ const rest = m[2];
729
+ // The `[x]` rewrite in BACKLOG keeps the ORIGINAL heading under a
730
+ // <details> block wrapped in `~~`. That archived copy must not be counted
731
+ // as a second, open item — skip a heading whose text starts with `~~`.
732
+ if (rest.trim().startsWith('~~'))
733
+ continue;
734
+ const done = /^\[x\]/i.test(rest.trim());
735
+ // The title is the heading text with the [x] marker and the `~~` strike
736
+ // wrappers removed, so a done item still has a clean label.
737
+ const title = rest
738
+ .replace(/^\[x\]\s*/i, '')
739
+ .replace(/~~/g, '')
740
+ .trim();
741
+ out.push({ id, title, status: done ? 'done' : 'open', priority });
742
+ }
743
+ return out;
744
+ }
745
+ // The next item to work on: the highest-priority OPEN item, in file order.
746
+ // P0 first, then P1..P3. An explicit id overrides the priority search.
747
+ export function nextBacklogItem(items, wanted) {
748
+ const open = items.filter((i) => i.status === 'open');
749
+ if (wanted) {
750
+ const w = wanted.trim().toUpperCase();
751
+ return items.find((i) => i.id.toUpperCase() === w) || null;
752
+ }
753
+ if (!open.length)
754
+ return null;
755
+ const rank = (p) => {
756
+ const m = /P(\d)/.exec(p || '');
757
+ return m ? Number(m[1]) : 9;
758
+ };
759
+ return open.slice().sort((a, b) => rank(a.priority) - rank(b.priority))[0];
760
+ }
761
+ // The task text sent to the agent for a chosen item. Agent-facing English,
762
+ // like the rest of the loop. Keeps the working rules explicit so a fresh agent
763
+ // does not need any external context.
764
+ export function buildImprovePrompt(item) {
765
+ return ('Implement BACKLOG item ' +
766
+ item.id +
767
+ ' (priority ' +
768
+ (item.priority || '?') +
769
+ '): ' +
770
+ item.title +
771
+ NL +
772
+ NL +
773
+ 'Read the full item text in BACKLOG.md for the details and the rationale. ' +
774
+ 'Then: (1) implement it with the smallest reasonable change; ' +
775
+ '(2) run `npm run typecheck`, `npm run lint`, `npm run format:check` and ' +
776
+ '`npm test` and fix any failure; (3) mark the item done in BACKLOG.md by ' +
777
+ 'prefixing its heading with `[x]` and wrapping the original heading text ' +
778
+ 'in `~~ ~~`; (4) add a CHANGELOG entry if the change is user-visible; ' +
779
+ '(5) do NOT commit or push — leave the changes staged in the working tree ' +
780
+ 'for the operator to review. When done, call respond with a short report ' +
781
+ '(files changed, tests run, whether BACKLOG/CHANGELOG were updated).');
782
+ }
658
783
  // ---------- /review ----------
659
784
  export function buildReviewPrompt(focus, hasStaged = false) {
660
785
  const scope = hasStaged ? 'staged' : 'uncommitted';