wrencode 0.3.0__tar.gz → 0.3.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,802 @@
1
+ Metadata-Version: 2.5
2
+ Name: wrencode
3
+ Version: 0.3.2
4
+ Summary: A minimal agent harness for coding: a readable agent loop and the modules it calls
5
+ Project-URL: Homepage, https://github.com/almostly/wrencode
6
+ Project-URL: Repository, https://github.com/almostly/wrencode
7
+ Project-URL: Issues, https://github.com/almostly/wrencode/issues
8
+ Author: Almostly
9
+ License-Expression: MIT
10
+ Keywords: agent,ai,anthropic,cli,coding-assistant,llm,mlx,ollama,openai
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Environment :: Console
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Topic :: Software Development :: Code Generators
16
+ Requires-Python: >=3.9
17
+ Provides-Extra: agent-sdk
18
+ Requires-Dist: claude-agent-sdk; (python_version >= '3.10') and extra == 'agent-sdk'
19
+ Provides-Extra: history
20
+ Requires-Dist: psycopg[binary]; extra == 'history'
21
+ Provides-Extra: mlx
22
+ Requires-Dist: mlx-lm; extra == 'mlx'
23
+ Provides-Extra: sandbox
24
+ Requires-Dist: pydantic-monty; extra == 'sandbox'
25
+ Provides-Extra: transformers
26
+ Requires-Dist: torch; extra == 'transformers'
27
+ Requires-Dist: transformers; extra == 'transformers'
28
+ Description-Content-Type: text/markdown
29
+
30
+ # 🐦 WrenCode
31
+
32
+ A minimal agent harness for coding. The agent loop is one readable Python file.
33
+
34
+ Named after Harold Wren - the alias of a genius who built a superintelligent AI and operated quietly in the background.
35
+
36
+ -----
37
+
38
+ ## What it is
39
+
40
+ WrenCode is a coding agent harness: everything around the model that turns it into an agent. It runs the tool-calling loop, executes tools, builds the system prompt, and manages context, locally or via API, giving an LLM the ability to read, write, and edit files, search codebases, and run shell commands - enough to autonomously navigate and modify a real project.
41
+
42
+ Where Claude Code is the batteries-included harness, WrenCode is the **"understand and own your agent" harness**: the entire agent loop fits in one readable file, runs against local or hosted models, and is yours to hack.
43
+
44
+ ## Code layout
45
+
46
+ The package is `src/wrencode/`. Read `app.py` top to bottom to understand the
47
+ agent; the modules beside it are what it calls. The tests are in `tests/`, one
48
+ file per module.
49
+
50
+ |File |What's in it |
51
+ |------------------|------------------------------------------------------------------|
52
+ |`app.py` |The harness: the tools, the system prompt, the turn loop, parallel subagents, context compaction, headless mode, `main()`|
53
+ |`backends.py` |Talking to models: backend tables and state, HTTP with retries, streaming, request/response formats (Anthropic, OpenAI, Bedrock Converse, local), usage and prices, `get_response()`|
54
+ |`configure.py` |Picking a backend and model: the first-run chooser, `/configure` and `/model`, API-key prompts and verification, model lists, saved config|
55
+ |`ui.py` |The terminal: colors, the line editor with slash-command completion, approvals, the stream printer, Escape-to-cancel, tagged output from parallel subagents|
56
+ |`permissions.py` |Permission rules: allow and deny by tool and pattern, user and project files|
57
+ |`mcp.py` |MCP client: stdio and HTTP servers, their tools as `mcp__server__tool`|
58
+ |`web.py` |The `fetch` tool: a URL as readable text|
59
+ |`sdk.py` |The `claude-agent-sdk` backend |
60
+ |`sandbox.py` |The `python` tool's sandbox, on pydantic-monty |
61
+ |`history.py` |Conversation history in Postgres: a server or embedded PGlite |
62
+ |`synthesize.py` |The `synthesize` subcommand |
63
+ |`__main__.py` |`python -m wrencode`, and what the release binary is built from |
64
+
65
+ Each module imports only the ones below it in this table's dependency order (`app` → backends/configure/sdk/synthesize → ui → permissions), so the loop can be read without the rest.
66
+
67
+ ## Backends
68
+
69
+ On first run WrenCode asks you to pick a backend and saves the choice to
70
+ `~/.wrencode/config.json`. Run `wrencode --configure` any time to change it.
71
+ Set `BACKEND` (and the matching API key) in the environment to override the
72
+ saved choice, e.g. for CI.
73
+
74
+ |Backend |Description |Availability |
75
+ |--------------|----------------------------------------|----------------------|
76
+ |`anthropic` |Claude via Anthropic API |binary + source |
77
+ |`claude-agent-sdk`|Claude Code's agent loop and tools via the Claude Agent SDK|source install, Python 3.10+|
78
+ |`openai` |GPT models via OpenAI API |binary + source |
79
+ |`openrouter` |Any model via OpenRouter |binary + source |
80
+ |`nanogpt` |Any model via NanoGPT |binary + source |
81
+ |`ollama` |Local models via a running `ollama serve`|binary + source |
82
+ |`openai-compatible`|vLLM, llama.cpp, Hugging Face, any OpenAI-compatible server|binary + source|
83
+ |`bedrock` |Any model via AWS Bedrock Converse (AWS credentials)|binary + source|
84
+ |`local` |Local proxy via Anthropic-compatible API|binary + source |
85
+ |`transformers`|HuggingFace Transformers (CPU/MPS/GPU) |source install only |
86
+ |`mlx` |Apple Silicon via MLX |source install, macOS |
87
+
88
+ The standalone binary can't bundle the heavy ML stack, so the local-weights
89
+ backends (`mlx`, `transformers`) are only offered when running from source.
90
+
91
+ The default local models are
92
+ [`deburky/gpt-oss-claude-code`](https://huggingface.co/deburky/gpt-oss-claude-code)
93
+ (transformers) and
94
+ [`deburky/gpt-oss-claude-mlx`](https://huggingface.co/deburky/gpt-oss-claude-mlx)
95
+ (MLX) — override either with `MODEL=...`.
96
+
97
+ ### Claude Agent SDK
98
+
99
+ The `claude-agent-sdk` backend hands each prompt to the
100
+ [Claude Agent SDK](https://code.claude.com/docs/en/agent-sdk/overview), which
101
+ runs Claude Code's own agent loop, tools and subagents. WrenCode shows the
102
+ stream, asks before edits and commands, and prints the cost of each turn.
103
+ Install the extra, then pick the backend in `/configure`:
104
+
105
+ ```bash
106
+ pip install 'wrencode[agent-sdk]' # or: pip install claude-agent-sdk
107
+ ```
108
+
109
+ It always authenticates with `ANTHROPIC_API_KEY`, so usage bills to your
110
+ Console credits, including the monthly API credits that come with Max and Team
111
+ plans. Subscription logins are never used. Multi-workspace keys need
112
+ `ANTHROPIC_WORKSPACE_ID`. The conversation resumes per project across restarts;
113
+ `/clear` starts a new one. Your `~/.claude` hooks, plugins and MCP servers are
114
+ not loaded; project instructions come from `AGENTS.md` / `CLAUDE.md`.
115
+
116
+ To run several prompts in parallel, each as its own agent with a fresh
117
+ context, see `examples/agent_sdk_swarm.py`. It prints every answer and the
118
+ total cost.
119
+
120
+ ### OpenAI-compatible servers
121
+
122
+ `openai-compatible` talks to any server that implements OpenAI chat completions,
123
+ using native tool calls. Point it at the server with `OPENAI_COMPATIBLE_BASE_URL`
124
+ (default `http://localhost:8000/v1`). If the server serves exactly one model,
125
+ WrenCode uses it; otherwise set `MODEL`.
126
+
127
+ ```bash
128
+ # vLLM (tool calling needs these flags; pick the parser for your model)
129
+ vllm serve Qwen/Qwen2.5-Coder-7B-Instruct --enable-auto-tool-choice --tool-call-parser hermes
130
+ BACKEND=openai-compatible wrencode
131
+
132
+ # llama.cpp (--jinja enables tool calling)
133
+ llama-server -m qwen2.5-coder-7b-instruct-q4_k_m.gguf --jinja --port 8080
134
+ BACKEND=openai-compatible OPENAI_COMPATIBLE_BASE_URL=http://localhost:8080/v1 wrencode
135
+
136
+ # Hugging Face Inference Providers
137
+ BACKEND=openai-compatible OPENAI_COMPATIBLE_BASE_URL=https://router.huggingface.co/v1 \
138
+ OPENAI_COMPATIBLE_API_KEY=$HF_TOKEN MODEL=Qwen/Qwen2.5-Coder-32B-Instruct wrencode
139
+ ```
140
+
141
+ ## Tools
142
+
143
+ The agent has access to seven tools:
144
+
145
+ - **read** - read a file with line numbers, or list a directory
146
+ - **write** - write content to a file
147
+ - **edit** - replace a unique string in a file. If the text only matches with
148
+ its indentation shifted by a consistent amount (a common slip when quoting a
149
+ method), the edit is applied with the replacement shifted to match; otherwise
150
+ the error shows the closest lines in the file
151
+ - **glob** - find files by pattern, sorted by modification time
152
+ - **grep** - search files for a regex pattern using `rg` when available, falling back to `grep`
153
+ - **bash** - run a shell command with timeout and streaming output
154
+ - **fetch** - read a web page or URL as text, with approval (see Web access)
155
+ - **task** - delegate a self-contained subtask to a fresh subagent (its own context, same tools) that returns only its final result
156
+ - **python** - run a snippet of Python in a sandbox, with `pydantic-monty` installed (see below)
157
+
158
+ All file operations are sandboxed to the workspace root by default.
159
+
160
+ ### Subagents
161
+
162
+ The `task` tool runs a nested agent loop on a fresh message history, so the
163
+ parent's context only grows by the returned summary — useful for context-heavy
164
+ subtasks.
165
+
166
+ Task calls made in the same reply run in parallel, up to
167
+ `WRENCODE_MAX_PARALLEL_SUBAGENTS` at a time (default 4), on every backend
168
+ except the in-process `mlx` and `transformers` ones. Ask for it in the prompt,
169
+ for example "analyze each file in docs/ with its own subagent, in parallel".
170
+ Each subagent's output is tagged `[1]`, `[2]`, and so on; approval prompts
171
+ take turns and pause the other agents' output until you answer. Escape stops
172
+ the whole batch. Recursion is capped by `WRENCODE_MAX_SUBAGENT_DEPTH` (default 2), and
173
+ each subagent round is bounded. For autonomous subagent runs, enable
174
+ `--yes` / `WRENCODE_AUTO_APPROVE` so sub-tool calls don't block on confirmation.
175
+
176
+ ### Python sandbox
177
+
178
+ With [pydantic-monty](https://github.com/pydantic/monty) installed, the agent gets
179
+ an eighth tool, `python`. It runs a snippet of model-written code in Monty, a
180
+ Python interpreter built as a sandbox: each call is a fresh interpreter with no
181
+ network, no shell, no environment variables, and a read-only view of the
182
+ workspace at `/workspace`, which is also the working directory, so
183
+ `open("foo.py")` works. wrencode's `read(path)`, `glob(pat)` and `grep(pat)` are
184
+ callable inside the snippet. Printed output and the value of a trailing
185
+ expression come back to the model; an exception comes back as its traceback.
186
+ Since the snippet can't change anything, it runs without an approval prompt;
187
+ changes still go through `write` and `edit`.
188
+
189
+ ```bash
190
+ pip install 'wrencode[sandbox]' # or: pip install pydantic-monty
191
+ ```
192
+
193
+ Monty runs a subset of Python: no class inheritance, generators or third-party
194
+ packages, and a curated standard library (`json`, `re`, `math`, `datetime`,
195
+ `pathlib`, ...). `WRENCODE_SANDBOX_TIMEOUT` (default 30 seconds) and
196
+ `WRENCODE_SANDBOX_MEMORY_MB` (default 256) bound each run. The standalone binary
197
+ doesn't bundle Monty, so the tool is a source-install feature.
198
+
199
+ ## Project instructions (AGENTS.md)
200
+
201
+ WrenCode reads [`AGENTS.md`](https://agents.md) files and adds them to the
202
+ system prompt, so conventions you've written for other agents apply here too.
203
+ It looks in `~/.wrencode/`, then in every directory from the git root down to
204
+ the workspace (outside a git repo, only the workspace). A directory without an
205
+ `AGENTS.md` falls back to `CLAUDE.md`. Files closer to the workspace come later
206
+ and take precedence. The total is capped at 32,000 characters, and the files
207
+ loaded are listed at startup.
208
+
209
+ ## Headless mode
210
+
211
+ `-p` / `--print` runs a single prompt without the interactive UI, for scripts,
212
+ CI, and evals:
213
+
214
+ ```bash
215
+ wrencode -p "Why is test_parse failing?"
216
+ git diff | wrencode -p "Review this diff" # prompt from stdin
217
+ wrencode --yes -p "Fix the lint errors" --max-turns 20
218
+ wrencode -p "List the TODOs" --output-format json | jq -r .result
219
+ wrencode --yes -p "Make the tests pass" --verify "python3 -m unittest -q"
220
+ ```
221
+
222
+ - stdout carries only the final answer (or one JSON object with
223
+ `--output-format json`: `result`, `is_error`, `stop_reason`, `num_turns`,
224
+ `backend`, `model`); progress and tool output go to stderr.
225
+ - Each run starts from a fresh history and doesn't touch the saved one.
226
+ - Without `--yes`, writes and shell commands are declined (the model is told
227
+ why) instead of waiting for approval. Read-only tools always work.
228
+ - The exit code is `0` when the agent finishes, `1` if it errors, hits
229
+ `--max-turns`, or stops on repeated tool errors, and `2` for bad arguments.
230
+ - `--verify CMD` checks the agent's claim of being done: WrenCode runs `CMD`
231
+ in the workspace when the agent finishes, and if it fails, sends the output
232
+ back and lets the agent continue (up to 3 attempts in all). The result says
233
+ `verified: true/false`, and a final failure exits `1` with
234
+ `stop_reason: "verify_failed"`. `--max-turns` applies to each attempt.
235
+
236
+ ### Structured output
237
+
238
+ `--json-schema` makes the answer a JSON value that matches a schema, given as
239
+ a file or inline:
240
+
241
+ ```bash
242
+ wrencode -p "Review this repo for bugs" --json-schema bugs.schema.json
243
+ wrencode -p "Is the build green?" --json-schema '{"type": "object", "properties": {"green": {"type": "boolean"}}, "required": ["green"]}'
244
+ ```
245
+
246
+ The agent gets a `respond` tool whose arguments are your schema, and the run
247
+ ends when it calls `respond` with a valid answer. If the answer doesn't match,
248
+ the validation errors go back to the model so it can fix them; if it never
249
+ calls `respond`, the run fails with `stop_reason: "no_structured_output"`.
250
+ stdout is the JSON value (or, with `--output-format json`, the usual object
251
+ with a `structured_output` field). Validation is built in and covers the
252
+ common keywords: `type`, `enum`, `const`, `properties`, `required`,
253
+ `additionalProperties`, `items`, length and numeric bounds, `pattern`, and
254
+ `anyOf`/`oneOf`/`allOf`.
255
+
256
+ ## Synthesize
257
+
258
+ `wrencode synthesize` fuses several agent chat transcripts into one document: a
259
+ semantic git-merge for conversations. Each transcript is normalized to user and
260
+ assistant turns, the model extracts its decisions, problems solved, files touched
261
+ and open questions, and those fact sets are then reconciled across chats. Every
262
+ claim cites the chat it came from: the first eight characters of the file name,
263
+ or more when two names would clash, so ids are unique within a run.
264
+
265
+ ```bash
266
+ wrencode synthesize # pick from this project's Claude Code history
267
+ wrencode synthesize a.jsonl b.jsonl # fuse these transcripts
268
+ wrencode synthesize ~/.claude/projects/-home-me-app --all # a whole directory, no picker
269
+ wrencode synthesize diff a.jsonl b.jsonl # only where the chats diverge
270
+ wrencode synthesize log --out DECISIONS.md # a decision timeline, written to a file
271
+ ```
272
+
273
+ Three modes, chosen by the first word after `synthesize`:
274
+
275
+ - **merge** (default) — `Reinforced decisions` that two or more chats agree on,
276
+ `Unique contributions`, `⚠ Conflicts` where chats contradict each other (a
277
+ later chat that overrode an earlier one is marked resolved), and
278
+ `Open questions`.
279
+ - **diff** — only the divergences: conflicts and what appears in just one chat,
280
+ like `git diff`.
281
+ - **log** — one chronological timeline of decisions, oldest first, noting where
282
+ a later chat supersedes an earlier one, and ending with the net state.
283
+
284
+ Inputs can be files, directories (every `*.jsonl` inside, newest first), or
285
+ nothing, which lists this project's Claude Code transcripts from
286
+ `~/.claude/projects/`. With a directory or no paths, an interactive picker lets
287
+ you choose: `↑`/`↓` move, space toggles, `a` selects all, Enter confirms, Esc
288
+ cancels. `--all` skips the picker. Transcripts are recognized in three formats:
289
+ Claude Code JSONL (tool calls, tool results and thinking are dropped), generic
290
+ JSONL message logs (Codex/OpenAI style), and a JSON list of messages or a
291
+ `{"messages": [...]}` object. Anything else is read as one block of text.
292
+
293
+ The synthesis uses the configured backend at temperature 0; local models are
294
+ loaded on demand. A transcript longer than the model's context window
295
+ (`WRENCODE_CONTEXT_TOKENS`) is sent with its start and, mostly, its end, since the
296
+ latest decisions override earlier ones. A chat whose extraction doesn't come back
297
+ as JSON is reported and contributes nothing. With a single transcript the result
298
+ degrades to a structured summary.
299
+
300
+ ## Token usage and spend
301
+
302
+ After each turn wrencode prints one dim line with what the backend reported:
303
+
304
+ ```
305
+ ↑ 1.6k ↓ 108 42 tok/s ↗ ▰▱▱▱▱▱▱▱▱▱ 1% $0.0042 · total $0.21
306
+ ```
307
+
308
+ Up is the turn's input tokens, down its output, then the speed: output tokens
309
+ per second over the turn's requests, with an arrow against the previous turn
310
+ (`↗` at least a tenth faster, `↘` a tenth slower, `→` about the same). The rate is measured over the whole request, so the network and the wait
311
+ for the first token are in it; a drop usually means the provider is busy. The
312
+ meter is the context fill: the size of the latest request against the model's
313
+ window, which is what the next turn starts from and what auto-compaction watches.
314
+ It turns yellow at the compaction threshold. Then the turn's cost and, after the
315
+ first turn, the session total. The terminal tab title shows the context fill, the
316
+ session totals, the spend and the speed. `/usage` prints the full numbers:
317
+
318
+ ```
319
+ input cached written output calls cost
320
+ this turn 1.6k 1.6k 54 108 1 $0.0042
321
+ session 4.7k 3.9k 1.6k 236 3 $0.213
322
+ context: 1.6k of 128k (1%); input = uncached + cached (read) + written
323
+ price: $2/$10 per MTok (cache read $0.20, write $2.50; built-in)
324
+ speed: 42 tok/s this turn (↗ from 36 tok/s last turn); 39 tok/s this session; output tokens over the request's wall time
325
+ ```
326
+
327
+ While you type, a dim line under the prompt estimates what sending the message
328
+ will cost in input tokens: the context the model already holds (at the cache-read
329
+ rate for the part it served from the cache last time), its last reply, and your
330
+ text at about four characters per token. Output can't be known ahead, so it isn't
331
+ counted.
332
+
333
+ ```
334
+ ❯ fix the bug in the parser
335
+ ≈ $0.0031 input
336
+ ```
337
+
338
+ **Where the prices come from.** Each call is priced at the four rates in effect
339
+ for the model: input, output, cache read and cache write, in USD per million
340
+ tokens. Anthropic's Models API lists models but not prices, so Claude models (and
341
+ the OpenAI, Amazon Nova and Bedrock ids wrencode knows) come from a built-in
342
+ table, checked October 2026; `/usage` says `built-in` and the startup line shows
343
+ the rate. OpenRouter and NanoGPT list prices in their model catalogs, and
344
+ `/configure` saves those to `~/.wrencode/prices.json` when it fetches the model
345
+ list, keyed `backend/model id`; you can add your own entries there as
346
+ `[input, output, cache_read, cache_write]`. `WRENCODE_PRICE=input,output[,cache_read[,cache_write]]`
347
+ overrides everything for the current model (`WRENCODE_PRICE=0,0` marks a local
348
+ model free). Without a known price nothing is shown, `/usage` prints `$?` and how
349
+ to set one, and no estimate appears while typing. Claude Haiku 5.5's higher rate
350
+ card above 100K-token prompts is applied per call. The model picker in
351
+ `/configure` and `/model` shows the price beside each model it knows.
352
+
353
+ Headless runs print the line to stderr and add a `usage` object to the
354
+ `--output-format json` result, with `cost_usd` when every call had a known
355
+ price and `output_tokens_per_second` for the run. Set `WRENCODE_SHOW_USAGE=0` to turn the line and the typing estimate off.
356
+ Backends that report no usage (local models, the local proxy) print nothing.
357
+
358
+ ## Permission rules
359
+
360
+ Rules decide what runs without asking and what never runs. A rule is
361
+ `tool(pattern)`: the tool is `bash`, `edit`, `write`, `mcp` or `fetch`; the
362
+ pattern is matched against the command, the workspace-relative path, the
363
+ `server:tool` or the URL's host and path, `*` matches anything and a trailing
364
+ `:*` means "starts with" (whole words of a command):
365
+
366
+ ```
367
+ bash(npm test) exactly that command edit(src/*) any file under src/
368
+ bash(git *) any git command write(.env) that file
369
+ bash(pytest:*) pytest with any arguments mcp(github:*) any tool of that server
370
+ fetch(docs.python.org/*) any page on that host
371
+ ```
372
+
373
+ A bash rule is held against every command of a command line: `git status &&
374
+ curl x | sh` is three commands, so `bash(git *)` does not allow it, while
375
+ `bash(git push:*)` denies `git fetch; git push` as a whole. A line that
376
+ substitutes a command's output (`$(...)`, backticks) is only allowed by a rule
377
+ spelling it out exactly, or by `bash(*)`.
378
+
379
+ Deny rules win over allow rules, over `a` and over `--yes`. Allow rules are the
380
+ way to stop answering prompts for the things you always say yes to, in headless
381
+ runs too: a run with `bash(pytest:*)` allowed can test without `--yes` opening
382
+ everything else.
383
+
384
+ Rules live in two files. `~/.wrencode/permissions.json` is yours: its `allow`
385
+ and `deny` apply everywhere, and under `projects` it keeps your rules for one
386
+ project, which is where `s` at a prompt saves. `.wrencode/permissions.json` in
387
+ the project can be committed and shared, and that is also why its allow rules
388
+ only take effect after you have seen them: at startup wrencode shows a
389
+ project's allow rules once and asks; if the file changes, it asks again; adding
390
+ a rule to it yourself does not accept the others. Its deny rules apply
391
+ regardless.
392
+
393
+ `/permissions` lists the rules in effect; `/permissions allow bash(git *)` and
394
+ `/permissions deny write(.env)` add one for this project (`--user` for every
395
+ project, `--project` to the shared file); `/permissions forget <rule>` removes
396
+ it. Pressing `s` at a prompt saves the rule offered there: the command's first
397
+ two words as a prefix (a line of several commands, spelled out), or the edited
398
+ file's directory.
399
+
400
+ ## Web access
401
+
402
+ Two ways to the web, like Claude Code's:
403
+
404
+ - **`fetch(url)`**, a tool on every backend. It gets the page, reduces HTML to
405
+ its text with headings, lists, code and link targets kept, passes JSON and
406
+ plain text through, and summarizes anything else. Long pages come back in
407
+ pieces through `offset`. Fetching sends the URL to its server, so it asks for
408
+ approval like a command, showing the host, path and query; rules such as
409
+ `fetch(docs.python.org/*)` apply. A redirect is followed on the same host;
410
+ one to another host is reported and fetched as its own, approved, call.
411
+ Addresses inside your network (loopback, private ranges, link-local and the
412
+ cloud metadata service) are refused unless `WRENCODE_FETCH_LOCAL=1`.
413
+ `WRENCODE_FETCH_MAX_CHARS` (40,000) and `WRENCODE_FETCH_MAX_BYTES` (4 MB)
414
+ bound a page, compressed or not.
415
+ - **Web search** on the Anthropic backend, through Anthropic's server-side
416
+ search tool. Claude searches and reads results on Anthropic's side; the
417
+ transcript shows each search and its results under it, and each search is
418
+ billed at $10 per 1,000 on top of tokens, counted in the usage line, `/usage`
419
+ and the headless result. `WRENCODE_WEB_SEARCH=0` turns it off. Other backends
420
+ have no search unless an MCP server provides one.
421
+
422
+ ## MCP servers
423
+
424
+ Tools from [Model Context Protocol](https://modelcontextprotocol.io) servers
425
+ join the tool list. Declare them in `.wrencode/mcp.json` in the project (Claude
426
+ Code's `.mcp.json` is read too, same format) or `~/.wrencode/mcp.json`:
427
+
428
+ ```json
429
+ {"mcpServers": {
430
+ "github": {"command": "npx", "args": ["-y", "@modelcontextprotocol/server-github"],
431
+ "env": {"GITHUB_TOKEN": "..."}},
432
+ "docs": {"url": "https://example.com/mcp", "headers": {"Authorization": "Bearer ..."}}
433
+ }}
434
+ ```
435
+
436
+ A `command` entry runs as a subprocess spoken to over stdio; a `url` entry is
437
+ Streamable HTTP. Each server's tools are offered to the model as
438
+ `mcp__<server>__<tool>`; permission rules such as `mcp(github:*)` apply to
439
+ every call, and a call asks for approval unless the server marks the tool
440
+ read-only. `/mcp` lists the servers, their tools and any connection error;
441
+ `/mcp reload` re-reads the files and reconnects.
442
+
443
+ A server's command is looked up on your PATH, it gets your environment without
444
+ the backend API keys and `WRENCODE_*` settings (its own `env` can pass
445
+ anything), and an HTTP server's headers are sent only to the URL configured,
446
+ never across a redirect. A project's servers run on your machine, so, like its
447
+ permission rules, they are shown in full once at startup and start only after
448
+ you accept them; a change to the file asks again, and their `env` cannot set
449
+ loader variables such as `LD_PRELOAD`, `NODE_OPTIONS` or `PYTHONPATH`.
450
+ `WRENCODE_MCP_TIMEOUT` (default 120s) bounds a tool call and
451
+ `WRENCODE_MCP_CONNECT_TIMEOUT` (default 20s) a connection.
452
+
453
+ ## Context management
454
+
455
+ Long sessions are compacted automatically. Before each model call WrenCode
456
+ estimates the prompt size (about 4 characters per token), and once it passes
457
+ `WRENCODE_COMPACT_AT` (default 75%) of `WRENCODE_CONTEXT_TOKENS` (default
458
+ 128,000) it has the model summarize the older messages: the request, files
459
+ touched, commands and results, decisions, and what's left to do. The most recent
460
+ messages, about a quarter of the window, are kept verbatim, along with the
461
+ user's latest request, so it works mid-task, between tool calls. If a request still
462
+ fails with a context-length error, WrenCode compacts and retries once.
463
+
464
+ Set `WRENCODE_CONTEXT_TOKENS` to your model's window, especially for local
465
+ models with small ones. `/compact` summarizes on demand.
466
+
467
+ ## Security
468
+
469
+ wrencode reads untrusted files and runs commands on your machine, so the goal is
470
+ narrower than "can't be attacked": nothing changes outside the sandbox without you
471
+ seeing and approving the real action, and opening an untrusted repository is safe.
472
+
473
+ - **Approvals show what will run.** Every write, edit, shell command, fetch and MCP
474
+ call asks first. Control characters, escape sequences and Unicode direction
475
+ overrides in a command, a file, a URL or a model reply are displayed as `^[`,
476
+ `^M`, `\u202e` and so on, never interpreted, so nothing can redraw the screen or
477
+ hide part of a command. Writes under a hidden path (`.git/hooks`,
478
+ `.github/workflows`, dotfiles) are flagged, after resolving the path. `--yes`
479
+ turns the prompts off; use it only in a sandbox you can throw away. Permission
480
+ rules narrow that: deny rules hold even under `--yes` and for read-only MCP
481
+ tools, a bash rule must cover every command of a command line, and a project's
482
+ allow rules apply only after you accept them, so a repository cannot grant
483
+ itself anything.
484
+ - **The web stays at arm's length.** `fetch` reaches public addresses only, follows
485
+ redirects on the same host only, and shows the full URL for approval. MCP servers
486
+ do not get the backend keys, a project's cannot preload code through the
487
+ environment, and an HTTP server's credentials never follow a redirect.
488
+ - **File tools stay in the workspace.** Paths are resolved (symlinks followed) and
489
+ must land inside the workspace root unless `WRENCODE_UNRESTRICTED_PATHS=1`;
490
+ `glob` drops matches that lead outside. `grep` passes the pattern and path as
491
+ arguments, never as flags.
492
+ - **A project's `.env` can't reconfigure the agent.** It may set `*_API_KEY` and
493
+ `ANTHROPIC_WORKSPACE_ID` only. The backend, any server URL, auto-approve, and the
494
+ config, history and workspace locations come from your shell or the `.env` at the
495
+ root of a source checkout; the names a project `.env` set, and the ones it tried
496
+ to, are reported at startup.
497
+ - **`AGENTS.md` / `CLAUDE.md` are prompt input.** A repository's instructions go into
498
+ the system prompt by design, which means a repository can steer the agent. The
499
+ approval prompts are the control; the files loaded are listed at startup.
500
+ - **Keys go only to their backend.** Fixed hosts for Anthropic, OpenAI, OpenRouter,
501
+ NanoGPT and Bedrock; the URL you configured for openai-compatible, Ollama and the
502
+ local proxy. Saved keys, the conversation history and model caches are owner-only
503
+ files (`0600`) under `~/.wrencode`; the embedded Postgres data and socket
504
+ directories are owner-only (`0700`). The Agent SDK backend runs with the API key
505
+ only, subscription credentials blanked.
506
+ - **The `python` tool is sandboxed** in pydantic-monty: no network, shell or
507
+ environment, a read-only workspace, and time and memory limits.
508
+ - **Releases are verifiable.** Each binary is published with a SHA-256 checksum that
509
+ `install.sh` checks. Hosted backends are reached over TLS with certificate
510
+ verification. wrencode has no runtime dependencies beyond the standard library.
511
+
512
+ Please report security issues privately through the repository's GitHub security
513
+ advisories rather than in a public issue.
514
+
515
+ ## Installation
516
+
517
+ ### Option 1: Standalone binary (recommended)
518
+
519
+ Run the guided installer:
520
+
521
+ ```bash
522
+ curl -fsSL https://raw.githubusercontent.com/almostly/wrencode/main/install.sh | sh
523
+ ```
524
+
525
+ It detects your OS/arch, downloads the matching binary from the latest GitHub
526
+ Release, and installs it to `~/.local/bin` (no sudo). Override the location with
527
+ `WRENCODE_INSTALL_DIR`, or pin a release with `WRENCODE_VERSION`:
528
+
529
+ ```bash
530
+ curl -fsSL https://raw.githubusercontent.com/almostly/wrencode/main/install.sh \
531
+ | WRENCODE_INSTALL_DIR=/usr/local/bin WRENCODE_VERSION=0.1.3 sh
532
+ ```
533
+
534
+ Or download and run it locally:
535
+
536
+ ```bash
537
+ curl -fsSL https://raw.githubusercontent.com/almostly/wrencode/main/install.sh -o install.sh
538
+ chmod +x install.sh
539
+ ./install.sh
540
+ ```
541
+
542
+ Manual install (fallback): download the right binary from GitHub Releases, make it executable, and move it into your `PATH`.
543
+
544
+ Windows: download `wrencode-windows-x64.exe` from GitHub Releases and put it on your `PATH`.
545
+
546
+ macOS Apple Silicon:
547
+
548
+ ```bash
549
+ curl -L https://github.com/almostly/wrencode/releases/latest/download/wrencode-macos-arm64 -o wrencode
550
+ chmod +x wrencode
551
+ sudo mv wrencode /usr/local/bin/wrencode
552
+ ```
553
+
554
+ macOS Intel:
555
+
556
+ ```bash
557
+ curl -L https://github.com/almostly/wrencode/releases/latest/download/wrencode-macos-x64 -o wrencode
558
+ chmod +x wrencode
559
+ sudo mv wrencode /usr/local/bin/wrencode
560
+ ```
561
+
562
+ Linux x64:
563
+
564
+ ```bash
565
+ curl -L https://github.com/almostly/wrencode/releases/latest/download/wrencode-linux-x64 -o wrencode
566
+ chmod +x wrencode
567
+ sudo mv wrencode /usr/local/bin/wrencode
568
+ ```
569
+
570
+ ### Option 2: Run from source
571
+
572
+ Standard library only (except the backend you choose). An editable install puts
573
+ the `wrencode` command on your PATH and pulls in nothing else; `python3 -m
574
+ wrencode` with `PYTHONPATH=src` works without installing.
575
+
576
+ ```bash
577
+ git clone https://github.com/almostly/wrencode
578
+ cd wrencode
579
+ pip install -e .
580
+ ```
581
+
582
+ For MLX (Mac Silicon):
583
+
584
+ ```bash
585
+ pip install mlx-lm
586
+ ```
587
+
588
+ For Anthropic:
589
+
590
+ ```bash
591
+ pip install anthropic # not required - uses urllib directly
592
+ export ANTHROPIC_API_KEY=your_key
593
+ ```
594
+
595
+ For OpenAI:
596
+
597
+ ```bash
598
+ export OPENAI_API_KEY=your_key
599
+ ```
600
+
601
+ For OpenRouter:
602
+
603
+ ```bash
604
+ export OPENROUTER_API_KEY=your_key
605
+ ```
606
+
607
+ For NanoGPT:
608
+
609
+ ```bash
610
+ export NANOGPT_API_KEY=your_key
611
+ ```
612
+
613
+ For HuggingFace Transformers:
614
+
615
+ ```bash
616
+ pip install transformers torch
617
+ ```
618
+
619
+ For the `python` sandbox tool:
620
+
621
+ ```bash
622
+ pip install pydantic-monty
623
+ ```
624
+
625
+ ## Usage
626
+
627
+ ```bash
628
+ # Standalone binary — prompts for a backend on first run
629
+ wrencode
630
+
631
+ # Re-pick the backend at any time
632
+ wrencode --configure
633
+
634
+ # Or from a source checkout (after pip install -e .) — also prompts on first run
635
+ python3 -m wrencode
636
+
637
+ # Anthropic Claude (model list is fetched live from the API during /configure)
638
+ BACKEND=anthropic python3 -m wrencode
639
+ # Multi-workspace Anthropic keys also need a workspace id:
640
+ # ANTHROPIC_WORKSPACE_ID=wrkspc_... BACKEND=anthropic python3 -m wrencode
641
+
642
+ # OpenAI (model list fetched live from the API during /configure)
643
+ BACKEND=openai MODEL=gpt-4o python3 -m wrencode
644
+
645
+ # OpenRouter
646
+ BACKEND=openrouter MODEL=anthropic/claude-3-haiku python3 -m wrencode
647
+
648
+ # NanoGPT
649
+ BACKEND=nanogpt MODEL=z-ai/glm-5.3-flash-uncensored python3 -m wrencode
650
+
651
+ # Ollama (needs `ollama serve` running and the model pulled)
652
+ BACKEND=ollama MODEL=llama3.2 python3 -m wrencode
653
+
654
+ # HuggingFace model
655
+ BACKEND=transformers MODEL=deburky/gpt-oss-claude-code python3 -m wrencode
656
+
657
+ # Local proxy
658
+ BACKEND=local LOCAL_PORT=8082 python3 -m wrencode
659
+ ```
660
+
661
+ ## Developing
662
+
663
+ ```bash
664
+ python3 -m unittest discover -s tests -t . # the test suite (stdlib unittest, no install needed)
665
+ uvx ruff check . && uvx ruff format . # lint and format; the rule set is in pyproject.toml
666
+ uvx ty check src # type check
667
+ ```
668
+
669
+ There are no `# noqa` markers: a rule the design contradicts is turned off in
670
+ `pyproject.toml` with its reason, and a test keeps it that way.
671
+
672
+ ## Releasing
673
+
674
+ Versions and [`CHANGELOG.md`](CHANGELOG.md) are managed with
675
+ [commitizen](https://commitizen-tools.github.io/commitizen/), so write commit
676
+ messages as [conventional commits](https://www.conventionalcommits.org/)
677
+ (`feat: ...`, `fix(edit): ...`, `refactor: ...`). To cut a release:
678
+
679
+ ```bash
680
+ uvx --from commitizen cz bump # bumps WRENCODE_VERSION, updates CHANGELOG.md, tags
681
+ git push origin main --tags
682
+ ```
683
+
684
+ Preview the next changelog entry with `uvx --from commitizen cz changelog --dry-run`.
685
+
686
+ Binaries are built automatically by GitHub Actions when a version tag is pushed,
687
+ and the package is published to PyPI. Without a terminal to push a tag from, run
688
+ the `build-and-release` workflow from the Actions tab on `main` with the version
689
+ as its input: it builds that commit, checks the binaries report that version,
690
+ creates the tag and the release. Then run `publish-pypi` with target `pypi`,
691
+ since a tag created by a workflow does not trigger the others.
692
+
693
+ This publishes release assets:
694
+ - `wrencode-linux-x64`
695
+ - `wrencode-macos-x64`
696
+ - `wrencode-macos-arm64`
697
+ - `wrencode-windows-x64.exe`
698
+
699
+ ## Slash Commands
700
+
701
+ |Command |Description |
702
+ |--------------|----------------------------------------------|
703
+ |`/help` |Show available commands |
704
+ |`/model` |Switch model, or `/model <id>` to set it directly|
705
+ |`/backend`, `/configure`|Switch backend, model and API key |
706
+ |`/clear` or `/c`|Clear the conversation (a new session with Postgres history)|
707
+ |`/sessions` |List this project's conversations (Postgres history)|
708
+ |`/resume [id]`|Continue an earlier session: a picker, or by id|
709
+ |`/search <text>`|Find past sessions by their words, then `/resume` one|
710
+ |`/sync` |Copy this project's history to the mirror now (Postgres history)|
711
+ |`/usage` |Token usage and spend for this turn and the session, and the price in effect|
712
+ |`/permissions`|Rules that allow or deny actions without asking: list, `allow`, `deny`, `forget`|
713
+ |`/mcp` |MCP servers and their tools; `/mcp reload` reconnects|
714
+ |`/compact` |Summarize history to reduce context |
715
+ |`/quit`, `/q` or `/exit`|Quit |
716
+
717
+ Type `/` to see matching commands: ↑↓ pick, Tab completes, Enter runs.
718
+
719
+ ## In the terminal
720
+
721
+ What a session looks like, and the keys that drive it.
722
+
723
+ - **Who is speaking.** Your line keeps the `❯` prompt. The model's reply opens
724
+ with a cyan dot; a tool call opens with a green dot, and what it returned
725
+ hangs under it after `⎿`: one line when that says it all (`ok`, `3 lines
726
+ read`, an error), a short excerpt otherwise. Code the model quotes is drawn
727
+ in its own tint with a gutter.
728
+ - **Approvals.** Before a write, edit or shell command runs, you see exactly
729
+ what it does: a unified diff of the file with a few lines of context, red
730
+ and green, headed by the file and line; then one question, `Apply to
731
+ app.py? Enter yes · a always · s allow bash(npm test:*) · n no`. `a`
732
+ stops asking for the rest of the session, `s` saves a rule so this kind of
733
+ action never asks again, `n` asks what to do differently.
734
+ - **Waiting.** The spinner says what is happening (`thinking`, `running
735
+ grep`, `running bash`), how long it has been, and that Escape cancels the
736
+ turn; a tool that finishes within half a second shows nothing, and a prompt
737
+ takes the spinner off the line while it waits for you. On the
738
+ Anthropic and OpenAI-style backends the reply then streams in as the model
739
+ writes it, rendered line by line; Escape stops it mid-sentence. The usage
740
+ line still counts the whole request. `WRENCODE_STREAM=0` waits for whole
741
+ replies instead; headless runs and Bedrock always do.
742
+ - **The prompt.** `←` `→` move, `Home`/`End` or `Ctrl-A`/`Ctrl-E` jump,
743
+ `Ctrl-W` deletes the word before the cursor, `Ctrl-U` to the start of the
744
+ line, `Ctrl-K` to the end, `↑`/`↓` walk the input history. A message can
745
+ span lines: end a line with `\` and press Enter, press `Alt-Enter`, or
746
+ paste text with line breaks (they are kept). `↑`/`↓` then move between the
747
+ lines, and Enter sends the whole message. While you type, a dim line
748
+ estimates the input cost of sending it.
749
+ - **Pickers.** Models and sessions are picked with `↑`/`↓` and Enter; Escape
750
+ cancels, a number jumps. Without a terminal they fall back to a numbered
751
+ prompt.
752
+ - **Errors.** A failed request is reported as a sentence and a next step
753
+ (`The API key was rejected. Run /configure to enter a new one.`), with the
754
+ raw message under it; `WRENCODE_DEBUG=1` prints it in full.
755
+ - **Colors.** wrencode draws in [Baseline](https://github.com/xRiskLab/vscode-themes),
756
+ dark or light to match the terminal's background: prose in the editor text
757
+ color, code the way the editor colors it (identifiers blue, keywords red,
758
+ strings, calls and attributes orange, constants magenta, comments muted), the
759
+ accent blue on the assistant's dot, the prompt and the banner. The background
760
+ is read from `COLORFGBG` or asked of the terminal itself (most answer); when
761
+ neither says, the first run asks you once and `/theme light|dark|auto`
762
+ changes the saved answer. `WRENCODE_THEME=light` or `dark` forces one, `ansi` draws with the terminal's
763
+ own colors instead, and a [Zed](https://zed.dev) theme file
764
+ (`~/.config/zed/themes/mine.json#Mine Light`) replaces the palette with its
765
+ own.
766
+ - **Embedded PGlite** (the default): Postgres compiled to WebAssembly, run by Node.js
767
+ with a persistent data directory under `~/.wrencode/pglite`. The first launch runs
768
+ `npm install` there for the pinned `@electric-sql/pglite` packages; after that it
769
+ starts in about a second, serves a Unix socket in an owner-only directory, and
770
+ stops when wrencode exits. It is a single-user database: one wrencode at a time
771
+ holds it, and a second one started meanwhile says so and uses `history.json`
772
+ for that run.
773
+ - **A Postgres server**: set `WRENCODE_DATABASE_URL=postgres://user:pass@host/db`
774
+ and the same schema is created there.
775
+
776
+ Each session stores its message list exactly as the backend format needs it
777
+ (JSONB), replaced whole on every save, so switching backends mid-history behaves as
778
+ it always has. Headless runs (`-p`) never read or write history.
779
+
780
+ ### Mirroring to another Postgres
781
+
782
+ PGlite has a real write-ahead log but, as a single-user engine, no replication
783
+ protocol: nothing can subscribe to it. wrencode replicates at the application level
784
+ instead, which is exact because it owns every write and saves each session whole.
785
+ Set `WRENCODE_MIRROR_URL=postgres://user:pass@host/db` and every saved session is
786
+ copied there, matched by a stable session uid, from a background thread so a slow or
787
+ unreachable mirror never holds up the loop. Only the latest snapshot per session is
788
+ kept pending; an outage is reported once, retried with backoff, and `/sync` queues
789
+ the whole project's history again and waits for it. The mirror has the same schema,
790
+ so it can serve as `WRENCODE_DATABASE_URL` for another machine. Like the other
791
+ settings that steer wrencode, the mirror URL is read from your shell, never from a
792
+ project's `.env`.
793
+
794
+ ## License
795
+
796
+ MIT - Copyright 2026 Almostly.
797
+
798
+ -----
799
+
800
+ <p align="center">
801
+ <img src="assets/almostly-badge.svg" alt="Almostly" />
802
+ </p>