closecode-ai 0.2.1__tar.gz → 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (31) hide show
  1. closecode_ai-0.4.0/PKG-INFO +252 -0
  2. closecode_ai-0.4.0/README.md +229 -0
  3. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/agent.py +10 -1
  4. closecode_ai-0.4.0/closecode_ai.egg-info/PKG-INFO +252 -0
  5. closecode_ai-0.4.0/harness.py +529 -0
  6. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/main.py +9 -0
  7. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/modes.py +3 -1
  8. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/pyproject.toml +1 -1
  9. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/session.py +22 -2
  10. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/tools.py +33 -1
  11. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/tui.py +11 -5
  12. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/ui.py +43 -2
  13. closecode_ai-0.2.1/PKG-INFO +0 -252
  14. closecode_ai-0.2.1/README.md +0 -229
  15. closecode_ai-0.2.1/closecode_ai.egg-info/PKG-INFO +0 -252
  16. closecode_ai-0.2.1/harness.py +0 -200
  17. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/closecode_ai.egg-info/SOURCES.txt +0 -0
  18. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/closecode_ai.egg-info/dependency_links.txt +0 -0
  19. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/closecode_ai.egg-info/entry_points.txt +0 -0
  20. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/closecode_ai.egg-info/requires.txt +0 -0
  21. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/closecode_ai.egg-info/top_level.txt +0 -0
  22. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/config.py +0 -0
  23. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/debug_response.py +0 -0
  24. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/guardrails.py +0 -0
  25. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/llm.py +0 -0
  26. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/mcp_tools.py +0 -0
  27. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/render.py +0 -0
  28. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/search.py +0 -0
  29. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/setup.cfg +0 -0
  30. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/todos.py +0 -0
  31. {closecode_ai-0.2.1 → closecode_ai-0.4.0}/token_tracker.py +0 -0
@@ -0,0 +1,252 @@
1
+ Metadata-Version: 2.4
2
+ Name: closecode-ai
3
+ Version: 0.4.0
4
+ Summary: CloseCode — an agentic terminal coding assistant (LangGraph + OpenRouter + MCP)
5
+ Author: Om Gite
6
+ License-Expression: MIT
7
+ Requires-Python: >=3.10
8
+ Description-Content-Type: text/markdown
9
+ Requires-Dist: langchain
10
+ Requires-Dist: langgraph
11
+ Requires-Dist: langchain-huggingface
12
+ Requires-Dist: langsmith
13
+ Requires-Dist: huggingface_hub
14
+ Requires-Dist: python-dotenv
15
+ Requires-Dist: requests
16
+ Requires-Dist: langchain-openai
17
+ Requires-Dist: langchain-mcp-adapters
18
+ Requires-Dist: mcp-server-git
19
+ Requires-Dist: langchain-openrouter
20
+ Requires-Dist: rich
21
+ Requires-Dist: pyfiglet
22
+ Requires-Dist: textual
23
+
24
+ <div align="center">
25
+
26
+ # CloseCode
27
+
28
+ **An agentic terminal coding assistant — LangGraph loop, sandboxed execution, and a full-screen TUI.**
29
+
30
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
31
+ [![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
32
+ <!-- Once CI exists, add: [![CI](https://github.com/omgite333/CLOSECODE/actions/workflows/ci.yml/badge.svg)](https://github.com/omgite333/CLOSECODE/actions) -->
33
+
34
+ <!-- SCREENSHOT: full-screen TUI on startup — banner, model name, workdir, tool list.
35
+ Capture: run `closecode`, wait for the banner, screenshot the terminal.
36
+ Save as: docs/banner.png -->
37
+ ![CloseCode banner](docs/banner.png)
38
+
39
+ </div>
40
+
41
+ ---
42
+
43
+ CloseCode reads a task, decides what to do, runs a tool, looks at the result, and repeats —
44
+ the same loop as OpenCode or Claude Code, built from scratch on LangGraph (agent loop),
45
+ LangChain (tool + model abstraction), and OpenRouter (model access). It runs entirely in
46
+ your terminal, in a full-screen TUI or a classic line-based mode, and every action it takes
47
+ on your files or shell goes through a sandboxed harness with permission prompts.
48
+
49
+
50
+ ## Why
51
+
52
+ Most people can't `pip install openai` and get an agent — the hard part isn't calling a
53
+ model, it's the loop around it: binding the right tools per mode, confirming risky actions,
54
+ persisting sessions, and stopping the model from doing something destructive. CloseCode is
55
+ that scaffolding, built openly, with open models via OpenRouter instead of a closed API.
56
+
57
+ ## Features
58
+
59
+ - **Full-screen TUI** (Textual) or a classic line-based REPL — same agent loop underneath,
60
+ switchable with `--no-tui`
61
+ - **Sandboxed execution** — file and shell operations are confined to a working directory;
62
+ path traversal (`../../etc/passwd`, absolute paths, Windows drive prefixes) is blocked
63
+ at the harness level, not just by convention
64
+ - **Plan / Build modes** — Plan mode literally never binds write/edit/bash tools to the
65
+ model, so it can explore and propose a plan with no possibility of a side effect
66
+ - **Guardrails** — four layers: input-scope filtering, destructive-command blocking
67
+ (`rm -rf /`, fork bombs, `curl | sh`), malicious-write scanning, and output redaction
68
+ - **Persistent sessions** — SQLite-backed conversation history; `/resume`, `/sessions`,
69
+ `--continue`
70
+ - **Todo tracking** — the agent maintains a visible task list for multi-step work
71
+ (`todo_write` / `todo_read`)
72
+ - **Code-aware search** — dedicated `grep`/`glob` tools (not raw shell), sandboxed and
73
+ available even in Plan mode since they're read-only
74
+ - **Git tools via MCP** — status, diff, log, commit, branches, through `mcp-server-git`
75
+ - **Any OpenRouter model** — free-tier models by default; switch with `/model`, browse
76
+ with `/models`
77
+ - **Per-user config** — your API key and default model persist in `~/.closecode/config.json`
78
+ (0600 permissions), independent of which directory you launch from
79
+
80
+ ## Install
81
+
82
+ ```bash
83
+ pipx install closecode-ai
84
+ closecode
85
+ ```
86
+
87
+ <!-- Adjust this section once published — this is the target state, not necessarily
88
+ live yet. Until it's on PyPI, use the git-based install below instead. -->
89
+
90
+ **From source, right now:**
91
+
92
+ ```bash
93
+ git clone https://github.com/omgite333/CLOSECODE.git
94
+ cd CLOSECODE
95
+ pip install -e .
96
+ closecode
97
+ ```
98
+
99
+ `pipx` is recommended over `pip` for the packaged version since it installs CLI tools into
100
+ an isolated environment and puts them straight on your `PATH`.
101
+
102
+ ## Quick start
103
+
104
+ 1. Get a free API key at [openrouter.ai/settings/keys](https://openrouter.ai/settings/keys)
105
+ 2. Run `closecode` — on first launch it prompts for the key (input hidden) and offers to
106
+ save it to `~/.closecode/config.json` so you're not asked again
107
+ 3. Type a task:
108
+
109
+ ```
110
+ > find every place we call the old auth API and list the files
111
+ ```
112
+
113
+ <!-- SCREENSHOT: a permission prompt in action — "Allow agent to run: `grep -r ...`?"
114
+ This is worth showing on its own since it's the project's core safety story.
115
+ Save as: docs/permission-prompt.png -->
116
+ ![Permission prompt](docs/permission-prompt.png)
117
+
118
+ ## How it works
119
+
120
+ ```
121
+ user input
122
+ │
123
+ ▼
124
+ ┌──────────┐ binds tools for current mode ┌───────────────┐
125
+ │ agent │ ───────────────────────────────▶ │ LangGraph loop│
126
+ │ (llm.py) │ │ (agent.py) │
127
+ └──────────┘ └──────┬────────┘
128
+ │ tool call
129
+ ▼
130
+ ┌───────────────────────┐
131
+ │ Harness │
132
+ │ sandboxed fs + shell │──▶ guardrails.py
133
+ └───────────────────────┘ (blocks/scans)
134
+ │ result
135
+ ▼
136
+ ┌──────────────────────┐
137
+ │ Renderer interface │
138
+ │ (render.py) │
139
+ └──────┬──────────┬────┘
140
+ ▼ ▼
141
+ tui.py (Textual) ui.py (classic)
142
+ ```
143
+
144
+ A single `Renderer` interface (`render.py`) decouples the agent loop from presentation, so
145
+ the Textual TUI and the classic REPL are two implementations of the same contract rather
146
+ than two copies of the agent logic.
147
+
148
+ ## Modes
149
+
150
+ | Mode | Command | Tools available | Use for |
151
+ |---|---|---|---|
152
+ | **Plan** | `/plan` | Read-only: `read_file`, `list_dir`, `grep`, `glob`, `tavily_search` | Exploring a codebase, proposing an approach with zero risk of a side effect |
153
+ | **Build** | `/build` | Everything, including `write_file`, `edit_file`, `bash`, `run_tests`, git tools | Actually making changes |
154
+
155
+ Plan mode isn't a prompt instruction the model can ignore — the write/edit/bash tools are
156
+ never bound to the model in the first place.
157
+
158
+ ## Tools
159
+
160
+ | Tool | Description |
161
+ |---|---|
162
+ | `read_file` / `write_file` / `edit_file` | Read, overwrite, or targeted find-and-replace on a file |
163
+ | `list_dir` | List a directory's contents |
164
+ | `glob` / `grep` | Find files by pattern / search file contents by regex — sandboxed, available in Plan mode |
165
+ | `bash` | Run a shell command in the sandbox (capped timeout, guardrail-checked) |
166
+ | `run_tests` | Run the project's test command and report pass/fail |
167
+ | `todo_write` / `todo_read` | Maintain a visible multi-step task list |
168
+ | `tavily_search` | Web search for current docs/APIs (requires `TAVILY_API_KEY`) |
169
+ | git tools (`status`, `diff`, `log`, `commit`, branches) | Via `mcp-server-git`, enabled with `AGENT_ENABLE_GIT=true` |
170
+
171
+ ## Commands
172
+
173
+ ```
174
+ /plan, /build switch modes
175
+ /models [filter] list OpenRouter models (free-tier first)
176
+ /model <id|number> switch model
177
+ /key update your OpenRouter API key
178
+ /sessions list saved conversations
179
+ /resume <id> switch to a saved session
180
+ /delete <id> delete a saved session
181
+ /usage token usage for this session
182
+ /clear start a new session
183
+ /help show all commands
184
+ ```
185
+
186
+ ## Configuration
187
+
188
+ All settings are environment variables (`.env`, or exported in your shell) — see
189
+ `.env.example` for the full list. Key ones:
190
+
191
+ | Variable | Default | Purpose |
192
+ |---|---|---|
193
+ | `OPENROUTER_API_KEY` | — | Required (or saved via `/key` into `~/.closecode/config.json`) |
194
+ | `OPENROUTER_MODEL` | `nvidia/nemotron-3.5-lightning:free` | Default model |
195
+ | `AGENT_WORKDIR` | `./sandbox` | Directory the agent is confined to |
196
+ | `AGENT_AUTO_APPROVE` | `false` | Skip permission prompts (guardrails still apply) |
197
+ | `AGENT_ENABLE_GIT` | `false` | Load git tools via MCP |
198
+ | `AGENT_DISABLE_GUARDRAILS` | `false` | Disable guardrails — trusted/isolated testing only |
199
+ | `TAVILY_API_KEY` | — | Enables `tavily_search` |
200
+ | `LANGCHAIN_API_KEY` | — | Enables LangSmith tracing |
201
+
202
+ ## Safety model
203
+
204
+ This is layered defense, not a single mechanism:
205
+
206
+ 1. **Input scope** — off-topic requests are redirected; clearly malicious requests
207
+ (keyloggers, phishing kits, account-hacking) are refused before reaching the model
208
+ 2. **Command blocking** — destructive shell commands are blocked before execution, even
209
+ with auto-approve on
210
+ 3. **Write scanning** — file writes/edits are scanned for malware indicators before
211
+ they're applied
212
+ 4. **Output redaction** — flagged content is scrubbed from conversation history
213
+ 5. **Sandboxed paths** — every file operation resolves through the harness, which refuses
214
+ to write outside the configured working directory regardless of how the path is phrased
215
+
216
+ These are conservative heuristics layered on top of the sandbox and per-action permission
217
+ prompts — not a formal guarantee. Shell commands currently run on the host inside a
218
+ path-restricted directory, not inside a container; see [Roadmap](#roadmap).
219
+
220
+ ## A note on model choice
221
+
222
+ Tool-calling reliability varies a lot across open models — this is the single biggest
223
+ factor in how well the agent performs. Frontier closed models are heavily trained for
224
+ reliable tool use; open models are improving but inconsistent. Roughly in order of
225
+ reliability, worth trying via `/model`:
226
+
227
+ - `qwen/qwen-2.5-coder-32b-instruct` — code-specialized, solid tool use, free tier available
228
+ - `qwen/qwen-2.5-72b-instruct` — strong, reliable, paid
229
+ - `meta-llama/llama-3.1-70b-instruct` — strong, reliable, paid
230
+ - `meta-llama/llama-3.1-8b-instruct` — fastest/cheapest, least reliable
231
+
232
+ If a smaller model frequently fails to call tools or hallucinates arguments, that's a
233
+ known gap between open and closed models on agentic tasks, not a bug here. Switching
234
+ models is the first thing to try before changing anything else.
235
+
236
+ ## Roadmap
237
+
238
+ - [ ] Test suite + CI (guardrails, sandbox path resolution, todo-store invariants)
239
+ - [ ] Docker-based sandbox for shell execution, not just path restriction
240
+ - [ ] Client/server split — `build_graph()` behind FastAPI/WebSocket, thin streaming client
241
+ - [ ] PyPI release + prebuilt binaries (PyInstaller) for no-Python-required installs
242
+ - [ ] Homebrew tap
243
+
244
+ ## Contributing
245
+
246
+ Issues and PRs welcome. If you're adding a tool, follow the pattern in `tools.py` /
247
+ `search.py`: bind state via a module-level `bind_*()` function, keep it sandboxed to the
248
+ harness root, and add it to the Plan-mode allowlist only if it's genuinely read-only.
249
+
250
+ ## License
251
+
252
+ MIT — see [LICENSE](LICENSE).
@@ -0,0 +1,229 @@
1
+ <div align="center">
2
+
3
+ # CloseCode
4
+
5
+ **An agentic terminal coding assistant — LangGraph loop, sandboxed execution, and a full-screen TUI.**
6
+
7
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
8
+ [![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
9
+ <!-- Once CI exists, add: [![CI](https://github.com/omgite333/CLOSECODE/actions/workflows/ci.yml/badge.svg)](https://github.com/omgite333/CLOSECODE/actions) -->
10
+
11
+ <!-- SCREENSHOT: full-screen TUI on startup — banner, model name, workdir, tool list.
12
+ Capture: run `closecode`, wait for the banner, screenshot the terminal.
13
+ Save as: docs/banner.png -->
14
+ ![CloseCode banner](docs/banner.png)
15
+
16
+ </div>
17
+
18
+ ---
19
+
20
+ CloseCode reads a task, decides what to do, runs a tool, looks at the result, and repeats —
21
+ the same loop as OpenCode or Claude Code, built from scratch on LangGraph (agent loop),
22
+ LangChain (tool + model abstraction), and OpenRouter (model access). It runs entirely in
23
+ your terminal, in a full-screen TUI or a classic line-based mode, and every action it takes
24
+ on your files or shell goes through a sandboxed harness with permission prompts.
25
+
26
+
27
+ ## Why
28
+
29
+ Most people can't `pip install openai` and get an agent — the hard part isn't calling a
30
+ model, it's the loop around it: binding the right tools per mode, confirming risky actions,
31
+ persisting sessions, and stopping the model from doing something destructive. CloseCode is
32
+ that scaffolding, built openly, with open models via OpenRouter instead of a closed API.
33
+
34
+ ## Features
35
+
36
+ - **Full-screen TUI** (Textual) or a classic line-based REPL — same agent loop underneath,
37
+ switchable with `--no-tui`
38
+ - **Sandboxed execution** — file and shell operations are confined to a working directory;
39
+ path traversal (`../../etc/passwd`, absolute paths, Windows drive prefixes) is blocked
40
+ at the harness level, not just by convention
41
+ - **Plan / Build modes** — Plan mode literally never binds write/edit/bash tools to the
42
+ model, so it can explore and propose a plan with no possibility of a side effect
43
+ - **Guardrails** — four layers: input-scope filtering, destructive-command blocking
44
+ (`rm -rf /`, fork bombs, `curl | sh`), malicious-write scanning, and output redaction
45
+ - **Persistent sessions** — SQLite-backed conversation history; `/resume`, `/sessions`,
46
+ `--continue`
47
+ - **Todo tracking** — the agent maintains a visible task list for multi-step work
48
+ (`todo_write` / `todo_read`)
49
+ - **Code-aware search** — dedicated `grep`/`glob` tools (not raw shell), sandboxed and
50
+ available even in Plan mode since they're read-only
51
+ - **Git tools via MCP** — status, diff, log, commit, branches, through `mcp-server-git`
52
+ - **Any OpenRouter model** — free-tier models by default; switch with `/model`, browse
53
+ with `/models`
54
+ - **Per-user config** — your API key and default model persist in `~/.closecode/config.json`
55
+ (0600 permissions), independent of which directory you launch from
56
+
57
+ ## Install
58
+
59
+ ```bash
60
+ pipx install closecode-ai
61
+ closecode
62
+ ```
63
+
64
+ <!-- Adjust this section once published — this is the target state, not necessarily
65
+ live yet. Until it's on PyPI, use the git-based install below instead. -->
66
+
67
+ **From source, right now:**
68
+
69
+ ```bash
70
+ git clone https://github.com/omgite333/CLOSECODE.git
71
+ cd CLOSECODE
72
+ pip install -e .
73
+ closecode
74
+ ```
75
+
76
+ `pipx` is recommended over `pip` for the packaged version since it installs CLI tools into
77
+ an isolated environment and puts them straight on your `PATH`.
78
+
79
+ ## Quick start
80
+
81
+ 1. Get a free API key at [openrouter.ai/settings/keys](https://openrouter.ai/settings/keys)
82
+ 2. Run `closecode` — on first launch it prompts for the key (input hidden) and offers to
83
+ save it to `~/.closecode/config.json` so you're not asked again
84
+ 3. Type a task:
85
+
86
+ ```
87
+ > find every place we call the old auth API and list the files
88
+ ```
89
+
90
+ <!-- SCREENSHOT: a permission prompt in action — "Allow agent to run: `grep -r ...`?"
91
+ This is worth showing on its own since it's the project's core safety story.
92
+ Save as: docs/permission-prompt.png -->
93
+ ![Permission prompt](docs/permission-prompt.png)
94
+
95
+ ## How it works
96
+
97
+ ```
98
+ user input
99
+ │
100
+ ▼
101
+ ┌──────────┐ binds tools for current mode ┌───────────────┐
102
+ │ agent │ ───────────────────────────────▶ │ LangGraph loop│
103
+ │ (llm.py) │ │ (agent.py) │
104
+ └──────────┘ └──────┬────────┘
105
+ │ tool call
106
+ ▼
107
+ ┌───────────────────────┐
108
+ │ Harness │
109
+ │ sandboxed fs + shell │──▶ guardrails.py
110
+ └───────────────────────┘ (blocks/scans)
111
+ │ result
112
+ ▼
113
+ ┌──────────────────────┐
114
+ │ Renderer interface │
115
+ │ (render.py) │
116
+ └──────┬──────────┬────┘
117
+ ▼ ▼
118
+ tui.py (Textual) ui.py (classic)
119
+ ```
120
+
121
+ A single `Renderer` interface (`render.py`) decouples the agent loop from presentation, so
122
+ the Textual TUI and the classic REPL are two implementations of the same contract rather
123
+ than two copies of the agent logic.
124
+
125
+ ## Modes
126
+
127
+ | Mode | Command | Tools available | Use for |
128
+ |---|---|---|---|
129
+ | **Plan** | `/plan` | Read-only: `read_file`, `list_dir`, `grep`, `glob`, `tavily_search` | Exploring a codebase, proposing an approach with zero risk of a side effect |
130
+ | **Build** | `/build` | Everything, including `write_file`, `edit_file`, `bash`, `run_tests`, git tools | Actually making changes |
131
+
132
+ Plan mode isn't a prompt instruction the model can ignore — the write/edit/bash tools are
133
+ never bound to the model in the first place.
134
+
135
+ ## Tools
136
+
137
+ | Tool | Description |
138
+ |---|---|
139
+ | `read_file` / `write_file` / `edit_file` | Read, overwrite, or targeted find-and-replace on a file |
140
+ | `list_dir` | List a directory's contents |
141
+ | `glob` / `grep` | Find files by pattern / search file contents by regex — sandboxed, available in Plan mode |
142
+ | `bash` | Run a shell command in the sandbox (capped timeout, guardrail-checked) |
143
+ | `run_tests` | Run the project's test command and report pass/fail |
144
+ | `todo_write` / `todo_read` | Maintain a visible multi-step task list |
145
+ | `tavily_search` | Web search for current docs/APIs (requires `TAVILY_API_KEY`) |
146
+ | git tools (`status`, `diff`, `log`, `commit`, branches) | Via `mcp-server-git`, enabled with `AGENT_ENABLE_GIT=true` |
147
+
148
+ ## Commands
149
+
150
+ ```
151
+ /plan, /build switch modes
152
+ /models [filter] list OpenRouter models (free-tier first)
153
+ /model <id|number> switch model
154
+ /key update your OpenRouter API key
155
+ /sessions list saved conversations
156
+ /resume <id> switch to a saved session
157
+ /delete <id> delete a saved session
158
+ /usage token usage for this session
159
+ /clear start a new session
160
+ /help show all commands
161
+ ```
162
+
163
+ ## Configuration
164
+
165
+ All settings are environment variables (`.env`, or exported in your shell) — see
166
+ `.env.example` for the full list. Key ones:
167
+
168
+ | Variable | Default | Purpose |
169
+ |---|---|---|
170
+ | `OPENROUTER_API_KEY` | — | Required (or saved via `/key` into `~/.closecode/config.json`) |
171
+ | `OPENROUTER_MODEL` | `nvidia/nemotron-3.5-lightning:free` | Default model |
172
+ | `AGENT_WORKDIR` | `./sandbox` | Directory the agent is confined to |
173
+ | `AGENT_AUTO_APPROVE` | `false` | Skip permission prompts (guardrails still apply) |
174
+ | `AGENT_ENABLE_GIT` | `false` | Load git tools via MCP |
175
+ | `AGENT_DISABLE_GUARDRAILS` | `false` | Disable guardrails — trusted/isolated testing only |
176
+ | `TAVILY_API_KEY` | — | Enables `tavily_search` |
177
+ | `LANGCHAIN_API_KEY` | — | Enables LangSmith tracing |
178
+
179
+ ## Safety model
180
+
181
+ This is layered defense, not a single mechanism:
182
+
183
+ 1. **Input scope** — off-topic requests are redirected; clearly malicious requests
184
+ (keyloggers, phishing kits, account-hacking) are refused before reaching the model
185
+ 2. **Command blocking** — destructive shell commands are blocked before execution, even
186
+ with auto-approve on
187
+ 3. **Write scanning** — file writes/edits are scanned for malware indicators before
188
+ they're applied
189
+ 4. **Output redaction** — flagged content is scrubbed from conversation history
190
+ 5. **Sandboxed paths** — every file operation resolves through the harness, which refuses
191
+ to write outside the configured working directory regardless of how the path is phrased
192
+
193
+ These are conservative heuristics layered on top of the sandbox and per-action permission
194
+ prompts — not a formal guarantee. Shell commands currently run on the host inside a
195
+ path-restricted directory, not inside a container; see [Roadmap](#roadmap).
196
+
197
+ ## A note on model choice
198
+
199
+ Tool-calling reliability varies a lot across open models — this is the single biggest
200
+ factor in how well the agent performs. Frontier closed models are heavily trained for
201
+ reliable tool use; open models are improving but inconsistent. Roughly in order of
202
+ reliability, worth trying via `/model`:
203
+
204
+ - `qwen/qwen-2.5-coder-32b-instruct` — code-specialized, solid tool use, free tier available
205
+ - `qwen/qwen-2.5-72b-instruct` — strong, reliable, paid
206
+ - `meta-llama/llama-3.1-70b-instruct` — strong, reliable, paid
207
+ - `meta-llama/llama-3.1-8b-instruct` — fastest/cheapest, least reliable
208
+
209
+ If a smaller model frequently fails to call tools or hallucinates arguments, that's a
210
+ known gap between open and closed models on agentic tasks, not a bug here. Switching
211
+ models is the first thing to try before changing anything else.
212
+
213
+ ## Roadmap
214
+
215
+ - [ ] Test suite + CI (guardrails, sandbox path resolution, todo-store invariants)
216
+ - [ ] Docker-based sandbox for shell execution, not just path restriction
217
+ - [ ] Client/server split — `build_graph()` behind FastAPI/WebSocket, thin streaming client
218
+ - [ ] PyPI release + prebuilt binaries (PyInstaller) for no-Python-required installs
219
+ - [ ] Homebrew tap
220
+
221
+ ## Contributing
222
+
223
+ Issues and PRs welcome. If you're adding a tool, follow the pattern in `tools.py` /
224
+ `search.py`: bind state via a module-level `bind_*()` function, keep it sandboxed to the
225
+ harness root, and add it to the Plan-mode allowlist only if it's genuinely read-only.
226
+
227
+ ## License
228
+
229
+ MIT — see [LICENSE](LICENSE).
@@ -9,7 +9,8 @@ from llm import get_llm
9
9
 
10
10
  SYSTEM_PROMPT = """You are a terminal coding agent running in a sandboxed working directory.
11
11
  You have tools for reading, writing, and editing files, running shell commands and tests,
12
- listing directories, and interacting with git (status, diff, log, commit, branches).
12
+ listing directories, managing background processes, undoing your own file changes,
13
+ and interacting with git (status, diff, log, commit, branches).
13
14
 
14
15
  Note: your available tools change depending on the current mode. In "plan" mode only
15
16
  read-only tools are bound to you (you literally cannot call write/edit/bash/commit tools
@@ -25,6 +26,14 @@ Rules:
25
26
  - Verify your work: after making a change, run a command, run tests, or read the
26
27
  file back to confirm it did what you intended. Never report a task complete
27
28
  without verifying — bugs are unacceptable, so test before you say "done".
29
+ - Long-running processes (dev servers, watchers, tunnels) go through
30
+ start_background, never a foreground bash call that would hang until timeout.
31
+ Use tail_logs to watch the process's output, fix what crashes, and restart
32
+ with kill_background + start_background. Always kill servers you started
33
+ when the task no longer needs them.
34
+ - Every write_file/edit_file is snapshotted before the change, and the user can
35
+ revert with /undo — treat that as a safety net, not a workflow: verify with
36
+ tests instead of writing sloppily and undoing.
28
37
  - Use git tools deliberately: check status/diff before committing, and never force-push
29
38
  or hard-reset unless the user explicitly asked for that specific action.
30
39
  - When the task is complete, reply with plain text summarizing what you did.