handcode 0.3.0rc1__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agentctl/__init__.py +0 -0
- agentctl/adapters/__init__.py +0 -0
- agentctl/adapters/litellm/__init__.py +9 -0
- agentctl/adapters/litellm/hook.py +49 -0
- agentctl/adapters/litellm/recorder.py +187 -0
- agentctl/adapters/openhands/__init__.py +169 -0
- agentctl/adapters/openhands/handoff.py +155 -0
- agentctl/adapters/openhands/seam_b.py +259 -0
- agentctl/adapters/openhands/seam_c.py +209 -0
- agentctl/cli.py +1450 -0
- agentctl/control/__init__.py +0 -0
- agentctl/control/cost/__init__.py +4 -0
- agentctl/control/cost/ledger.py +210 -0
- agentctl/control/dash.py +697 -0
- agentctl/control/keys.py +440 -0
- agentctl/control/matrix/__init__.py +0 -0
- agentctl/control/matrix/data/tools.yaml +149 -0
- agentctl/control/policy/__init__.py +10 -0
- agentctl/control/policy/compile.py +258 -0
- agentctl/control/policy/data/policy.compiled.json +38 -0
- agentctl/control/policy/data/policy.yaml +46 -0
- agentctl/control/probe.py +399 -0
- agentctl/control/providers.py +293 -0
- agentctl/control/proxy.py +536 -0
- agentctl/control/proxyenv.py +309 -0
- agentctl/control/replay/__init__.py +14 -0
- agentctl/control/replay/cassette.py +281 -0
- agentctl/control/replay/server.py +109 -0
- agentctl/demo/__init__.py +214 -0
- agentctl/demo/child.py +84 -0
- agentctl/demo/mock.py +79 -0
- agentctl/demo/tool.py +62 -0
- agentctl/gha.py +488 -0
- agentctl/kernel/__init__.py +0 -0
- agentctl/kernel/classify.py +170 -0
- agentctl/kernel/gate.py +391 -0
- agentctl/kernel/hook.py +229 -0
- agentctl/kernel/ledger/__init__.py +0 -0
- agentctl/kernel/ledger/models.py +160 -0
- agentctl/kernel/ledger/schema.sql +62 -0
- agentctl/kernel/ledger/store.py +596 -0
- agentctl/kernel/paths.py +203 -0
- agentctl/kernel/policy.py +160 -0
- agentctl/kernel/reconcile/__init__.py +31 -0
- agentctl/kernel/reconcile/base.py +106 -0
- agentctl/kernel/reconcile/external.py +137 -0
- agentctl/kernel/reconcile/filesystem.py +162 -0
- agentctl/kernel/reconcile/git.py +162 -0
- agentctl/runtime/__init__.py +20 -0
- agentctl/runtime/citations.py +179 -0
- agentctl/runtime/config.py +97 -0
- agentctl/runtime/doctor.py +335 -0
- agentctl/runtime/init.py +148 -0
- agentctl/runtime/lease.py +143 -0
- agentctl/runtime/orchestrate.py +187 -0
- agentctl/runtime/plugins.py +130 -0
- agentctl/runtime/report.py +361 -0
- agentctl/runtime/runner.py +787 -0
- agentctl/runtime/runs.py +191 -0
- agentctl/runtime/subagent.py +274 -0
- agentctl/runtime/tools.py +350 -0
- handcode-0.3.0rc1.dist-info/METADATA +659 -0
- handcode-0.3.0rc1.dist-info/RECORD +67 -0
- handcode-0.3.0rc1.dist-info/WHEEL +5 -0
- handcode-0.3.0rc1.dist-info/entry_points.txt +3 -0
- handcode-0.3.0rc1.dist-info/licenses/LICENSE +21 -0
- handcode-0.3.0rc1.dist-info/top_level.txt +1 -0
|
@@ -0,0 +1,659 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: handcode
|
|
3
|
+
Version: 0.3.0rc1
|
|
4
|
+
Summary: A coding agent whose work survives crashes, restarts and provider switches
|
|
5
|
+
Author: csdeepak
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/csdeepak/HandCode
|
|
8
|
+
Project-URL: Documentation, https://github.com/csdeepak/HandCode/tree/main/guide
|
|
9
|
+
Project-URL: Changelog, https://github.com/csdeepak/HandCode/blob/main/CHANGELOG.md
|
|
10
|
+
Project-URL: Issues, https://github.com/csdeepak/HandCode/issues
|
|
11
|
+
Keywords: agent,coding-agent,openhands,litellm,llm,crash-recovery,idempotency
|
|
12
|
+
Classifier: Development Status :: 3 - Alpha
|
|
13
|
+
Classifier: Environment :: Console
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Operating System :: OS Independent
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
19
|
+
Classifier: Topic :: Software Development
|
|
20
|
+
Requires-Python: >=3.12
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
License-File: LICENSE
|
|
23
|
+
Requires-Dist: pyyaml>=6.0
|
|
24
|
+
Provides-Extra: openhands
|
|
25
|
+
Requires-Dist: openhands-sdk>=1.45.0; extra == "openhands"
|
|
26
|
+
Provides-Extra: proxy
|
|
27
|
+
Requires-Dist: litellm[proxy]>=1.100.0; extra == "proxy"
|
|
28
|
+
Provides-Extra: dev
|
|
29
|
+
Requires-Dist: pytest>=8.0; extra == "dev"
|
|
30
|
+
Requires-Dist: markdown-it-py>=3; extra == "dev"
|
|
31
|
+
Dynamic: license-file
|
|
32
|
+
|
|
33
|
+
# HandCode
|
|
34
|
+
|
|
35
|
+
[](https://github.com/csdeepak/HandCode/actions/workflows/ci.yml)
|
|
36
|
+
|
|
37
|
+
**HandCode keeps a coding agent's work safe across crashes, restarts and
|
|
38
|
+
provider switches.** Its command is `agentctl`.
|
|
39
|
+
|
|
40
|
+
When an agent's process dies in the middle of a task and you resume it, the
|
|
41
|
+
agent framework re-runs whatever was in flight, so a commit can happen twice.
|
|
42
|
+
agentctl records every action before it runs. A resumed run never repeats one
|
|
43
|
+
that already happened, and when agentctl cannot tell whether it did, it stops
|
|
44
|
+
and asks you instead of guessing.
|
|
45
|
+
|
|
46
|
+
It wraps the [OpenHands](https://github.com/OpenHands/software-agent-sdk)
|
|
47
|
+
agent SDK and works with any provider LiteLLM supports, free tiers included.
|
|
48
|
+
|
|
49
|
+
**Documentation:** <https://csdeepak.github.io/HandCode/>
|
|
50
|
+
|
|
51
|
+
## See it in a minute
|
|
52
|
+
|
|
53
|
+
```bash
|
|
54
|
+
agentctl demo # no API key, no network, $0
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
```
|
|
58
|
+
1. Plain OpenHands, no agentctl
|
|
59
|
+
the agent committed, then its process was killed ... 1 new commit
|
|
60
|
+
resumed ............................................ 2 new commits <- the same commit, twice
|
|
61
|
+
|
|
62
|
+
2. With agentctl
|
|
63
|
+
the agent committed, then its process was killed ... 1 new commit
|
|
64
|
+
resumed ............................................ 1 new commit <- once
|
|
65
|
+
the ledger: the commit is OBSERVED (git probe: LANDED) · 0 waiting on you
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Real git, real process death and the real SDK. Only the model is scripted,
|
|
69
|
+
which is what makes it free and identical on every OS.
|
|
70
|
+
|
|
71
|
+
## Quickstart
|
|
72
|
+
|
|
73
|
+
Python 3.12+, `git`, and one API key (a free OpenRouter or Gemini key works).
|
|
74
|
+
|
|
75
|
+
```bash
|
|
76
|
+
git clone https://github.com/csdeepak/HandCode && cd HandCode
|
|
77
|
+
python3 -m venv .venv && source .venv/bin/activate # Windows: see guide/quickstart.md
|
|
78
|
+
pip install uv && uv pip install -e ".[openhands]" # uv: seconds, where pip takes minutes
|
|
79
|
+
|
|
80
|
+
agentctl demo # see what it is for
|
|
81
|
+
agentctl init # one key, one checked default model
|
|
82
|
+
cd ../your-project
|
|
83
|
+
agentctl run "fix the failing date test" --accept "python -m pytest -q"
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Every run ends with a report of what was checked, what changed, what it cost,
|
|
87
|
+
and what needs you:
|
|
88
|
+
|
|
89
|
+
```
|
|
90
|
+
outcome PASS `python -m pytest -q` exited 0 (run by agentctl after the agent finished)
|
|
91
|
+
changed 1 file +5 -1
|
|
92
|
+
used 9 requests · 43.7K tokens · $0.00 (free-tier model) · 12s
|
|
93
|
+
actions 8 actions: 3 reads, 3 file writes, 2 commands
|
|
94
|
+
needs you nothing
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
The **[user guide](https://github.com/csdeepak/HandCode/blob/main/guide/README.md)** covers the
|
|
98
|
+
[quickstart](https://github.com/csdeepak/HandCode/blob/main/guide/quickstart.md) step by step (Windows included),
|
|
99
|
+
[concepts in plain words](https://github.com/csdeepak/HandCode/blob/main/guide/concepts.md),
|
|
100
|
+
[troubleshooting by the exact text you see](https://github.com/csdeepak/HandCode/blob/main/guide/troubleshooting.md), and an
|
|
101
|
+
[FAQ](https://github.com/csdeepak/HandCode/blob/main/guide/faq.md).
|
|
102
|
+
|
|
103
|
+
## What it protects you from
|
|
104
|
+
|
|
105
|
+
| | Without agentctl | With it |
|
|
106
|
+
|---|---|---|
|
|
107
|
+
| The process dies mid-action, then you resume | The in-flight action runs again | It runs at most once. If it already landed, the agent gets its result back |
|
|
108
|
+
| You cannot tell whether something happened | It is guessed | It is checked (did HEAD move?), or you are asked |
|
|
109
|
+
| Ctrl-C, a closed laptop, a killed terminal | Start again | `agentctl resume` |
|
|
110
|
+
| A rate limit mid-task | The run dies | `--wait 30m` waits it out and resumes, or `--pool` routes to another provider |
|
|
111
|
+
| A dangerous command with nobody watching | It runs, or the run hangs on a prompt | It is queued for `agentctl approve` / `deny` |
|
|
112
|
+
| "Done!" from the agent | Taken on trust | `--accept` runs your tests and reports PASS or FAIL |
|
|
113
|
+
| Two terminals resuming the same run | Both drive it | The second is refused and told which process holds it |
|
|
114
|
+
|
|
115
|
+
**What it does not do:** it is not a sandbox. Commands run on your machine.
|
|
116
|
+
agentctl stops actions repeating and asks before dangerous ones, but it does
|
|
117
|
+
not contain what an allowed command does. For a task you would not trust your
|
|
118
|
+
shell with, run it in the [container image](https://github.com/csdeepak/HandCode/blob/main/guide/quickstart.md#6-run-it-in-a-container-recommended-for-tasks-you-would-not-trust-your-shell-with),
|
|
119
|
+
where only the mounted repository is reachable
|
|
120
|
+
([concepts](https://github.com/csdeepak/HandCode/blob/main/guide/concepts.md#what-agentctl-does-not-do)).
|
|
121
|
+
|
|
122
|
+
**On GitHub.** Type a task in your repository's Actions tab and get a pull
|
|
123
|
+
request back, with the report as its description. Your key stays in your
|
|
124
|
+
repository's secrets, and the job that runs the agent holds no token that can
|
|
125
|
+
write ([the GitHub Action](https://github.com/csdeepak/HandCode/blob/main/guide/github-action.md)).
|
|
126
|
+
It has run on GitHub with a real model and opened a real pull request.
|
|
127
|
+
|
|
128
|
+
**Status.** 887 tests, including a chaos suite that kills the process at ten
|
|
129
|
+
points in the protocol, plus `verify.py`'s end-to-end crash experiments. All
|
|
130
|
+
run at no cost, on Linux and Windows in CI. Every feature here has also been
|
|
131
|
+
run at least once against a real provider. The design history, with what was
|
|
132
|
+
measured and what was found, is the numbered [`docs/`](https://github.com/csdeepak/HandCode/blob/main/INDEX.md) stream.
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
# Reference
|
|
137
|
+
|
|
138
|
+
## The problem, in one example
|
|
139
|
+
|
|
140
|
+
An agent runs `git commit`. The process dies before the result is recorded. On
|
|
141
|
+
resume, OpenHands re-drives the pending action through the real executor — and
|
|
142
|
+
commits again.
|
|
143
|
+
|
|
144
|
+
That is not a hypothesis. `docs/0014` reproduces it: **1 effect before the
|
|
145
|
+
crash, 2 after resume.**
|
|
146
|
+
|
|
147
|
+
With this layer installed:
|
|
148
|
+
|
|
149
|
+
| | M2a | M4 | M2b |
|
|
150
|
+
|---|---|---|---|
|
|
151
|
+
| Duplicate effect | none | none | none |
|
|
152
|
+
| Ledger state | `BLOCKED` | `COMMITTED` | `OBSERVED` |
|
|
153
|
+
| Human needed | yes | no | no |
|
|
154
|
+
| Agent receives | rejection | rejection | **the result** |
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## Verify it yourself
|
|
159
|
+
|
|
160
|
+
Requires **Python ≥ 3.12** (the OpenHands SDK does) and `git` on PATH. On an
|
|
161
|
+
older Python the install fails with a long list of `Requires-Python >=3.12`
|
|
162
|
+
lines that never names the cause, so check first.
|
|
163
|
+
|
|
164
|
+
**Windows** (PowerShell):
|
|
165
|
+
|
|
166
|
+
```powershell
|
|
167
|
+
py -3.12 -m venv .venv
|
|
168
|
+
.venv\Scripts\Activate.ps1
|
|
169
|
+
pip install -e ".[dev,openhands]"
|
|
170
|
+
python verify.py
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
Clone into a short path such as `C:\src\HandCode`. One file the install
|
|
174
|
+
unpacks sits 139 characters deep inside `.venv`, so a clone path longer than
|
|
175
|
+
about 120 characters hits Windows' 260-character limit and pip fails partway
|
|
176
|
+
(`docs/0044` N3).
|
|
177
|
+
|
|
178
|
+
**macOS / Linux:**
|
|
179
|
+
|
|
180
|
+
```bash
|
|
181
|
+
python3 -m venv .venv # Ubuntu: sudo apt install python3.12-venv first
|
|
182
|
+
source .venv/bin/activate
|
|
183
|
+
pip install -e ".[dev,openhands]"
|
|
184
|
+
python verify.py
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
**Zero cost** — everything runs against a local mock provider. No API key, no
|
|
188
|
+
network, no tokens. Takes about five minutes. With this install, two of the
|
|
189
|
+
ten checks (Seam A and the full stack) are **SKIPPED**, because they need the
|
|
190
|
+
proxy extra below; that is expected, and `verify.py` says how to run them.
|
|
191
|
+
|
|
192
|
+
If a dependency has since shipped something incompatible, install the exact set
|
|
193
|
+
that is known to pass:
|
|
194
|
+
|
|
195
|
+
```bash
|
|
196
|
+
pip install -e ".[dev,openhands]" -c constraints.txt
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
CI runs both forms on Linux and Windows, and runs `verify.py` itself — so the
|
|
200
|
+
badge above means the correctness claim reproduces on a machine that is not
|
|
201
|
+
mine, which is the only version of that claim worth anything.
|
|
202
|
+
|
|
203
|
+
### The proxy extra needs two steps (only if you run the proxy yourself)
|
|
204
|
+
|
|
205
|
+
`agentctl run --pool` and `agentctl proxy up` do not need any of this: they
|
|
206
|
+
keep LiteLLM in a separate environment of their own. The steps below are for
|
|
207
|
+
running the full-stack checks in `verify.py`, or a proxy by hand.
|
|
208
|
+
|
|
209
|
+
`litellm[proxy]` declares `mcp<2.0`; the OpenHands SDK needs `fastmcp` and so
|
|
210
|
+
needs `mcp>=2`. They are incompatible on paper and work in practice, so the
|
|
211
|
+
override is explicit rather than hidden in a version range pip would refuse:
|
|
212
|
+
|
|
213
|
+
```bash
|
|
214
|
+
pip install -e ".[dev,openhands,proxy]"
|
|
215
|
+
pip install --upgrade "mcp>=2.2.0" "fastmcp>=4.0.3"
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
`fastmcp` is named explicitly because the proxy install leaves only
|
|
219
|
+
`fastmcp-slim`, which has no client support — and the SDK does
|
|
220
|
+
`from fastmcp import Client` on nearly every import path.
|
|
221
|
+
|
|
222
|
+
Only Seam A and the full-stack demo need this. Everything else — including the
|
|
223
|
+
whole correctness core — runs without the proxy.
|
|
224
|
+
|
|
225
|
+
---
|
|
226
|
+
|
|
227
|
+
## Set up your keys
|
|
228
|
+
|
|
229
|
+
```bash
|
|
230
|
+
agentctl keys --init # writes keys.env listing every provider + how to get one
|
|
231
|
+
agentctl keys # what is set, and where to get the rest
|
|
232
|
+
agentctl keys --install-hook # refuse any commit containing one of your keys
|
|
233
|
+
agentctl keys --check # does every key actually work? costs no tokens
|
|
234
|
+
agentctl dash # failover, effects, spend, policy on one screen
|
|
235
|
+
```
|
|
236
|
+
|
|
237
|
+
`keys.env` is gitignored **before** it is written, and no command ever prints a
|
|
238
|
+
value. If it ever becomes tracked by git, every command says so loudly and
|
|
239
|
+
tells you to rotate — a `.gitignore` entry added after a file is tracked does
|
|
240
|
+
nothing (`docs/0032`).
|
|
241
|
+
|
|
242
|
+
**One key is enough to start.** A daily cap on it stops work until it resets;
|
|
243
|
+
a key at a **second provider** survives that, and survives an outage too:
|
|
244
|
+
|
|
245
|
+
```
|
|
246
|
+
OPENROUTER_API_KEY=sk-or-v1-... # one provider
|
|
247
|
+
GEMINI_API_KEY=... # a second provider
|
|
248
|
+
```
|
|
249
|
+
|
|
250
|
+
Extra keys at the *same* provider are accepted with any suffix (`_2`, `_WORK`)
|
|
251
|
+
and each becomes its own deployment in the pool (`docs/0033`). **Whether that
|
|
252
|
+
is a separate quota is unverified**: OpenRouter's own limits page says extra
|
|
253
|
+
accounts do not change rate limits. Pooling several free accounts may also be
|
|
254
|
+
against a provider's terms. Check both before relying on it (`docs/0042`
|
|
255
|
+
§4.C). `agentctl dash` says which state you are actually in.
|
|
256
|
+
|
|
257
|
+
## Use it
|
|
258
|
+
|
|
259
|
+
```bash
|
|
260
|
+
agentctl init # once: one key, one checked default model
|
|
261
|
+
cd myproject
|
|
262
|
+
agentctl run "add type hints to utils.py" # no flags needed
|
|
263
|
+
```
|
|
264
|
+
|
|
265
|
+
`init` uses a key you already have in the environment or a keys file, or asks
|
|
266
|
+
for one and stores it in `~/.agentctl/keys.env`, outside every repository and
|
|
267
|
+
never echoed. It then sends **one** completion and records the model that
|
|
268
|
+
actually answered in `~/.agentctl/config.toml`. A paid provider is not called
|
|
269
|
+
unless you pass `--check-paid`.
|
|
270
|
+
|
|
271
|
+
`run` takes its model from `--model`, then `AGENTCTL_MODEL`, then that config,
|
|
272
|
+
then the first provider you hold a key for, and prints which one it used.
|
|
273
|
+
|
|
274
|
+
Every run ends with a report:
|
|
275
|
+
|
|
276
|
+
```
|
|
277
|
+
outcome PASS `python -m pytest -q` exited 0 (run by agentctl after the agent finished)
|
|
278
|
+
changed 1 file +5 -1
|
|
279
|
+
stats.py +5 -1
|
|
280
|
+
agent said "The median function in stats.py has been fixed to correctly handle..."
|
|
281
|
+
used 9 requests · 43.7K tokens · $0.00 (free-tier model) · 12s
|
|
282
|
+
actions 8 actions: 3 reads, 3 file writes, 2 commands
|
|
283
|
+
needs you nothing
|
|
284
|
+
```
|
|
285
|
+
|
|
286
|
+
- **`outcome` is only what was checked.** Pass `--accept "<command>"` (your
|
|
287
|
+
test suite, usually) and agentctl runs it itself once the agent has
|
|
288
|
+
finished, outside the agent's loop: exit 0 is PASS. Without it the outcome
|
|
289
|
+
reads `not checked`, never an implied success.
|
|
290
|
+
- **`changed`** is measured with git against where the run started, commits
|
|
291
|
+
included.
|
|
292
|
+
- **`used`** says what a cost figure can be trusted for. Free-tier models read
|
|
293
|
+
`$0.00`, and list prices that a free key is not billed are labelled as such.
|
|
294
|
+
- **The exit code is 0** when nothing failed a check and nothing is waiting on
|
|
295
|
+
you, and 1 otherwise.
|
|
296
|
+
|
|
297
|
+
```bash
|
|
298
|
+
agentctl doctor --workspace ./myproject # check everything BEFORE you spend
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
`doctor` checks packages, the SDK import chain, git, your provider keys, the
|
|
302
|
+
policy, and whether the workspace has uncommitted changes the agent is about
|
|
303
|
+
to edit.
|
|
304
|
+
|
|
305
|
+
> **Correction (2026-09-21).** This paragraph used to say OpenRouter exposes no
|
|
306
|
+
> remaining-free-request counter on any endpoint (`docs/0031`). That is false.
|
|
307
|
+
> `GET /api/v1/key` returns `free_model_daily_requests` with `used`, `limit`
|
|
308
|
+
> and `remaining`, and it costs nothing — measured against six live accounts,
|
|
309
|
+
> all reporting `limit=50`. `probe.py` was already calling the sibling endpoint
|
|
310
|
+
> `/api/v1/auth/key` the whole time. `docs/0038` §9.4 records the measurement;
|
|
311
|
+
> surfacing it is scheduled work.
|
|
312
|
+
|
|
313
|
+
Real tools (bash, read, write), a real model, every effect classified and
|
|
314
|
+
ledgered. Crash it and re-run with the conversation id the run printed at the
|
|
315
|
+
start. Work already done is not repeated:
|
|
316
|
+
|
|
317
|
+
```bash
|
|
318
|
+
agentctl run "" --workspace ./myproject --resume <conversation-id>
|
|
319
|
+
```
|
|
320
|
+
|
|
321
|
+
The task is `""` because the conversation already holds it.
|
|
322
|
+
|
|
323
|
+
Only one process may drive a conversation. If the run that held it crashed on
|
|
324
|
+
this machine, `--resume` sees that it is gone and takes over. If it is still
|
|
325
|
+
running, `--resume` refuses and names it, because two drivers of one
|
|
326
|
+
conversation is exactly what the ledger exists to prevent. `--takeover`
|
|
327
|
+
overrides that, for a holder you know is dead but this machine cannot check
|
|
328
|
+
(`docs/0046`).
|
|
329
|
+
|
|
330
|
+
### Run it again for nothing
|
|
331
|
+
|
|
332
|
+
```bash
|
|
333
|
+
agentctl run "..." --workspace ./app --record session.jsonl # once, for real
|
|
334
|
+
agentctl run "" --workspace ./app --replay session.jsonl # $0.00, offline
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
`--replay` serves every completion from the recording: no API key, no network,
|
|
338
|
+
no tokens, no sampling. A request the cassette does not contain is answered
|
|
339
|
+
with a 502 naming the turn that diverged, never with a plausible-looking
|
|
340
|
+
completion — a replay that cannot fail would not be worth running.
|
|
341
|
+
|
|
342
|
+
Replay pins the model, not the world: tool output feeds the next request, so
|
|
343
|
+
the workspace has to start where the recording did (`docs/0029`).
|
|
344
|
+
|
|
345
|
+
Two safety properties, deliberately separate — the **gate** stops an effect
|
|
346
|
+
happening *twice*; `--confirm-destructive` (on by default) stops one happening
|
|
347
|
+
*at all* without a human saying yes. It asks on two grounds: the effect is
|
|
348
|
+
destructive, **or** it writes outside the workspace. `echo x > ~/.bashrc` is an
|
|
349
|
+
ordinary idempotent write that is simply none of the agent's business
|
|
350
|
+
(`docs/0027`).
|
|
351
|
+
|
|
352
|
+
"Twice" means *by replay*. Once a result — success or failure — is in the
|
|
353
|
+
history the model reads, an identical later call is the model deciding to run
|
|
354
|
+
it again: the edit, test, re-test loop. It runs. A crash still never causes a
|
|
355
|
+
repeat (`docs/0045`).
|
|
356
|
+
|
|
357
|
+
### Or embed it
|
|
358
|
+
|
|
359
|
+
```python
|
|
360
|
+
from agentctl.adapters.openhands import protect
|
|
361
|
+
|
|
362
|
+
guard = protect(
|
|
363
|
+
ledger="./ledger.db",
|
|
364
|
+
conversation_id=str(conversation_id),
|
|
365
|
+
tools={"commit": CommitTool}, # gated at Seam C
|
|
366
|
+
repo_root="./workspace", # for the git probe
|
|
367
|
+
takeover=previous_holder_is_dead, # NEVER for a live one (docs/0046)
|
|
368
|
+
)
|
|
369
|
+
|
|
370
|
+
conv = Conversation(agent=agent, callbacks=[guard.seam_b], ...)
|
|
371
|
+
guard.attach(conv) # required, or the gate is inert
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
Omit `tools` to run Seam B only: still correct, but an already-landed effect is
|
|
375
|
+
blocked rather than resumed cleanly.
|
|
376
|
+
|
|
377
|
+
### Failing over
|
|
378
|
+
|
|
379
|
+
```bash
|
|
380
|
+
agentctl run "..." --pool # starts the managed pool if it is not running
|
|
381
|
+
agentctl proxy status # running? answering? where is its log?
|
|
382
|
+
agentctl proxy down # stop it
|
|
383
|
+
```
|
|
384
|
+
|
|
385
|
+
`--pool` (or `agentctl proxy up`) runs LiteLLM in **its own environment**,
|
|
386
|
+
`~/.agentctl/proxy-env`, created the first time (a few minutes) and reused
|
|
387
|
+
after that. Your install never needs the proxy extra.
|
|
388
|
+
- It verifies every model id with one completion and leaves out what cannot
|
|
389
|
+
serve, naming it.
|
|
390
|
+
- It runs in the background and outlives the command that started it.
|
|
391
|
+
|
|
392
|
+
`agentctl proxy --out ./proxy` still writes a config and start scripts, if you
|
|
393
|
+
would rather run the proxy yourself (`docs/0047`).
|
|
394
|
+
|
|
395
|
+
A pool over **one** provider key survives a transient upstream overload and a
|
|
396
|
+
per-model limit. It does **not** survive an account-wide daily cap — three
|
|
397
|
+
`:free` models on one key share one quota. The command counts credentials, not
|
|
398
|
+
deployments, and says so.
|
|
399
|
+
|
|
400
|
+
### Choosing one source
|
|
401
|
+
|
|
402
|
+
```bash
|
|
403
|
+
agentctl models # every source, and what choosing it costs
|
|
404
|
+
agentctl run "..." --source gemini --base-url http://localhost:4000
|
|
405
|
+
```
|
|
406
|
+
|
|
407
|
+
`pool` is the default and the widest thing you can ask for. A source group is
|
|
408
|
+
narrower **on purpose**, so every row says what it gives up:
|
|
409
|
+
|
|
410
|
+
```
|
|
411
|
+
pool 48 deployments 24 accounts the default
|
|
412
|
+
pool-openrouter 18 deployments 6 accounts gives up 30 of 48
|
|
413
|
+
pool-gemini 6 deployments 6 accounts gives up 42 of 48
|
|
414
|
+
```
|
|
415
|
+
|
|
416
|
+
That line is the feature. Narrowing to one source means an account-wide daily
|
|
417
|
+
cap has less to fail over to — the exact failure the multi-account design
|
|
418
|
+
exists to escape (`docs/0033`). `--verify` marks which sources can actually
|
|
419
|
+
serve, which is not the same as which ones you hold keys for.
|
|
420
|
+
|
|
421
|
+
### Delegating a read
|
|
422
|
+
|
|
423
|
+
```bash
|
|
424
|
+
agentctl subagent --init # writes an example definition
|
|
425
|
+
agentctl subagent reviewer "what does the gate do?"
|
|
426
|
+
```
|
|
427
|
+
|
|
428
|
+
Claude Code's Markdown frontmatter format, loaded through the SDK's own
|
|
429
|
+
`AgentDefinition`. A subagent may hold **`read_file` and nothing else** — no
|
|
430
|
+
shell, no writes, no MCP.
|
|
431
|
+
|
|
432
|
+
That is not a starter limitation, it is the whole safety argument. Multi-agent
|
|
433
|
+
is skipped because four correctness mechanisms assume a single writer
|
|
434
|
+
(`docs/0038` §4.2); a subagent that cannot produce an effect needs none of
|
|
435
|
+
them. So the restriction is enforced twice — the definition is validated, and
|
|
436
|
+
the tool list handed to the agent is built from a constant rather than from
|
|
437
|
+
the definition, because a property that depends on one function returning
|
|
438
|
+
correctly is one refactor from gone.
|
|
439
|
+
|
|
440
|
+
### Installing a plugin, one capability at a time
|
|
441
|
+
|
|
442
|
+
```bash
|
|
443
|
+
agentctl plugins ./some-plugin # the audit. Installs nothing.
|
|
444
|
+
```
|
|
445
|
+
|
|
446
|
+
A Claude Code plugin can carry agents, hooks, commands, skills and MCP
|
|
447
|
+
servers. **Five of those six are refused**, and each refusal is counted:
|
|
448
|
+
|
|
449
|
+
```
|
|
450
|
+
ADMITTED 1 read-only agent(s)
|
|
451
|
+
REJECTED 1 agent(s) that could change the world
|
|
452
|
+
REFUSED mcp_config 2 declared
|
|
453
|
+
commands 1 declared
|
|
454
|
+
```
|
|
455
|
+
|
|
456
|
+
"declines MCP servers" is a policy; "2 declared" is a fact about the thing in
|
|
457
|
+
front of you, and only the second tells you whether refusing it matters. A
|
|
458
|
+
broker that silently dropped half a plugin would leave you believing you had
|
|
459
|
+
installed something you had not.
|
|
460
|
+
|
|
461
|
+
### Policy
|
|
462
|
+
|
|
463
|
+
```bash
|
|
464
|
+
agentctl policy policy.yaml # compile, then show what it says
|
|
465
|
+
agentctl run "..." --policy policy.yaml
|
|
466
|
+
```
|
|
467
|
+
|
|
468
|
+
Compiling is a separate step on purpose. Every error a policy can contain —
|
|
469
|
+
an undefined pool, a daily cap below the per-task cap, a misspelled effect
|
|
470
|
+
class — surfaces there, where you are watching, rather than mid-run where the
|
|
471
|
+
only safe response is to stop. A typo like `desctructive` would otherwise
|
|
472
|
+
compile into an artifact where `DESTRUCTIVE` has no rule at all, and the file
|
|
473
|
+
would still read like protection.
|
|
474
|
+
|
|
475
|
+
The cap is enforced *before* the run starts, because a budget check that runs
|
|
476
|
+
after the work is an audit:
|
|
477
|
+
|
|
478
|
+
```
|
|
479
|
+
refusing to start: budget exceeded: $1.5000 of $1.00 (daily)
|
|
480
|
+
```
|
|
481
|
+
|
|
482
|
+
### Stopping, coming back, and deciding
|
|
483
|
+
|
|
484
|
+
Every run is recorded in `~/.agentctl/runs.db`, so none of these needs a path
|
|
485
|
+
or an id copied out of the scrollback (`docs/0049`):
|
|
486
|
+
|
|
487
|
+
```bash
|
|
488
|
+
agentctl status # recent runs: done, died, paused, waiting on you
|
|
489
|
+
agentctl resume # continue the last run here (or: resume <id prefix>)
|
|
490
|
+
agentctl blocked # everything waiting on you, in every workspace
|
|
491
|
+
```
|
|
492
|
+
|
|
493
|
+
- **Ctrl-C** pauses after the current step, and the report says how to
|
|
494
|
+
continue. A second Ctrl-C stops at once. Either way nothing is lost: the
|
|
495
|
+
ledger reconciles whatever was in flight on resume.
|
|
496
|
+
- **A run that died** (a killed terminal, a closed laptop) shows as `died`,
|
|
497
|
+
and `agentctl resume` continues it.
|
|
498
|
+
- **`--wait 30m`** on `run` or `resume` waits out a rate limit and resumes
|
|
499
|
+
the same conversation, up to that long and at most five times. Without it,
|
|
500
|
+
the run ends with the command that continues it.
|
|
501
|
+
|
|
502
|
+
**Approvals.** A dangerous action (`rm -rf`, or a write outside the
|
|
503
|
+
workspace) asks first. With nobody at a terminal, it is **queued** rather
|
|
504
|
+
than refused: the agent is told it is waiting and not to repeat it.
|
|
505
|
+
|
|
506
|
+
```bash
|
|
507
|
+
agentctl approve <id> # let it run: the agent is told on resume
|
|
508
|
+
agentctl deny <id> # refuse it: likewise
|
|
509
|
+
agentctl resume # the approved action runs without asking again
|
|
510
|
+
```
|
|
511
|
+
|
|
512
|
+
**Unknown outcomes.** The gate fails closed when it cannot tell whether an
|
|
513
|
+
effect happened. That is a different question, with a different command:
|
|
514
|
+
|
|
515
|
+
```bash
|
|
516
|
+
agentctl show <id> # everything known about one
|
|
517
|
+
agentctl resolve <id> --landed # it did happen; do not re-run it
|
|
518
|
+
agentctl resolve <id> --retry # it did not; allow a retry
|
|
519
|
+
```
|
|
520
|
+
|
|
521
|
+
### Checking usage
|
|
522
|
+
|
|
523
|
+
```bash
|
|
524
|
+
agentctl ingest hook_telemetry.json # load Seam A telemetry
|
|
525
|
+
agentctl cost --by-deployment # spend, with pricing coverage
|
|
526
|
+
agentctl cost --conversation <id> # cost per completed task
|
|
527
|
+
```
|
|
528
|
+
|
|
529
|
+
LiteLLM reports `0.0` for endpoints it cannot price, so a total is never shown
|
|
530
|
+
without the share of calls it actually covers. `$0.0042 + unknown (1/3 priced)`
|
|
531
|
+
is the honest answer, and the ledger will not print a bare number instead.
|
|
532
|
+
|
|
533
|
+
There is no "probably fine" — `resolve` requires an explicit choice.
|
|
534
|
+
|
|
535
|
+
---
|
|
536
|
+
|
|
537
|
+
## How it works
|
|
538
|
+
|
|
539
|
+
Three seams with unequal powers, and that inequality is the whole design:
|
|
540
|
+
|
|
541
|
+
| Seam | Where | Can block | Can substitute |
|
|
542
|
+
|---|---|---|---|
|
|
543
|
+
| **A** | LiteLLM `CustomLogger` | yes | n/a |
|
|
544
|
+
| **B** | Event callback + `block_action` | yes | **no** |
|
|
545
|
+
| **C** | `ToolDefinition.executor` wrap | yes | **yes** |
|
|
546
|
+
|
|
547
|
+
Seam C also stamps idempotency keys — it is the only place a call can be
|
|
548
|
+
modified before it executes, which is what makes remote effects retry-safe.
|
|
549
|
+
|
|
550
|
+
Every decision is made by the gate and enforced at Seam B. **Seam C is not a
|
|
551
|
+
second gate** — it exists only to honour the one verdict Seam B cannot deliver:
|
|
552
|
+
handing back a recorded result instead of a refusal.
|
|
553
|
+
|
|
554
|
+
The consequence is graceful degradation. Lose Seam C and the system still fails
|
|
555
|
+
closed correctly; it only loses clean resume.
|
|
556
|
+
|
|
557
|
+
```
|
|
558
|
+
AGENT HARNESS ──▶ [Seam B: block] ──▶ [Seam C: substitute] ──▶ tool
|
|
559
|
+
│ │
|
|
560
|
+
│ EFFECT LEDGER SQLite · WAL · fsync · fenced
|
|
561
|
+
▼
|
|
562
|
+
LLM DATA PLANE ──▶ providers
|
|
563
|
+
│ telemetry
|
|
564
|
+
▼
|
|
565
|
+
CONTROL PLANE out-of-band · may fail
|
|
566
|
+
```
|
|
567
|
+
|
|
568
|
+
Full reasoning: [`docs/0008`](https://github.com/csdeepak/HandCode/blob/main/docs/0008-system-architecture-v1.md).
|
|
569
|
+
Diagrams: [`docs/0011`](https://github.com/csdeepak/HandCode/blob/main/docs/0011-request-flow-architecture.md).
|
|
570
|
+
|
|
571
|
+
---
|
|
572
|
+
|
|
573
|
+
## What it does not do yet
|
|
574
|
+
|
|
575
|
+
Stated plainly, because a safety layer that oversells itself is worse than none:
|
|
576
|
+
|
|
577
|
+
- **No sandbox.** `execute_bash` runs on the host. The capability matrix stands
|
|
578
|
+
in for one, and a 75-command corpus keeps it honest (`docs/0026`) — but it is
|
|
579
|
+
regex over a command string, not a shell parser, and an interpreter
|
|
580
|
+
(`python -c "..."`) is opaque to it by construction. Not enough for untrusted
|
|
581
|
+
tasks. Listed first because it is the limitation the others assume away.
|
|
582
|
+
- **`EXTERNAL` effects need an idempotency key.** With one declared, a retry is
|
|
583
|
+
safe (`docs/0020`). Without one — `send_email` and friends — a crash with the
|
|
584
|
+
outcome unknown still fails closed: there is no safe retry. But a model that
|
|
585
|
+
*saw* a failure and retries it is not stopped (`docs/0045` §4).
|
|
586
|
+
- **Only two effect kinds are chaos-tested** (git commit, file append). The
|
|
587
|
+
ten crash points are covered for those; `EXTERNAL` is tested separately.
|
|
588
|
+
- **A non-compliant remote voids the guarantee.** We trust the server to honour
|
|
589
|
+
the key, and nothing local can detect that it did not.
|
|
590
|
+
- **Single process.** Fencing is implemented and tested; multi-host is not
|
|
591
|
+
exercised.
|
|
592
|
+
- **No cache affinity, no dashboard.** The policy compiler landed in M7;
|
|
593
|
+
`pools` and `tiering` are declarations the LiteLLM proxy would act on,
|
|
594
|
+
not things `agentctl run` routes by (`docs/0030` §6).
|
|
595
|
+
- **One real provider only.** Verified against OpenRouter free-tier models
|
|
596
|
+
(`docs/0023`); it found two real bugs on the first attempt. Anthropic's
|
|
597
|
+
`/v1/messages` path and paid pricing coverage remain untested.
|
|
598
|
+
|
|
599
|
+
---
|
|
600
|
+
|
|
601
|
+
## Repository
|
|
602
|
+
|
|
603
|
+
```
|
|
604
|
+
agentctl/ the code
|
|
605
|
+
kernel/ in-band, must not fail. No network, no harness imports.
|
|
606
|
+
control/ out-of-band, may fail. Never imported by the kernel.
|
|
607
|
+
adapters/ harness-specific. The portability cost lives here.
|
|
608
|
+
gha.py the GitHub Action's logic; standard library only
|
|
609
|
+
action.yml the GitHub Action's run job; publish/ is its second job
|
|
610
|
+
docs/ the numbered document stream. Highest number is newest.
|
|
611
|
+
examples/ a workflow to copy into your repository
|
|
612
|
+
experiments/ reproducible crash experiments, zero cost
|
|
613
|
+
guide/ the user guide; website/ builds it into the docs site
|
|
614
|
+
tests/ 887 tests, including the ten-point chaos suite
|
|
615
|
+
verify.py one command that proves all of the above
|
|
616
|
+
```
|
|
617
|
+
|
|
618
|
+
The kernel/control boundary is enforced by a test
|
|
619
|
+
([`tests/test_boundaries.py`](https://github.com/csdeepak/HandCode/blob/main/tests/test_boundaries.py)). That single test is
|
|
620
|
+
what stops the architecture rotting.
|
|
621
|
+
|
|
622
|
+
**Start with [INDEX.md](https://github.com/csdeepak/HandCode/blob/main/INDEX.md)** — every document, numbered, newest last.
|
|
623
|
+
[`docs/0001`](https://github.com/csdeepak/HandCode/blob/main/docs/0001-project-charter.md) is the charter,
|
|
624
|
+
[`docs/0009`](https://github.com/csdeepak/HandCode/blob/main/docs/0009-open-questions-register.md) says what is still open.
|
|
625
|
+
|
|
626
|
+
---
|
|
627
|
+
|
|
628
|
+
## Method
|
|
629
|
+
|
|
630
|
+
Research before building. Falsify before committing. Every component got a
|
|
631
|
+
BUILD / CONFIGURE / SKIP verdict backed by primary sources before any code was
|
|
632
|
+
written — and Phase 0 deleted two components and downgraded two more, which was
|
|
633
|
+
the point of running it.
|
|
634
|
+
|
|
635
|
+
Every milestone so far has been finished by a bug only *execution* could find:
|
|
636
|
+
cp1252 output on Windows, a crashed process holding its own lease, CRLF
|
|
637
|
+
breaking the append probe, and an observation shape that validated on
|
|
638
|
+
assignment then failed three components later. None were visible by reading.
|
|
639
|
+
|
|
640
|
+
That is why `verify.py` costs nothing to run.
|
|
641
|
+
|
|
642
|
+
---
|
|
643
|
+
|
|
644
|
+
## A note on this repository's location
|
|
645
|
+
|
|
646
|
+
This is deliberately its own git repository. It was developed inside a
|
|
647
|
+
directory whose *parent* was already a repo with an unrelated remote, and
|
|
648
|
+
keeping it separate is what stopped its files being swept into that one.
|
|
649
|
+
|
|
650
|
+
If you clone into a similar layout, check `git rev-parse --show-toplevel`
|
|
651
|
+
before your first commit. The git probe learned the same lesson the hard way
|
|
652
|
+
(`docs/0019`): it now records the toplevel it was configured with and refuses
|
|
653
|
+
to act on a different one.
|
|
654
|
+
|
|
655
|
+
---
|
|
656
|
+
|
|
657
|
+
## License
|
|
658
|
+
|
|
659
|
+
MIT. See [LICENSE](https://github.com/csdeepak/HandCode/blob/main/LICENSE).
|