handcode 0.3.0rc1__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (67) hide show
  1. agentctl/__init__.py +0 -0
  2. agentctl/adapters/__init__.py +0 -0
  3. agentctl/adapters/litellm/__init__.py +9 -0
  4. agentctl/adapters/litellm/hook.py +49 -0
  5. agentctl/adapters/litellm/recorder.py +187 -0
  6. agentctl/adapters/openhands/__init__.py +169 -0
  7. agentctl/adapters/openhands/handoff.py +155 -0
  8. agentctl/adapters/openhands/seam_b.py +259 -0
  9. agentctl/adapters/openhands/seam_c.py +209 -0
  10. agentctl/cli.py +1450 -0
  11. agentctl/control/__init__.py +0 -0
  12. agentctl/control/cost/__init__.py +4 -0
  13. agentctl/control/cost/ledger.py +210 -0
  14. agentctl/control/dash.py +697 -0
  15. agentctl/control/keys.py +440 -0
  16. agentctl/control/matrix/__init__.py +0 -0
  17. agentctl/control/matrix/data/tools.yaml +149 -0
  18. agentctl/control/policy/__init__.py +10 -0
  19. agentctl/control/policy/compile.py +258 -0
  20. agentctl/control/policy/data/policy.compiled.json +38 -0
  21. agentctl/control/policy/data/policy.yaml +46 -0
  22. agentctl/control/probe.py +399 -0
  23. agentctl/control/providers.py +293 -0
  24. agentctl/control/proxy.py +536 -0
  25. agentctl/control/proxyenv.py +309 -0
  26. agentctl/control/replay/__init__.py +14 -0
  27. agentctl/control/replay/cassette.py +281 -0
  28. agentctl/control/replay/server.py +109 -0
  29. agentctl/demo/__init__.py +214 -0
  30. agentctl/demo/child.py +84 -0
  31. agentctl/demo/mock.py +79 -0
  32. agentctl/demo/tool.py +62 -0
  33. agentctl/gha.py +488 -0
  34. agentctl/kernel/__init__.py +0 -0
  35. agentctl/kernel/classify.py +170 -0
  36. agentctl/kernel/gate.py +391 -0
  37. agentctl/kernel/hook.py +229 -0
  38. agentctl/kernel/ledger/__init__.py +0 -0
  39. agentctl/kernel/ledger/models.py +160 -0
  40. agentctl/kernel/ledger/schema.sql +62 -0
  41. agentctl/kernel/ledger/store.py +596 -0
  42. agentctl/kernel/paths.py +203 -0
  43. agentctl/kernel/policy.py +160 -0
  44. agentctl/kernel/reconcile/__init__.py +31 -0
  45. agentctl/kernel/reconcile/base.py +106 -0
  46. agentctl/kernel/reconcile/external.py +137 -0
  47. agentctl/kernel/reconcile/filesystem.py +162 -0
  48. agentctl/kernel/reconcile/git.py +162 -0
  49. agentctl/runtime/__init__.py +20 -0
  50. agentctl/runtime/citations.py +179 -0
  51. agentctl/runtime/config.py +97 -0
  52. agentctl/runtime/doctor.py +335 -0
  53. agentctl/runtime/init.py +148 -0
  54. agentctl/runtime/lease.py +143 -0
  55. agentctl/runtime/orchestrate.py +187 -0
  56. agentctl/runtime/plugins.py +130 -0
  57. agentctl/runtime/report.py +361 -0
  58. agentctl/runtime/runner.py +787 -0
  59. agentctl/runtime/runs.py +191 -0
  60. agentctl/runtime/subagent.py +274 -0
  61. agentctl/runtime/tools.py +350 -0
  62. handcode-0.3.0rc1.dist-info/METADATA +659 -0
  63. handcode-0.3.0rc1.dist-info/RECORD +67 -0
  64. handcode-0.3.0rc1.dist-info/WHEEL +5 -0
  65. handcode-0.3.0rc1.dist-info/entry_points.txt +3 -0
  66. handcode-0.3.0rc1.dist-info/licenses/LICENSE +21 -0
  67. handcode-0.3.0rc1.dist-info/top_level.txt +1 -0
@@ -0,0 +1,659 @@
1
+ Metadata-Version: 2.4
2
+ Name: handcode
3
+ Version: 0.3.0rc1
4
+ Summary: A coding agent whose work survives crashes, restarts and provider switches
5
+ Author: csdeepak
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/csdeepak/HandCode
8
+ Project-URL: Documentation, https://github.com/csdeepak/HandCode/tree/main/guide
9
+ Project-URL: Changelog, https://github.com/csdeepak/HandCode/blob/main/CHANGELOG.md
10
+ Project-URL: Issues, https://github.com/csdeepak/HandCode/issues
11
+ Keywords: agent,coding-agent,openhands,litellm,llm,crash-recovery,idempotency
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Environment :: Console
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Operating System :: OS Independent
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: Programming Language :: Python :: 3.13
19
+ Classifier: Topic :: Software Development
20
+ Requires-Python: >=3.12
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Requires-Dist: pyyaml>=6.0
24
+ Provides-Extra: openhands
25
+ Requires-Dist: openhands-sdk>=1.45.0; extra == "openhands"
26
+ Provides-Extra: proxy
27
+ Requires-Dist: litellm[proxy]>=1.100.0; extra == "proxy"
28
+ Provides-Extra: dev
29
+ Requires-Dist: pytest>=8.0; extra == "dev"
30
+ Requires-Dist: markdown-it-py>=3; extra == "dev"
31
+ Dynamic: license-file
32
+
33
+ # HandCode
34
+
35
+ [![CI](https://github.com/csdeepak/HandCode/actions/workflows/ci.yml/badge.svg)](https://github.com/csdeepak/HandCode/actions/workflows/ci.yml)
36
+
37
+ **HandCode keeps a coding agent's work safe across crashes, restarts and
38
+ provider switches.** Its command is `agentctl`.
39
+
40
+ When an agent's process dies in the middle of a task and you resume it, the
41
+ agent framework re-runs whatever was in flight, so a commit can happen twice.
42
+ agentctl records every action before it runs. A resumed run never repeats one
43
+ that already happened, and when agentctl cannot tell whether it did, it stops
44
+ and asks you instead of guessing.
45
+
46
+ It wraps the [OpenHands](https://github.com/OpenHands/software-agent-sdk)
47
+ agent SDK and works with any provider LiteLLM supports, free tiers included.
48
+
49
+ **Documentation:** <https://csdeepak.github.io/HandCode/>
50
+
51
+ ## See it in a minute
52
+
53
+ ```bash
54
+ agentctl demo # no API key, no network, $0
55
+ ```
56
+
57
+ ```
58
+ 1. Plain OpenHands, no agentctl
59
+ the agent committed, then its process was killed ... 1 new commit
60
+ resumed ............................................ 2 new commits <- the same commit, twice
61
+
62
+ 2. With agentctl
63
+ the agent committed, then its process was killed ... 1 new commit
64
+ resumed ............................................ 1 new commit <- once
65
+ the ledger: the commit is OBSERVED (git probe: LANDED) · 0 waiting on you
66
+ ```
67
+
68
+ Real git, real process death and the real SDK. Only the model is scripted,
69
+ which is what makes it free and identical on every OS.
70
+
71
+ ## Quickstart
72
+
73
+ Python 3.12+, `git`, and one API key (a free OpenRouter or Gemini key works).
74
+
75
+ ```bash
76
+ git clone https://github.com/csdeepak/HandCode && cd HandCode
77
+ python3 -m venv .venv && source .venv/bin/activate # Windows: see guide/quickstart.md
78
+ pip install uv && uv pip install -e ".[openhands]" # uv: seconds, where pip takes minutes
79
+
80
+ agentctl demo # see what it is for
81
+ agentctl init # one key, one checked default model
82
+ cd ../your-project
83
+ agentctl run "fix the failing date test" --accept "python -m pytest -q"
84
+ ```
85
+
86
+ Every run ends with a report of what was checked, what changed, what it cost,
87
+ and what needs you:
88
+
89
+ ```
90
+ outcome PASS `python -m pytest -q` exited 0 (run by agentctl after the agent finished)
91
+ changed 1 file +5 -1
92
+ used 9 requests · 43.7K tokens · $0.00 (free-tier model) · 12s
93
+ actions 8 actions: 3 reads, 3 file writes, 2 commands
94
+ needs you nothing
95
+ ```
96
+
97
+ The **[user guide](https://github.com/csdeepak/HandCode/blob/main/guide/README.md)** covers the
98
+ [quickstart](https://github.com/csdeepak/HandCode/blob/main/guide/quickstart.md) step by step (Windows included),
99
+ [concepts in plain words](https://github.com/csdeepak/HandCode/blob/main/guide/concepts.md),
100
+ [troubleshooting by the exact text you see](https://github.com/csdeepak/HandCode/blob/main/guide/troubleshooting.md), and an
101
+ [FAQ](https://github.com/csdeepak/HandCode/blob/main/guide/faq.md).
102
+
103
+ ## What it protects you from
104
+
105
+ | | Without agentctl | With it |
106
+ |---|---|---|
107
+ | The process dies mid-action, then you resume | The in-flight action runs again | It runs at most once. If it already landed, the agent gets its result back |
108
+ | You cannot tell whether something happened | It is guessed | It is checked (did HEAD move?), or you are asked |
109
+ | Ctrl-C, a closed laptop, a killed terminal | Start again | `agentctl resume` |
110
+ | A rate limit mid-task | The run dies | `--wait 30m` waits it out and resumes, or `--pool` routes to another provider |
111
+ | A dangerous command with nobody watching | It runs, or the run hangs on a prompt | It is queued for `agentctl approve` / `deny` |
112
+ | "Done!" from the agent | Taken on trust | `--accept` runs your tests and reports PASS or FAIL |
113
+ | Two terminals resuming the same run | Both drive it | The second is refused and told which process holds it |
114
+
115
+ **What it does not do:** it is not a sandbox. Commands run on your machine.
116
+ agentctl stops actions repeating and asks before dangerous ones, but it does
117
+ not contain what an allowed command does. For a task you would not trust your
118
+ shell with, run it in the [container image](https://github.com/csdeepak/HandCode/blob/main/guide/quickstart.md#6-run-it-in-a-container-recommended-for-tasks-you-would-not-trust-your-shell-with),
119
+ where only the mounted repository is reachable
120
+ ([concepts](https://github.com/csdeepak/HandCode/blob/main/guide/concepts.md#what-agentctl-does-not-do)).
121
+
122
+ **On GitHub.** Type a task in your repository's Actions tab and get a pull
123
+ request back, with the report as its description. Your key stays in your
124
+ repository's secrets, and the job that runs the agent holds no token that can
125
+ write ([the GitHub Action](https://github.com/csdeepak/HandCode/blob/main/guide/github-action.md)).
126
+ It has run on GitHub with a real model and opened a real pull request.
127
+
128
+ **Status.** 887 tests, including a chaos suite that kills the process at ten
129
+ points in the protocol, plus `verify.py`'s end-to-end crash experiments. All
130
+ run at no cost, on Linux and Windows in CI. Every feature here has also been
131
+ run at least once against a real provider. The design history, with what was
132
+ measured and what was found, is the numbered [`docs/`](https://github.com/csdeepak/HandCode/blob/main/INDEX.md) stream.
133
+
134
+ ---
135
+
136
+ # Reference
137
+
138
+ ## The problem, in one example
139
+
140
+ An agent runs `git commit`. The process dies before the result is recorded. On
141
+ resume, OpenHands re-drives the pending action through the real executor — and
142
+ commits again.
143
+
144
+ That is not a hypothesis. `docs/0014` reproduces it: **1 effect before the
145
+ crash, 2 after resume.**
146
+
147
+ With this layer installed:
148
+
149
+ | | M2a | M4 | M2b |
150
+ |---|---|---|---|
151
+ | Duplicate effect | none | none | none |
152
+ | Ledger state | `BLOCKED` | `COMMITTED` | `OBSERVED` |
153
+ | Human needed | yes | no | no |
154
+ | Agent receives | rejection | rejection | **the result** |
155
+
156
+ ---
157
+
158
+ ## Verify it yourself
159
+
160
+ Requires **Python ≥ 3.12** (the OpenHands SDK does) and `git` on PATH. On an
161
+ older Python the install fails with a long list of `Requires-Python >=3.12`
162
+ lines that never names the cause, so check first.
163
+
164
+ **Windows** (PowerShell):
165
+
166
+ ```powershell
167
+ py -3.12 -m venv .venv
168
+ .venv\Scripts\Activate.ps1
169
+ pip install -e ".[dev,openhands]"
170
+ python verify.py
171
+ ```
172
+
173
+ Clone into a short path such as `C:\src\HandCode`. One file the install
174
+ unpacks sits 139 characters deep inside `.venv`, so a clone path longer than
175
+ about 120 characters hits Windows' 260-character limit and pip fails partway
176
+ (`docs/0044` N3).
177
+
178
+ **macOS / Linux:**
179
+
180
+ ```bash
181
+ python3 -m venv .venv # Ubuntu: sudo apt install python3.12-venv first
182
+ source .venv/bin/activate
183
+ pip install -e ".[dev,openhands]"
184
+ python verify.py
185
+ ```
186
+
187
+ **Zero cost** — everything runs against a local mock provider. No API key, no
188
+ network, no tokens. Takes about five minutes. With this install, two of the
189
+ ten checks (Seam A and the full stack) are **SKIPPED**, because they need the
190
+ proxy extra below; that is expected, and `verify.py` says how to run them.
191
+
192
+ If a dependency has since shipped something incompatible, install the exact set
193
+ that is known to pass:
194
+
195
+ ```bash
196
+ pip install -e ".[dev,openhands]" -c constraints.txt
197
+ ```
198
+
199
+ CI runs both forms on Linux and Windows, and runs `verify.py` itself — so the
200
+ badge above means the correctness claim reproduces on a machine that is not
201
+ mine, which is the only version of that claim worth anything.
202
+
203
+ ### The proxy extra needs two steps (only if you run the proxy yourself)
204
+
205
+ `agentctl run --pool` and `agentctl proxy up` do not need any of this: they
206
+ keep LiteLLM in a separate environment of their own. The steps below are for
207
+ running the full-stack checks in `verify.py`, or a proxy by hand.
208
+
209
+ `litellm[proxy]` declares `mcp<2.0`; the OpenHands SDK needs `fastmcp` and so
210
+ needs `mcp>=2`. They are incompatible on paper and work in practice, so the
211
+ override is explicit rather than hidden in a version range pip would refuse:
212
+
213
+ ```bash
214
+ pip install -e ".[dev,openhands,proxy]"
215
+ pip install --upgrade "mcp>=2.2.0" "fastmcp>=4.0.3"
216
+ ```
217
+
218
+ `fastmcp` is named explicitly because the proxy install leaves only
219
+ `fastmcp-slim`, which has no client support — and the SDK does
220
+ `from fastmcp import Client` on nearly every import path.
221
+
222
+ Only Seam A and the full-stack demo need this. Everything else — including the
223
+ whole correctness core — runs without the proxy.
224
+
225
+ ---
226
+
227
+ ## Set up your keys
228
+
229
+ ```bash
230
+ agentctl keys --init # writes keys.env listing every provider + how to get one
231
+ agentctl keys # what is set, and where to get the rest
232
+ agentctl keys --install-hook # refuse any commit containing one of your keys
233
+ agentctl keys --check # does every key actually work? costs no tokens
234
+ agentctl dash # failover, effects, spend, policy on one screen
235
+ ```
236
+
237
+ `keys.env` is gitignored **before** it is written, and no command ever prints a
238
+ value. If it ever becomes tracked by git, every command says so loudly and
239
+ tells you to rotate — a `.gitignore` entry added after a file is tracked does
240
+ nothing (`docs/0032`).
241
+
242
+ **One key is enough to start.** A daily cap on it stops work until it resets;
243
+ a key at a **second provider** survives that, and survives an outage too:
244
+
245
+ ```
246
+ OPENROUTER_API_KEY=sk-or-v1-... # one provider
247
+ GEMINI_API_KEY=... # a second provider
248
+ ```
249
+
250
+ Extra keys at the *same* provider are accepted with any suffix (`_2`, `_WORK`)
251
+ and each becomes its own deployment in the pool (`docs/0033`). **Whether that
252
+ is a separate quota is unverified**: OpenRouter's own limits page says extra
253
+ accounts do not change rate limits. Pooling several free accounts may also be
254
+ against a provider's terms. Check both before relying on it (`docs/0042`
255
+ §4.C). `agentctl dash` says which state you are actually in.
256
+
257
+ ## Use it
258
+
259
+ ```bash
260
+ agentctl init # once: one key, one checked default model
261
+ cd myproject
262
+ agentctl run "add type hints to utils.py" # no flags needed
263
+ ```
264
+
265
+ `init` uses a key you already have in the environment or a keys file, or asks
266
+ for one and stores it in `~/.agentctl/keys.env`, outside every repository and
267
+ never echoed. It then sends **one** completion and records the model that
268
+ actually answered in `~/.agentctl/config.toml`. A paid provider is not called
269
+ unless you pass `--check-paid`.
270
+
271
+ `run` takes its model from `--model`, then `AGENTCTL_MODEL`, then that config,
272
+ then the first provider you hold a key for, and prints which one it used.
273
+
274
+ Every run ends with a report:
275
+
276
+ ```
277
+ outcome PASS `python -m pytest -q` exited 0 (run by agentctl after the agent finished)
278
+ changed 1 file +5 -1
279
+ stats.py +5 -1
280
+ agent said "The median function in stats.py has been fixed to correctly handle..."
281
+ used 9 requests · 43.7K tokens · $0.00 (free-tier model) · 12s
282
+ actions 8 actions: 3 reads, 3 file writes, 2 commands
283
+ needs you nothing
284
+ ```
285
+
286
+ - **`outcome` is only what was checked.** Pass `--accept "<command>"` (your
287
+ test suite, usually) and agentctl runs it itself once the agent has
288
+ finished, outside the agent's loop: exit 0 is PASS. Without it the outcome
289
+ reads `not checked`, never an implied success.
290
+ - **`changed`** is measured with git against where the run started, commits
291
+ included.
292
+ - **`used`** says what a cost figure can be trusted for. Free-tier models read
293
+ `$0.00`, and list prices that a free key is not billed are labelled as such.
294
+ - **The exit code is 0** when nothing failed a check and nothing is waiting on
295
+ you, and 1 otherwise.
296
+
297
+ ```bash
298
+ agentctl doctor --workspace ./myproject # check everything BEFORE you spend
299
+ ```
300
+
301
+ `doctor` checks packages, the SDK import chain, git, your provider keys, the
302
+ policy, and whether the workspace has uncommitted changes the agent is about
303
+ to edit.
304
+
305
+ > **Correction (2026-09-21).** This paragraph used to say OpenRouter exposes no
306
+ > remaining-free-request counter on any endpoint (`docs/0031`). That is false.
307
+ > `GET /api/v1/key` returns `free_model_daily_requests` with `used`, `limit`
308
+ > and `remaining`, and it costs nothing — measured against six live accounts,
309
+ > all reporting `limit=50`. `probe.py` was already calling the sibling endpoint
310
+ > `/api/v1/auth/key` the whole time. `docs/0038` §9.4 records the measurement;
311
+ > surfacing it is scheduled work.
312
+
313
+ Real tools (bash, read, write), a real model, every effect classified and
314
+ ledgered. Crash it and re-run with the conversation id the run printed at the
315
+ start. Work already done is not repeated:
316
+
317
+ ```bash
318
+ agentctl run "" --workspace ./myproject --resume <conversation-id>
319
+ ```
320
+
321
+ The task is `""` because the conversation already holds it.
322
+
323
+ Only one process may drive a conversation. If the run that held it crashed on
324
+ this machine, `--resume` sees that it is gone and takes over. If it is still
325
+ running, `--resume` refuses and names it, because two drivers of one
326
+ conversation is exactly what the ledger exists to prevent. `--takeover`
327
+ overrides that, for a holder you know is dead but this machine cannot check
328
+ (`docs/0046`).
329
+
330
+ ### Run it again for nothing
331
+
332
+ ```bash
333
+ agentctl run "..." --workspace ./app --record session.jsonl # once, for real
334
+ agentctl run "" --workspace ./app --replay session.jsonl # $0.00, offline
335
+ ```
336
+
337
+ `--replay` serves every completion from the recording: no API key, no network,
338
+ no tokens, no sampling. A request the cassette does not contain is answered
339
+ with a 502 naming the turn that diverged, never with a plausible-looking
340
+ completion — a replay that cannot fail would not be worth running.
341
+
342
+ Replay pins the model, not the world: tool output feeds the next request, so
343
+ the workspace has to start where the recording did (`docs/0029`).
344
+
345
+ Two safety properties, deliberately separate — the **gate** stops an effect
346
+ happening *twice*; `--confirm-destructive` (on by default) stops one happening
347
+ *at all* without a human saying yes. It asks on two grounds: the effect is
348
+ destructive, **or** it writes outside the workspace. `echo x > ~/.bashrc` is an
349
+ ordinary idempotent write that is simply none of the agent's business
350
+ (`docs/0027`).
351
+
352
+ "Twice" means *by replay*. Once a result — success or failure — is in the
353
+ history the model reads, an identical later call is the model deciding to run
354
+ it again: the edit, test, re-test loop. It runs. A crash still never causes a
355
+ repeat (`docs/0045`).
356
+
357
+ ### Or embed it
358
+
359
+ ```python
360
+ from agentctl.adapters.openhands import protect
361
+
362
+ guard = protect(
363
+ ledger="./ledger.db",
364
+ conversation_id=str(conversation_id),
365
+ tools={"commit": CommitTool}, # gated at Seam C
366
+ repo_root="./workspace", # for the git probe
367
+ takeover=previous_holder_is_dead, # NEVER for a live one (docs/0046)
368
+ )
369
+
370
+ conv = Conversation(agent=agent, callbacks=[guard.seam_b], ...)
371
+ guard.attach(conv) # required, or the gate is inert
372
+ ```
373
+
374
+ Omit `tools` to run Seam B only: still correct, but an already-landed effect is
375
+ blocked rather than resumed cleanly.
376
+
377
+ ### Failing over
378
+
379
+ ```bash
380
+ agentctl run "..." --pool # starts the managed pool if it is not running
381
+ agentctl proxy status # running? answering? where is its log?
382
+ agentctl proxy down # stop it
383
+ ```
384
+
385
+ `--pool` (or `agentctl proxy up`) runs LiteLLM in **its own environment**,
386
+ `~/.agentctl/proxy-env`, created the first time (a few minutes) and reused
387
+ after that. Your install never needs the proxy extra.
388
+ - It verifies every model id with one completion and leaves out what cannot
389
+ serve, naming it.
390
+ - It runs in the background and outlives the command that started it.
391
+
392
+ `agentctl proxy --out ./proxy` still writes a config and start scripts, if you
393
+ would rather run the proxy yourself (`docs/0047`).
394
+
395
+ A pool over **one** provider key survives a transient upstream overload and a
396
+ per-model limit. It does **not** survive an account-wide daily cap — three
397
+ `:free` models on one key share one quota. The command counts credentials, not
398
+ deployments, and says so.
399
+
400
+ ### Choosing one source
401
+
402
+ ```bash
403
+ agentctl models # every source, and what choosing it costs
404
+ agentctl run "..." --source gemini --base-url http://localhost:4000
405
+ ```
406
+
407
+ `pool` is the default and the widest thing you can ask for. A source group is
408
+ narrower **on purpose**, so every row says what it gives up:
409
+
410
+ ```
411
+ pool 48 deployments 24 accounts the default
412
+ pool-openrouter 18 deployments 6 accounts gives up 30 of 48
413
+ pool-gemini 6 deployments 6 accounts gives up 42 of 48
414
+ ```
415
+
416
+ That line is the feature. Narrowing to one source means an account-wide daily
417
+ cap has less to fail over to — the exact failure the multi-account design
418
+ exists to escape (`docs/0033`). `--verify` marks which sources can actually
419
+ serve, which is not the same as which ones you hold keys for.
420
+
421
+ ### Delegating a read
422
+
423
+ ```bash
424
+ agentctl subagent --init # writes an example definition
425
+ agentctl subagent reviewer "what does the gate do?"
426
+ ```
427
+
428
+ Claude Code's Markdown frontmatter format, loaded through the SDK's own
429
+ `AgentDefinition`. A subagent may hold **`read_file` and nothing else** — no
430
+ shell, no writes, no MCP.
431
+
432
+ That is not a starter limitation, it is the whole safety argument. Multi-agent
433
+ is skipped because four correctness mechanisms assume a single writer
434
+ (`docs/0038` §4.2); a subagent that cannot produce an effect needs none of
435
+ them. So the restriction is enforced twice — the definition is validated, and
436
+ the tool list handed to the agent is built from a constant rather than from
437
+ the definition, because a property that depends on one function returning
438
+ correctly is one refactor from gone.
439
+
440
+ ### Installing a plugin, one capability at a time
441
+
442
+ ```bash
443
+ agentctl plugins ./some-plugin # the audit. Installs nothing.
444
+ ```
445
+
446
+ A Claude Code plugin can carry agents, hooks, commands, skills and MCP
447
+ servers. **Five of those six are refused**, and each refusal is counted:
448
+
449
+ ```
450
+ ADMITTED 1 read-only agent(s)
451
+ REJECTED 1 agent(s) that could change the world
452
+ REFUSED mcp_config 2 declared
453
+ commands 1 declared
454
+ ```
455
+
456
+ "declines MCP servers" is a policy; "2 declared" is a fact about the thing in
457
+ front of you, and only the second tells you whether refusing it matters. A
458
+ broker that silently dropped half a plugin would leave you believing you had
459
+ installed something you had not.
460
+
461
+ ### Policy
462
+
463
+ ```bash
464
+ agentctl policy policy.yaml # compile, then show what it says
465
+ agentctl run "..." --policy policy.yaml
466
+ ```
467
+
468
+ Compiling is a separate step on purpose. Every error a policy can contain —
469
+ an undefined pool, a daily cap below the per-task cap, a misspelled effect
470
+ class — surfaces there, where you are watching, rather than mid-run where the
471
+ only safe response is to stop. A typo like `desctructive` would otherwise
472
+ compile into an artifact where `DESTRUCTIVE` has no rule at all, and the file
473
+ would still read like protection.
474
+
475
+ The cap is enforced *before* the run starts, because a budget check that runs
476
+ after the work is an audit:
477
+
478
+ ```
479
+ refusing to start: budget exceeded: $1.5000 of $1.00 (daily)
480
+ ```
481
+
482
+ ### Stopping, coming back, and deciding
483
+
484
+ Every run is recorded in `~/.agentctl/runs.db`, so none of these needs a path
485
+ or an id copied out of the scrollback (`docs/0049`):
486
+
487
+ ```bash
488
+ agentctl status # recent runs: done, died, paused, waiting on you
489
+ agentctl resume # continue the last run here (or: resume <id prefix>)
490
+ agentctl blocked # everything waiting on you, in every workspace
491
+ ```
492
+
493
+ - **Ctrl-C** pauses after the current step, and the report says how to
494
+ continue. A second Ctrl-C stops at once. Either way nothing is lost: the
495
+ ledger reconciles whatever was in flight on resume.
496
+ - **A run that died** (a killed terminal, a closed laptop) shows as `died`,
497
+ and `agentctl resume` continues it.
498
+ - **`--wait 30m`** on `run` or `resume` waits out a rate limit and resumes
499
+ the same conversation, up to that long and at most five times. Without it,
500
+ the run ends with the command that continues it.
501
+
502
+ **Approvals.** A dangerous action (`rm -rf`, or a write outside the
503
+ workspace) asks first. With nobody at a terminal, it is **queued** rather
504
+ than refused: the agent is told it is waiting and not to repeat it.
505
+
506
+ ```bash
507
+ agentctl approve <id> # let it run: the agent is told on resume
508
+ agentctl deny <id> # refuse it: likewise
509
+ agentctl resume # the approved action runs without asking again
510
+ ```
511
+
512
+ **Unknown outcomes.** The gate fails closed when it cannot tell whether an
513
+ effect happened. That is a different question, with a different command:
514
+
515
+ ```bash
516
+ agentctl show <id> # everything known about one
517
+ agentctl resolve <id> --landed # it did happen; do not re-run it
518
+ agentctl resolve <id> --retry # it did not; allow a retry
519
+ ```
520
+
521
+ ### Checking usage
522
+
523
+ ```bash
524
+ agentctl ingest hook_telemetry.json # load Seam A telemetry
525
+ agentctl cost --by-deployment # spend, with pricing coverage
526
+ agentctl cost --conversation <id> # cost per completed task
527
+ ```
528
+
529
+ LiteLLM reports `0.0` for endpoints it cannot price, so a total is never shown
530
+ without the share of calls it actually covers. `$0.0042 + unknown (1/3 priced)`
531
+ is the honest answer, and the ledger will not print a bare number instead.
532
+
533
+ There is no "probably fine" — `resolve` requires an explicit choice.
534
+
535
+ ---
536
+
537
+ ## How it works
538
+
539
+ Three seams with unequal powers, and that inequality is the whole design:
540
+
541
+ | Seam | Where | Can block | Can substitute |
542
+ |---|---|---|---|
543
+ | **A** | LiteLLM `CustomLogger` | yes | n/a |
544
+ | **B** | Event callback + `block_action` | yes | **no** |
545
+ | **C** | `ToolDefinition.executor` wrap | yes | **yes** |
546
+
547
+ Seam C also stamps idempotency keys — it is the only place a call can be
548
+ modified before it executes, which is what makes remote effects retry-safe.
549
+
550
+ Every decision is made by the gate and enforced at Seam B. **Seam C is not a
551
+ second gate** — it exists only to honour the one verdict Seam B cannot deliver:
552
+ handing back a recorded result instead of a refusal.
553
+
554
+ The consequence is graceful degradation. Lose Seam C and the system still fails
555
+ closed correctly; it only loses clean resume.
556
+
557
+ ```
558
+ AGENT HARNESS ──▶ [Seam B: block] ──▶ [Seam C: substitute] ──▶ tool
559
+ │ │
560
+ │ EFFECT LEDGER SQLite · WAL · fsync · fenced
561
+ ▼
562
+ LLM DATA PLANE ──▶ providers
563
+ │ telemetry
564
+ ▼
565
+ CONTROL PLANE out-of-band · may fail
566
+ ```
567
+
568
+ Full reasoning: [`docs/0008`](https://github.com/csdeepak/HandCode/blob/main/docs/0008-system-architecture-v1.md).
569
+ Diagrams: [`docs/0011`](https://github.com/csdeepak/HandCode/blob/main/docs/0011-request-flow-architecture.md).
570
+
571
+ ---
572
+
573
+ ## What it does not do yet
574
+
575
+ Stated plainly, because a safety layer that oversells itself is worse than none:
576
+
577
+ - **No sandbox.** `execute_bash` runs on the host. The capability matrix stands
578
+ in for one, and a 75-command corpus keeps it honest (`docs/0026`) — but it is
579
+ regex over a command string, not a shell parser, and an interpreter
580
+ (`python -c "..."`) is opaque to it by construction. Not enough for untrusted
581
+ tasks. Listed first because it is the limitation the others assume away.
582
+ - **`EXTERNAL` effects need an idempotency key.** With one declared, a retry is
583
+ safe (`docs/0020`). Without one — `send_email` and friends — a crash with the
584
+ outcome unknown still fails closed: there is no safe retry. But a model that
585
+ *saw* a failure and retries it is not stopped (`docs/0045` §4).
586
+ - **Only two effect kinds are chaos-tested** (git commit, file append). The
587
+ ten crash points are covered for those; `EXTERNAL` is tested separately.
588
+ - **A non-compliant remote voids the guarantee.** We trust the server to honour
589
+ the key, and nothing local can detect that it did not.
590
+ - **Single process.** Fencing is implemented and tested; multi-host is not
591
+ exercised.
592
+ - **No cache affinity, no dashboard.** The policy compiler landed in M7;
593
+ `pools` and `tiering` are declarations the LiteLLM proxy would act on,
594
+ not things `agentctl run` routes by (`docs/0030` §6).
595
+ - **One real provider only.** Verified against OpenRouter free-tier models
596
+ (`docs/0023`); it found two real bugs on the first attempt. Anthropic's
597
+ `/v1/messages` path and paid pricing coverage remain untested.
598
+
599
+ ---
600
+
601
+ ## Repository
602
+
603
+ ```
604
+ agentctl/ the code
605
+ kernel/ in-band, must not fail. No network, no harness imports.
606
+ control/ out-of-band, may fail. Never imported by the kernel.
607
+ adapters/ harness-specific. The portability cost lives here.
608
+ gha.py the GitHub Action's logic; standard library only
609
+ action.yml the GitHub Action's run job; publish/ is its second job
610
+ docs/ the numbered document stream. Highest number is newest.
611
+ examples/ a workflow to copy into your repository
612
+ experiments/ reproducible crash experiments, zero cost
613
+ guide/ the user guide; website/ builds it into the docs site
614
+ tests/ 887 tests, including the ten-point chaos suite
615
+ verify.py one command that proves all of the above
616
+ ```
617
+
618
+ The kernel/control boundary is enforced by a test
619
+ ([`tests/test_boundaries.py`](https://github.com/csdeepak/HandCode/blob/main/tests/test_boundaries.py)). That single test is
620
+ what stops the architecture rotting.
621
+
622
+ **Start with [INDEX.md](https://github.com/csdeepak/HandCode/blob/main/INDEX.md)** — every document, numbered, newest last.
623
+ [`docs/0001`](https://github.com/csdeepak/HandCode/blob/main/docs/0001-project-charter.md) is the charter,
624
+ [`docs/0009`](https://github.com/csdeepak/HandCode/blob/main/docs/0009-open-questions-register.md) says what is still open.
625
+
626
+ ---
627
+
628
+ ## Method
629
+
630
+ Research before building. Falsify before committing. Every component got a
631
+ BUILD / CONFIGURE / SKIP verdict backed by primary sources before any code was
632
+ written — and Phase 0 deleted two components and downgraded two more, which was
633
+ the point of running it.
634
+
635
+ Every milestone so far has been finished by a bug only *execution* could find:
636
+ cp1252 output on Windows, a crashed process holding its own lease, CRLF
637
+ breaking the append probe, and an observation shape that validated on
638
+ assignment then failed three components later. None were visible by reading.
639
+
640
+ That is why `verify.py` costs nothing to run.
641
+
642
+ ---
643
+
644
+ ## A note on this repository's location
645
+
646
+ This is deliberately its own git repository. It was developed inside a
647
+ directory whose *parent* was already a repo with an unrelated remote, and
648
+ keeping it separate is what stopped its files being swept into that one.
649
+
650
+ If you clone into a similar layout, check `git rev-parse --show-toplevel`
651
+ before your first commit. The git probe learned the same lesson the hard way
652
+ (`docs/0019`): it now records the toplevel it was configured with and refuses
653
+ to act on a different one.
654
+
655
+ ---
656
+
657
+ ## License
658
+
659
+ MIT. See [LICENSE](https://github.com/csdeepak/HandCode/blob/main/LICENSE).