corecoder 0.7.0__tar.gz → 0.8.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (73) hide show
  1. {corecoder-0.7.0 → corecoder-0.8.0}/.github/workflows/ci.yml +6 -0
  2. {corecoder-0.7.0 → corecoder-0.8.0}/PKG-INFO +20 -20
  3. {corecoder-0.7.0 → corecoder-0.8.0}/README.md +18 -18
  4. {corecoder-0.7.0 → corecoder-0.8.0}/README_CN.md +18 -18
  5. {corecoder-0.7.0 → corecoder-0.8.0}/article/00-index.md +3 -3
  6. {corecoder-0.7.0 → corecoder-0.8.0}/article/00-index_EN.md +3 -3
  7. {corecoder-0.7.0 → corecoder-0.8.0}/article/01-the-loop.md +1 -1
  8. {corecoder-0.7.0 → corecoder-0.8.0}/article/01-the-loop_EN.md +1 -1
  9. {corecoder-0.7.0 → corecoder-0.8.0}/article/02-tools.md +5 -5
  10. {corecoder-0.7.0 → corecoder-0.8.0}/article/02-tools_EN.md +5 -5
  11. {corecoder-0.7.0 → corecoder-0.8.0}/article/03-llm-and-cost.md +2 -2
  12. {corecoder-0.7.0 → corecoder-0.8.0}/article/03-llm-and-cost_EN.md +2 -2
  13. {corecoder-0.7.0 → corecoder-0.8.0}/article/04-context.md +1 -1
  14. {corecoder-0.7.0 → corecoder-0.8.0}/article/04-context_EN.md +1 -1
  15. {corecoder-0.7.0 → corecoder-0.8.0}/article/05-parallel-and-subagents.md +1 -1
  16. {corecoder-0.7.0 → corecoder-0.8.0}/article/05-parallel-and-subagents_EN.md +1 -1
  17. {corecoder-0.7.0 → corecoder-0.8.0}/article/06-session-and-cli.md +1 -1
  18. {corecoder-0.7.0 → corecoder-0.8.0}/article/06-session-and-cli_EN.md +1 -1
  19. {corecoder-0.7.0 → corecoder-0.8.0}/article/08-extensibility.md +1 -1
  20. {corecoder-0.7.0 → corecoder-0.8.0}/article/08-extensibility_EN.md +1 -1
  21. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/__init__.py +1 -1
  22. corecoder-0.8.0/corecoder/checkpoints.py +93 -0
  23. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/cli.py +2 -2
  24. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/demo.py +1 -0
  25. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/permissions.py +31 -4
  26. {corecoder-0.7.0 → corecoder-0.8.0}/examples/plan_hooks_demo.py +0 -0
  27. {corecoder-0.7.0 → corecoder-0.8.0}/pyproject.toml +2 -2
  28. corecoder-0.8.0/tests/conftest.py +32 -0
  29. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_checkpoints.py +36 -2
  30. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_core.py +40 -0
  31. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_permissions.py +26 -0
  32. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_plan_mode.py +5 -12
  33. corecoder-0.8.0/tests/test_repl.py +140 -0
  34. corecoder-0.8.0/tests/test_slash_commands.py +130 -0
  35. corecoder-0.7.0/corecoder/checkpoints.py +0 -44
  36. corecoder-0.7.0/tests/conftest.py +0 -11
  37. {corecoder-0.7.0 → corecoder-0.8.0}/.github/workflows/publish.yml +0 -0
  38. {corecoder-0.7.0 → corecoder-0.8.0}/.gitignore +0 -0
  39. {corecoder-0.7.0 → corecoder-0.8.0}/LICENSE +0 -0
  40. {corecoder-0.7.0 → corecoder-0.8.0}/article/07-build-your-own.md +0 -0
  41. {corecoder-0.7.0 → corecoder-0.8.0}/article/07-build-your-own_EN.md +0 -0
  42. {corecoder-0.7.0 → corecoder-0.8.0}/assets/demo-plan-hooks.gif +0 -0
  43. {corecoder-0.7.0 → corecoder-0.8.0}/assets/demo.png +0 -0
  44. {corecoder-0.7.0 → corecoder-0.8.0}/assets/demo_en.png +0 -0
  45. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/__main__.py +0 -0
  46. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/agent.py +0 -0
  47. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/config.py +0 -0
  48. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/context.py +0 -0
  49. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/hooks.py +0 -0
  50. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/llm.py +0 -0
  51. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/mcp.py +0 -0
  52. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/prompt.py +0 -0
  53. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/session.py +0 -0
  54. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/shell.py +0 -0
  55. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/__init__.py +0 -0
  56. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/agent.py +0 -0
  57. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/base.py +0 -0
  58. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/bash.py +0 -0
  59. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/edit.py +0 -0
  60. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/glob_tool.py +0 -0
  61. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/grep.py +0 -0
  62. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/read.py +0 -0
  63. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/todo.py +0 -0
  64. {corecoder-0.7.0 → corecoder-0.8.0}/corecoder/tools/write.py +0 -0
  65. {corecoder-0.7.0 → corecoder-0.8.0}/tests/__init__.py +0 -0
  66. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_demo.py +0 -0
  67. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_hooks.py +0 -0
  68. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_litellm.py +0 -0
  69. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_mcp.py +0 -0
  70. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_safety_matrix.py +0 -0
  71. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_session.py +0 -0
  72. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_shell.py +0 -0
  73. {corecoder-0.7.0 → corecoder-0.8.0}/tests/test_tools.py +0 -0
@@ -3,8 +3,14 @@ name: CI
3
3
  on:
4
4
  push:
5
5
  branches: [main]
6
+ paths-ignore:
7
+ - "article/**"
8
+ - "*.md"
6
9
  pull_request:
7
10
  branches: [main]
11
+ paths-ignore:
12
+ - "article/**"
13
+ - "*.md"
8
14
 
9
15
  jobs:
10
16
  test:
@@ -1,7 +1,7 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: corecoder
3
- Version: 0.7.0
4
- Summary: Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. (formerly NanoCoder)
3
+ Version: 0.8.0
4
+ Summary: Minimal AI coding agent (2,735 lines of Python) inspired by Claude Code. Works with any LLM. (formerly NanoCoder)
5
5
  Project-URL: Homepage, https://github.com/he-yufeng/CoreCoder
6
6
  Project-URL: Repository, https://github.com/he-yufeng/CoreCoder
7
7
  Project-URL: Issues, https://github.com/he-yufeng/CoreCoder/issues
@@ -37,7 +37,7 @@ Description-Content-Type: text/markdown
37
37
 
38
38
  # CoreCoder
39
39
 
40
- **The nanoGPT of coding agents. A 1.3k-line engine inside 2,658 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
40
+ **The nanoGPT of coding agents. A 1.3k-line engine inside 2,735 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
41
41
 
42
42
  *learn from it · fork it · ship something better*
43
43
 
@@ -60,7 +60,7 @@ Description-Content-Type: text/markdown
60
60
 
61
61
  | | CoreCoder | Claude Code | aider | nanoGPT |
62
62
  |---|---|---|---|---|
63
- | Lines of code | ~1,309 engine / 2,658 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
63
+ | Lines of code | ~1,309 engine / 2,735 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
64
64
  | Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
65
65
  | Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
66
66
  | What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
@@ -71,9 +71,9 @@ The nanoGPT column is there as a reference point: minimal, readable, but it teac
71
71
 
72
72
  I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
73
73
 
74
- The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,658 physical lines, 2,138 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
74
+ The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,735 physical lines, 2,205 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
75
75
 
76
- And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 171 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
76
+ And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 215 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
77
77
 
78
78
  The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
79
79
 
@@ -126,15 +126,15 @@ corecoder/
126
126
  ├── llm.py streaming client + retry + cost 332 lines
127
127
  ├── context.py three-tier context compaction 220 lines
128
128
  ├── session.py save / resume + path-traversal guard 97 lines
129
- ├── permissions.py consent for mutating tools 48 lines
129
+ ├── permissions.py consent for mutating tools 75 lines
130
130
  ├── hooks.py Pre/PostToolUse shell hooks 87 lines
131
131
  ├── shell.py POSIX shell routing (Git Bash on Windows) 61 lines
132
132
  ├── mcp.py MCP stdio client for external tools 208 lines
133
133
  ├── prompt.py system prompt 41 lines
134
134
  ├── cli.py REPL + slash commands + one-shot 358 lines
135
135
  ├── config.py env-var config 55 lines
136
- ├── checkpoints.py /undo snapshot and restore 44 lines
137
- ├── demo.py offline end-to-end demo 100 lines
136
+ ├── checkpoints.py /undo snapshot and restore 93 lines
137
+ ├── demo.py offline end-to-end demo 101 lines
138
138
  └── tools/
139
139
  ├── bash.py shell + dangerous-command gate + cd 203 lines
140
140
  ├── edit.py unique-match search/replace + diff 99 lines
@@ -170,7 +170,7 @@ def chat(self, user_input):
170
170
  return "(hit the round limit)"
171
171
  ```
172
172
 
173
- That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
173
+ That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the engine, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
174
174
 
175
175
  Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
176
176
 
@@ -188,7 +188,7 @@ I also wrote a bilingual source-reading series, one intro plus eight parts, each
188
188
 
189
189
  - **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
190
190
  - **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
191
- - **[02 · The tool system: letting the model act, safely](article/02-tools_EN.md)** — the seven tools in `tools/` and the bash safety gate
191
+ - **[02 · The tool system: letting the model act, safely](article/02-tools_EN.md)** — the eight tools in `tools/` and the bash safety gate
192
192
  - **[03 · Plug in any LLM, and keep the bill honest](article/03-llm-and-cost_EN.md)** — `llm.py`'s provider wrapper, retries, and cost accounting
193
193
  - **[04 · Surviving a long task on a finite window](article/04-context_EN.md)** — `context.py`'s three-tier compaction and orphaned tool messages
194
194
  - **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
@@ -200,8 +200,8 @@ I also wrote a bilingual source-reading series, one intro plus eight parts, each
200
200
 
201
201
  Once you understand it, the natural next step is to fork. Getting started doesn't take much:
202
202
 
203
- - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (267 lines) is the entry point for all provider adaptation.
204
- - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (27 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
203
+ - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (332 lines) is the entry point for all provider adaptation.
204
+ - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (32 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
205
205
  - **Rewrite the system prompt.** `prompt.py` is all of 41 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
206
206
  - **Import it as a library.** The top level exports `Agent`, `LLM`, and `Config`, ready to embed in your own program:
207
207
 
@@ -236,13 +236,13 @@ Inside the REPL, `/help` lists everything; these are the ones you'll reach for:
236
236
  quit / exit exit (Ctrl+C cancels the current round)
237
237
  ```
238
238
 
239
- Session IDs are sanitized to safe characters before they become filenames, every archive lands under `~/.corecoder/sessions`, and a malicious session name can't traverse out.
239
+ Session IDs are sanitized to safe characters before they become filenames, every archive lands under `~/.corecoder/sessions`, and a malicious session name can't traverse out. Undo history persists the same way: checkpoints land in `~/.corecoder/checkpoints.json`, so `/undo` still reaches back after a restart.
240
240
 
241
241
  ## Permissions
242
242
 
243
243
  Read-only tools (`read_file`, `glob`, `grep`, `todo_write`) run the moment the model asks. The mutating ones (`edit_file`, `write_file`, `bash`, and spawning a sub-agent) stop for consent first, and the REPL banner shows which mode you're in:
244
244
 
245
- - In the REPL you get one prompt per call: allow once, always allow this tool, or deny. "Always" is remembered per tool for the rest of the session, and a sub-agent inherits the same layer, so consent follows the work wherever it happens.
245
+ - In the REPL you get one prompt per call: allow once, always allow this tool, or deny. "Always" is remembered per tool and persists in `~/.corecoder/permissions.json` across restarts; a sub-agent inherits the same layer, so consent follows the work wherever it happens.
246
246
  - In one-shot mode (`-p`) there is nobody to ask, so a mutating call is refused on the spot and the refusal goes back to the model as an ordinary tool result: the loop never hangs on input that can't arrive. Pass `--yes` to approve everything up front (scripts, CI).
247
247
  - The decision itself is pure logic in `permissions.py`, with the terminal only supplying the prompt callback. You can unit-test consent without a TTY, or reuse the layer in your own embedding.
248
248
 
@@ -271,8 +271,8 @@ Two worth stealing (the commands lean on `jq`, the usual suspect):
271
271
  # the same turn instead of waiting for CI.
272
272
  {
273
273
  "PostToolUse": [{
274
- "matcher": "edit",
275
- "command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
274
+ "matcher": "edit_file",
275
+ "command": "f=$(jq -r .tool_input.file_path); ruff check \"$f\" 2>&1 | head -20"
276
276
  }]
277
277
  }
278
278
 
@@ -280,8 +280,8 @@ Two worth stealing (the commands lean on `jq`, the usual suspect):
280
280
  # touch. Exit code 2 vetoes the call and the message reaches the model.
281
281
  {
282
282
  "PreToolUse": [{
283
- "matcher": "edit",
284
- "command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
283
+ "matcher": "edit_file",
284
+ "command": "case \"$(jq -r .tool_input.file_path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
285
285
  }]
286
286
  }
287
287
  ```
@@ -314,7 +314,7 @@ If working through CoreCoder was useful, here are a few other tools I've built a
314
314
 
315
315
  ## Contributing / License
316
316
 
317
- Before you send anything, run `pytest tests/ -q` (171 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
317
+ Before you send anything, run `pytest tests/ -q` (215 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
318
318
 
319
319
  ---
320
320
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  # CoreCoder
4
4
 
5
- **The nanoGPT of coding agents. A 1.3k-line engine inside 2,658 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
5
+ **The nanoGPT of coding agents. A 1.3k-line engine inside 2,735 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
6
6
 
7
7
  *learn from it · fork it · ship something better*
8
8
 
@@ -25,7 +25,7 @@
25
25
 
26
26
  | | CoreCoder | Claude Code | aider | nanoGPT |
27
27
  |---|---|---|---|---|
28
- | Lines of code | ~1,309 engine / 2,658 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
28
+ | Lines of code | ~1,309 engine / 2,735 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
29
29
  | Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
30
30
  | Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
31
31
  | What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
@@ -36,9 +36,9 @@ The nanoGPT column is there as a reference point: minimal, readable, but it teac
36
36
 
37
37
  I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
38
38
 
39
- The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,658 physical lines, 2,138 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
39
+ The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,735 physical lines, 2,205 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
40
40
 
41
- And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 171 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
41
+ And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 215 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
42
42
 
43
43
  The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
44
44
 
@@ -91,15 +91,15 @@ corecoder/
91
91
  ├── llm.py streaming client + retry + cost 332 lines
92
92
  ├── context.py three-tier context compaction 220 lines
93
93
  ├── session.py save / resume + path-traversal guard 97 lines
94
- ├── permissions.py consent for mutating tools 48 lines
94
+ ├── permissions.py consent for mutating tools 75 lines
95
95
  ├── hooks.py Pre/PostToolUse shell hooks 87 lines
96
96
  ├── shell.py POSIX shell routing (Git Bash on Windows) 61 lines
97
97
  ├── mcp.py MCP stdio client for external tools 208 lines
98
98
  ├── prompt.py system prompt 41 lines
99
99
  ├── cli.py REPL + slash commands + one-shot 358 lines
100
100
  ├── config.py env-var config 55 lines
101
- ├── checkpoints.py /undo snapshot and restore 44 lines
102
- ├── demo.py offline end-to-end demo 100 lines
101
+ ├── checkpoints.py /undo snapshot and restore 93 lines
102
+ ├── demo.py offline end-to-end demo 101 lines
103
103
  └── tools/
104
104
  ├── bash.py shell + dangerous-command gate + cd 203 lines
105
105
  ├── edit.py unique-match search/replace + diff 99 lines
@@ -135,7 +135,7 @@ def chat(self, user_input):
135
135
  return "(hit the round limit)"
136
136
  ```
137
137
 
138
- That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
138
+ That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the engine, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
139
139
 
140
140
  Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
141
141
 
@@ -153,7 +153,7 @@ I also wrote a bilingual source-reading series, one intro plus eight parts, each
153
153
 
154
154
  - **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
155
155
  - **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
156
- - **[02 · The tool system: letting the model act, safely](article/02-tools_EN.md)** — the seven tools in `tools/` and the bash safety gate
156
+ - **[02 · The tool system: letting the model act, safely](article/02-tools_EN.md)** — the eight tools in `tools/` and the bash safety gate
157
157
  - **[03 · Plug in any LLM, and keep the bill honest](article/03-llm-and-cost_EN.md)** — `llm.py`'s provider wrapper, retries, and cost accounting
158
158
  - **[04 · Surviving a long task on a finite window](article/04-context_EN.md)** — `context.py`'s three-tier compaction and orphaned tool messages
159
159
  - **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
@@ -165,8 +165,8 @@ I also wrote a bilingual source-reading series, one intro plus eight parts, each
165
165
 
166
166
  Once you understand it, the natural next step is to fork. Getting started doesn't take much:
167
167
 
168
- - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (267 lines) is the entry point for all provider adaptation.
169
- - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (27 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
168
+ - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (332 lines) is the entry point for all provider adaptation.
169
+ - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (32 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
170
170
  - **Rewrite the system prompt.** `prompt.py` is all of 41 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
171
171
  - **Import it as a library.** The top level exports `Agent`, `LLM`, and `Config`, ready to embed in your own program:
172
172
 
@@ -201,13 +201,13 @@ Inside the REPL, `/help` lists everything; these are the ones you'll reach for:
201
201
  quit / exit exit (Ctrl+C cancels the current round)
202
202
  ```
203
203
 
204
- Session IDs are sanitized to safe characters before they become filenames, every archive lands under `~/.corecoder/sessions`, and a malicious session name can't traverse out.
204
+ Session IDs are sanitized to safe characters before they become filenames, every archive lands under `~/.corecoder/sessions`, and a malicious session name can't traverse out. Undo history persists the same way: checkpoints land in `~/.corecoder/checkpoints.json`, so `/undo` still reaches back after a restart.
205
205
 
206
206
  ## Permissions
207
207
 
208
208
  Read-only tools (`read_file`, `glob`, `grep`, `todo_write`) run the moment the model asks. The mutating ones (`edit_file`, `write_file`, `bash`, and spawning a sub-agent) stop for consent first, and the REPL banner shows which mode you're in:
209
209
 
210
- - In the REPL you get one prompt per call: allow once, always allow this tool, or deny. "Always" is remembered per tool for the rest of the session, and a sub-agent inherits the same layer, so consent follows the work wherever it happens.
210
+ - In the REPL you get one prompt per call: allow once, always allow this tool, or deny. "Always" is remembered per tool and persists in `~/.corecoder/permissions.json` across restarts; a sub-agent inherits the same layer, so consent follows the work wherever it happens.
211
211
  - In one-shot mode (`-p`) there is nobody to ask, so a mutating call is refused on the spot and the refusal goes back to the model as an ordinary tool result: the loop never hangs on input that can't arrive. Pass `--yes` to approve everything up front (scripts, CI).
212
212
  - The decision itself is pure logic in `permissions.py`, with the terminal only supplying the prompt callback. You can unit-test consent without a TTY, or reuse the layer in your own embedding.
213
213
 
@@ -236,8 +236,8 @@ Two worth stealing (the commands lean on `jq`, the usual suspect):
236
236
  # the same turn instead of waiting for CI.
237
237
  {
238
238
  "PostToolUse": [{
239
- "matcher": "edit",
240
- "command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
239
+ "matcher": "edit_file",
240
+ "command": "f=$(jq -r .tool_input.file_path); ruff check \"$f\" 2>&1 | head -20"
241
241
  }]
242
242
  }
243
243
 
@@ -245,8 +245,8 @@ Two worth stealing (the commands lean on `jq`, the usual suspect):
245
245
  # touch. Exit code 2 vetoes the call and the message reaches the model.
246
246
  {
247
247
  "PreToolUse": [{
248
- "matcher": "edit",
249
- "command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
248
+ "matcher": "edit_file",
249
+ "command": "case \"$(jq -r .tool_input.file_path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
250
250
  }]
251
251
  }
252
252
  ```
@@ -279,7 +279,7 @@ If working through CoreCoder was useful, here are a few other tools I've built a
279
279
 
280
280
  ## Contributing / License
281
281
 
282
- Before you send anything, run `pytest tests/ -q` (171 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
282
+ Before you send anything, run `pytest tests/ -q` (215 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
283
283
 
284
284
  ---
285
285
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  # CoreCoder
4
4
 
5
- **编程 agent 里的 nanoGPT。1.3k 行引擎、整包 2594 行纯 Python 全部一口气可读,读懂一个 coding agent 到底怎么运作,再 fork 出你自己的。**
5
+ **编程 agent 里的 nanoGPT。1.3k 行引擎、整包 2735 行纯 Python 全部一口气可读,读懂一个 coding agent 到底怎么运作,再 fork 出你自己的。**
6
6
 
7
7
  *learn from it · fork it · ship something better*
8
8
 
@@ -25,7 +25,7 @@
25
25
 
26
26
  | | CoreCoder | Claude Code | aider | nanoGPT |
27
27
  |---|---|---|---|---|
28
- | 代码量 | 引擎约 1308 行 / 整包 2594 行 | 几十万行(闭源) | 数万行 Python | 约 600 行(两个文件) |
28
+ | 代码量 | 引擎约 1309 行 / 整包 2735 行 | 几十万行(闭源) | 数万行 Python | 约 600 行(两个文件) |
29
29
  | 读完要多久 | 一个下午 | 读不了(闭源) | 得啃几天 | 一个下午 |
30
30
  | 能不能下断点改了再跑 | 能,每一行 | 不能 | 能,但量大 | 能 |
31
31
  | 定位 | 读懂并 fork 出你自己的 agent | 生产级编程助手 | 终端结对编程 | 教学用最小 GPT |
@@ -36,9 +36,9 @@ nanoGPT 那一列是拿来对照的:它最小、可读,但教的是训一个
36
36
 
37
37
  我一直觉得 coding agent 被讲得太玄了。把 Claude Code、Cursor 这类工具扒到底,核心是一个 while 循环套着一个大模型,外加七八个让它能真正动手的工具。难的从来不是这个循环,而是循环跑进真实世界以后要兜的那些底。CoreCoder 就是把这个核心老老实实写出来的最小版本。
38
38
 
39
- 引擎部分(循环、模型接口、上下文、工具、会话)去掉空行和注释是 1309 行。连最外层的 CLI、配置、打包一起算,整个包 25 个文件、物理 2658 行、净 2138 行,每个文件都短到能一口气读完。自 1161 行快照之后的增长都花在了看得见的功能上:plan mode、hooks、checkpoints,下文各有交代。
39
+ 引擎部分(循环、模型接口、上下文、工具、会话)去掉空行和注释是 1309 行。连最外层的 CLI、配置、打包一起算,整个包 25 个文件、物理 2735 行、净 2205 行,每个文件都短到能一口气读完。自 1161 行快照之后的增长都花在了看得见的功能上:plan mode、hooks、checkpoints,下文各有交代。
40
40
 
41
- 它真能跑:读写文件、执行 shell、派子 agent、分三层压上下文,还能随时把这趟烧掉的 token 和美元数报给你。任何要动你磁盘、要跑命令的调用,都会先停下来等你点头,171 个测试是绿的。但能跑不是为了劝你拿去日用,而是为了让这份「注释」不撒谎:一个解释 agent 怎么运作的范例,自己得真能运作。
41
+ 它真能跑:读写文件、执行 shell、派子 agent、分三层压上下文,还能随时把这趟烧掉的 token 和美元数报给你。任何要动你磁盘、要跑命令的调用,都会先停下来等你点头,215 个测试是绿的。但能跑不是为了劝你拿去日用,而是为了让这份「注释」不撒谎:一个解释 agent 怎么运作的范例,自己得真能运作。
42
42
 
43
43
  代码来自一次公开拆解。公开的源码分析里,Claude Code 这类生产级 agent 暴露出不少关键架构,我挑出最核心的一层,用尽量少的代码诚实地复写了一遍。所以读 CoreCoder,约等于读一份基于公开源码分析的「可运行注释版」:讲的是这类 agent 的核心思路,而它本身只是最小复写,就摆在你机器上,随你拆、随你改。
44
44
 
@@ -91,15 +91,15 @@ corecoder/
91
91
  ├── llm.py 流式客户端 + 重试 + 成本统计 332 行
92
92
  ├── context.py 三层上下文压缩 220 行
93
93
  ├── session.py 会话存盘 / 续聊 + 路径穿越防护 97 行
94
- ├── permissions.py 改动类工具的用户授权 48 行
94
+ ├── permissions.py 改动类工具的用户授权 75 行
95
95
  ├── hooks.py 工具调用前后的用户 shell 钩子 87 行
96
96
  ├── shell.py POSIX shell 路由(Windows 走 Git Bash) 61 行
97
97
  ├── mcp.py MCP stdio 客户端,接外部工具 208 行
98
98
  ├── prompt.py 系统提示词 41 行
99
99
  ├── cli.py REPL + 斜杠命令 + 一次性模式 358 行
100
100
  ├── config.py 环境变量配置 55 行
101
- ├── checkpoints.py /undo 快照与回滚 44 行
102
- ├── demo.py 离线端到端演示 100 行
101
+ ├── checkpoints.py /undo 快照与回滚 93 行
102
+ ├── demo.py 离线端到端演示 101 行
103
103
  └── tools/
104
104
  ├── bash.py shell + 危险命令闸 + cd 追踪 203 行
105
105
  ├── edit.py 唯一匹配搜索替换 + diff 99 行
@@ -135,7 +135,7 @@ def chat(self, user_input):
135
135
  return "(已达轮次上限)"
136
136
  ```
137
137
 
138
- 就这么点。这个循环的核心骨架就二十来行,把并行执行和被 Ctrl+C 打断后的回填都算上,也才四十多行。CoreCoder 一千多行里剩下的,几乎全在收拾它真跑起来之后冒出来的岔子。`llm.py` 最后成了全项目最大的文件,不是因为调模型有多难,而是流式返回里一个工具调用的参数会被切成好几段先后送到、得按顺序拼回去,provider 偶尔吐半截 JSON 或把 usage 填成 null,限流(429)、超时、连接中断和 5xx 都得退避重试,其余 4xx 该直接抛就别硬试。已经开始出字的流也可能半路断掉,所以重试包住的是整个请求,不只是建连。这些不起眼的脏活,而不是那个循环,才是一个 agent 从能演示走到能交付真正吃工程功夫的地方;第三篇文章顺着它拆到每一行。
138
+ 就这么点。这个循环的核心骨架就二十来行,把并行执行和被 Ctrl+C 打断后的回填都算上,也才四十多行。CoreCoder 一千多行里剩下的,几乎全在收拾它真跑起来之后冒出来的岔子。`llm.py` 最后成了引擎里最大的文件,不是因为调模型有多难,而是流式返回里一个工具调用的参数会被切成好几段先后送到、得按顺序拼回去,provider 偶尔吐半截 JSON 或把 usage 填成 null,限流(429)、超时、连接中断和 5xx 都得退避重试,其余 4xx 该直接抛就别硬试。已经开始出字的流也可能半路断掉,所以重试包住的是整个请求,不只是建连。这些不起眼的脏活,而不是那个循环,才是一个 agent 从能演示走到能交付真正吃工程功夫的地方;第三篇文章顺着它拆到每一行。
139
139
 
140
140
  有三个决定值得单独看,因为它们是「先读懂别人怎么做」之后才做得出的取舍,也是你 fork 自己 agent 时可以直接抄走的判断。
141
141
 
@@ -153,7 +153,7 @@ def chat(self, user_input):
153
153
 
154
154
  - **[导言 · 用 CoreCoder 读懂 Claude Code,再造一个你自己的](article/00-index.md)**
155
155
  - **[01 一个 agent 的本体,是一个 while 循环](article/01-the-loop.md)** — `agent.py` 的主循环、打断与轮次上限
156
- - **[02 工具系统:让模型安全地动手](article/02-tools.md)** — `tools/` 七个工具与 bash 安全闸
156
+ - **[02 工具系统:让模型安全地动手](article/02-tools.md)** — `tools/` 八个工具与 bash 安全闸
157
157
  - **[03 接入任意大模型,顺便把账算清楚](article/03-llm-and-cost.md)** — `llm.py` 的 provider 包装、重试与成本统计
158
158
  - **[04 用有限的窗口扛住一个长任务](article/04-context.md)** — `context.py` 的三层压缩与孤儿 tool 消息
159
159
  - **[05 并行执行与子 agent](article/05-parallel-and-subagents.md)** — 线程池并发与子 agent 隔离
@@ -165,8 +165,8 @@ def chat(self, user_input):
165
165
 
166
166
  读懂之后,最自然的下一步就是 fork。起手不用伤筋动骨:
167
167
 
168
- - **换个你常用的模型。** 就是上面那两个环境变量,`llm.py`(331 行)是所有 provider 适配的入口。
169
- - **加一件你自己的工具。** 照 `tools/base.py`(27 行)的工具基类写个新文件,跑测试、抓网页、调 LSP 都行,第二篇文章末尾手把手带你写第一个。
168
+ - **换个你常用的模型。** 就是上面那两个环境变量,`llm.py`(332 行)是所有 provider 适配的入口。
169
+ - **加一件你自己的工具。** 照 `tools/base.py`(32 行)的工具基类写个新文件,跑测试、抓网页、调 LSP 都行,第二篇文章末尾手把手带你写第一个。
170
170
  - **改系统提示词。** `prompt.py` 才 41 行,改一句就能看到 agent 的脾气变了,是门槛最低的「改一处就有反馈」。
171
171
  - **直接当库 import。** 顶层导出了 `Agent`、`LLM`、`Config`,能嵌进你自己的程序:
172
172
 
@@ -201,13 +201,13 @@ README 只给方向,每条的代码细节第七篇接着讲。挑一个动手
201
201
  quit / exit 退出(Ctrl+C 取消当前回合)
202
202
  ```
203
203
 
204
- 会话 ID 会先清洗成安全字符再拿去当文件名,存档统统落在 `~/.corecoder/sessions` 里,恶意会话名穿越不出去。
204
+ 会话 ID 会先清洗成安全字符再拿去当文件名,存档统统落在 `~/.corecoder/sessions` 里,恶意会话名穿越不出去。undo 历史也一样落盘:快照写在 `~/.corecoder/checkpoints.json`,重启之后 `/undo` 照样能往回撤。
205
205
 
206
206
  ## 权限
207
207
 
208
208
  只读工具(`read_file`、`glob`、`grep`、`todo_write`)模型一调就跑。会动手的那些(`edit_file`、`write_file`、`bash`,以及派生子 agent)先停下来等你点头,REPL 启动横幅里能看到当前是哪种模式:
209
209
 
210
- - REPL 里每次调用问一次:允许这一次、本工具本次会话都允许、或者拒绝。「都允许」按工具记到会话结束;子 agent 继承同一层授权,活走到哪,许可跟到哪。
210
+ - REPL 里每次调用问一次:允许这一次、本工具一直允许、或者拒绝。「一直允许」按工具记下,写进 `~/.corecoder/permissions.json`,重启也有效;子 agent 继承同一层授权,活走到哪,许可跟到哪。
211
211
  - 一次性模式(`-p`)没人可问,改动类调用当场被拒,拒绝理由作为普通工具结果回给模型:循环绝不会卡在等一个永远不会来的输入上。要全部预授权就加 `--yes`(脚本、CI 场景)。
212
212
  - 判断本身是 `permissions.py` 里的纯逻辑,终端只是塞进来一个提问回调。不用 TTY 也能单测授权逻辑,或者直接搬进你自己的嵌入场景。
213
213
 
@@ -235,8 +235,8 @@ REPL 里 `/plan` 开关计划模式。开着的时候,提示符变成 `(plan)`
235
235
  # 模型同一回合就能看到输出,自己把低级错误修了,不用等 CI 回来。
236
236
  {
237
237
  "PostToolUse": [{
238
- "matcher": "edit",
239
- "command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
238
+ "matcher": "edit_file",
239
+ "command": "f=$(jq -r .tool_input.file_path); ruff check \"$f\" 2>&1 | head -20"
240
240
  }]
241
241
  }
242
242
 
@@ -244,8 +244,8 @@ REPL 里 `/plan` 开关计划模式。开着的时候,提示符变成 `(plan)`
244
244
  # 拒绝原因会送到模型那边。
245
245
  {
246
246
  "PreToolUse": [{
247
- "matcher": "edit",
248
- "command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
247
+ "matcher": "edit_file",
248
+ "command": "case \"$(jq -r .tool_input.file_path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
249
249
  }]
250
250
  }
251
251
  ```
@@ -278,7 +278,7 @@ REPL 里 `/plan` 开关计划模式。开着的时候,提示符变成 `(plan)`
278
278
 
279
279
  ## 贡献 / License
280
280
 
281
- 动手之前先跑一遍 `pytest tests/ -q`(171 个测试)、`ruff check` 和 `compileall`,绿了再提。MIT License,欢迎 fork 拿去造更好的东西,能在 README 里留一句出处就更好。
281
+ 动手之前先跑一遍 `pytest tests/ -q`(215 个测试)、`ruff check` 和 `compileall`,绿了再提。MIT License,欢迎 fork 拿去造更好的东西,能在 README 里留一句出处就更好。
282
282
 
283
283
  ---
284
284
 
@@ -22,10 +22,10 @@ CoreCoder 做的事,是把这套骨架压到一千行出头的纯 Python。准
22
22
 
23
23
  前六篇是「读懂」。每篇盯住 agent 的一个子系统,先讲 Claude Code 在这件事上的做法和取舍,再翻到 CoreCoder 对应的真实代码,看同一个想法被压缩成几十行后长什么样。最后一篇是「自己做」,把前面所有零件接起来,从 fork 到加一个自定义工具到换模型,落到一个能跑的成品。
24
24
 
25
- 1. [一个 agent 的本体,是一个 while 循环](01-the-loop.md)。整个 agent 最核心的东西,是一个「问模型、跑工具、把结果喂回去、再问」的循环。我们看 CoreCoder 的 `agent.py`(150 行)怎么把它写明白,以及打断、轮次上限、半截工具调用回填这些真实世界的麻烦各自怎么收场。
25
+ 1. [一个 agent 的本体,是一个 while 循环](01-the-loop.md)。整个 agent 最核心的东西,是一个「问模型、跑工具、把结果喂回去、再问」的循环。我们看 CoreCoder 的 `agent.py`(240 行)怎么把它写明白,以及打断、轮次上限、半截工具调用回填这些真实世界的麻烦各自怎么收场。
26
26
  2. [工具系统:让模型安全地动手](02-tools.md)。模型本身只会吐字,是工具让它能读文件、写文件、跑命令。这篇讲 CoreCoder 的七个工具,重点是那个看似平平无奇、实则是 Claude Code 关键创新的「唯一性搜索替换」编辑,以及 bash 的安全闸。v0.6.0 之后还补了一节「门之外」:plan 模式、hooks、MCP 这三个进阶件怎么挂在一次调用的前后。末尾教你写第一个自己的工具。
27
- 3. [接入任意大模型,顺便把钱算清楚](03-llm-and-cost.md)。`llm.py`(336 行,全系列最大的文件)怎么用一套 OpenAI 兼容接口接住 DeepSeek、Qwen、Kimi、本地 Ollama,怎么做指数退避重试,怎么在流式输出里顺手把 token 和美元成本统计出来。
28
- 4. [用有限的窗口扛住一个长任务](04-context.md)。上下文窗口是 agent 的硬约束。`context.py`(210 行)实现了三层压缩,从轻到重。这篇还会讲一个特别容易踩、API 一定报错的坑:孤儿 tool 消息。这是我做这个项目时真改过的 bug。
27
+ 3. [接入任意大模型,顺便把钱算清楚](03-llm-and-cost.md)。`llm.py`(332 行,引擎里最大的文件)怎么用一套 OpenAI 兼容接口接住 DeepSeek、Qwen、Kimi、本地 Ollama,怎么做指数退避重试,怎么在流式输出里顺手把 token 和美元成本统计出来。
28
+ 4. [用有限的窗口扛住一个长任务](04-context.md)。上下文窗口是 agent 的硬约束。`context.py`(220 行)实现了三层压缩,从轻到重。这篇还会讲一个特别容易踩、API 一定报错的坑:孤儿 tool 消息。这是我做这个项目时真改过的 bug。
29
29
  5. [并行执行与子 agent](05-parallel-and-subagents.md)。模型一次返回多个工具调用时,CoreCoder 用线程池并发跑。这篇老实讲这个简化版相对 Claude Code 的流式执行器差在哪,并发又会引入什么新麻烦,以及子 agent 为什么不准递归。
30
30
  6. [把它跑成一个真正的命令行工具](06-session-and-cli.md)。会话存盘、断点续聊、斜杠命令、一次性模式。`session.py` 里有个不起眼但很要命的安全细节:怎么防住用恶意会话名做路径穿越。
31
31
  7. [Fork CoreCoder,搭一个你自己的 coding agent](07-build-your-own.md)。收尾的实操篇。从 clone 到换成你常用的模型,到加一个真正有用的自定义工具,到改系统提示词调教它的风格,到打包发布。读完前六篇你已经懂了原理,这篇让你真有一个东西。
@@ -22,10 +22,10 @@ The answer is engineering. And the kind that, once you've read it, makes you thi
22
22
 
23
23
  The first six are about understanding. Each one fixes on a single subsystem of the agent, first describing how Claude Code handles that thing and the tradeoffs it makes, then turning to the real code in CoreCoder to see what the same idea looks like once it's compressed into a few dozen lines. The last piece is about building it yourself, wiring all the parts back together, from fork to a custom tool to swapping the model, landing on something that runs.
24
24
 
25
- 1. [An agent is, at heart, a while loop](01-the-loop_EN.md). The most central thing in the whole agent is a loop: ask the model, run a tool, feed the result back, ask again. We look at how CoreCoder's `agent.py` (150 lines) writes it out plainly, and how real-world headaches like interruption, a round cap, and backfilling half-finished tool calls each get handled.
25
+ 1. [An agent is, at heart, a while loop](01-the-loop_EN.md). The most central thing in the whole agent is a loop: ask the model, run a tool, feed the result back, ask again. We look at how CoreCoder's `agent.py` (240 lines) writes it out plainly, and how real-world headaches like interruption, a round cap, and backfilling half-finished tool calls each get handled.
26
26
  2. [The tool system: letting the model act, safely](02-tools_EN.md). The model on its own only emits text. Tools are what let it read files, write files, run commands. This piece covers CoreCoder's seven tools, with the spotlight on the seemingly unremarkable unique search-and-replace edit that is in fact one of Claude Code's key innovations, plus bash's safety gate. Since v0.6.0 it also carries a "beyond the gates" section on how the three advanced pieces — plan mode, hooks, and MCP — hang off the moments around a call. At the end I'll have you write your first tool.
27
- 3. [Plug in any model, and get the bill right while you're at it](03-llm-and-cost_EN.md). How `llm.py` (336 lines, the largest single file in the project) uses one OpenAI-compatible interface to catch DeepSeek, Qwen, Kimi, and local Ollama, how it does exponential-backoff retry, and how it tallies tokens and dollar cost right inside the streaming output.
28
- 4. [Surviving a long task in a finite window](04-context_EN.md). The context window is the agent's hard constraint. `context.py` (210 lines) implements three layers of compression, lightest to heaviest. This piece also covers a trap that's easy to hit and that the API will always reject: the orphaned tool message. That's a bug I actually fixed while building this project.
27
+ 3. [Plug in any model, and get the bill right while you're at it](03-llm-and-cost_EN.md). How `llm.py` (332 lines, the largest single file in the engine) uses one OpenAI-compatible interface to catch DeepSeek, Qwen, Kimi, and local Ollama, how it does exponential-backoff retry, and how it tallies tokens and dollar cost right inside the streaming output.
28
+ 4. [Surviving a long task in a finite window](04-context_EN.md). The context window is the agent's hard constraint. `context.py` (220 lines) implements three layers of compression, lightest to heaviest. This piece also covers a trap that's easy to hit and that the API will always reject: the orphaned tool message. That's a bug I actually fixed while building this project.
29
29
  5. [Parallel execution and sub-agents](05-parallel-and-subagents_EN.md). When the model returns several tool calls at once, CoreCoder runs them concurrently on a thread pool. This piece is honest about where this simplified version falls short of Claude Code's streaming executor, what new trouble concurrency brings in, and why a sub-agent is not allowed to recurse.
30
30
  6. [Turning it into a real command-line tool](06-session-and-cli_EN.md). Session save, resume, slash commands, one-shot mode. `session.py` holds an unremarkable but critical security detail: how to stop a malicious session name from turning into a path traversal.
31
31
  7. [Fork CoreCoder and build your own coding agent](07-build-your-own_EN.md). The hands-on finale. From clone, to switching to the model you actually use, to adding a genuinely useful custom tool, to tuning the system prompt to shape its style, to packaging and release. After the first six pieces you understand the principles; this one leaves you with something real.
@@ -2,7 +2,7 @@
2
2
 
3
3
  如果只能用一句话解释编码 agent,我会这么说:它是一个循环,反复地问模型「下一步干什么」,照着模型说的去动手,把结果再讲给模型听,直到模型说「不用动手了,我有答案了」。
4
4
 
5
- 听起来朴素到有点失望。但这就是真相。Claude Code 把这件事做到了几十万行,可那个最中心的东西,公开拆解里叫 `query.ts`,它的主体是一个一千七百行上下的 `while` 循环。CoreCoder 把同一个循环写在 `corecoder/agent.py` 里,连空行带注释一共 150 行。两者形状一模一样,区别只是后者你能一眼看完。
5
+ 听起来朴素到有点失望。但这就是真相。Claude Code 把这件事做到了几十万行,可那个最中心的东西,公开拆解里叫 `query.ts`,它的主体是一个一千七百行上下的 `while` 循环。CoreCoder 把同一个循环写在 `corecoder/agent.py`(240 行,连空行带注释)里。两者形状一模一样,区别只是后者你能一眼看完。
6
6
 
7
7
  这一篇,我们就把这个循环逐段读透。
8
8
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  If I had only one sentence to explain a coding agent, I'd put it like this: it's a loop that keeps asking the model "what's next," does what the model says, reports the result back to the model, and repeats until the model says "no need to act, I have the answer."
4
4
 
5
- It sounds almost disappointingly plain. But that is the truth of it. Claude Code took this same thing and built it out to hundreds of thousands of lines, yet the most central piece, the part public teardowns call `query.ts`, is at its core a `while` loop of around seventeen hundred lines. CoreCoder writes the same loop in `corecoder/agent.py`, 150 lines including blanks and comments. The two have an identical shape; the only difference is you can read the latter in a single glance.
5
+ It sounds almost disappointingly plain. But that is the truth of it. Claude Code took this same thing and built it out to hundreds of thousands of lines, yet the most central piece, the part public teardowns call `query.ts`, is at its core a `while` loop of around seventeen hundred lines. CoreCoder writes the same loop in `corecoder/agent.py` (240 lines including blanks and comments). The two have an identical shape; the only difference is you can read the latter in a single glance.
6
6
 
7
7
  In this piece we read that loop closely, section by section.
8
8
 
@@ -4,11 +4,11 @@
4
4
 
5
5
  模型本身只会做一件事,根据上文吐出下文。它不能读你的文件,不能跑你的测试,不能往磁盘写一个字节。让它从「会说」变成「会做」的,是工具。工具是 agent 真正接触世界的那只手。所以一个 agent 强不强,很大程度上取决于它的工具设计得好不好:接口是否清晰、错误反馈是否到位、危险操作是否拦得住。
6
6
 
7
- CoreCoder 给了模型七个工具:`bash`、`read_file`、`write_file`、`edit_file`、`glob`、`grep`、`agent`。这一篇我们先看它们共同的骨架,再细抠其中两个最值得说的,最后我带你写一个自己的。
7
+ CoreCoder 给了模型八个工具:`bash`、`read_file`、`write_file`、`edit_file`、`glob`、`grep`、`todo_write`、`agent`。这一篇我们先看它们共同的骨架,再细抠其中两个最值得说的,最后我带你写一个自己的。
8
8
 
9
9
  ## 一个工具长什么样
10
10
 
11
- 所有工具继承自 `tools/base.py` 里的 `Tool`,整个基类 27 行:
11
+ 所有工具继承自 `tools/base.py` 里的 `Tool`,整个基类 32 行:
12
12
 
13
13
  ```python
14
14
  class Tool(ABC):
@@ -57,7 +57,7 @@ ALL_TOOLS = [
57
57
 
58
58
  ## edit_file:一个看着平平无奇的关键创新
59
59
 
60
- 七个工具里,如果只能挑一个讲,我会挑 `edit_file`。因为「让模型修改一个已有文件」这个看似简单的需求,背后死过好几条路。
60
+ 八个工具里,如果只能挑一个讲,我会挑 `edit_file`。因为「让模型修改一个已有文件」这个看似简单的需求,背后死过好几条路。
61
61
 
62
62
  第一条死路是让模型按行号打补丁,比如「把第 42 行换成这样」。问题是模型对行号的感知极不可靠,它脑子里的第 42 行和文件里真实的第 42 行经常对不上,差一行就改错地方。而且只要文件在前面被动过一次,后面所有行号全部漂移。
63
63
 
@@ -115,7 +115,7 @@ except UnicodeDecodeError:
115
115
 
116
116
  `read_file`、`edit_file` 这些工具能造成的破坏有限。`bash` 不一样,它能跑任意 shell 命令,模型一旦写出 `rm -rf /`,后果是真实的。
117
117
 
118
- Claude Code 的 `BashTool` 公开拆解里是 1143 行,里头有命令分类器、有基于 `sandbox-exec` 和 `seccomp` 的真沙箱、有输出截断、有交互式命令拦截。CoreCoder 的 `bash.py` 是 127 行的蒸馏版,保留了四件最要紧的事:危险命令检测、输出截断、超时、工作目录跟踪。
118
+ Claude Code 的 `BashTool` 公开拆解里是 1143 行,里头有命令分类器、有基于 `sandbox-exec` 和 `seccomp` 的真沙箱、有输出截断、有交互式命令拦截。CoreCoder 的 `bash.py` 是 203 行的蒸馏版,保留了四件最要紧的事:危险命令检测、输出截断、超时、工作目录跟踪。
119
119
 
120
120
  危险命令检测是一张正则黑名单:
121
121
 
@@ -161,7 +161,7 @@ v0.6.0 在「调用前后」这条缝上又加了三样东西。它们不改变
161
161
 
162
162
  第三个是 MCP。`mcp.py` 里每个 server 是一个子进程,走 stdio 上一行一条的 JSON-RPC:启动握手、`tools/list` 拉清单,然后每个远程工具以 `mcp__server__tool` 的名字登记进工具表,从此授权、hooks、plan 模式对它们和内置工具一视同仁。值得学的是这个「一视同」是怎么来的:不是靠 MCP 代码里写多少特殊分支,而是因为工具边界本身(`Tool` 基类加上面三道闸)定义得够干净,外部工具只是又一个实现类。一个挂着死掉的 server 最坏也就是那一次调用返回个错误字符串,循环照转。
163
163
 
164
- 这三块都属于「进阶件」:它们不在 agent 的骨架上,骨架仍然是那个循环加七件工具。但它们回答的是同一个问题——一次工具调用前后,还有谁能说话。答案从「两道闸」变成了「两道闸加一圈可插拔的旁路」,而旁路的每一环都被设计成「坏了也不拖累主循环」。
164
+ 这三块都属于「进阶件」:它们不在 agent 的骨架上,骨架仍然是那个循环加八件工具。但它们回答的是同一个问题——一次工具调用前后,还有谁能说话。答案从「两道闸」变成了「两道闸加一圈可插拔的旁路」,而旁路的每一环都被设计成「坏了也不拖累主循环」。
165
165
 
166
166
  ## 动手:写你自己的第一个工具
167
167
 
@@ -4,11 +4,11 @@ In the loop from the last piece, one step got glossed over: executing tools. Thi
4
4
 
5
5
  The model itself does only one thing, emitting the next text given the text so far. It can't read your files, can't run your tests, can't write a single byte to disk. What turns it from "able to talk" into "able to do" is tools. A tool is the hand through which an agent actually touches the world. So how strong an agent is depends largely on how well its tools are designed: whether the interface is clear, whether the error feedback lands, whether dangerous operations get stopped.
6
6
 
7
- CoreCoder gives the model seven tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `agent`. In this piece we first look at the skeleton they share, then dig into the two most worth discussing, and finally I'll have you write one of your own.
7
+ CoreCoder gives the model eight tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `todo_write`, `agent`. In this piece we first look at the skeleton they share, then dig into the two most worth discussing, and finally I'll have you write one of your own.
8
8
 
9
9
  ## What a tool looks like
10
10
 
11
- Every tool inherits from `Tool` in `tools/base.py`, the whole base class being 27 lines:
11
+ Every tool inherits from `Tool` in `tools/base.py`, the whole base class being 32 lines:
12
12
 
13
13
  ```python
14
14
  class Tool(ABC):
@@ -57,7 +57,7 @@ To add a tool, drop an instance into this list. We'll actually do that at the en
57
57
 
58
58
  ## edit_file: a key innovation that looks unremarkable
59
59
 
60
- If I could only discuss one of the seven tools, I'd pick `edit_file`. Because "let the model modify an existing file," a need that looks simple, has several dead bodies behind it.
60
+ If I could only discuss one of the eight tools, I'd pick `edit_file`. Because "let the model modify an existing file," a need that looks simple, has several dead bodies behind it.
61
61
 
62
62
  The first dead end is having the model patch by line number, say "replace line 42 with this." The problem is the model's sense of line numbers is wildly unreliable; the line 42 in its head and the real line 42 in the file often don't match, and being off by one means editing the wrong place. Worse, the moment something earlier in the file gets touched, every line number after it shifts.
63
63
 
@@ -115,7 +115,7 @@ Without this check, the moment the model accidentally runs `edit_file` on a bina
115
115
 
116
116
  Tools like `read_file` and `edit_file` can only do limited damage. `bash` is different; it runs arbitrary shell commands, and the moment the model writes `rm -rf /`, the consequences are real.
117
117
 
118
- Claude Code's `BashTool` is 1,143 lines in public teardowns, with a command classifier, a real sandbox built on `sandbox-exec` and `seccomp`, output truncation, and interactive-command interception. CoreCoder's `bash.py` is a 127-line distillation that keeps the four most essential things: dangerous-command detection, output truncation, timeout, and working-directory tracking.
118
+ Claude Code's `BashTool` is 1,143 lines in public teardowns, with a command classifier, a real sandbox built on `sandbox-exec` and `seccomp`, output truncation, and interactive-command interception. CoreCoder's `bash.py` is a 203-line distillation that keeps the four most essential things: dangerous-command detection, output truncation, timeout, and working-directory tracking.
119
119
 
120
120
  Dangerous-command detection is a regex blocklist:
121
121
 
@@ -161,7 +161,7 @@ The second is hooks. `~/.corecoder/hooks.json` lets users hang shell commands on
161
161
 
162
162
  The third is MCP. In `mcp.py`, each server is a subprocess speaking newline-delimited JSON-RPC over stdio: handshake, `tools/list`, and then every remote tool registers as `mcp__server__tool`, after which consent, hooks, and plan mode treat it exactly like the built-ins. The lesson is where that "exactly like" comes from: not from special-casing in the MCP code, but from a tool boundary (the `Tool` base class plus the three gates) clean enough that an external tool is just one more implementation. A wedged or dead server costs one error string on that one call; the loop keeps turning.
163
163
 
164
- All three are advanced pieces. They are not on the agent's skeleton — the skeleton is still the loop plus seven tools. But they answer the same question: around a tool call, who else gets to speak? The answer went from "two gates" to "two gates plus a ring of pluggable bypasses," and every bypass is designed to fail without dragging the main loop down with it.
164
+ All three are advanced pieces. They are not on the agent's skeleton — the skeleton is still the loop plus eight tools. But they answer the same question: around a tool call, who else gets to speak? The answer went from "two gates" to "two gates plus a ring of pluggable bypasses," and every bypass is designed to fail without dragging the main loop down with it.
165
165
 
166
166
  ## Hands-on: write your first tool
167
167
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  前两篇讲的是循环和工具,也就是 agent 的手脚。这一篇讲大脑接口:模型怎么接进来,流式输出怎么处理,provider 抽风了怎么扛,以及一个被很多教程跳过、但你上线后第一天就会关心的问题,这一轮到底花了多少钱。
4
4
 
5
- 对应的文件是 `corecoder/llm.py`,336 行,是整个项目最大的单文件。它大,是因为它替你扛下了和真实 API 打交道时所有不优雅的部分。
5
+ 对应的文件是 `corecoder/llm.py`,332 行,是引擎里最大的单文件。它大,是因为它替你扛下了和真实 API 打交道时所有不优雅的部分。
6
6
 
7
7
  ## 一个赌注:大家都长得像 OpenAI
8
8
 
@@ -183,7 +183,7 @@ Tokens: 12043 prompt + 3201 completion = 15244 total (~$0.0621)
183
183
 
184
184
  ## 配置从哪来
185
185
 
186
- 最后串一下 `config.py`(57 行)。它从环境变量读配置,带一个合理的优先级:
186
+ 最后串一下 `config.py`(55 行)。它从环境变量读配置,带一个合理的优先级:
187
187
 
188
188
  ```python
189
189
  api_key = (
@@ -2,7 +2,7 @@
2
2
 
3
3
  The last two pieces covered the loop and the tools, the agent's hands and feet. This piece covers the brain's interface: how the model gets plugged in, how streaming output is handled, how to survive a provider acting up, and a question many tutorials skip but that you'll care about on day one after going live, namely how much this round actually cost.
4
4
 
5
- The file is `corecoder/llm.py`, 336 lines, the largest single file in the whole project. It's large because it carries, on your behalf, all the inelegant parts of dealing with a real API.
5
+ The file is `corecoder/llm.py`, 332 lines, the largest single file in the engine. It's large because it carries, on your behalf, all the inelegant parts of dealing with a real API.
6
6
 
7
7
  ## A bet: everyone looks like OpenAI
8
8
 
@@ -183,7 +183,7 @@ Of course this is only an estimate; the price table goes stale, and cache discou
183
183
 
184
184
  ## Where config comes from
185
185
 
186
- Finally, a thread through `config.py` (57 lines). It reads config from environment variables with a sensible priority:
186
+ Finally, a thread through `config.py` (55 lines). It reads config from environment variables with a sensible priority:
187
187
 
188
188
  ```python
189
189
  api_key = (