corecoder 0.6.0__tar.gz → 0.7.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {corecoder-0.6.0 → corecoder-0.7.0}/PKG-INFO +55 -22
- {corecoder-0.6.0 → corecoder-0.7.0}/README.md +54 -21
- {corecoder-0.6.0 → corecoder-0.7.0}/README_CN.md +54 -22
- {corecoder-0.6.0 → corecoder-0.7.0}/article/00-index.md +2 -1
- {corecoder-0.6.0 → corecoder-0.7.0}/article/00-index_EN.md +2 -1
- {corecoder-0.6.0 → corecoder-0.7.0}/article/02-tools.md +12 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/02-tools_EN.md +12 -0
- corecoder-0.7.0/article/08-extensibility.md +98 -0
- corecoder-0.7.0/article/08-extensibility_EN.md +98 -0
- corecoder-0.7.0/assets/demo-plan-hooks.gif +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/__init__.py +1 -1
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/agent.py +30 -3
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/checkpoints.py +6 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/cli.py +14 -2
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/context.py +12 -2
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/demo.py +1 -1
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/hooks.py +4 -2
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/llm.py +77 -12
- corecoder-0.7.0/corecoder/shell.py +61 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/agent.py +8 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/base.py +5 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/bash.py +86 -14
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/edit.py +26 -23
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/write.py +8 -5
- corecoder-0.7.0/examples/plan_hooks_demo.py +140 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/pyproject.toml +1 -1
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_checkpoints.py +17 -0
- corecoder-0.7.0/tests/test_core.py +734 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_demo.py +18 -1
- corecoder-0.7.0/tests/test_safety_matrix.py +208 -0
- corecoder-0.7.0/tests/test_shell.py +71 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_tools.py +76 -0
- corecoder-0.6.0/tests/test_core.py +0 -311
- {corecoder-0.6.0 → corecoder-0.7.0}/.github/workflows/ci.yml +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/.github/workflows/publish.yml +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/.gitignore +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/LICENSE +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/01-the-loop.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/01-the-loop_EN.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/03-llm-and-cost.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/03-llm-and-cost_EN.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/04-context.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/04-context_EN.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/05-parallel-and-subagents.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/05-parallel-and-subagents_EN.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/06-session-and-cli.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/06-session-and-cli_EN.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/07-build-your-own.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/article/07-build-your-own_EN.md +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/assets/demo.png +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/assets/demo_en.png +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/__main__.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/config.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/mcp.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/permissions.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/prompt.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/session.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/__init__.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/glob_tool.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/grep.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/read.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/corecoder/tools/todo.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/__init__.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/conftest.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_hooks.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_litellm.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_mcp.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_permissions.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_plan_mode.py +0 -0
- {corecoder-0.6.0 → corecoder-0.7.0}/tests/test_session.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: corecoder
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.7.0
|
|
4
4
|
Summary: Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. (formerly NanoCoder)
|
|
5
5
|
Project-URL: Homepage, https://github.com/he-yufeng/CoreCoder
|
|
6
6
|
Project-URL: Repository, https://github.com/he-yufeng/CoreCoder
|
|
@@ -37,7 +37,7 @@ Description-Content-Type: text/markdown
|
|
|
37
37
|
|
|
38
38
|
# CoreCoder
|
|
39
39
|
|
|
40
|
-
**The nanoGPT of coding agents. 1,
|
|
40
|
+
**The nanoGPT of coding agents. A 1.3k-line engine inside 2,658 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
|
|
41
41
|
|
|
42
42
|
*learn from it · fork it · ship something better*
|
|
43
43
|
|
|
@@ -47,7 +47,7 @@ Description-Content-Type: text/markdown
|
|
|
47
47
|
[](https://python.org)
|
|
48
48
|
[](LICENSE)
|
|
49
49
|
[](https://github.com/he-yufeng/CoreCoder/actions)
|
|
50
|
-
[](article/00-index_EN.md)
|
|
51
51
|
[](article/00-index_EN.md)
|
|
52
52
|
|
|
53
53
|
</div>
|
|
@@ -60,7 +60,7 @@ Description-Content-Type: text/markdown
|
|
|
60
60
|
|
|
61
61
|
| | CoreCoder | Claude Code | aider | nanoGPT |
|
|
62
62
|
|---|---|---|---|---|
|
|
63
|
-
| Lines of code | ~1,
|
|
63
|
+
| Lines of code | ~1,309 engine / 2,658 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
|
|
64
64
|
| Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
|
|
65
65
|
| Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
|
|
66
66
|
| What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
|
|
@@ -71,15 +71,15 @@ The nanoGPT column is there as a reference point: minimal, readable, but it teac
|
|
|
71
71
|
|
|
72
72
|
I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
|
|
73
73
|
|
|
74
|
-
The engine (loop, model interface, context, tools, sessions) is 1,
|
|
74
|
+
The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,658 physical lines, 2,138 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
|
|
75
75
|
|
|
76
|
-
And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first.
|
|
76
|
+
And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 171 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
|
|
77
77
|
|
|
78
78
|
The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
|
|
79
79
|
|
|
80
80
|
<p align="center">
|
|
81
|
-
<img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/
|
|
82
|
-
alt="
|
|
81
|
+
<img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo-plan-hooks.gif" width="760"
|
|
82
|
+
alt="Plan mode in action: the agent reads fib.py, gets its edit refused while plan mode is on, presents a plan, and only after approval edits, tests, and reports — with Pre/PostToolUse hooks firing around every call.">
|
|
83
83
|
</p>
|
|
84
84
|
|
|
85
85
|
<p align="center"><sub><i>These thousand lines really do run a full loop end to end: ask it to fix buggy.py and it reads the file, edits the code, runs it once to confirm, then reports back on its own. Watch it, then come back and read the code.</i></sub></p>
|
|
@@ -107,7 +107,9 @@ Give it a model and a key and it goes. It speaks the OpenAI-compatible API by de
|
|
|
107
107
|
| OmniRoute | `OPENAI_API_KEY=your-key OPENAI_BASE_URL=http://localhost:20128/v1 CORECODER_MODEL=auto` |
|
|
108
108
|
| Local Ollama | `OPENAI_API_KEY=ollama OPENAI_BASE_URL=http://localhost:11434/v1 CORECODER_MODEL=qwen2.5-coder` |
|
|
109
109
|
|
|
110
|
-
Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
|
|
110
|
+
Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. Thinking models are first-class too: deepseek-reasoner, kimi-k3 and friends stream their chain-of-thought, and CoreCoder shows it dimmed as it works, kept out of the conversation history so providers never see it come back. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
|
|
111
|
+
|
|
112
|
+
Smoke-tested end to end (read the file, edit it, run it, report back) against DeepSeek, Qwen3 and Kimi K2 via a single OpenRouter-compatible endpoint; each completed the full loop. One note for one-shot scripts: `-p` refuses mutating tools unless you pass `--yes`, by design.
|
|
111
113
|
|
|
112
114
|
```bash
|
|
113
115
|
corecoder # interactive REPL
|
|
@@ -120,26 +122,31 @@ Laid out flat, the whole project is this big. Skim it before you clone and you'l
|
|
|
120
122
|
|
|
121
123
|
```
|
|
122
124
|
corecoder/
|
|
123
|
-
├── agent.py agent loop + parallel tool exec
|
|
124
|
-
├── llm.py streaming client + retry + cost
|
|
125
|
-
├── context.py three-tier context compaction
|
|
125
|
+
├── agent.py agent loop + parallel tool exec 240 lines ← start here
|
|
126
|
+
├── llm.py streaming client + retry + cost 332 lines
|
|
127
|
+
├── context.py three-tier context compaction 220 lines
|
|
126
128
|
├── session.py save / resume + path-traversal guard 97 lines
|
|
127
129
|
├── permissions.py consent for mutating tools 48 lines
|
|
128
|
-
├── hooks.py Pre/PostToolUse shell hooks
|
|
130
|
+
├── hooks.py Pre/PostToolUse shell hooks 87 lines
|
|
131
|
+
├── shell.py POSIX shell routing (Git Bash on Windows) 61 lines
|
|
129
132
|
├── mcp.py MCP stdio client for external tools 208 lines
|
|
130
133
|
├── prompt.py system prompt 41 lines
|
|
131
|
-
├── cli.py REPL + slash commands + one-shot
|
|
134
|
+
├── cli.py REPL + slash commands + one-shot 358 lines
|
|
132
135
|
├── config.py env-var config 55 lines
|
|
136
|
+
├── checkpoints.py /undo snapshot and restore 44 lines
|
|
137
|
+
├── demo.py offline end-to-end demo 100 lines
|
|
133
138
|
└── tools/
|
|
134
|
-
├── bash.py shell + dangerous-command gate + cd
|
|
135
|
-
├── edit.py unique-match search/replace + diff
|
|
139
|
+
├── bash.py shell + dangerous-command gate + cd 203 lines
|
|
140
|
+
├── edit.py unique-match search/replace + diff 99 lines
|
|
136
141
|
├── grep.py content search 93 lines
|
|
137
142
|
├── glob_tool.py filename matching 52 lines
|
|
138
143
|
├── read.py file read 56 lines
|
|
139
|
-
├── write.py file write
|
|
144
|
+
├── write.py file write 46 lines
|
|
140
145
|
├── todo.py agent-maintained task checklist 79 lines
|
|
141
|
-
├── agent.py sub-agent spawning
|
|
142
|
-
└── base.py tool base class
|
|
146
|
+
├── agent.py sub-agent spawning 72 lines
|
|
147
|
+
└── base.py tool base class 32 lines
|
|
148
|
+
examples/
|
|
149
|
+
└── plan_hooks_demo.py offline plan mode + hooks demo (no API key)
|
|
143
150
|
```
|
|
144
151
|
|
|
145
152
|
Eight tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `todo_write` (a task checklist the agent maintains for itself), and `agent` (which spawns a sub-agent). Everything else is the CLI shell, config, and packaging wrapped around that engine core. If `~/.corecoder/mcp.json` exists, its MCP servers join the eight as extra `mcp__*` tools; the MCP section below covers it.
|
|
@@ -163,7 +170,7 @@ def chat(self, user_input):
|
|
|
163
170
|
return "(hit the round limit)"
|
|
164
171
|
```
|
|
165
172
|
|
|
166
|
-
That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
|
|
173
|
+
That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
|
|
167
174
|
|
|
168
175
|
Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
|
|
169
176
|
|
|
@@ -177,7 +184,7 @@ Every one of these *whys* is traced down to the actual lines of code in the seri
|
|
|
177
184
|
|
|
178
185
|
## The source-reading series · 8 bilingual essays
|
|
179
186
|
|
|
180
|
-
I also wrote a bilingual source-reading series, one intro plus
|
|
187
|
+
I also wrote a bilingual source-reading series, one intro plus eight parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking, and the eighth is about extending it without touching the loop; read them in any order.
|
|
181
188
|
|
|
182
189
|
- **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
|
|
183
190
|
- **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
|
|
@@ -187,6 +194,7 @@ I also wrote a bilingual source-reading series, one intro plus seven parts, each
|
|
|
187
194
|
- **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
|
|
188
195
|
- **[06 · Turning it into a real command-line tool](article/06-session-and-cli_EN.md)** — `session.py` and path-traversal defense
|
|
189
196
|
- **[07 · Fork CoreCoder into your own coding agent](article/07-build-your-own_EN.md)** — from fork to custom tools to swapping models
|
|
197
|
+
- **[08 · Three ways to extend without touching the loop: MCP, hooks, and plan mode](article/08-extensibility_EN.md)** — the v0.6.0 extensibility trio and the contract that makes them safe
|
|
190
198
|
|
|
191
199
|
## Fork it, build something better
|
|
192
200
|
|
|
@@ -255,6 +263,31 @@ Drop a `hooks.json` under `~/.corecoder` and your own shell commands run around
|
|
|
255
263
|
|
|
256
264
|
Each hook gets the call as JSON on stdin (`tool_name`, `tool_input`; post hooks also get `tool_response`). The matcher is an exact tool name; empty or `*` fires on every tool. A pre hook can veto the call with exit code 2, and its stderr travels back to the model as the reason so it can route around the block. Post hooks only observe and can never block. A hook that errors or runs past ten seconds is skipped with a warning: hooks assist the loop, they never get to kill it. The whole mechanism is `hooks.py`, and the REPL banner shows how many hooks loaded.
|
|
257
265
|
|
|
266
|
+
Two worth stealing (the commands lean on `jq`, the usual suspect):
|
|
267
|
+
|
|
268
|
+
```bash
|
|
269
|
+
# 1. lint gate: after every edit/write, run the project's fast linter on the
|
|
270
|
+
# touched file. The model sees the output and fixes its own mistakes in
|
|
271
|
+
# the same turn instead of waiting for CI.
|
|
272
|
+
{
|
|
273
|
+
"PostToolUse": [{
|
|
274
|
+
"matcher": "edit",
|
|
275
|
+
"command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
|
|
276
|
+
}]
|
|
277
|
+
}
|
|
278
|
+
|
|
279
|
+
# 2. write protect: refuse edits under paths you never want an agent to
|
|
280
|
+
# touch. Exit code 2 vetoes the call and the message reaches the model.
|
|
281
|
+
{
|
|
282
|
+
"PreToolUse": [{
|
|
283
|
+
"matcher": "edit",
|
|
284
|
+
"command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
|
|
285
|
+
}]
|
|
286
|
+
}
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Both are plain shell; nothing here is CoreCoder-specific syntax beyond the JSON shape and the exit-code-2 veto.
|
|
290
|
+
|
|
258
291
|
## MCP servers
|
|
259
292
|
|
|
260
293
|
Drop a `mcp.json` under `~/.corecoder` and tools from any MCP server join the agent over stdio, the same config shape as Claude Code's:
|
|
@@ -281,7 +314,7 @@ If working through CoreCoder was useful, here are a few other tools I've built a
|
|
|
281
314
|
|
|
282
315
|
## Contributing / License
|
|
283
316
|
|
|
284
|
-
Before you send anything, run `pytest tests/ -q` (
|
|
317
|
+
Before you send anything, run `pytest tests/ -q` (171 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
|
|
285
318
|
|
|
286
319
|
---
|
|
287
320
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# CoreCoder
|
|
4
4
|
|
|
5
|
-
**The nanoGPT of coding agents. 1,
|
|
5
|
+
**The nanoGPT of coding agents. A 1.3k-line engine inside 2,658 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
|
|
6
6
|
|
|
7
7
|
*learn from it · fork it · ship something better*
|
|
8
8
|
|
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
[](https://python.org)
|
|
13
13
|
[](LICENSE)
|
|
14
14
|
[](https://github.com/he-yufeng/CoreCoder/actions)
|
|
15
|
-
[](article/00-index_EN.md)
|
|
16
16
|
[](article/00-index_EN.md)
|
|
17
17
|
|
|
18
18
|
</div>
|
|
@@ -25,7 +25,7 @@
|
|
|
25
25
|
|
|
26
26
|
| | CoreCoder | Claude Code | aider | nanoGPT |
|
|
27
27
|
|---|---|---|---|---|
|
|
28
|
-
| Lines of code | ~1,
|
|
28
|
+
| Lines of code | ~1,309 engine / 2,658 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
|
|
29
29
|
| Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
|
|
30
30
|
| Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
|
|
31
31
|
| What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
|
|
@@ -36,15 +36,15 @@ The nanoGPT column is there as a reference point: minimal, readable, but it teac
|
|
|
36
36
|
|
|
37
37
|
I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
|
|
38
38
|
|
|
39
|
-
The engine (loop, model interface, context, tools, sessions) is 1,
|
|
39
|
+
The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,658 physical lines, 2,138 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
|
|
40
40
|
|
|
41
|
-
And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first.
|
|
41
|
+
And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 171 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
|
|
42
42
|
|
|
43
43
|
The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
|
|
44
44
|
|
|
45
45
|
<p align="center">
|
|
46
|
-
<img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/
|
|
47
|
-
alt="
|
|
46
|
+
<img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo-plan-hooks.gif" width="760"
|
|
47
|
+
alt="Plan mode in action: the agent reads fib.py, gets its edit refused while plan mode is on, presents a plan, and only after approval edits, tests, and reports — with Pre/PostToolUse hooks firing around every call.">
|
|
48
48
|
</p>
|
|
49
49
|
|
|
50
50
|
<p align="center"><sub><i>These thousand lines really do run a full loop end to end: ask it to fix buggy.py and it reads the file, edits the code, runs it once to confirm, then reports back on its own. Watch it, then come back and read the code.</i></sub></p>
|
|
@@ -72,7 +72,9 @@ Give it a model and a key and it goes. It speaks the OpenAI-compatible API by de
|
|
|
72
72
|
| OmniRoute | `OPENAI_API_KEY=your-key OPENAI_BASE_URL=http://localhost:20128/v1 CORECODER_MODEL=auto` |
|
|
73
73
|
| Local Ollama | `OPENAI_API_KEY=ollama OPENAI_BASE_URL=http://localhost:11434/v1 CORECODER_MODEL=qwen2.5-coder` |
|
|
74
74
|
|
|
75
|
-
Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
|
|
75
|
+
Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. Thinking models are first-class too: deepseek-reasoner, kimi-k3 and friends stream their chain-of-thought, and CoreCoder shows it dimmed as it works, kept out of the conversation history so providers never see it come back. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
|
|
76
|
+
|
|
77
|
+
Smoke-tested end to end (read the file, edit it, run it, report back) against DeepSeek, Qwen3 and Kimi K2 via a single OpenRouter-compatible endpoint; each completed the full loop. One note for one-shot scripts: `-p` refuses mutating tools unless you pass `--yes`, by design.
|
|
76
78
|
|
|
77
79
|
```bash
|
|
78
80
|
corecoder # interactive REPL
|
|
@@ -85,26 +87,31 @@ Laid out flat, the whole project is this big. Skim it before you clone and you'l
|
|
|
85
87
|
|
|
86
88
|
```
|
|
87
89
|
corecoder/
|
|
88
|
-
├── agent.py agent loop + parallel tool exec
|
|
89
|
-
├── llm.py streaming client + retry + cost
|
|
90
|
-
├── context.py three-tier context compaction
|
|
90
|
+
├── agent.py agent loop + parallel tool exec 240 lines ← start here
|
|
91
|
+
├── llm.py streaming client + retry + cost 332 lines
|
|
92
|
+
├── context.py three-tier context compaction 220 lines
|
|
91
93
|
├── session.py save / resume + path-traversal guard 97 lines
|
|
92
94
|
├── permissions.py consent for mutating tools 48 lines
|
|
93
|
-
├── hooks.py Pre/PostToolUse shell hooks
|
|
95
|
+
├── hooks.py Pre/PostToolUse shell hooks 87 lines
|
|
96
|
+
├── shell.py POSIX shell routing (Git Bash on Windows) 61 lines
|
|
94
97
|
├── mcp.py MCP stdio client for external tools 208 lines
|
|
95
98
|
├── prompt.py system prompt 41 lines
|
|
96
|
-
├── cli.py REPL + slash commands + one-shot
|
|
99
|
+
├── cli.py REPL + slash commands + one-shot 358 lines
|
|
97
100
|
├── config.py env-var config 55 lines
|
|
101
|
+
├── checkpoints.py /undo snapshot and restore 44 lines
|
|
102
|
+
├── demo.py offline end-to-end demo 100 lines
|
|
98
103
|
└── tools/
|
|
99
|
-
├── bash.py shell + dangerous-command gate + cd
|
|
100
|
-
├── edit.py unique-match search/replace + diff
|
|
104
|
+
├── bash.py shell + dangerous-command gate + cd 203 lines
|
|
105
|
+
├── edit.py unique-match search/replace + diff 99 lines
|
|
101
106
|
├── grep.py content search 93 lines
|
|
102
107
|
├── glob_tool.py filename matching 52 lines
|
|
103
108
|
├── read.py file read 56 lines
|
|
104
|
-
├── write.py file write
|
|
109
|
+
├── write.py file write 46 lines
|
|
105
110
|
├── todo.py agent-maintained task checklist 79 lines
|
|
106
|
-
├── agent.py sub-agent spawning
|
|
107
|
-
└── base.py tool base class
|
|
111
|
+
├── agent.py sub-agent spawning 72 lines
|
|
112
|
+
└── base.py tool base class 32 lines
|
|
113
|
+
examples/
|
|
114
|
+
└── plan_hooks_demo.py offline plan mode + hooks demo (no API key)
|
|
108
115
|
```
|
|
109
116
|
|
|
110
117
|
Eight tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `todo_write` (a task checklist the agent maintains for itself), and `agent` (which spawns a sub-agent). Everything else is the CLI shell, config, and packaging wrapped around that engine core. If `~/.corecoder/mcp.json` exists, its MCP servers join the eight as extra `mcp__*` tools; the MCP section below covers it.
|
|
@@ -128,7 +135,7 @@ def chat(self, user_input):
|
|
|
128
135
|
return "(hit the round limit)"
|
|
129
136
|
```
|
|
130
137
|
|
|
131
|
-
That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
|
|
138
|
+
That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
|
|
132
139
|
|
|
133
140
|
Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
|
|
134
141
|
|
|
@@ -142,7 +149,7 @@ Every one of these *whys* is traced down to the actual lines of code in the seri
|
|
|
142
149
|
|
|
143
150
|
## The source-reading series · 8 bilingual essays
|
|
144
151
|
|
|
145
|
-
I also wrote a bilingual source-reading series, one intro plus
|
|
152
|
+
I also wrote a bilingual source-reading series, one intro plus eight parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking, and the eighth is about extending it without touching the loop; read them in any order.
|
|
146
153
|
|
|
147
154
|
- **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
|
|
148
155
|
- **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
|
|
@@ -152,6 +159,7 @@ I also wrote a bilingual source-reading series, one intro plus seven parts, each
|
|
|
152
159
|
- **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
|
|
153
160
|
- **[06 · Turning it into a real command-line tool](article/06-session-and-cli_EN.md)** — `session.py` and path-traversal defense
|
|
154
161
|
- **[07 · Fork CoreCoder into your own coding agent](article/07-build-your-own_EN.md)** — from fork to custom tools to swapping models
|
|
162
|
+
- **[08 · Three ways to extend without touching the loop: MCP, hooks, and plan mode](article/08-extensibility_EN.md)** — the v0.6.0 extensibility trio and the contract that makes them safe
|
|
155
163
|
|
|
156
164
|
## Fork it, build something better
|
|
157
165
|
|
|
@@ -220,6 +228,31 @@ Drop a `hooks.json` under `~/.corecoder` and your own shell commands run around
|
|
|
220
228
|
|
|
221
229
|
Each hook gets the call as JSON on stdin (`tool_name`, `tool_input`; post hooks also get `tool_response`). The matcher is an exact tool name; empty or `*` fires on every tool. A pre hook can veto the call with exit code 2, and its stderr travels back to the model as the reason so it can route around the block. Post hooks only observe and can never block. A hook that errors or runs past ten seconds is skipped with a warning: hooks assist the loop, they never get to kill it. The whole mechanism is `hooks.py`, and the REPL banner shows how many hooks loaded.
|
|
222
230
|
|
|
231
|
+
Two worth stealing (the commands lean on `jq`, the usual suspect):
|
|
232
|
+
|
|
233
|
+
```bash
|
|
234
|
+
# 1. lint gate: after every edit/write, run the project's fast linter on the
|
|
235
|
+
# touched file. The model sees the output and fixes its own mistakes in
|
|
236
|
+
# the same turn instead of waiting for CI.
|
|
237
|
+
{
|
|
238
|
+
"PostToolUse": [{
|
|
239
|
+
"matcher": "edit",
|
|
240
|
+
"command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
|
|
241
|
+
}]
|
|
242
|
+
}
|
|
243
|
+
|
|
244
|
+
# 2. write protect: refuse edits under paths you never want an agent to
|
|
245
|
+
# touch. Exit code 2 vetoes the call and the message reaches the model.
|
|
246
|
+
{
|
|
247
|
+
"PreToolUse": [{
|
|
248
|
+
"matcher": "edit",
|
|
249
|
+
"command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
|
|
250
|
+
}]
|
|
251
|
+
}
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
Both are plain shell; nothing here is CoreCoder-specific syntax beyond the JSON shape and the exit-code-2 veto.
|
|
255
|
+
|
|
223
256
|
## MCP servers
|
|
224
257
|
|
|
225
258
|
Drop a `mcp.json` under `~/.corecoder` and tools from any MCP server join the agent over stdio, the same config shape as Claude Code's:
|
|
@@ -246,7 +279,7 @@ If working through CoreCoder was useful, here are a few other tools I've built a
|
|
|
246
279
|
|
|
247
280
|
## Contributing / License
|
|
248
281
|
|
|
249
|
-
Before you send anything, run `pytest tests/ -q` (
|
|
282
|
+
Before you send anything, run `pytest tests/ -q` (171 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
|
|
250
283
|
|
|
251
284
|
---
|
|
252
285
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# CoreCoder
|
|
4
4
|
|
|
5
|
-
**编程 agent 里的 nanoGPT。
|
|
5
|
+
**编程 agent 里的 nanoGPT。1.3k 行引擎、整包 2594 行纯 Python 全部一口气可读,读懂一个 coding agent 到底怎么运作,再 fork 出你自己的。**
|
|
6
6
|
|
|
7
7
|
*learn from it · fork it · ship something better*
|
|
8
8
|
|
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
[](https://python.org)
|
|
13
13
|
[](LICENSE)
|
|
14
14
|
[](https://github.com/he-yufeng/CoreCoder/actions)
|
|
15
|
-
[](article/)
|
|
16
16
|
[](article/)
|
|
17
17
|
|
|
18
18
|
</div>
|
|
@@ -25,7 +25,7 @@
|
|
|
25
25
|
|
|
26
26
|
| | CoreCoder | Claude Code | aider | nanoGPT |
|
|
27
27
|
|---|---|---|---|---|
|
|
28
|
-
| 代码量 | 引擎约
|
|
28
|
+
| 代码量 | 引擎约 1308 行 / 整包 2594 行 | 几十万行(闭源) | 数万行 Python | 约 600 行(两个文件) |
|
|
29
29
|
| 读完要多久 | 一个下午 | 读不了(闭源) | 得啃几天 | 一个下午 |
|
|
30
30
|
| 能不能下断点改了再跑 | 能,每一行 | 不能 | 能,但量大 | 能 |
|
|
31
31
|
| 定位 | 读懂并 fork 出你自己的 agent | 生产级编程助手 | 终端结对编程 | 教学用最小 GPT |
|
|
@@ -36,15 +36,15 @@ nanoGPT 那一列是拿来对照的:它最小、可读,但教的是训一个
|
|
|
36
36
|
|
|
37
37
|
我一直觉得 coding agent 被讲得太玄了。把 Claude Code、Cursor 这类工具扒到底,核心是一个 while 循环套着一个大模型,外加七八个让它能真正动手的工具。难的从来不是这个循环,而是循环跑进真实世界以后要兜的那些底。CoreCoder 就是把这个核心老老实实写出来的最小版本。
|
|
38
38
|
|
|
39
|
-
引擎部分(循环、模型接口、上下文、工具、会话)去掉空行和注释是
|
|
39
|
+
引擎部分(循环、模型接口、上下文、工具、会话)去掉空行和注释是 1309 行。连最外层的 CLI、配置、打包一起算,整个包 25 个文件、物理 2658 行、净 2138 行,每个文件都短到能一口气读完。自 1161 行快照之后的增长都花在了看得见的功能上:plan mode、hooks、checkpoints,下文各有交代。
|
|
40
40
|
|
|
41
|
-
它真能跑:读写文件、执行 shell、派子 agent、分三层压上下文,还能随时把这趟烧掉的 token 和美元数报给你。任何要动你磁盘、要跑命令的调用,都会先停下来等你点头,
|
|
41
|
+
它真能跑:读写文件、执行 shell、派子 agent、分三层压上下文,还能随时把这趟烧掉的 token 和美元数报给你。任何要动你磁盘、要跑命令的调用,都会先停下来等你点头,171 个测试是绿的。但能跑不是为了劝你拿去日用,而是为了让这份「注释」不撒谎:一个解释 agent 怎么运作的范例,自己得真能运作。
|
|
42
42
|
|
|
43
43
|
代码来自一次公开拆解。公开的源码分析里,Claude Code 这类生产级 agent 暴露出不少关键架构,我挑出最核心的一层,用尽量少的代码诚实地复写了一遍。所以读 CoreCoder,约等于读一份基于公开源码分析的「可运行注释版」:讲的是这类 agent 的核心思路,而它本身只是最小复写,就摆在你机器上,随你拆、随你改。
|
|
44
44
|
|
|
45
45
|
<p align="center">
|
|
46
|
-
<img src="assets/demo.
|
|
47
|
-
alt="
|
|
46
|
+
<img src="assets/demo-plan-hooks.gif" width="760"
|
|
47
|
+
alt="plan mode 实战:agent 先读 fib.py,写入被 plan mode 拦下,先给计划;批准之后才动手改文件、跑测试,Pre/PostToolUse hooks 在每次调用前后触发">
|
|
48
48
|
</p>
|
|
49
49
|
|
|
50
50
|
<p align="center"><sub><i>这一千行真能跑通一个完整回合:让它修 buggy.py,它自己读文件、改代码、跑一遍确认、再给结论。看完就回来读代码。</i></sub></p>
|
|
@@ -72,7 +72,9 @@ pip install -e .
|
|
|
72
72
|
| OmniRoute | `OPENAI_API_KEY=your-key OPENAI_BASE_URL=http://localhost:20128/v1 CORECODER_MODEL=auto` |
|
|
73
73
|
| 本地 Ollama | `OPENAI_API_KEY=ollama OPENAI_BASE_URL=http://localhost:11434/v1 CORECODER_MODEL=qwen2.5-coder` |
|
|
74
74
|
|
|
75
|
-
Kimi、Qwen 这些同样是改这两个变量;连 OpenAI 兼容接口都不给的 provider,装上可选的 LiteLLM 后端(`pip install "corecoder[litellm]"
|
|
75
|
+
Kimi、Qwen 这些同样是改这两个变量;连 OpenAI 兼容接口都不给的 provider,装上可选的 LiteLLM 后端(`pip install "corecoder[litellm]"`)能路由一百多家。第三篇文章把这块讲得更细。思考模型也是一等公民:deepseek-reasoner、kimi-k3 这类模型的思考过程会实时流出来,CoreCoder 把它用暗色显示出来,但不进对话历史,provider 永远不会在回包里看到它。key 可以直接 `export`,也可以在项目根目录扔个 `.env`,启动时自动加载。然后:
|
|
76
|
+
|
|
77
|
+
端到端真机冒烟过三家(读文件、改代码、跑一次确认、自己报告):DeepSeek、Qwen3、Kimi K2,走同一个 OpenRouter 兼容端点,各自完整跑完全循环。写脚本用 one-shot 的留意:`-p` 默认拒绝一切改动类工具,要加 `--yes`,这是设计如此。
|
|
76
78
|
|
|
77
79
|
```bash
|
|
78
80
|
corecoder # 交互式 REPL
|
|
@@ -85,26 +87,31 @@ corecoder -p "给 parse_config() 加错误处理" # 一次性模式,干完
|
|
|
85
87
|
|
|
86
88
|
```
|
|
87
89
|
corecoder/
|
|
88
|
-
├── agent.py agent 主循环 + 并行工具执行
|
|
89
|
-
├── llm.py 流式客户端 + 重试 + 成本统计
|
|
90
|
-
├── context.py 三层上下文压缩
|
|
90
|
+
├── agent.py agent 主循环 + 并行工具执行 240 行 ← 从这里开始读
|
|
91
|
+
├── llm.py 流式客户端 + 重试 + 成本统计 332 行
|
|
92
|
+
├── context.py 三层上下文压缩 220 行
|
|
91
93
|
├── session.py 会话存盘 / 续聊 + 路径穿越防护 97 行
|
|
92
94
|
├── permissions.py 改动类工具的用户授权 48 行
|
|
93
|
-
├── hooks.py 工具调用前后的用户 shell 钩子
|
|
95
|
+
├── hooks.py 工具调用前后的用户 shell 钩子 87 行
|
|
96
|
+
├── shell.py POSIX shell 路由(Windows 走 Git Bash) 61 行
|
|
94
97
|
├── mcp.py MCP stdio 客户端,接外部工具 208 行
|
|
95
98
|
├── prompt.py 系统提示词 41 行
|
|
96
|
-
├── cli.py REPL + 斜杠命令 + 一次性模式
|
|
99
|
+
├── cli.py REPL + 斜杠命令 + 一次性模式 358 行
|
|
97
100
|
├── config.py 环境变量配置 55 行
|
|
101
|
+
├── checkpoints.py /undo 快照与回滚 44 行
|
|
102
|
+
├── demo.py 离线端到端演示 100 行
|
|
98
103
|
└── tools/
|
|
99
|
-
├── bash.py shell + 危险命令闸 + cd 追踪
|
|
100
|
-
├── edit.py 唯一匹配搜索替换 + diff
|
|
104
|
+
├── bash.py shell + 危险命令闸 + cd 追踪 203 行
|
|
105
|
+
├── edit.py 唯一匹配搜索替换 + diff 99 行
|
|
101
106
|
├── grep.py 内容搜索 93 行
|
|
102
107
|
├── glob_tool.py 文件名匹配 52 行
|
|
103
108
|
├── read.py 文件读取 56 行
|
|
104
|
-
├── write.py 文件写入
|
|
109
|
+
├── write.py 文件写入 46 行
|
|
105
110
|
├── todo.py agent 自维护的任务清单 79 行
|
|
106
|
-
├── agent.py 子 agent 派生
|
|
107
|
-
└── base.py 工具基类
|
|
111
|
+
├── agent.py 子 agent 派生 72 行
|
|
112
|
+
└── base.py 工具基类 32 行
|
|
113
|
+
examples/
|
|
114
|
+
└── plan_hooks_demo.py 离线 plan mode + hooks 演示(免 API key)
|
|
108
115
|
```
|
|
109
116
|
|
|
110
117
|
八个工具:`bash`、`read_file`、`write_file`、`edit_file`、`glob`、`grep`、`todo_write`(agent 自己维护的任务清单)、`agent`(派子 agent)。其余都是包在引擎核心外面的 CLI 外壳、配置和打包。存在 `~/.corecoder/mcp.json` 时,里面的 MCP 服务器会以 `mcp__*` 工具的身份并进这八件里,下面有专门一节讲。
|
|
@@ -128,7 +135,7 @@ def chat(self, user_input):
|
|
|
128
135
|
return "(已达轮次上限)"
|
|
129
136
|
```
|
|
130
137
|
|
|
131
|
-
就这么点。这个循环的核心骨架就二十来行,把并行执行和被 Ctrl+C 打断后的回填都算上,也才四十多行。CoreCoder 一千多行里剩下的,几乎全在收拾它真跑起来之后冒出来的岔子。`llm.py` 最后成了全项目最大的文件,不是因为调模型有多难,而是流式返回里一个工具调用的参数会被切成好几段先后送到、得按顺序拼回去,provider 偶尔吐半截 JSON 或把 usage 填成 null,限流(429)、超时、连接中断和 5xx 都得退避重试,其余 4xx
|
|
138
|
+
就这么点。这个循环的核心骨架就二十来行,把并行执行和被 Ctrl+C 打断后的回填都算上,也才四十多行。CoreCoder 一千多行里剩下的,几乎全在收拾它真跑起来之后冒出来的岔子。`llm.py` 最后成了全项目最大的文件,不是因为调模型有多难,而是流式返回里一个工具调用的参数会被切成好几段先后送到、得按顺序拼回去,provider 偶尔吐半截 JSON 或把 usage 填成 null,限流(429)、超时、连接中断和 5xx 都得退避重试,其余 4xx 该直接抛就别硬试。已经开始出字的流也可能半路断掉,所以重试包住的是整个请求,不只是建连。这些不起眼的脏活,而不是那个循环,才是一个 agent 从能演示走到能交付真正吃工程功夫的地方;第三篇文章顺着它拆到每一行。
|
|
132
139
|
|
|
133
140
|
有三个决定值得单独看,因为它们是「先读懂别人怎么做」之后才做得出的取舍,也是你 fork 自己 agent 时可以直接抄走的判断。
|
|
134
141
|
|
|
@@ -142,7 +149,7 @@ def chat(self, user_input):
|
|
|
142
149
|
|
|
143
150
|
## 配套源码导读 · 八篇双语
|
|
144
151
|
|
|
145
|
-
|
|
152
|
+
我还写了一套双语源码导读,一篇导言加八篇正文,每篇都配英文镜像(`_EN.md`)。它对着 CoreCoder 的真实代码,讲 Claude Code 这类 agent 的内部构造。有一条给自己立的硬规矩:每一处行数、每一段代码都从仓库里现读现核,绝不凭印象编。前六篇带你读懂,第七篇带你 fork,第八篇讲怎么不动主循环地扩展它,哪篇先读都行。
|
|
146
153
|
|
|
147
154
|
- **[导言 · 用 CoreCoder 读懂 Claude Code,再造一个你自己的](article/00-index.md)**
|
|
148
155
|
- **[01 一个 agent 的本体,是一个 while 循环](article/01-the-loop.md)** — `agent.py` 的主循环、打断与轮次上限
|
|
@@ -152,12 +159,13 @@ def chat(self, user_input):
|
|
|
152
159
|
- **[05 并行执行与子 agent](article/05-parallel-and-subagents.md)** — 线程池并发与子 agent 隔离
|
|
153
160
|
- **[06 把它跑成一个真正的命令行工具](article/06-session-and-cli.md)** — `session.py` 与路径穿越防护
|
|
154
161
|
- **[07 Fork CoreCoder,搭一个你自己的 coding agent](article/07-build-your-own.md)** — 从 fork 到加自定义工具到换模型
|
|
162
|
+
- **[08 不动主循环的三种加法:MCP、钩子与计划模式](article/08-extensibility.md)** — v0.6.0 扩展三件套,以及让它们成立的那条契约
|
|
155
163
|
|
|
156
164
|
## Fork 它,造个更好的
|
|
157
165
|
|
|
158
166
|
读懂之后,最自然的下一步就是 fork。起手不用伤筋动骨:
|
|
159
167
|
|
|
160
|
-
- **换个你常用的模型。** 就是上面那两个环境变量,`llm.py`(
|
|
168
|
+
- **换个你常用的模型。** 就是上面那两个环境变量,`llm.py`(331 行)是所有 provider 适配的入口。
|
|
161
169
|
- **加一件你自己的工具。** 照 `tools/base.py`(27 行)的工具基类写个新文件,跑测试、抓网页、调 LSP 都行,第二篇文章末尾手把手带你写第一个。
|
|
162
170
|
- **改系统提示词。** `prompt.py` 才 41 行,改一句就能看到 agent 的脾气变了,是门槛最低的「改一处就有反馈」。
|
|
163
171
|
- **直接当库 import。** 顶层导出了 `Agent`、`LLM`、`Config`,能嵌进你自己的程序:
|
|
@@ -220,6 +228,30 @@ REPL 里 `/plan` 开关计划模式。开着的时候,提示符变成 `(plan)`
|
|
|
220
228
|
|
|
221
229
|
每个钩子从 stdin 拿到这次调用的 JSON(`tool_name`、`tool_input`,post 钩子还带 `tool_response`)。matcher 是精确的工具名,留空或写 `*` 表示对所有工具生效。pre 钩子可以否决这次调用:退出码 2,它的 stderr 会作为理由回给模型,让它换条路走。post 钩子只观察,永远拦不住。钩子报错或超过十秒会被跳过并记一条警告:钩子是来帮忙的,没权力弄死循环。整个机制就是 `hooks.py` 一个文件,REPL 启动横幅会显示加载了几条。
|
|
222
230
|
|
|
231
|
+
两个能直接抄走的(命令里用了 `jq`,装一下就有):
|
|
232
|
+
|
|
233
|
+
```bash
|
|
234
|
+
# 1. lint 门:每次 edit/write 之后立刻对改动的文件跑快速 lint。
|
|
235
|
+
# 模型同一回合就能看到输出,自己把低级错误修了,不用等 CI 回来。
|
|
236
|
+
{
|
|
237
|
+
"PostToolUse": [{
|
|
238
|
+
"matcher": "edit",
|
|
239
|
+
"command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
|
|
240
|
+
}]
|
|
241
|
+
}
|
|
242
|
+
|
|
243
|
+
# 2. 写保护:指定路径一律不让 agent 碰。退出码 2 当场否决,
|
|
244
|
+
# 拒绝原因会送到模型那边。
|
|
245
|
+
{
|
|
246
|
+
"PreToolUse": [{
|
|
247
|
+
"matcher": "edit",
|
|
248
|
+
"command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
|
|
249
|
+
}]
|
|
250
|
+
}
|
|
251
|
+
```
|
|
252
|
+
|
|
253
|
+
两个例子都是纯 shell;除了 JSON 形状和退出码 2 否决这两个约定,没有任何 CoreCoder 特有的语法。
|
|
254
|
+
|
|
223
255
|
## MCP 服务器
|
|
224
256
|
|
|
225
257
|
在 `~/.corecoder` 下放一个 `mcp.json`,任何 MCP 服务器的工具就能通过 stdio 接进 agent,配置形状和 Claude Code 的一样:
|
|
@@ -246,7 +278,7 @@ REPL 里 `/plan` 开关计划模式。开着的时候,提示符变成 `(plan)`
|
|
|
246
278
|
|
|
247
279
|
## 贡献 / License
|
|
248
280
|
|
|
249
|
-
动手之前先跑一遍 `pytest tests/ -q`(
|
|
281
|
+
动手之前先跑一遍 `pytest tests/ -q`(171 个测试)、`ruff check` 和 `compileall`,绿了再提。MIT License,欢迎 fork 拿去造更好的东西,能在 README 里留一句出处就更好。
|
|
250
282
|
|
|
251
283
|
---
|
|
252
284
|
|
|
@@ -23,12 +23,13 @@ CoreCoder 做的事,是把这套骨架压到一千行出头的纯 Python。准
|
|
|
23
23
|
前六篇是「读懂」。每篇盯住 agent 的一个子系统,先讲 Claude Code 在这件事上的做法和取舍,再翻到 CoreCoder 对应的真实代码,看同一个想法被压缩成几十行后长什么样。最后一篇是「自己做」,把前面所有零件接起来,从 fork 到加一个自定义工具到换模型,落到一个能跑的成品。
|
|
24
24
|
|
|
25
25
|
1. [一个 agent 的本体,是一个 while 循环](01-the-loop.md)。整个 agent 最核心的东西,是一个「问模型、跑工具、把结果喂回去、再问」的循环。我们看 CoreCoder 的 `agent.py`(150 行)怎么把它写明白,以及打断、轮次上限、半截工具调用回填这些真实世界的麻烦各自怎么收场。
|
|
26
|
-
2. [工具系统:让模型安全地动手](02-tools.md)。模型本身只会吐字,是工具让它能读文件、写文件、跑命令。这篇讲 CoreCoder 的七个工具,重点是那个看似平平无奇、实则是 Claude Code 关键创新的「唯一性搜索替换」编辑,以及 bash
|
|
26
|
+
2. [工具系统:让模型安全地动手](02-tools.md)。模型本身只会吐字,是工具让它能读文件、写文件、跑命令。这篇讲 CoreCoder 的七个工具,重点是那个看似平平无奇、实则是 Claude Code 关键创新的「唯一性搜索替换」编辑,以及 bash 的安全闸。v0.6.0 之后还补了一节「门之外」:plan 模式、hooks、MCP 这三个进阶件怎么挂在一次调用的前后。末尾教你写第一个自己的工具。
|
|
27
27
|
3. [接入任意大模型,顺便把钱算清楚](03-llm-and-cost.md)。`llm.py`(336 行,全系列最大的文件)怎么用一套 OpenAI 兼容接口接住 DeepSeek、Qwen、Kimi、本地 Ollama,怎么做指数退避重试,怎么在流式输出里顺手把 token 和美元成本统计出来。
|
|
28
28
|
4. [用有限的窗口扛住一个长任务](04-context.md)。上下文窗口是 agent 的硬约束。`context.py`(210 行)实现了三层压缩,从轻到重。这篇还会讲一个特别容易踩、API 一定报错的坑:孤儿 tool 消息。这是我做这个项目时真改过的 bug。
|
|
29
29
|
5. [并行执行与子 agent](05-parallel-and-subagents.md)。模型一次返回多个工具调用时,CoreCoder 用线程池并发跑。这篇老实讲这个简化版相对 Claude Code 的流式执行器差在哪,并发又会引入什么新麻烦,以及子 agent 为什么不准递归。
|
|
30
30
|
6. [把它跑成一个真正的命令行工具](06-session-and-cli.md)。会话存盘、断点续聊、斜杠命令、一次性模式。`session.py` 里有个不起眼但很要命的安全细节:怎么防住用恶意会话名做路径穿越。
|
|
31
31
|
7. [Fork CoreCoder,搭一个你自己的 coding agent](07-build-your-own.md)。收尾的实操篇。从 clone 到换成你常用的模型,到加一个真正有用的自定义工具,到改系统提示词调教它的风格,到打包发布。读完前六篇你已经懂了原理,这篇让你真有一个东西。
|
|
32
|
+
8. [不动主循环的三种加法:MCP、钩子与计划模式](08-extensibility.md)。给 agent 加东西的三条正路:`mcp.py`(208 行)把远程工具接成与内建同权,hooks 用退出码 2 否决一次调用却永远杀不死循环,计划模式用一个布尔量压住包括 `--yes` 在内的整个权限层。最后一节把三块收拢回同一条契约:工具同形、错误是普通返回值、循环不被外围杀死。
|
|
32
33
|
|
|
33
34
|
不想按顺序也行。想搞懂「它凭什么敢自动跑命令」直接跳第二篇;想接自己的模型直接跳第三篇;只想赶紧 fork 出个能用的,第七篇是自洽的。
|
|
34
35
|
|
|
@@ -23,12 +23,13 @@ The answer is engineering. And the kind that, once you've read it, makes you thi
|
|
|
23
23
|
The first six are about understanding. Each one fixes on a single subsystem of the agent, first describing how Claude Code handles that thing and the tradeoffs it makes, then turning to the real code in CoreCoder to see what the same idea looks like once it's compressed into a few dozen lines. The last piece is about building it yourself, wiring all the parts back together, from fork to a custom tool to swapping the model, landing on something that runs.
|
|
24
24
|
|
|
25
25
|
1. [An agent is, at heart, a while loop](01-the-loop_EN.md). The most central thing in the whole agent is a loop: ask the model, run a tool, feed the result back, ask again. We look at how CoreCoder's `agent.py` (150 lines) writes it out plainly, and how real-world headaches like interruption, a round cap, and backfilling half-finished tool calls each get handled.
|
|
26
|
-
2. [The tool system: letting the model act, safely](02-tools_EN.md). The model on its own only emits text. Tools are what let it read files, write files, run commands. This piece covers CoreCoder's seven tools, with the spotlight on the seemingly unremarkable unique search-and-replace edit that is in fact one of Claude Code's key innovations, plus bash's safety gate. At the end I'll have you write your first tool.
|
|
26
|
+
2. [The tool system: letting the model act, safely](02-tools_EN.md). The model on its own only emits text. Tools are what let it read files, write files, run commands. This piece covers CoreCoder's seven tools, with the spotlight on the seemingly unremarkable unique search-and-replace edit that is in fact one of Claude Code's key innovations, plus bash's safety gate. Since v0.6.0 it also carries a "beyond the gates" section on how the three advanced pieces — plan mode, hooks, and MCP — hang off the moments around a call. At the end I'll have you write your first tool.
|
|
27
27
|
3. [Plug in any model, and get the bill right while you're at it](03-llm-and-cost_EN.md). How `llm.py` (336 lines, the largest single file in the project) uses one OpenAI-compatible interface to catch DeepSeek, Qwen, Kimi, and local Ollama, how it does exponential-backoff retry, and how it tallies tokens and dollar cost right inside the streaming output.
|
|
28
28
|
4. [Surviving a long task in a finite window](04-context_EN.md). The context window is the agent's hard constraint. `context.py` (210 lines) implements three layers of compression, lightest to heaviest. This piece also covers a trap that's easy to hit and that the API will always reject: the orphaned tool message. That's a bug I actually fixed while building this project.
|
|
29
29
|
5. [Parallel execution and sub-agents](05-parallel-and-subagents_EN.md). When the model returns several tool calls at once, CoreCoder runs them concurrently on a thread pool. This piece is honest about where this simplified version falls short of Claude Code's streaming executor, what new trouble concurrency brings in, and why a sub-agent is not allowed to recurse.
|
|
30
30
|
6. [Turning it into a real command-line tool](06-session-and-cli_EN.md). Session save, resume, slash commands, one-shot mode. `session.py` holds an unremarkable but critical security detail: how to stop a malicious session name from turning into a path traversal.
|
|
31
31
|
7. [Fork CoreCoder and build your own coding agent](07-build-your-own_EN.md). The hands-on finale. From clone, to switching to the model you actually use, to adding a genuinely useful custom tool, to tuning the system prompt to shape its style, to packaging and release. After the first six pieces you understand the principles; this one leaves you with something real.
|
|
32
|
+
8. [Three ways to extend without touching the loop: MCP, hooks, and plan mode](08-extensibility_EN.md). The three right ways to grow an agent: `mcp.py` (208 lines) wires remote tools in as first-class citizens, hooks veto a single call with exit code 2 yet can never kill the loop, and plan mode is one boolean that outranks the entire permission layer including `--yes`. The coda folds all three back into one contract: uniform tools, errors as ordinary return values, and a loop its perimeter cannot kill.
|
|
32
33
|
|
|
33
34
|
You don't have to read in order. Want to understand how it dares to run commands on its own? Jump to piece two. Want to plug in your own model? Jump to piece three. Just want to fork something usable fast? Piece seven stands on its own.
|
|
34
35
|
|
|
@@ -151,6 +151,18 @@ if warning:
|
|
|
151
151
|
|
|
152
152
|
这正对应 Claude Code 的两阶段门控,公开拆解里叫 `validateInput` 和 `checkPermissions`:一个验输入合不合法,一个验这个操作允不允许。把「格式对不对」和「该不该做」分成两关,好处是各自的失败能给出各自精准的反馈,模型也能针对性地修正。一个混在一起的大 try-except 做不到这种精度。
|
|
153
153
|
|
|
154
|
+
## 门之外:plan 模式、hooks,和住在别的进程里的工具
|
|
155
|
+
|
|
156
|
+
v0.6.0 在「调用前后」这条缝上又加了三样东西。它们不改变工具本身,改的是一次调用从「模型想调」到「真的执行」之间还经过谁。
|
|
157
|
+
|
|
158
|
+
第一个是 plan 模式,第三道闸,而且权限最高。`agent.py` 里它就一行判断:`self.plan_mode and tc.name not in Permission.READ_ONLY`,命中就拒绝,连 `--yes` 都压不过它。这个次序是刻意的:用户开 plan 模式就是想让 agent 先只读地读一遍代码再说,这时候任何「我已经授权过了」都不该生效。模型收到的拒绝消息也不是一句冷冰冰的报错,而是一段解释 plan 模式是什么的话,让它知道自己该先把计划讲清楚。只读先行、改动靠后,这个模式值得记住,因为它几乎是零成本的:一个布尔值,加在白名单判断前面。
|
|
159
|
+
|
|
160
|
+
第二个是 hooks。`~/.corecoder/hooks.json` 里用户可以挂 shell 命令,分 PreToolUse 和 PostToolUse 两种。Pre 钩子在授权检查之前触发(`agent.py` 里就是 `_pre_hooks(tc) or self._permit(tc)` 这一个 `or` 决定先后),返回码 2 表示否决,否决理由从 stderr 原样喂回模型,模型拿到的是一个能读懂、能修正的文本,而不是一个异常栈。Post 钩子只能旁观,永远不能拦。还有一个更关键的取舍:钩子挂了、超时了、返回码不对,都只是记一条 warning 然后跳过。钩子是建议者,不是法官,主循环的命不能交给用户随手写的一段 shell。
|
|
161
|
+
|
|
162
|
+
第三个是 MCP。`mcp.py` 里每个 server 是一个子进程,走 stdio 上一行一条的 JSON-RPC:启动握手、`tools/list` 拉清单,然后每个远程工具以 `mcp__server__tool` 的名字登记进工具表,从此授权、hooks、plan 模式对它们和内置工具一视同仁。值得学的是这个「一视同」是怎么来的:不是靠 MCP 代码里写多少特殊分支,而是因为工具边界本身(`Tool` 基类加上面三道闸)定义得够干净,外部工具只是又一个实现类。一个挂着死掉的 server 最坏也就是那一次调用返回个错误字符串,循环照转。
|
|
163
|
+
|
|
164
|
+
这三块都属于「进阶件」:它们不在 agent 的骨架上,骨架仍然是那个循环加七件工具。但它们回答的是同一个问题——一次工具调用前后,还有谁能说话。答案从「两道闸」变成了「两道闸加一圈可插拔的旁路」,而旁路的每一环都被设计成「坏了也不拖累主循环」。
|
|
165
|
+
|
|
154
166
|
## 动手:写你自己的第一个工具
|
|
155
167
|
|
|
156
168
|
讲了这么多,不如真加一个。假设我们想给 agent 一个查当前时间的能力(模型自己是不知道现在几点的)。新建 `corecoder/tools/now.py`:
|