corecoder 0.5.0__tar.gz → 0.7.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (70) hide show
  1. {corecoder-0.5.0 → corecoder-0.7.0}/PKG-INFO +97 -30
  2. {corecoder-0.5.0 → corecoder-0.7.0}/README.md +96 -29
  3. {corecoder-0.5.0 → corecoder-0.7.0}/README_CN.md +96 -30
  4. {corecoder-0.5.0 → corecoder-0.7.0}/article/00-index.md +2 -1
  5. {corecoder-0.5.0 → corecoder-0.7.0}/article/00-index_EN.md +2 -1
  6. {corecoder-0.5.0 → corecoder-0.7.0}/article/02-tools.md +12 -0
  7. {corecoder-0.5.0 → corecoder-0.7.0}/article/02-tools_EN.md +12 -0
  8. corecoder-0.7.0/article/08-extensibility.md +98 -0
  9. corecoder-0.7.0/article/08-extensibility_EN.md +98 -0
  10. corecoder-0.7.0/assets/demo-plan-hooks.gif +0 -0
  11. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/__init__.py +1 -1
  12. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/agent.py +68 -8
  13. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/checkpoints.py +6 -0
  14. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/cli.py +47 -6
  15. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/config.py +15 -17
  16. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/context.py +12 -2
  17. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/demo.py +28 -12
  18. corecoder-0.7.0/corecoder/hooks.py +87 -0
  19. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/llm.py +82 -111
  20. corecoder-0.7.0/corecoder/mcp.py +208 -0
  21. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/prompt.py +8 -0
  22. corecoder-0.7.0/corecoder/shell.py +61 -0
  23. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/__init__.py +0 -7
  24. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/agent.py +10 -1
  25. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/base.py +5 -0
  26. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/bash.py +86 -14
  27. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/edit.py +26 -23
  28. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/write.py +8 -5
  29. corecoder-0.7.0/examples/plan_hooks_demo.py +140 -0
  30. {corecoder-0.5.0 → corecoder-0.7.0}/pyproject.toml +1 -1
  31. corecoder-0.7.0/tests/conftest.py +11 -0
  32. {corecoder-0.5.0 → corecoder-0.7.0}/tests/test_checkpoints.py +17 -0
  33. corecoder-0.7.0/tests/test_core.py +734 -0
  34. {corecoder-0.5.0 → corecoder-0.7.0}/tests/test_demo.py +20 -3
  35. corecoder-0.7.0/tests/test_hooks.py +145 -0
  36. corecoder-0.7.0/tests/test_mcp.py +198 -0
  37. {corecoder-0.5.0 → corecoder-0.7.0}/tests/test_permissions.py +3 -2
  38. corecoder-0.7.0/tests/test_plan_mode.py +100 -0
  39. corecoder-0.7.0/tests/test_safety_matrix.py +208 -0
  40. corecoder-0.7.0/tests/test_shell.py +71 -0
  41. {corecoder-0.5.0 → corecoder-0.7.0}/tests/test_tools.py +78 -1
  42. corecoder-0.5.0/tests/test_core.py +0 -311
  43. {corecoder-0.5.0 → corecoder-0.7.0}/.github/workflows/ci.yml +0 -0
  44. {corecoder-0.5.0 → corecoder-0.7.0}/.github/workflows/publish.yml +0 -0
  45. {corecoder-0.5.0 → corecoder-0.7.0}/.gitignore +0 -0
  46. {corecoder-0.5.0 → corecoder-0.7.0}/LICENSE +0 -0
  47. {corecoder-0.5.0 → corecoder-0.7.0}/article/01-the-loop.md +0 -0
  48. {corecoder-0.5.0 → corecoder-0.7.0}/article/01-the-loop_EN.md +0 -0
  49. {corecoder-0.5.0 → corecoder-0.7.0}/article/03-llm-and-cost.md +0 -0
  50. {corecoder-0.5.0 → corecoder-0.7.0}/article/03-llm-and-cost_EN.md +0 -0
  51. {corecoder-0.5.0 → corecoder-0.7.0}/article/04-context.md +0 -0
  52. {corecoder-0.5.0 → corecoder-0.7.0}/article/04-context_EN.md +0 -0
  53. {corecoder-0.5.0 → corecoder-0.7.0}/article/05-parallel-and-subagents.md +0 -0
  54. {corecoder-0.5.0 → corecoder-0.7.0}/article/05-parallel-and-subagents_EN.md +0 -0
  55. {corecoder-0.5.0 → corecoder-0.7.0}/article/06-session-and-cli.md +0 -0
  56. {corecoder-0.5.0 → corecoder-0.7.0}/article/06-session-and-cli_EN.md +0 -0
  57. {corecoder-0.5.0 → corecoder-0.7.0}/article/07-build-your-own.md +0 -0
  58. {corecoder-0.5.0 → corecoder-0.7.0}/article/07-build-your-own_EN.md +0 -0
  59. {corecoder-0.5.0 → corecoder-0.7.0}/assets/demo.png +0 -0
  60. {corecoder-0.5.0 → corecoder-0.7.0}/assets/demo_en.png +0 -0
  61. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/__main__.py +0 -0
  62. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/permissions.py +0 -0
  63. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/session.py +0 -0
  64. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/glob_tool.py +0 -0
  65. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/grep.py +0 -0
  66. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/read.py +0 -0
  67. {corecoder-0.5.0 → corecoder-0.7.0}/corecoder/tools/todo.py +0 -0
  68. {corecoder-0.5.0 → corecoder-0.7.0}/tests/__init__.py +0 -0
  69. {corecoder-0.5.0 → corecoder-0.7.0}/tests/test_litellm.py +0 -0
  70. {corecoder-0.5.0 → corecoder-0.7.0}/tests/test_session.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: corecoder
3
- Version: 0.5.0
3
+ Version: 0.7.0
4
4
  Summary: Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. (formerly NanoCoder)
5
5
  Project-URL: Homepage, https://github.com/he-yufeng/CoreCoder
6
6
  Project-URL: Repository, https://github.com/he-yufeng/CoreCoder
@@ -37,7 +37,7 @@ Description-Content-Type: text/markdown
37
37
 
38
38
  # CoreCoder
39
39
 
40
- **The nanoGPT of coding agents. 1,217 lines of pure Python — understand how a coding agent actually works, then fork your own.**
40
+ **The nanoGPT of coding agents. A 1.3k-line engine inside 2,658 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
41
41
 
42
42
  *learn from it · fork it · ship something better*
43
43
 
@@ -47,7 +47,7 @@ Description-Content-Type: text/markdown
47
47
  [![Python](https://img.shields.io/badge/python-3.10+-blue)](https://python.org)
48
48
  [![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)
49
49
  [![Tests](https://github.com/he-yufeng/CoreCoder/actions/workflows/ci.yml/badge.svg)](https://github.com/he-yufeng/CoreCoder/actions)
50
- [![engine](https://img.shields.io/badge/engine-1217_LoC-blue)](article/00-index_EN.md)
50
+ [![engine](https://img.shields.io/badge/engine-1309_LoC-blue)](article/00-index_EN.md)
51
51
  [![essays](https://img.shields.io/badge/source--reading-8_bilingual-orange)](article/00-index_EN.md)
52
52
 
53
53
  </div>
@@ -60,7 +60,7 @@ Description-Content-Type: text/markdown
60
60
 
61
61
  | | CoreCoder | Claude Code | aider | nanoGPT |
62
62
  |---|---|---|---|---|
63
- | Lines of code | ~1,217 engine / 2,107 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
63
+ | Lines of code | ~1,309 engine / 2,658 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
64
64
  | Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
65
65
  | Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
66
66
  | What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
@@ -71,15 +71,15 @@ The nanoGPT column is there as a reference point: minimal, readable, but it teac
71
71
 
72
72
  I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
73
73
 
74
- The engine (loop, model interface, context, tools, sessions) is 1,217 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 22 files: 2,107 physical lines, 1,697 net, every one short enough to read in a single sitting.
74
+ The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,658 physical lines, 2,138 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
75
75
 
76
- And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 119 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
76
+ And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 171 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
77
77
 
78
78
  The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
79
79
 
80
80
  <p align="center">
81
- <img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo_en.png" width="760"
82
- alt="A real CoreCoder run: corecoder -p asks it to fix buggy.py; the agent reads the file, edits the code, runs it to confirm, and reports what it changed.">
81
+ <img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo-plan-hooks.gif" width="760"
82
+ alt="Plan mode in action: the agent reads fib.py, gets its edit refused while plan mode is on, presents a plan, and only after approval edits, tests, and reports — with Pre/PostToolUse hooks firing around every call.">
83
83
  </p>
84
84
 
85
85
  <p align="center"><sub><i>These thousand lines really do run a full loop end to end: ask it to fix buggy.py and it reads the file, edits the code, runs it once to confirm, then reports back on its own. Watch it, then come back and read the code.</i></sub></p>
@@ -107,7 +107,9 @@ Give it a model and a key and it goes. It speaks the OpenAI-compatible API by de
107
107
  | OmniRoute | `OPENAI_API_KEY=your-key OPENAI_BASE_URL=http://localhost:20128/v1 CORECODER_MODEL=auto` |
108
108
  | Local Ollama | `OPENAI_API_KEY=ollama OPENAI_BASE_URL=http://localhost:11434/v1 CORECODER_MODEL=qwen2.5-coder` |
109
109
 
110
- Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
110
+ Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. Thinking models are first-class too: deepseek-reasoner, kimi-k3 and friends stream their chain-of-thought, and CoreCoder shows it dimmed as it works, kept out of the conversation history so providers never see it come back. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
111
+
112
+ Smoke-tested end to end (read the file, edit it, run it, report back) against DeepSeek, Qwen3 and Kimi K2 via a single OpenRouter-compatible endpoint; each completed the full loop. One note for one-shot scripts: `-p` refuses mutating tools unless you pass `--yes`, by design.
111
113
 
112
114
  ```bash
113
115
  corecoder # interactive REPL
@@ -120,27 +122,34 @@ Laid out flat, the whole project is this big. Skim it before you clone and you'l
120
122
 
121
123
  ```
122
124
  corecoder/
123
- ├── agent.py agent loop + parallel tool exec 180 lines ← start here
124
- ├── llm.py streaming client + retry + cost 336 lines
125
- ├── context.py three-tier context compaction 210 lines
125
+ ├── agent.py agent loop + parallel tool exec 240 lines ← start here
126
+ ├── llm.py streaming client + retry + cost 332 lines
127
+ ├── context.py three-tier context compaction 220 lines
126
128
  ├── session.py save / resume + path-traversal guard 97 lines
127
129
  ├── permissions.py consent for mutating tools 48 lines
128
- ├── prompt.py system prompt 33 lines
129
- ├── cli.py REPL + slash commands + one-shot 317 lines
130
- ├── config.py env-var config 57 lines
130
+ ├── hooks.py Pre/PostToolUse shell hooks 87 lines
131
+ ├── shell.py POSIX shell routing (Git Bash on Windows) 61 lines
132
+ ├── mcp.py MCP stdio client for external tools 208 lines
133
+ ├── prompt.py system prompt 41 lines
134
+ ├── cli.py REPL + slash commands + one-shot 358 lines
135
+ ├── config.py env-var config 55 lines
136
+ ├── checkpoints.py /undo snapshot and restore 44 lines
137
+ ├── demo.py offline end-to-end demo 100 lines
131
138
  └── tools/
132
- ├── bash.py shell + dangerous-command gate + cd 127 lines
133
- ├── edit.py unique-match search/replace + diff 92 lines
134
- ├── grep.py content search 79 lines
135
- ├── glob_tool.py filename matching 47 lines
136
- ├── read.py file read 53 lines
137
- ├── write.py file write 38 lines
139
+ ├── bash.py shell + dangerous-command gate + cd 203 lines
140
+ ├── edit.py unique-match search/replace + diff 99 lines
141
+ ├── grep.py content search 93 lines
142
+ ├── glob_tool.py filename matching 52 lines
143
+ ├── read.py file read 56 lines
144
+ ├── write.py file write 46 lines
138
145
  ├── todo.py agent-maintained task checklist 79 lines
139
- ├── agent.py sub-agent spawning 63 lines
140
- └── base.py tool base class 27 lines
146
+ ├── agent.py sub-agent spawning 72 lines
147
+ └── base.py tool base class 32 lines
148
+ examples/
149
+ └── plan_hooks_demo.py offline plan mode + hooks demo (no API key)
141
150
  ```
142
151
 
143
- Eight tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `todo_write` (a task checklist the agent maintains for itself), and `agent` (which spawns a sub-agent). Everything else is the CLI shell, config, and packaging wrapped around that engine core.
152
+ Eight tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `todo_write` (a task checklist the agent maintains for itself), and `agent` (which spawns a sub-agent). Everything else is the CLI shell, config, and packaging wrapped around that engine core. If `~/.corecoder/mcp.json` exists, its MCP servers join the eight as extra `mcp__*` tools; the MCP section below covers it.
144
153
 
145
154
  ## A `while` loop is the whole agent
146
155
 
@@ -161,7 +170,7 @@ def chat(self, user_input):
161
170
  return "(hit the round limit)"
162
171
  ```
163
172
 
164
- That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
173
+ That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
165
174
 
166
175
  Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
167
176
 
@@ -175,7 +184,7 @@ Every one of these *whys* is traced down to the actual lines of code in the seri
175
184
 
176
185
  ## The source-reading series · 8 bilingual essays
177
186
 
178
- I also wrote a bilingual source-reading series, one intro plus seven parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking; read them in any order.
187
+ I also wrote a bilingual source-reading series, one intro plus eight parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking, and the eighth is about extending it without touching the loop; read them in any order.
179
188
 
180
189
  - **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
181
190
  - **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
@@ -185,14 +194,15 @@ I also wrote a bilingual source-reading series, one intro plus seven parts, each
185
194
  - **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
186
195
  - **[06 · Turning it into a real command-line tool](article/06-session-and-cli_EN.md)** — `session.py` and path-traversal defense
187
196
  - **[07 · Fork CoreCoder into your own coding agent](article/07-build-your-own_EN.md)** — from fork to custom tools to swapping models
197
+ - **[08 · Three ways to extend without touching the loop: MCP, hooks, and plan mode](article/08-extensibility_EN.md)** — the v0.6.0 extensibility trio and the contract that makes them safe
188
198
 
189
199
  ## Fork it, build something better
190
200
 
191
201
  Once you understand it, the natural next step is to fork. Getting started doesn't take much:
192
202
 
193
- - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (336 lines) is the entry point for all provider adaptation.
203
+ - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (267 lines) is the entry point for all provider adaptation.
194
204
  - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (27 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
195
- - **Rewrite the system prompt.** `prompt.py` is all of 33 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
205
+ - **Rewrite the system prompt.** `prompt.py` is all of 41 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
196
206
  - **Import it as a library.** The top level exports `Agent`, `LLM`, and `Config`, ready to embed in your own program:
197
207
 
198
208
  ```python
@@ -207,7 +217,7 @@ Going deeper, the directions are out in the open too. None of the following is i
207
217
  - **The dangerous-command blocking in bash is just a regex blacklist.** It guards against slips, not a security sandbox. Facing untrusted input means reaching for seccomp or container isolation. This is the hardest of the four; it goes all the way down to the syscall and isolation layer.
208
218
  - **Retry is only exponential backoff.** No fallback model, no hard dollar budget. Follow `llm.py` down and add a fallback model chain plus a stop-on-over-budget gate; the change stays mostly inside that one file.
209
219
  - **Sub-agents only run the plainest synchronous execution.** Make it async or a streaming executor and you close the exact gap the fifth essay identifies between this and how production agents stream execution.
210
- - **No MCP, no RAG.** Wire up MCP to give it the external tool ecosystem, or add retrieval-based code location for big repos. Both are real ways to grow from a minimal core into your own stronger agent.
220
+ - **No RAG, and the MCP client speaks tools only.** Retrieval-based code location for big repos is still open, and `mcp.py` leaves resources and prompts unimplemented on purpose. Either one is a real way to grow from a minimal core into your own stronger agent.
211
221
 
212
222
  The README only points; the seventh essay picks up the code details for each. Pick one and start; that's the whole reason the core is kept this small.
213
223
 
@@ -221,6 +231,7 @@ Inside the REPL, `/help` lists everything; these are the ones you'll reach for:
221
231
  /tokens token usage and cost estimate
222
232
  /diff files changed this session
223
233
  /undo revert the most recent file change
234
+ /plan toggle plan mode (read-only, then a plan to approve)
224
235
  /save /sessions save / list sessions
225
236
  quit / exit exit (Ctrl+C cancels the current round)
226
237
  ```
@@ -235,6 +246,62 @@ Read-only tools (`read_file`, `glob`, `grep`, `todo_write`) run the moment the m
235
246
  - In one-shot mode (`-p`) there is nobody to ask, so a mutating call is refused on the spot and the refusal goes back to the model as an ordinary tool result: the loop never hangs on input that can't arrive. Pass `--yes` to approve everything up front (scripts, CI).
236
247
  - The decision itself is pure logic in `permissions.py`, with the terminal only supplying the prompt callback. You can unit-test consent without a TTY, or reuse the layer in your own embedding.
237
248
 
249
+ ## Plan mode
250
+
251
+ `/plan` toggles plan mode in the REPL. While it's on, the prompt shows `(plan)` and every mutating call (writes, edits, bash, MCP tools, sub-agents) is refused on the spot: the refusal goes back to the model as an ordinary tool result, telling it to keep investigating read-only and present a numbered plan instead. When the plan looks right, `approve` (or `/plan` again) hands control back and the agent executes. Mechanically it is one flag on the `Agent` plus one refusal branch ahead of the consent gate, which itself stays untouched; there is no plan file and nothing is remembered between sessions.
252
+
253
+ ## Hooks
254
+
255
+ Drop a `hooks.json` under `~/.corecoder` and your own shell commands run around every tool call, the same idea as Claude Code's hooks:
256
+
257
+ ```json
258
+ {
259
+ "PreToolUse": [{"matcher": "bash", "command": "cat >> ~/.corecoder/audit.jsonl"}],
260
+ "PostToolUse": [{"matcher": "*", "command": "cat >> ~/.corecoder/trace.jsonl"}]
261
+ }
262
+ ```
263
+
264
+ Each hook gets the call as JSON on stdin (`tool_name`, `tool_input`; post hooks also get `tool_response`). The matcher is an exact tool name; empty or `*` fires on every tool. A pre hook can veto the call with exit code 2, and its stderr travels back to the model as the reason so it can route around the block. Post hooks only observe and can never block. A hook that errors or runs past ten seconds is skipped with a warning: hooks assist the loop, they never get to kill it. The whole mechanism is `hooks.py`, and the REPL banner shows how many hooks loaded.
265
+
266
+ Two worth stealing (the commands lean on `jq`, the usual suspect):
267
+
268
+ ```bash
269
+ # 1. lint gate: after every edit/write, run the project's fast linter on the
270
+ # touched file. The model sees the output and fixes its own mistakes in
271
+ # the same turn instead of waiting for CI.
272
+ {
273
+ "PostToolUse": [{
274
+ "matcher": "edit",
275
+ "command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
276
+ }]
277
+ }
278
+
279
+ # 2. write protect: refuse edits under paths you never want an agent to
280
+ # touch. Exit code 2 vetoes the call and the message reaches the model.
281
+ {
282
+ "PreToolUse": [{
283
+ "matcher": "edit",
284
+ "command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
285
+ }]
286
+ }
287
+ ```
288
+
289
+ Both are plain shell; nothing here is CoreCoder-specific syntax beyond the JSON shape and the exit-code-2 veto.
290
+
291
+ ## MCP servers
292
+
293
+ Drop a `mcp.json` under `~/.corecoder` and tools from any MCP server join the agent over stdio, the same config shape as Claude Code's:
294
+
295
+ ```json
296
+ {
297
+ "mcpServers": {
298
+ "filesystem": {"command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "/tmp"]}
299
+ }
300
+ }
301
+ ```
302
+
303
+ Each configured server starts as a subprocess at launch, handshakes, and lists its tools; every one is registered as `mcp__<server>__<tool>`, so hook matchers and the consent gate treat it exactly like a built-in. MCP tools stay out of the read-only set, meaning the agent asks before running one. The handshake gets fifteen seconds, a call gets sixty, and a server that dies or never answers fails that one call as an ordinary tool result instead of killing the loop. The client speaks the tools slice of the protocol (initialize, tools/list, tools/call) and nothing else, which keeps the whole thing inside `mcp.py` at about 200 lines. With no `mcp.json` there is no MCP and nothing changes.
304
+
238
305
  ## Related Projects
239
306
 
240
307
  If working through CoreCoder was useful, here are a few other tools I've built around agents and LLM systems:
@@ -247,7 +314,7 @@ If working through CoreCoder was useful, here are a few other tools I've built a
247
314
 
248
315
  ## Contributing / License
249
316
 
250
- Before you send anything, run `pytest tests/ -q` (119 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
317
+ Before you send anything, run `pytest tests/ -q` (171 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
251
318
 
252
319
  ---
253
320
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  # CoreCoder
4
4
 
5
- **The nanoGPT of coding agents. 1,217 lines of pure Python — understand how a coding agent actually works, then fork your own.**
5
+ **The nanoGPT of coding agents. A 1.3k-line engine inside 2,658 readable lines of pure Python: understand how a coding agent actually works, then fork your own.**
6
6
 
7
7
  *learn from it · fork it · ship something better*
8
8
 
@@ -12,7 +12,7 @@
12
12
  [![Python](https://img.shields.io/badge/python-3.10+-blue)](https://python.org)
13
13
  [![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)
14
14
  [![Tests](https://github.com/he-yufeng/CoreCoder/actions/workflows/ci.yml/badge.svg)](https://github.com/he-yufeng/CoreCoder/actions)
15
- [![engine](https://img.shields.io/badge/engine-1217_LoC-blue)](article/00-index_EN.md)
15
+ [![engine](https://img.shields.io/badge/engine-1309_LoC-blue)](article/00-index_EN.md)
16
16
  [![essays](https://img.shields.io/badge/source--reading-8_bilingual-orange)](article/00-index_EN.md)
17
17
 
18
18
  </div>
@@ -25,7 +25,7 @@
25
25
 
26
26
  | | CoreCoder | Claude Code | aider | nanoGPT |
27
27
  |---|---|---|---|---|
28
- | Lines of code | ~1,217 engine / 2,107 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
28
+ | Lines of code | ~1,309 engine / 2,658 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
29
29
  | Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
30
30
  | Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
31
31
  | What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
@@ -36,15 +36,15 @@ The nanoGPT column is there as a reference point: minimal, readable, but it teac
36
36
 
37
37
  I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
38
38
 
39
- The engine (loop, model interface, context, tools, sessions) is 1,217 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 22 files: 2,107 physical lines, 1,697 net, every one short enough to read in a single sitting.
39
+ The engine (loop, model interface, context, tools, sessions) is 1,309 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 25 files: 2,658 physical lines, 2,138 net, every one short enough to read in a single sitting. The growth since the original 1,161-line snapshot went into visible features: plan mode, hooks and checkpoints, each documented below.
40
40
 
41
- And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 119 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
41
+ And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. Anything that would mutate your disk or run a command stops for your consent first. 171 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
42
42
 
43
43
  The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
44
44
 
45
45
  <p align="center">
46
- <img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo_en.png" width="760"
47
- alt="A real CoreCoder run: corecoder -p asks it to fix buggy.py; the agent reads the file, edits the code, runs it to confirm, and reports what it changed.">
46
+ <img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo-plan-hooks.gif" width="760"
47
+ alt="Plan mode in action: the agent reads fib.py, gets its edit refused while plan mode is on, presents a plan, and only after approval edits, tests, and reports — with Pre/PostToolUse hooks firing around every call.">
48
48
  </p>
49
49
 
50
50
  <p align="center"><sub><i>These thousand lines really do run a full loop end to end: ask it to fix buggy.py and it reads the file, edits the code, runs it once to confirm, then reports back on its own. Watch it, then come back and read the code.</i></sub></p>
@@ -72,7 +72,9 @@ Give it a model and a key and it goes. It speaks the OpenAI-compatible API by de
72
72
  | OmniRoute | `OPENAI_API_KEY=your-key OPENAI_BASE_URL=http://localhost:20128/v1 CORECODER_MODEL=auto` |
73
73
  | Local Ollama | `OPENAI_API_KEY=ollama OPENAI_BASE_URL=http://localhost:11434/v1 CORECODER_MODEL=qwen2.5-coder` |
74
74
 
75
- Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
75
+ Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. Thinking models are first-class too: deepseek-reasoner, kimi-k3 and friends stream their chain-of-thought, and CoreCoder shows it dimmed as it works, kept out of the conversation history so providers never see it come back. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
76
+
77
+ Smoke-tested end to end (read the file, edit it, run it, report back) against DeepSeek, Qwen3 and Kimi K2 via a single OpenRouter-compatible endpoint; each completed the full loop. One note for one-shot scripts: `-p` refuses mutating tools unless you pass `--yes`, by design.
76
78
 
77
79
  ```bash
78
80
  corecoder # interactive REPL
@@ -85,27 +87,34 @@ Laid out flat, the whole project is this big. Skim it before you clone and you'l
85
87
 
86
88
  ```
87
89
  corecoder/
88
- ├── agent.py agent loop + parallel tool exec 180 lines ← start here
89
- ├── llm.py streaming client + retry + cost 336 lines
90
- ├── context.py three-tier context compaction 210 lines
90
+ ├── agent.py agent loop + parallel tool exec 240 lines ← start here
91
+ ├── llm.py streaming client + retry + cost 332 lines
92
+ ├── context.py three-tier context compaction 220 lines
91
93
  ├── session.py save / resume + path-traversal guard 97 lines
92
94
  ├── permissions.py consent for mutating tools 48 lines
93
- ├── prompt.py system prompt 33 lines
94
- ├── cli.py REPL + slash commands + one-shot 317 lines
95
- ├── config.py env-var config 57 lines
95
+ ├── hooks.py Pre/PostToolUse shell hooks 87 lines
96
+ ├── shell.py POSIX shell routing (Git Bash on Windows) 61 lines
97
+ ├── mcp.py MCP stdio client for external tools 208 lines
98
+ ├── prompt.py system prompt 41 lines
99
+ ├── cli.py REPL + slash commands + one-shot 358 lines
100
+ ├── config.py env-var config 55 lines
101
+ ├── checkpoints.py /undo snapshot and restore 44 lines
102
+ ├── demo.py offline end-to-end demo 100 lines
96
103
  └── tools/
97
- ├── bash.py shell + dangerous-command gate + cd 127 lines
98
- ├── edit.py unique-match search/replace + diff 92 lines
99
- ├── grep.py content search 79 lines
100
- ├── glob_tool.py filename matching 47 lines
101
- ├── read.py file read 53 lines
102
- ├── write.py file write 38 lines
104
+ ├── bash.py shell + dangerous-command gate + cd 203 lines
105
+ ├── edit.py unique-match search/replace + diff 99 lines
106
+ ├── grep.py content search 93 lines
107
+ ├── glob_tool.py filename matching 52 lines
108
+ ├── read.py file read 56 lines
109
+ ├── write.py file write 46 lines
103
110
  ├── todo.py agent-maintained task checklist 79 lines
104
- ├── agent.py sub-agent spawning 63 lines
105
- └── base.py tool base class 27 lines
111
+ ├── agent.py sub-agent spawning 72 lines
112
+ └── base.py tool base class 32 lines
113
+ examples/
114
+ └── plan_hooks_demo.py offline plan mode + hooks demo (no API key)
106
115
  ```
107
116
 
108
- Eight tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `todo_write` (a task checklist the agent maintains for itself), and `agent` (which spawns a sub-agent). Everything else is the CLI shell, config, and packaging wrapped around that engine core.
117
+ Eight tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, `todo_write` (a task checklist the agent maintains for itself), and `agent` (which spawns a sub-agent). Everything else is the CLI shell, config, and packaging wrapped around that engine core. If `~/.corecoder/mcp.json` exists, its MCP servers join the eight as extra `mcp__*` tools; the MCP section below covers it.
109
118
 
110
119
  ## A `while` loop is the whole agent
111
120
 
@@ -126,7 +135,7 @@ def chat(self, user_input):
126
135
  return "(hit the round limit)"
127
136
  ```
128
137
 
129
- That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
138
+ That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess the loop runs into once it meets the real world. `llm.py` ends up the biggest file in the project, not because calling a model is hard, but because a streamed response splinters each tool call's arguments into fragments you have to restitch in order, a provider will hand you half a JSON object or a null `usage` field, and 429s, timeouts, dropped connections and 5xx all need backoff-and-retry while the other 4xx should just raise. A stream can also die after it has already started, so the retry wraps the whole request, not just the connect. That unglamorous grunt work, not the loop, is where the real engineering of taking an agent from demo to delivery actually lives; the third essay follows it down to the line.
130
139
 
131
140
  Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
132
141
 
@@ -140,7 +149,7 @@ Every one of these *whys* is traced down to the actual lines of code in the seri
140
149
 
141
150
  ## The source-reading series · 8 bilingual essays
142
151
 
143
- I also wrote a bilingual source-reading series, one intro plus seven parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking; read them in any order.
152
+ I also wrote a bilingual source-reading series, one intro plus eight parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking, and the eighth is about extending it without touching the loop; read them in any order.
144
153
 
145
154
  - **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
146
155
  - **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
@@ -150,14 +159,15 @@ I also wrote a bilingual source-reading series, one intro plus seven parts, each
150
159
  - **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
151
160
  - **[06 · Turning it into a real command-line tool](article/06-session-and-cli_EN.md)** — `session.py` and path-traversal defense
152
161
  - **[07 · Fork CoreCoder into your own coding agent](article/07-build-your-own_EN.md)** — from fork to custom tools to swapping models
162
+ - **[08 · Three ways to extend without touching the loop: MCP, hooks, and plan mode](article/08-extensibility_EN.md)** — the v0.6.0 extensibility trio and the contract that makes them safe
153
163
 
154
164
  ## Fork it, build something better
155
165
 
156
166
  Once you understand it, the natural next step is to fork. Getting started doesn't take much:
157
167
 
158
- - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (336 lines) is the entry point for all provider adaptation.
168
+ - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (267 lines) is the entry point for all provider adaptation.
159
169
  - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (27 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
160
- - **Rewrite the system prompt.** `prompt.py` is all of 33 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
170
+ - **Rewrite the system prompt.** `prompt.py` is all of 41 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
161
171
  - **Import it as a library.** The top level exports `Agent`, `LLM`, and `Config`, ready to embed in your own program:
162
172
 
163
173
  ```python
@@ -172,7 +182,7 @@ Going deeper, the directions are out in the open too. None of the following is i
172
182
  - **The dangerous-command blocking in bash is just a regex blacklist.** It guards against slips, not a security sandbox. Facing untrusted input means reaching for seccomp or container isolation. This is the hardest of the four; it goes all the way down to the syscall and isolation layer.
173
183
  - **Retry is only exponential backoff.** No fallback model, no hard dollar budget. Follow `llm.py` down and add a fallback model chain plus a stop-on-over-budget gate; the change stays mostly inside that one file.
174
184
  - **Sub-agents only run the plainest synchronous execution.** Make it async or a streaming executor and you close the exact gap the fifth essay identifies between this and how production agents stream execution.
175
- - **No MCP, no RAG.** Wire up MCP to give it the external tool ecosystem, or add retrieval-based code location for big repos. Both are real ways to grow from a minimal core into your own stronger agent.
185
+ - **No RAG, and the MCP client speaks tools only.** Retrieval-based code location for big repos is still open, and `mcp.py` leaves resources and prompts unimplemented on purpose. Either one is a real way to grow from a minimal core into your own stronger agent.
176
186
 
177
187
  The README only points; the seventh essay picks up the code details for each. Pick one and start; that's the whole reason the core is kept this small.
178
188
 
@@ -186,6 +196,7 @@ Inside the REPL, `/help` lists everything; these are the ones you'll reach for:
186
196
  /tokens token usage and cost estimate
187
197
  /diff files changed this session
188
198
  /undo revert the most recent file change
199
+ /plan toggle plan mode (read-only, then a plan to approve)
189
200
  /save /sessions save / list sessions
190
201
  quit / exit exit (Ctrl+C cancels the current round)
191
202
  ```
@@ -200,6 +211,62 @@ Read-only tools (`read_file`, `glob`, `grep`, `todo_write`) run the moment the m
200
211
  - In one-shot mode (`-p`) there is nobody to ask, so a mutating call is refused on the spot and the refusal goes back to the model as an ordinary tool result: the loop never hangs on input that can't arrive. Pass `--yes` to approve everything up front (scripts, CI).
201
212
  - The decision itself is pure logic in `permissions.py`, with the terminal only supplying the prompt callback. You can unit-test consent without a TTY, or reuse the layer in your own embedding.
202
213
 
214
+ ## Plan mode
215
+
216
+ `/plan` toggles plan mode in the REPL. While it's on, the prompt shows `(plan)` and every mutating call (writes, edits, bash, MCP tools, sub-agents) is refused on the spot: the refusal goes back to the model as an ordinary tool result, telling it to keep investigating read-only and present a numbered plan instead. When the plan looks right, `approve` (or `/plan` again) hands control back and the agent executes. Mechanically it is one flag on the `Agent` plus one refusal branch ahead of the consent gate, which itself stays untouched; there is no plan file and nothing is remembered between sessions.
217
+
218
+ ## Hooks
219
+
220
+ Drop a `hooks.json` under `~/.corecoder` and your own shell commands run around every tool call, the same idea as Claude Code's hooks:
221
+
222
+ ```json
223
+ {
224
+ "PreToolUse": [{"matcher": "bash", "command": "cat >> ~/.corecoder/audit.jsonl"}],
225
+ "PostToolUse": [{"matcher": "*", "command": "cat >> ~/.corecoder/trace.jsonl"}]
226
+ }
227
+ ```
228
+
229
+ Each hook gets the call as JSON on stdin (`tool_name`, `tool_input`; post hooks also get `tool_response`). The matcher is an exact tool name; empty or `*` fires on every tool. A pre hook can veto the call with exit code 2, and its stderr travels back to the model as the reason so it can route around the block. Post hooks only observe and can never block. A hook that errors or runs past ten seconds is skipped with a warning: hooks assist the loop, they never get to kill it. The whole mechanism is `hooks.py`, and the REPL banner shows how many hooks loaded.
230
+
231
+ Two worth stealing (the commands lean on `jq`, the usual suspect):
232
+
233
+ ```bash
234
+ # 1. lint gate: after every edit/write, run the project's fast linter on the
235
+ # touched file. The model sees the output and fixes its own mistakes in
236
+ # the same turn instead of waiting for CI.
237
+ {
238
+ "PostToolUse": [{
239
+ "matcher": "edit",
240
+ "command": "f=$(jq -r .tool_input.path); ruff check \"$f\" 2>&1 | head -20"
241
+ }]
242
+ }
243
+
244
+ # 2. write protect: refuse edits under paths you never want an agent to
245
+ # touch. Exit code 2 vetoes the call and the message reaches the model.
246
+ {
247
+ "PreToolUse": [{
248
+ "matcher": "edit",
249
+ "command": "case \"$(jq -r .tool_input.path)\" in .env*|*/secrets/*|*.pem) echo 'that path is off-limits' >&2; exit 2;; esac"
250
+ }]
251
+ }
252
+ ```
253
+
254
+ Both are plain shell; nothing here is CoreCoder-specific syntax beyond the JSON shape and the exit-code-2 veto.
255
+
256
+ ## MCP servers
257
+
258
+ Drop a `mcp.json` under `~/.corecoder` and tools from any MCP server join the agent over stdio, the same config shape as Claude Code's:
259
+
260
+ ```json
261
+ {
262
+ "mcpServers": {
263
+ "filesystem": {"command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "/tmp"]}
264
+ }
265
+ }
266
+ ```
267
+
268
+ Each configured server starts as a subprocess at launch, handshakes, and lists its tools; every one is registered as `mcp__<server>__<tool>`, so hook matchers and the consent gate treat it exactly like a built-in. MCP tools stay out of the read-only set, meaning the agent asks before running one. The handshake gets fifteen seconds, a call gets sixty, and a server that dies or never answers fails that one call as an ordinary tool result instead of killing the loop. The client speaks the tools slice of the protocol (initialize, tools/list, tools/call) and nothing else, which keeps the whole thing inside `mcp.py` at about 200 lines. With no `mcp.json` there is no MCP and nothing changes.
269
+
203
270
  ## Related Projects
204
271
 
205
272
  If working through CoreCoder was useful, here are a few other tools I've built around agents and LLM systems:
@@ -212,7 +279,7 @@ If working through CoreCoder was useful, here are a few other tools I've built a
212
279
 
213
280
  ## Contributing / License
214
281
 
215
- Before you send anything, run `pytest tests/ -q` (119 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
282
+ Before you send anything, run `pytest tests/ -q` (171 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
216
283
 
217
284
  ---
218
285