corecoder 0.3.0__tar.gz → 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (71) hide show
  1. corecoder-0.4.0/.github/workflows/ci.yml +64 -0
  2. {corecoder-0.3.0 → corecoder-0.4.0}/.github/workflows/publish.yml +7 -5
  3. {corecoder-0.3.0 → corecoder-0.4.0}/.gitignore +1 -1
  4. corecoder-0.4.0/PKG-INFO +234 -0
  5. corecoder-0.4.0/README.md +199 -0
  6. corecoder-0.4.0/README_CN.md +199 -0
  7. corecoder-0.4.0/article/00-index.md +70 -0
  8. corecoder-0.4.0/article/00-index_EN.md +70 -0
  9. corecoder-0.4.0/article/01-the-loop.md +149 -0
  10. corecoder-0.4.0/article/01-the-loop_EN.md +149 -0
  11. corecoder-0.4.0/article/02-tools.md +204 -0
  12. corecoder-0.4.0/article/02-tools_EN.md +204 -0
  13. corecoder-0.4.0/article/03-llm-and-cost.md +207 -0
  14. corecoder-0.4.0/article/03-llm-and-cost_EN.md +207 -0
  15. corecoder-0.4.0/article/04-context.md +154 -0
  16. corecoder-0.4.0/article/04-context_EN.md +154 -0
  17. corecoder-0.4.0/article/05-parallel-and-subagents.md +142 -0
  18. corecoder-0.4.0/article/05-parallel-and-subagents_EN.md +142 -0
  19. corecoder-0.4.0/article/06-session-and-cli.md +151 -0
  20. corecoder-0.4.0/article/06-session-and-cli_EN.md +151 -0
  21. corecoder-0.4.0/article/07-build-your-own.md +186 -0
  22. corecoder-0.4.0/article/07-build-your-own_EN.md +186 -0
  23. corecoder-0.4.0/assets/demo.png +0 -0
  24. corecoder-0.4.0/assets/demo_en.png +0 -0
  25. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/__init__.py +1 -1
  26. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/agent.py +45 -17
  27. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/cli.py +14 -2
  28. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/config.py +2 -2
  29. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/context.py +21 -7
  30. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/llm.py +18 -9
  31. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/session.py +38 -9
  32. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/bash.py +27 -15
  33. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/edit.py +5 -2
  34. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/grep.py +4 -3
  35. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/read.py +1 -1
  36. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/write.py +1 -1
  37. {corecoder-0.3.0 → corecoder-0.4.0}/pyproject.toml +7 -4
  38. corecoder-0.4.0/tests/test_core.py +233 -0
  39. {corecoder-0.3.0 → corecoder-0.4.0}/tests/test_litellm.py +8 -2
  40. corecoder-0.4.0/tests/test_session.py +75 -0
  41. corecoder-0.4.0/tests/test_tools.py +304 -0
  42. corecoder-0.3.0/.github/workflows/ci.yml +0 -27
  43. corecoder-0.3.0/PKG-INFO +0 -239
  44. corecoder-0.3.0/README.md +0 -204
  45. corecoder-0.3.0/README_CN.md +0 -182
  46. corecoder-0.3.0/article/00-index.md +0 -29
  47. corecoder-0.3.0/article/00-index_EN.md +0 -35
  48. corecoder-0.3.0/article/01-architecture-overview.md +0 -143
  49. corecoder-0.3.0/article/01-architecture-overview_EN.md +0 -144
  50. corecoder-0.3.0/article/02-agent-loop.md +0 -237
  51. corecoder-0.3.0/article/02-agent-loop_EN.md +0 -242
  52. corecoder-0.3.0/article/03-tool-system.md +0 -169
  53. corecoder-0.3.0/article/03-tool-system_EN.md +0 -163
  54. corecoder-0.3.0/article/04-context-compression.md +0 -146
  55. corecoder-0.3.0/article/04-context-compression_EN.md +0 -145
  56. corecoder-0.3.0/article/05-streaming-executor.md +0 -213
  57. corecoder-0.3.0/article/05-streaming-executor_EN.md +0 -213
  58. corecoder-0.3.0/article/06-multi-agent.md +0 -193
  59. corecoder-0.3.0/article/06-multi-agent_EN.md +0 -192
  60. corecoder-0.3.0/article/07-hidden-features.md +0 -98
  61. corecoder-0.3.0/article/07-hidden-features_EN.md +0 -98
  62. corecoder-0.3.0/tests/test_core.py +0 -146
  63. corecoder-0.3.0/tests/test_tools.py +0 -196
  64. {corecoder-0.3.0 → corecoder-0.4.0}/LICENSE +0 -0
  65. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/__main__.py +0 -0
  66. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/prompt.py +0 -0
  67. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/__init__.py +0 -0
  68. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/agent.py +0 -0
  69. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/base.py +0 -0
  70. {corecoder-0.3.0 → corecoder-0.4.0}/corecoder/tools/glob_tool.py +0 -0
  71. {corecoder-0.3.0 → corecoder-0.4.0}/tests/__init__.py +0 -0
@@ -0,0 +1,64 @@
1
+ name: CI
2
+
3
+ on:
4
+ push:
5
+ branches: [main]
6
+ pull_request:
7
+ branches: [main]
8
+
9
+ jobs:
10
+ test:
11
+ runs-on: ${{ matrix.os }}
12
+ strategy:
13
+ fail-fast: false
14
+ matrix:
15
+ os: [ubuntu-latest, macos-latest, windows-latest]
16
+ python-version: ["3.10", "3.11", "3.12", "3.13"]
17
+
18
+ steps:
19
+ - uses: actions/checkout@v6
20
+ - uses: actions/setup-python@v6
21
+ with:
22
+ python-version: ${{ matrix.python-version }}
23
+ cache: pip
24
+
25
+ - name: Install dependencies
26
+ run: |
27
+ python -m pip install -U pip
28
+ python -m pip install -e ".[dev]"
29
+
30
+ - name: Run tests
31
+ run: python -m pytest tests/ -v
32
+
33
+ - name: Compile
34
+ run: python -m compileall -q corecoder tests
35
+
36
+ package:
37
+ runs-on: ubuntu-latest
38
+ steps:
39
+ - uses: actions/checkout@v6
40
+ - uses: actions/setup-python@v6
41
+ with:
42
+ python-version: "3.13"
43
+ cache: pip
44
+
45
+ - name: Build
46
+ run: |
47
+ python -m pip install -U pip build twine
48
+ python -m build
49
+ python -m twine check dist/*
50
+
51
+ lint:
52
+ runs-on: ubuntu-latest
53
+ steps:
54
+ - uses: actions/checkout@v6
55
+ - uses: actions/setup-python@v6
56
+ with:
57
+ python-version: "3.13"
58
+ cache: pip
59
+
60
+ - name: Install ruff
61
+ run: python -m pip install -U pip ruff
62
+
63
+ - name: Ruff check
64
+ run: ruff check corecoder tests
@@ -12,16 +12,18 @@ jobs:
12
12
  runs-on: ubuntu-latest
13
13
  environment: pypi
14
14
  steps:
15
- - uses: actions/checkout@v4
16
- - uses: actions/setup-python@v5
15
+ - uses: actions/checkout@v6
16
+ - uses: actions/setup-python@v6
17
17
  with:
18
- python-version: "3.12"
18
+ python-version: "3.13"
19
19
 
20
20
  - name: Install build tools
21
- run: pip install build
21
+ run: python -m pip install -U pip build twine
22
22
 
23
23
  - name: Build package
24
- run: python -m build
24
+ run: |
25
+ python -m build
26
+ python -m twine check dist/*
25
27
 
26
28
  - name: Publish to PyPI
27
29
  uses: pypa/gh-action-pypi-publish@release/v1
@@ -13,5 +13,5 @@ venv/
13
13
  .ruff_cache/
14
14
  .corecoder_history
15
15
  .DS_Store
16
- dist/
17
16
  .pytest_cache/
17
+ .tmp/
@@ -0,0 +1,234 @@
1
+ Metadata-Version: 2.4
2
+ Name: corecoder
3
+ Version: 0.4.0
4
+ Summary: Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. (formerly NanoCoder)
5
+ Project-URL: Homepage, https://github.com/he-yufeng/CoreCoder
6
+ Project-URL: Repository, https://github.com/he-yufeng/CoreCoder
7
+ Project-URL: Issues, https://github.com/he-yufeng/CoreCoder/issues
8
+ Author-email: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
9
+ License-Expression: MIT
10
+ License-File: LICENSE
11
+ Keywords: agent,ai,claude-code,cli,coding,llm
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Environment :: Console
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Operating System :: OS Independent
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.10
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Topic :: Software Development
22
+ Classifier: Topic :: Software Development :: Code Generators
23
+ Classifier: Topic :: Terminals
24
+ Requires-Python: >=3.10
25
+ Requires-Dist: openai>=1.0
26
+ Requires-Dist: prompt-toolkit>=3.0
27
+ Requires-Dist: python-dotenv>=1.0
28
+ Requires-Dist: rich>=13.0
29
+ Provides-Extra: dev
30
+ Requires-Dist: pytest>=7.0; extra == 'dev'
31
+ Requires-Dist: ruff>=0.9.0; extra == 'dev'
32
+ Provides-Extra: litellm
33
+ Requires-Dist: litellm<2.0.0,>=1.60.0; extra == 'litellm'
34
+ Description-Content-Type: text/markdown
35
+
36
+ <div align="center">
37
+
38
+ # CoreCoder
39
+
40
+ **The nanoGPT of coding agents. 1,081 lines of pure Python — understand how a coding agent actually works, then fork your own.**
41
+
42
+ *learn from it · fork it · ship something better*
43
+
44
+ [中文](README_CN.md) | English | [Source-reading series · 8 bilingual essays](article/00-index_EN.md)
45
+
46
+ [![PyPI](https://img.shields.io/pypi/v/corecoder)](https://pypi.org/project/corecoder/)
47
+ [![Python](https://img.shields.io/badge/python-3.10+-blue)](https://python.org)
48
+ [![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)
49
+ [![Tests](https://github.com/he-yufeng/CoreCoder/actions/workflows/ci.yml/badge.svg)](https://github.com/he-yufeng/CoreCoder/actions)
50
+ [![engine](https://img.shields.io/badge/engine-1081_LoC-blue)](article/00-index_EN.md)
51
+ [![essays](https://img.shields.io/badge/source--reading-8_bilingual-orange)](article/00-index_EN.md)
52
+
53
+ </div>
54
+
55
+ - **Readable end to end.** Read the whole engine in an afternoon: 1,081 lines of pure Python, with no magic hidden anywhere you can't follow it.
56
+ - **Hackable.** Set a breakpoint on any line, change it, rerun, all on your own machine. It genuinely works, which makes this a living reference rather than a diagram. It just isn't meant to be your daily driver.
57
+ - **The gaps are the point.** It deliberately keeps only the minimal core; what's missing isn't half-finished, it's where you branch off and make it your own.
58
+
59
+ ## What this is
60
+
61
+ I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
62
+
63
+ The engine (loop, model interface, context, tools, sessions) is 1,081 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 18 files: 1,714 physical lines, 1,385 net, every one short enough to read in a single sitting.
64
+
65
+ And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. 86 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
66
+
67
+ The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
68
+
69
+ <p align="center">
70
+ <img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo_en.png" width="760"
71
+ alt="A real CoreCoder run: corecoder -p asks it to fix buggy.py; the agent reads the file, edits the code, runs it to confirm, and reports what it changed.">
72
+ </p>
73
+
74
+ <p align="center"><sub><i>These thousand lines really do run a full loop end to end: ask it to fix buggy.py and it reads the file, edits the code, runs it once to confirm, then reports back on its own. Watch it, then come back and read the code.</i></sub></p>
75
+
76
+ This README follows the same arc: the first half helps you **read it** (the code map, the main loop, eight essays), the second half helps you **fork it** and points at a few directions worth pushing further.
77
+
78
+ ## Run it once first (five minutes before you read)
79
+
80
+ Before you read the source, get it running on your machine once to build some intuition. It's a foundation meant for forking, so the recommended path is to clone it and install editable, reading and changing as you go:
81
+
82
+ ```bash
83
+ git clone https://github.com/he-yufeng/CoreCoder
84
+ cd CoreCoder
85
+ pip install -e .
86
+ ```
87
+
88
+ If you just want to get it running first, `pip install corecoder` works too.
89
+
90
+ Give it a model and a key and it goes. It speaks the OpenAI-compatible API by default, and switching providers is usually just two environment variables:
91
+
92
+ | Provider | Example env vars |
93
+ |---|---|
94
+ | OpenAI (default `gpt-5.5`) | `OPENAI_API_KEY=sk-...` |
95
+ | DeepSeek | `OPENAI_API_KEY=sk-... OPENAI_BASE_URL=https://api.deepseek.com CORECODER_MODEL=deepseek-chat` |
96
+ | Local Ollama | `OPENAI_API_KEY=ollama OPENAI_BASE_URL=http://localhost:11434/v1 CORECODER_MODEL=qwen2.5-coder` |
97
+
98
+ Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
99
+
100
+ ```bash
101
+ corecoder # interactive REPL
102
+ corecoder -p "add error handling to parse_config()" # one-shot mode, exits when done
103
+ ```
104
+
105
+ ## Read it: the code map
106
+
107
+ Laid out flat, the whole project is this big. Skim it before you clone and you'll know where everything is. This is the most concrete difference from Claude Code's hundreds of thousands of lines: you can read it like the table of contents of a book. Start from the main loop in `agent.py`; that's the heart of the whole agent.
108
+
109
+ ```
110
+ corecoder/
111
+ ├── agent.py agent loop + parallel tool exec 150 lines ← start here
112
+ ├── llm.py streaming client + retry + cost 336 lines
113
+ ├── context.py three-tier context compaction 210 lines
114
+ ├── session.py save / resume + path-traversal guard 97 lines
115
+ ├── prompt.py system prompt 33 lines
116
+ ├── cli.py REPL + slash commands + one-shot 270 lines
117
+ ├── config.py env-var config 57 lines
118
+ └── tools/
119
+ ├── bash.py shell + dangerous-command gate + cd 127 lines
120
+ ├── edit.py unique-match search/replace + diff 92 lines
121
+ ├── grep.py content search 79 lines
122
+ ├── glob_tool.py filename matching 47 lines
123
+ ├── read.py file read 53 lines
124
+ ├── write.py file write 38 lines
125
+ ├── agent.py sub-agent spawning 58 lines
126
+ └── base.py tool base class 27 lines
127
+ ```
128
+
129
+ Seven tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, and `agent` (which spawns a sub-agent). Add the packaging files like `__init__` and `__main__` and the package is 18 files; strip the CLI shell and config and the engine itself is about 1,081 lines.
130
+
131
+ ## A `while` loop is the whole agent
132
+
133
+ The whole of an agent fits in one sentence: hand the user's words to the model, run whatever tools it asks for, stuff the results back into the context, ask again, and keep going until it stops asking for tools and gives an answer. In code, that's about a dozen lines:
134
+
135
+ ```python
136
+ # corecoder/agent.py · the main loop (trimmed skeleton)
137
+ def chat(self, user_input):
138
+ self.messages.append(user_input)
139
+
140
+ for _ in range(self.max_rounds): # bounded, so it can't run away
141
+ reply = self.llm.chat(self.messages, self.tools) # ask the model what to do next
142
+ if not reply.tool_calls: # model wants no more tools
143
+ return reply.text # -> done, hand the answer back
144
+ results = run_parallel(reply.tool_calls) # tools requested -> run in parallel
145
+ self.messages += results # feed results back, loop again
146
+
147
+ return "(hit the round limit)"
148
+ ```
149
+
150
+ That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess that shows up once the loop runs against the real world. `llm.py` is the biggest file in the project, not because calling a model is hard, but because in a streamed response a single tool call's arguments arrive in several fragments, one after another, and you have to stitch them back together in order. A provider will occasionally hand you half a JSON object, or fill the `usage` field with null; rate limits (429), timeouts, dropped connections and 5xx all need backoff-and-retry, while the other 4xx should just raise instead of being retried into the ground. Even an OpenAI extension like `stream_options` gets handled: some providers reject it outright with a 400, so the code strips it and resends exactly once, only on a 400, and never stacks that on top of the retry backoff. Lay all this grunt work out and it's the genuinely hard engineering part of taking an agent from demo to delivery. The unglamorous part turns out to be the loop itself; the thousand lines of fallback around it are where the real work lives.
151
+
152
+ Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
153
+
154
+ **`edit_file` does search-and-replace on a unique match, not line numbers.** Line numbers are a trap: the model only has to miscount by one and it quietly edits the wrong place. Anchor on a unique snippet of the original instead. If there's no match, it hands the start of the file back so the model can re-anchor; if there are several matches, it makes the model bring more surrounding context rather than gamble on one. On a successful edit it returns a diff. Recoverable on failure, verifiable on success: the whole loop stays inside the tool.
155
+
156
+ **Context isn't cut all at once when it's full; it gives ground in three tiers, cheapest first.** At half full (50%) it trims over-long tool outputs in place, a tier that's purely mechanical and costs no model call. If 70% still isn't enough, it has the model summarize the older turns into a single paragraph while keeping the most recent ones verbatim. Only at 90% does it hit the emergency tier and pull everything, summary and recent turns alike, down to its tightest form. Blunt truncation tends to throw away exactly the early decision a long task leans on most; tiering lets it surrender the least important things first instead of lopping off the oldest decisions wholesale from the start.
157
+
158
+ **You constrain a sub-agent by withholding the tool, not by writing rules and hoping it obeys.** A spawned sub-agent gets an isolated context and its own separate history, with a toolset exactly one item shorter than the parent's: the `agent` tool itself, so it can't recursively spawn more sub-agents. Handing it one fewer tool is cleaner than legislating a rule after the fact. It also reuses the parent's model connection (its spend folded into the same running total), truncates its output once it runs past 5,000 characters down to just the opening, and runs on a shorter round limit than the parent. The same restraint, end to end.
159
+
160
+ Every one of these *whys* is traced down to the actual lines of code in the series below.
161
+
162
+ ## The source-reading series · 8 bilingual essays
163
+
164
+ I also wrote a bilingual source-reading series, one intro plus seven parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking; read them in any order.
165
+
166
+ - **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
167
+ - **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
168
+ - **[02 · The tool system: letting the model act, safely](article/02-tools_EN.md)** — the seven tools in `tools/` and the bash safety gate
169
+ - **[03 · Plug in any LLM, and keep the bill honest](article/03-llm-and-cost_EN.md)** — `llm.py`'s provider wrapper, retries, and cost accounting
170
+ - **[04 · Surviving a long task on a finite window](article/04-context_EN.md)** — `context.py`'s three-tier compaction and orphaned tool messages
171
+ - **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
172
+ - **[06 · Turning it into a real command-line tool](article/06-session-and-cli_EN.md)** — `session.py` and path-traversal defense
173
+ - **[07 · Fork CoreCoder into your own coding agent](article/07-build-your-own_EN.md)** — from fork to custom tools to swapping models
174
+
175
+ ## Fork it, build something better
176
+
177
+ Once you understand it, the natural next step is to fork. Getting started doesn't take much:
178
+
179
+ - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (336 lines) is the entry point for all provider adaptation.
180
+ - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (27 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
181
+ - **Rewrite the system prompt.** `prompt.py` is all of 33 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
182
+ - **Import it as a library.** The top level exports `Agent`, `LLM`, and `Config`, ready to embed in your own program:
183
+
184
+ ```python
185
+ from corecoder import Agent, LLM
186
+
187
+ llm = LLM(model="deepseek-chat", api_key="sk-...", base_url="https://api.deepseek.com")
188
+ print(Agent(llm=llm).chat("find every TODO comment in this project and list them"))
189
+ ```
190
+
191
+ Going deeper, the directions are out in the open too. None of the following is in CoreCoder, by design, not because it's unfinished. Flip it around and each one is an entry point you can carry into a real tool of your own:
192
+
193
+ - **The dangerous-command blocking in bash is just a regex blacklist.** It guards against slips, not a security sandbox. Facing untrusted input means reaching for seccomp or container isolation. This is the hardest of the four; it goes all the way down to the syscall and isolation layer.
194
+ - **Retry is only exponential backoff.** No fallback model, no hard dollar budget. Follow `llm.py` down and add a fallback model chain plus a stop-on-over-budget gate; the change stays mostly inside that one file.
195
+ - **Sub-agents only run the plainest synchronous execution.** Make it async or a streaming executor and you close the exact gap the fifth essay identifies between this and how production agents stream execution.
196
+ - **No MCP, no RAG.** Wire up MCP to give it the external tool ecosystem, or add retrieval-based code location for big repos. Both are real ways to grow from a minimal core into your own stronger agent.
197
+
198
+ The README only points; the seventh essay picks up the code details for each. Pick one and start; that's the whole reason the core is kept this small.
199
+
200
+ ## How it compares
201
+
202
+ | | CoreCoder | Claude Code | aider | nanoGPT |
203
+ |---|---|---|---|---|
204
+ | Lines of code | ~1,081 engine / 1,714 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
205
+ | Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
206
+ | Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
207
+ | What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
208
+
209
+ The nanoGPT column is there as a reference point: minimal, readable, but it teaches you to train a GPT. CoreCoder is after the same thing, only the subject is an agent that actually edits code. Sitting it next to Claude Code and aider isn't about competing for their users. CoreCoder is the foundation you stand on while you learn from them and get going; it isn't in the same race.
210
+
211
+ ## Commands
212
+
213
+ Inside the REPL, `/help` lists everything; these are the ones you'll reach for:
214
+
215
+ ```
216
+ /model <name> switch model
217
+ /compact compact the context by hand
218
+ /tokens token usage and cost estimate
219
+ /diff files changed this session
220
+ /save /sessions save / list sessions
221
+ quit / exit exit (Ctrl+C cancels the current round)
222
+ ```
223
+
224
+ Session IDs are sanitized to safe characters before they become filenames, every archive lands under `~/.corecoder/sessions`, and a malicious session name can't traverse out.
225
+
226
+ ## Contributing / License
227
+
228
+ Before you send anything, run `pytest tests/ -q` (86 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
229
+
230
+ ---
231
+
232
+ By [Yufeng He](https://github.com/he-yufeng), formerly at Moonshot AI (Kimi). I earlier wrote a fairly complete [Claude Code source analysis](https://zhuanlan.zhihu.com/p/1898797658343862272) on Zhihu; this project is its hands-on counterpart: that one walks you through reading it, this one through rebuilding it.
233
+
234
+ > CoreCoder was formerly named NanoCoder; it was renamed to avoid confusion with [Nano-Collective/nanocoder](https://github.com/Nano-Collective/nanocoder), and old links redirect here automatically.
@@ -0,0 +1,199 @@
1
+ <div align="center">
2
+
3
+ # CoreCoder
4
+
5
+ **The nanoGPT of coding agents. 1,081 lines of pure Python — understand how a coding agent actually works, then fork your own.**
6
+
7
+ *learn from it · fork it · ship something better*
8
+
9
+ [中文](README_CN.md) | English | [Source-reading series · 8 bilingual essays](article/00-index_EN.md)
10
+
11
+ [![PyPI](https://img.shields.io/pypi/v/corecoder)](https://pypi.org/project/corecoder/)
12
+ [![Python](https://img.shields.io/badge/python-3.10+-blue)](https://python.org)
13
+ [![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)
14
+ [![Tests](https://github.com/he-yufeng/CoreCoder/actions/workflows/ci.yml/badge.svg)](https://github.com/he-yufeng/CoreCoder/actions)
15
+ [![engine](https://img.shields.io/badge/engine-1081_LoC-blue)](article/00-index_EN.md)
16
+ [![essays](https://img.shields.io/badge/source--reading-8_bilingual-orange)](article/00-index_EN.md)
17
+
18
+ </div>
19
+
20
+ - **Readable end to end.** Read the whole engine in an afternoon: 1,081 lines of pure Python, with no magic hidden anywhere you can't follow it.
21
+ - **Hackable.** Set a breakpoint on any line, change it, rerun, all on your own machine. It genuinely works, which makes this a living reference rather than a diagram. It just isn't meant to be your daily driver.
22
+ - **The gaps are the point.** It deliberately keeps only the minimal core; what's missing isn't half-finished, it's where you branch off and make it your own.
23
+
24
+ ## What this is
25
+
26
+ I've always felt coding agents get talked about as if they were arcane. Strip a tool like Claude Code or Cursor all the way down and the core is a `while` loop wrapped around a large model, plus seven or eight tools that let it actually do things. The hard part was never the loop; it's everything the loop has to cope with once it meets the real world. CoreCoder is the minimal version that writes that core out honestly.
27
+
28
+ The engine (loop, model interface, context, tools, sessions) is 1,081 lines once you drop blank lines and comments. Counting the outer CLI, config and packaging too, the whole package is 18 files: 1,714 physical lines, 1,385 net, every one short enough to read in a single sitting.
29
+
30
+ And it really runs: reads and writes files, executes shell, spawns sub-agents, compacts context in three tiers, and tells you the tokens and dollars a run burned whenever you ask. 86 tests, all green. But the point of it running isn't to become your daily driver. It runs so the walkthrough can't lie: a reference that shows how an agent works has to actually work.
31
+
32
+ The code came out of a public teardown: open analyses have already exposed a lot of the load-bearing architecture inside production agents like Claude Code. I took the most essential layer and rewrote it honestly, in as little code as I could. So reading CoreCoder is roughly like reading a runnable, annotated take on how that kind of agent works, except it's only a minimal reimplementation, sitting right there on your machine for you to take apart and change.
33
+
34
+ <p align="center">
35
+ <img src="https://raw.githubusercontent.com/he-yufeng/CoreCoder/main/assets/demo_en.png" width="760"
36
+ alt="A real CoreCoder run: corecoder -p asks it to fix buggy.py; the agent reads the file, edits the code, runs it to confirm, and reports what it changed.">
37
+ </p>
38
+
39
+ <p align="center"><sub><i>These thousand lines really do run a full loop end to end: ask it to fix buggy.py and it reads the file, edits the code, runs it once to confirm, then reports back on its own. Watch it, then come back and read the code.</i></sub></p>
40
+
41
+ This README follows the same arc: the first half helps you **read it** (the code map, the main loop, eight essays), the second half helps you **fork it** and points at a few directions worth pushing further.
42
+
43
+ ## Run it once first (five minutes before you read)
44
+
45
+ Before you read the source, get it running on your machine once to build some intuition. It's a foundation meant for forking, so the recommended path is to clone it and install editable, reading and changing as you go:
46
+
47
+ ```bash
48
+ git clone https://github.com/he-yufeng/CoreCoder
49
+ cd CoreCoder
50
+ pip install -e .
51
+ ```
52
+
53
+ If you just want to get it running first, `pip install corecoder` works too.
54
+
55
+ Give it a model and a key and it goes. It speaks the OpenAI-compatible API by default, and switching providers is usually just two environment variables:
56
+
57
+ | Provider | Example env vars |
58
+ |---|---|
59
+ | OpenAI (default `gpt-5.5`) | `OPENAI_API_KEY=sk-...` |
60
+ | DeepSeek | `OPENAI_API_KEY=sk-... OPENAI_BASE_URL=https://api.deepseek.com CORECODER_MODEL=deepseek-chat` |
61
+ | Local Ollama | `OPENAI_API_KEY=ollama OPENAI_BASE_URL=http://localhost:11434/v1 CORECODER_MODEL=qwen2.5-coder` |
62
+
63
+ Kimi, Qwen and the like are the same two variables; for providers that don't even offer an OpenAI-compatible endpoint, the optional LiteLLM backend (`pip install "corecoder[litellm]"`) routes to a hundred-plus of them. The third essay goes into this in detail. The key can be `export`ed directly or dropped into a `.env` at the project root, which is loaded on startup. Then:
64
+
65
+ ```bash
66
+ corecoder # interactive REPL
67
+ corecoder -p "add error handling to parse_config()" # one-shot mode, exits when done
68
+ ```
69
+
70
+ ## Read it: the code map
71
+
72
+ Laid out flat, the whole project is this big. Skim it before you clone and you'll know where everything is. This is the most concrete difference from Claude Code's hundreds of thousands of lines: you can read it like the table of contents of a book. Start from the main loop in `agent.py`; that's the heart of the whole agent.
73
+
74
+ ```
75
+ corecoder/
76
+ ├── agent.py agent loop + parallel tool exec 150 lines ← start here
77
+ ├── llm.py streaming client + retry + cost 336 lines
78
+ ├── context.py three-tier context compaction 210 lines
79
+ ├── session.py save / resume + path-traversal guard 97 lines
80
+ ├── prompt.py system prompt 33 lines
81
+ ├── cli.py REPL + slash commands + one-shot 270 lines
82
+ ├── config.py env-var config 57 lines
83
+ └── tools/
84
+ ├── bash.py shell + dangerous-command gate + cd 127 lines
85
+ ├── edit.py unique-match search/replace + diff 92 lines
86
+ ├── grep.py content search 79 lines
87
+ ├── glob_tool.py filename matching 47 lines
88
+ ├── read.py file read 53 lines
89
+ ├── write.py file write 38 lines
90
+ ├── agent.py sub-agent spawning 58 lines
91
+ └── base.py tool base class 27 lines
92
+ ```
93
+
94
+ Seven tools: `bash`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, and `agent` (which spawns a sub-agent). Add the packaging files like `__init__` and `__main__` and the package is 18 files; strip the CLI shell and config and the engine itself is about 1,081 lines.
95
+
96
+ ## A `while` loop is the whole agent
97
+
98
+ The whole of an agent fits in one sentence: hand the user's words to the model, run whatever tools it asks for, stuff the results back into the context, ask again, and keep going until it stops asking for tools and gives an answer. In code, that's about a dozen lines:
99
+
100
+ ```python
101
+ # corecoder/agent.py · the main loop (trimmed skeleton)
102
+ def chat(self, user_input):
103
+ self.messages.append(user_input)
104
+
105
+ for _ in range(self.max_rounds): # bounded, so it can't run away
106
+ reply = self.llm.chat(self.messages, self.tools) # ask the model what to do next
107
+ if not reply.tool_calls: # model wants no more tools
108
+ return reply.text # -> done, hand the answer back
109
+ results = run_parallel(reply.tool_calls) # tools requested -> run in parallel
110
+ self.messages += results # feed results back, loop again
111
+
112
+ return "(hit the round limit)"
113
+ ```
114
+
115
+ That's the whole thing. The core skeleton is about twenty lines; counting parallel execution and the bookkeeping after a Ctrl+C interrupt, maybe forty. Almost everything else in CoreCoder's thousand-odd lines is there to clean up the mess that shows up once the loop runs against the real world. `llm.py` is the biggest file in the project, not because calling a model is hard, but because in a streamed response a single tool call's arguments arrive in several fragments, one after another, and you have to stitch them back together in order. A provider will occasionally hand you half a JSON object, or fill the `usage` field with null; rate limits (429), timeouts, dropped connections and 5xx all need backoff-and-retry, while the other 4xx should just raise instead of being retried into the ground. Even an OpenAI extension like `stream_options` gets handled: some providers reject it outright with a 400, so the code strips it and resends exactly once, only on a 400, and never stacks that on top of the retry backoff. Lay all this grunt work out and it's the genuinely hard engineering part of taking an agent from demo to delivery. The unglamorous part turns out to be the loop itself; the thousand lines of fallback around it are where the real work lives.
116
+
117
+ Three decisions are worth a closer look, because they're the kind of call you can only make after you've understood how others did it, and they're judgments you can lift straight into your own fork.
118
+
119
+ **`edit_file` does search-and-replace on a unique match, not line numbers.** Line numbers are a trap: the model only has to miscount by one and it quietly edits the wrong place. Anchor on a unique snippet of the original instead. If there's no match, it hands the start of the file back so the model can re-anchor; if there are several matches, it makes the model bring more surrounding context rather than gamble on one. On a successful edit it returns a diff. Recoverable on failure, verifiable on success: the whole loop stays inside the tool.
120
+
121
+ **Context isn't cut all at once when it's full; it gives ground in three tiers, cheapest first.** At half full (50%) it trims over-long tool outputs in place, a tier that's purely mechanical and costs no model call. If 70% still isn't enough, it has the model summarize the older turns into a single paragraph while keeping the most recent ones verbatim. Only at 90% does it hit the emergency tier and pull everything, summary and recent turns alike, down to its tightest form. Blunt truncation tends to throw away exactly the early decision a long task leans on most; tiering lets it surrender the least important things first instead of lopping off the oldest decisions wholesale from the start.
122
+
123
+ **You constrain a sub-agent by withholding the tool, not by writing rules and hoping it obeys.** A spawned sub-agent gets an isolated context and its own separate history, with a toolset exactly one item shorter than the parent's: the `agent` tool itself, so it can't recursively spawn more sub-agents. Handing it one fewer tool is cleaner than legislating a rule after the fact. It also reuses the parent's model connection (its spend folded into the same running total), truncates its output once it runs past 5,000 characters down to just the opening, and runs on a shorter round limit than the parent. The same restraint, end to end.
124
+
125
+ Every one of these *whys* is traced down to the actual lines of code in the series below.
126
+
127
+ ## The source-reading series · 8 bilingual essays
128
+
129
+ I also wrote a bilingual source-reading series, one intro plus seven parts, each in Chinese with an English mirror. Against CoreCoder's actual code, it walks through how agents like Claude Code work under the hood. One hard rule I set myself: every line count and every snippet is re-read and re-checked from the repo, never written from memory. The first six get you reading, the seventh gets you forking; read them in any order.
130
+
131
+ - **[Intro · Read Claude Code through CoreCoder, then build your own](article/00-index_EN.md)**
132
+ - **[01 · An agent, at its core, is a `while` loop](article/01-the-loop_EN.md)** — the main loop in `agent.py`, interrupts, and the round limit
133
+ - **[02 · The tool system: letting the model act, safely](article/02-tools_EN.md)** — the seven tools in `tools/` and the bash safety gate
134
+ - **[03 · Plug in any LLM, and keep the bill honest](article/03-llm-and-cost_EN.md)** — `llm.py`'s provider wrapper, retries, and cost accounting
135
+ - **[04 · Surviving a long task on a finite window](article/04-context_EN.md)** — `context.py`'s three-tier compaction and orphaned tool messages
136
+ - **[05 · Parallel execution and sub-agents](article/05-parallel-and-subagents_EN.md)** — thread-pool concurrency and sub-agent isolation
137
+ - **[06 · Turning it into a real command-line tool](article/06-session-and-cli_EN.md)** — `session.py` and path-traversal defense
138
+ - **[07 · Fork CoreCoder into your own coding agent](article/07-build-your-own_EN.md)** — from fork to custom tools to swapping models
139
+
140
+ ## Fork it, build something better
141
+
142
+ Once you understand it, the natural next step is to fork. Getting started doesn't take much:
143
+
144
+ - **Swap in a model you actually use.** It's the two env vars from above; `llm.py` (336 lines) is the entry point for all provider adaptation.
145
+ - **Add a tool of your own.** Write a new file against the tool base class in `tools/base.py` (27 lines): run tests, fetch a page, call an LSP, whatever. The end of the second essay walks you through your first one by hand.
146
+ - **Rewrite the system prompt.** `prompt.py` is all of 33 lines; change one line and you'll watch the agent's temperament shift. It's the cheapest "change one thing, see a result" in the whole project.
147
+ - **Import it as a library.** The top level exports `Agent`, `LLM`, and `Config`, ready to embed in your own program:
148
+
149
+ ```python
150
+ from corecoder import Agent, LLM
151
+
152
+ llm = LLM(model="deepseek-chat", api_key="sk-...", base_url="https://api.deepseek.com")
153
+ print(Agent(llm=llm).chat("find every TODO comment in this project and list them"))
154
+ ```
155
+
156
+ Going deeper, the directions are out in the open too. None of the following is in CoreCoder, by design, not because it's unfinished. Flip it around and each one is an entry point you can carry into a real tool of your own:
157
+
158
+ - **The dangerous-command blocking in bash is just a regex blacklist.** It guards against slips, not a security sandbox. Facing untrusted input means reaching for seccomp or container isolation. This is the hardest of the four; it goes all the way down to the syscall and isolation layer.
159
+ - **Retry is only exponential backoff.** No fallback model, no hard dollar budget. Follow `llm.py` down and add a fallback model chain plus a stop-on-over-budget gate; the change stays mostly inside that one file.
160
+ - **Sub-agents only run the plainest synchronous execution.** Make it async or a streaming executor and you close the exact gap the fifth essay identifies between this and how production agents stream execution.
161
+ - **No MCP, no RAG.** Wire up MCP to give it the external tool ecosystem, or add retrieval-based code location for big repos. Both are real ways to grow from a minimal core into your own stronger agent.
162
+
163
+ The README only points; the seventh essay picks up the code details for each. Pick one and start; that's the whole reason the core is kept this small.
164
+
165
+ ## How it compares
166
+
167
+ | | CoreCoder | Claude Code | aider | nanoGPT |
168
+ |---|---|---|---|---|
169
+ | Lines of code | ~1,081 engine / 1,714 total | hundreds of thousands (closed) | tens of thousands of Python | ~600 (two files) |
170
+ | Time to read it all | one afternoon | can't (closed) | a few days of slogging | one afternoon |
171
+ | Breakpoint, change, rerun? | yes, every line | no | yes, but there's a lot | yes |
172
+ | What it's for | understand one, then fork your own | production coding assistant | terminal pair-programming | minimal GPT for teaching |
173
+
174
+ The nanoGPT column is there as a reference point: minimal, readable, but it teaches you to train a GPT. CoreCoder is after the same thing, only the subject is an agent that actually edits code. Sitting it next to Claude Code and aider isn't about competing for their users. CoreCoder is the foundation you stand on while you learn from them and get going; it isn't in the same race.
175
+
176
+ ## Commands
177
+
178
+ Inside the REPL, `/help` lists everything; these are the ones you'll reach for:
179
+
180
+ ```
181
+ /model <name> switch model
182
+ /compact compact the context by hand
183
+ /tokens token usage and cost estimate
184
+ /diff files changed this session
185
+ /save /sessions save / list sessions
186
+ quit / exit exit (Ctrl+C cancels the current round)
187
+ ```
188
+
189
+ Session IDs are sanitized to safe characters before they become filenames, every archive lands under `~/.corecoder/sessions`, and a malicious session name can't traverse out.
190
+
191
+ ## Contributing / License
192
+
193
+ Before you send anything, run `pytest tests/ -q` (86 tests), `ruff check`, and `compileall`, and make sure they're green. MIT licensed: fork it, learn from it, ship something better. A mention of this project is appreciated.
194
+
195
+ ---
196
+
197
+ By [Yufeng He](https://github.com/he-yufeng), formerly at Moonshot AI (Kimi). I earlier wrote a fairly complete [Claude Code source analysis](https://zhuanlan.zhihu.com/p/1898797658343862272) on Zhihu; this project is its hands-on counterpart: that one walks you through reading it, this one through rebuilding it.
198
+
199
+ > CoreCoder was formerly named NanoCoder; it was renamed to avoid confusion with [Nano-Collective/nanocoder](https://github.com/Nano-Collective/nanocoder), and old links redirect here automatically.