fapony 0.3.6 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,26 +6,26 @@
6
6
 
7
7
  [![npm](https://img.shields.io/npm/v/fapony.svg)](https://www.npmjs.com/package/fapony)
8
8
 
9
- **Where did your tokens go?** fapony reads the session logs Claude Code, Codex, OpenCode and ZCode
10
- already write, and puts them all on one yardstick — tokens, cost and time per model, per client,
11
- per workflow. Nothing to instrument, no per-project setup, no waiting for data to accumulate: it
12
- runs on the history already sitting on your disk.
9
+ **See what your coding agents actually cost.** fapony reads the session logs Claude Code, Codex,
10
+ OpenCode and ZCode already write, and puts them all on one yardstick — tokens, cost and time per
11
+ model, per client, per workflow. Nothing to instrument, no per-project setup: it runs on the
12
+ history already sitting on your disk.
13
13
 
14
14
  <p align="center">
15
15
  <img src="images/summary.webp" width="800" alt="fapony usage-web summary cards">
16
16
  </p>
17
17
 
18
18
  ```bash
19
- fapony usage-scan # read the session logs already on your disk
20
- fapony price-scan # fetch the price table (needed once, for cost)
21
- fapony usage-web # every session you already have, all clients, one page
19
+ npm install -g fapony # needs Bun — https://bun.sh
20
+ fapony usage-scan # read the session logs already on your disk
21
+ fapony price-scan # fetch the price table (needed once, for cost)
22
+ fapony usage-web # every session you already have, all clients, one page
22
23
  ```
23
24
 
24
25
  **Cost is the part your client probably isn't logging.** Of the four, only OpenCode writes a real
25
- dollar figure into its session log — Claude Code, ZCode and Codex record `0`. fapony prices those
26
- sessions at published list rates and labels the number `imputed`, so a figure you can compare
27
- across clients exists at all. A model it can't find a rate for stays `unpriced`: nothing is
28
- quietly counted as free.
26
+ dollar figure into its session log — the others record `0`. fapony prices those sessions at
27
+ published list rates and labels the number `imputed`; a model it can't find a rate for stays
28
+ `unpriced` — nothing is quietly counted as free.
29
29
 
30
30
  <details>
31
31
  <summary>full usage-web dashboard preview</summary>
@@ -36,151 +36,88 @@ quietly counted as free.
36
36
 
37
37
  </details>
38
38
 
39
- That is day one. Past that, fapony keeps what coding agents actually did — the frozen
40
- ledger of graded runs (rounds, pass/fail, cost per grade, readable via CLI, no new grades)
41
- plus the live mem log — through 3 MCP tools any agent can call. If you juggle more than
42
- one agent, this is the point: the numbers come from the same yardstick everywhere, so
43
- "which model earns its keep on which kind of task" becomes a data question instead of a
44
- vibe. On top of history it checks claims against git facts: handoff conformance and
45
- allowlisted evidence — with everything the agent claimed but couldn't prove marked as such.
46
-
47
- **What that question looks like answered, from one project's own (frozen — reads history,
48
- no new grades) ledger — the top of the `n≥5` frontier (`fapony stats --mode verdict --regime code`):**
49
-
50
- | model | tokens/pass | quality | n |
51
- |---|---|---|---|
52
- | `claude-opus-5` | 22.5M | 3.8 | 10 |
53
- | `claude-sonnet-5` | 5.6M | 3.5 | 11 |
54
- | `muse-spark-1.3-contributor-free` | 4.3M | 4.0 | 5 |
55
-
56
- Same quality band, an 8× token spread — the kind of answer a session log can't give (it has tokens,
57
- no grades) and a benchmark can't give either (it has grades, not your codebase). One caveat that's
58
- on you to hold: work isn't randomly assigned to models, so a gap this size is a strong prior, not a
59
- controlled trial — you likely route easy tasks to the cheap model already. `n≥5` is fapony's own
60
- floor before a model counts toward the frontier at all; below that it's a data point, not a pick.
61
-
62
- **The reason to keep it running is the third layer: knowledge accumulation — and the thing it
63
- accumulates is pain.** An agent has no memory of pain across sessions: it writes the 37th
64
- hand-rolled `try/catch` as cheerfully as the first, because every session starts new. Wrappers and
65
- shared libraries get built by *people* who were hurt by the same thing often enough to remember.
66
- That is why a codebase written with agents from day one tends not to grow a shared layer — nobody
67
- in the room remembers. fapony is the part that remembers: mem rows carry
68
- `files[]`, so the zones that keep coming back in re-done work are a query, not a hunch
69
- (the frozen ledger's old graded rows carry them too).
70
- Paired with `fapony debt`, which tracks how far the codebase has actually moved to a convention you
71
- already decided on, that is the loop: notice the repeated cost, name the shared thing, watch the
72
- migration finish. Finding dead code and duplication is *not* part of it — knip and friends already
73
- do that better, and a convention with a `checker` is deliberately left to the checker.
74
-
75
- **The measurement layer underneath it:** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: which model earns its keep on which kind of work **in this project**, at what token cost. The frozen ledger still answers that from history — every old verdict carries a `regime` (`code` / `fix` / `review` / `plan` / `inquiry` / `test`), and runs split by whether there was a plan at all — so "does planning beat diving in, and for which model" stays a table, not an argument. New accumulation goes to the mem log instead: decisions, bugs and notes with `files[]`, written by the agents doing the work.
76
-
77
- Three tiers, deliberately: **measurement ships today** and needs no per-project setup — raw facts nobody can call unfair. **Verification is the sharper edge** but stays beta until its evidence layer is hardened; fapony doesn't control your agent's flow, so it never promises "verified" as a headline. **Knowledge accumulation is the compounding one** — it's worthless on run 1 and gets more useful every run after, which is exactly why it's the layer competitors can't clone by copying a feature list.
78
-
79
- Adopting it doesn't change your workflow. There is no loop to join and no framework to learn: install the MCP server, point your agent at it, and read the reports. fapony also ships plans, skills and read-only seed commands from its own dogfooding — those are conveniences, kept in their own section below, and deleting all of them costs you nothing the ledger can measure.
39
+ Raw facts from logs are hard to argue with — a vendor can dispute a verdict as unfair; they can't
40
+ dispute their own token count. That is the whole measurement layer: tokens and cost, nothing
41
+ self-graded.
80
42
 
81
- ## What fapony is not
43
+ ## Past day one
82
44
 
83
- Stated up front, because the gap between these two things is where most tooling oversells:
45
+ Two more layers, both optional, both compounding:
84
46
 
85
- - **It does not run your test suite.** The evidence collector runs an allowlist *you* write in
86
- `.fapony/evidence.json`, and never a command an agent proposes. No allowlist, no evidence.
87
- - **It does not judge your code.** Mem rows *record* decisions, bugs and notes; a human or
88
- a working agent supplies them. fapony is the memory, not the judge. (The frozen
89
- ledger's old grades work the same way — *stored*, never computed.)
90
- - **It checks conformance, not correctness.** What it can verify is that a claim lines up with git
91
- facts and that uncertainty was declared — not that the code works. Those are different
92
- guarantees and fapony only offers the first.
93
- - **Almost nothing blocks.** No CI failure, no gate on your own commands. The one exception is the
94
- Stop hook, once per turn when a commit lands with no new mem row; the read/edit/commit hints only annotate.
95
- Skip the install of all of them and you are back to exactly the workflow you had.
96
- - **Model attribution is inferred, not declared.** A gate in the frozen ledger is attributed to whichever client
97
- session was live in that worktree at that moment. When one model writes the code and another
98
- reviews and files the verdict, the grade lands on the reviewer. Reports label it `inferred`;
99
- read it as such.
100
- - **The knowledge layer is empty on run 1.** It is worth something around run 5 and more every run
101
- after.
102
-
103
- ## Quick start (MCP)
47
+ **Memory — the mem log.** An agent has no memory of pain across sessions: it writes the 37th
48
+ hand-rolled `try/catch` as cheerfully as the first, because every session starts new. Wrappers
49
+ and shared libraries get built by *people* who were hurt often enough to remember. fapony
50
+ remembers instead: one MCP call per unit of work (`mem_add`) records the decision, bug or note
51
+ with the files it touched, and `mem_find` answers *"what was ever decided about this file?"*
52
+ before the next agent touches it. The log lives in your repo (`.fapony/.memory/`), so it crosses
53
+ machines over git for free.
54
+
55
+ **Convention debt.** `fapony debt` answers the question nothing else does: *we decided this six
56
+ months ago — how far along is the move?* ESLint says this line is wrong; nothing says 11 of 47
57
+ files have migrated. Dead code and duplication it deliberately leaves to knip and friends —
58
+ they already do that better.
59
+
60
+ ```mermaid
61
+ flowchart LR
62
+ A[Claude Code] --> F[fapony]
63
+ B[OpenCode] --> F
64
+ C[ZCode] --> F
65
+ D[Codex] --> F
66
+ F --> U[usage — tokens & cost]
67
+ F --> M[mem log — what was decided here]
68
+ F --> D[debt — how far the move has gone]
69
+ ```
70
+
71
+ Adopting it doesn't change your workflow: install it, point your agent at it, read the reports.
72
+ Both layers above are per-project (`fapony init`) and worthless on run 1 — they get more useful
73
+ every run after, which is exactly why they're retention, not the reason to install.
74
+
75
+ ## Quick start
104
76
 
105
77
  ```bash
106
78
  # 1. Install (needs Bun — https://bun.sh)
107
- npm install -g fapony # or: bun add -g fapony
79
+ npm install -g fapony
108
80
  # from source instead:
109
81
  # git clone https://github.com/kire21b/fapony.git && cd fapony && bun install && bun link
110
- # note: `bun link` claims the global `fapony` bin by package name, not path — running it
111
- # from a second checkout silently repoints the command there. Re-run it in the one you want.
112
-
113
- # 2. Wire it into your MCP client
114
- fapony install # detects installed clients, asks which to wire
115
- fapony install --all # skip the prompt, wire everything detected
116
- # use --platform <name> to force a specific client (bypasses detection)
117
- # zcode/codex need their config to exist first — open the app once if you never have
118
- # claude/opencode also symlink skill/<name>/ into ~/.claude/skills — an existing
119
- # skill of the same name is reported, never overwritten
120
- # …or add it manually to any MCP client (e.g. Claude Desktop):
121
- # { "mcpServers": { "fapony": { "command": "fapony", "args": ["mcp"] } } }
82
+ # (`bun link` claims the global `fapony` bin by package name, not path — re-run it in the
83
+ # checkout you want to be the one)
122
84
 
123
- # 3. Measure — zero per-project setup
124
- fapony usage-scan # scan the session logs already on disk → cache
125
- fapony price-scan # fetch the OpenRouter price table → ~/.config/fapony/prices.json
126
- fapony usage-web # dashboard; re-run the scans to refresh
85
+ # 2. Measure — zero per-project setup
86
+ fapony usage-scan # scan the session logs already on disk → cache
87
+ fapony price-scan # fetch the OpenRouter price table → ~/.config/fapony/prices.json
88
+ fapony usage-web # dashboard; re-run the scans to refresh
127
89
  # both scans are manual by design — nothing fetches or re-reads session logs behind your back
128
90
 
129
- # 4. Turn on the knowledge layer (per project you want it in)
130
- fapony init /path/to/your-worktree
131
- # creates .fapony/ — .memory/ (the mem log the 3 MCP tools read and write),
132
- # conventions.json for `fapony debt`, plan/spec/done, and evidence.json
133
- # conventions.json + evidence.json are shared rules: commit them
134
- ```
135
-
136
- With `.fapony/evidence.json` in place, any graded run from the frozen ledger can be replayed
137
- as a report. This one is a CLI command, not an MCP tool — the schemas cost every session of
138
- every client and no skill called them (see [The 3 tools](#the-3-tools) below). No new runs
139
- can be created; run ids come from `fapony stats` reading history:
91
+ # 3. Wire your clients
92
+ fapony install # detects installed clients, asks which to wire
93
+ fapony install --all # skip the prompt, wire everything detected
94
+ # claude/opencode also symlink skill/<name>/ into ~/.claude/skills — an existing
95
+ # skill of the same name is reported, never overwritten
140
96
 
141
- ```bash
142
- fapony report <run-id> # run ids come from `fapony stats`
97
+ # 4. Turn on the memory layer (per project you want it in)
98
+ fapony init /path/to/your-worktree
99
+ # creates .fapony/ — .memory/ (the mem log the 3 MCP tools read and write)
100
+ # and conventions.json for `fapony debt` (shared rules: commit them)
143
101
  ```
144
102
 
145
- You get one report: git facts (files, commits, branch), handoff conformance (claims vs. reality), evidence from the allowlisted commands (pass/fail/timeout/unverified), the frozen 6-grade verdict, and cost — with anything the agent claimed but couldn't prove marked as such.
146
-
147
- Sections that have nothing to report say so (`not_run`, `unavailable`) rather than disappearing — a report with no evidence must not read like a report that passed.
148
-
149
- Two details for the allowlist once it is under version control. If your `.gitignore` ignores `.fapony/` wholesale, re-include the file (dir before file): `**/.fapony/*`, then `!**/.fapony/evidence.json`. In a monorepo, give an app its own `apps/<app>/.fapony/evidence.json` — reports whose changed files all sit under that app use it; anything else uses the root one.
150
-
151
- Two things worth knowing about the report header and budget:
152
-
153
- - **`server_sha`** — every report is stamped with the git SHA of the fapony code that produced it, read once at server start. MCP servers are long-lived: after you edit fapony and don't restart the client, reports keep coming from the old build. Compare the stamp against `git log -1` in the fapony repo; if they differ, reconnect the server before trusting the result.
154
- - **Evidence budget** — each allowlisted command gets `timeout_ms` (default 30s), and the whole report is capped at 180s total. A command that doesn't fit is reported as `timeout`, never as a pass. Time your real suite and set `timeout_ms` accordingly.
155
-
156
- ## How it fits
157
-
158
- ```mermaid
159
- flowchart LR
160
- A[Claude Code] --> F[fapony MCP]
161
- B[OpenCode] --> F
162
- C[ZCode] --> F
163
- D[Codex] --> F
164
- E[Cursor] --> F
165
- F --> G[git facts + session logs]
166
- G --> S[stats / usage]
167
- G --> V[verification report]
168
- G --> M[mem log - what was decided here]
169
- ```
103
+ …or add it manually to any MCP client: `{ "mcpServers": { "fapony": { "command": "fapony", "args": ["mcp"] } } }`. Full protocol and adapter examples: [docs/mcp-handcheck.md](https://github.com/kire21b/fapony/blob/main/docs/mcp-handcheck.md).
170
104
 
171
- fapony never drives the agent. It sits on two sides of your work that never touch each
172
- other, and **the rest of this README is organised along that line**: the ledger below is the
173
- product, and everything under "The work side" after it is a convenience you can delete without
174
- losing a single number.
105
+ ## What fapony is not
175
106
 
176
- | | The ledger | The work side |
177
- |---|---|---|
178
- | What it is | 3 MCP tools (mem) + a frozen SQLite ledger (reads history) | plans, skills, read-only seed commands |
179
- | Needs | an MCP client | nothing — or your own tooling instead |
180
- | Writes | one mem row per unit of work, into the project's log | nothing |
181
- | Skip it and | there is no fapony | fapony still answers every question |
107
+ Stated up front, because the gap between these two things is where most tooling oversells:
182
108
 
183
- ### What runs where
109
+ - **It does not run your test suite.** The evidence collector runs an allowlist *you* write in
110
+ `.fapony/evidence.json`, never a command an agent proposes. No allowlist, no evidence.
111
+ - **It does not judge your code.** Mem rows *record* what a human or a working agent supplies.
112
+ fapony is the memory, not the judge.
113
+ - **It checks conformance, not correctness** — that a claim lines up with git facts and that
114
+ uncertainty was declared, not that the code works.
115
+ - **Almost nothing blocks.** The one exception is the Stop hook, once per turn when a commit
116
+ lands with no new mem row; everything else only annotates.
117
+ - **Model attribution is inferred, not declared** — reports label it `inferred`; read it as such.
118
+ - **The knowledge layer is empty on run 1** — worth something around run 5, more every run after.
119
+
120
+ ## What runs where
184
121
 
185
122
  `fapony install` wires five clients (Claude Code, OpenCode, Cursor, ZCode, Codex). MCP is the only
186
123
  piece all of them get — the hooks and in-process hints are per-client, and the read/edit hints
@@ -199,119 +136,46 @@ is `tool.execute.after`. Nothing here is required: skip the hooks and every MCP
199
136
  | Skills symlinked into `~/.agents/skills` | — | — | — | ✅ | ✅ |
200
137
  | `usage-scan` reads this client's session log | ✅ | ✅ | — | ✅ | ✅ |
201
138
 
202
- `—` means not wired, not impossible: Cursor has no PreToolUse hook, and ZCode/Codex expose no
203
- in-process hook surface for read/edit hints yet (Codex's `apply_patch` sends patch text, not
204
- resolved file paths). Codex hooks require trust via `/hooks` before they run — `fapony install`
205
- tells you when. The hints live on hooks rather than MCP on purpose — they must fire mid-turn
206
- without the agent deciding to call anything ([why](#when-to-call-what)).
139
+ `—` means not wired, not impossible. Codex hooks require trust via `/hooks` before they run —
140
+ `fapony install` tells you when. The hints live on hooks rather than MCP on purpose — they must
141
+ fire mid-turn without the agent deciding to call anything.
207
142
 
208
- ## The ledger — this is the product
143
+ ## The ledger — one habit, 3 tools
209
144
 
210
145
  One habit feeds it: record a mem row when a unit of work ends. Everything else on this page is
211
- optional around that. `mem_add` needs no plan file and no skill — any agent that speaks MCP
212
- can call it, and calling it is what turns a pile of session logs into an answer the next
213
- session can find.
214
-
215
- ### One turn, end to end
216
-
217
- ```mermaid
218
- sequenceDiagram
219
- autonumber
220
- participant A as Any MCP client
221
- participant F as fapony MCP
222
- participant M as project mem log (.fapony/.memory)
223
-
224
- Note over A,F: end a turn with a commit and no new mem row → the Stop hook blocks it once
225
- A->>F: mem_add (kind + files + text)
226
- F->>M: one mem row in the project's log, stamped with the model that did it
227
- opt proof, not just a claim — CLI, for runs from the frozen ledger
228
- A->>F: fapony report <run-id>
229
- F-->>A: git facts + evidence from .fapony/evidence.json, stamped with server_sha
230
- end
231
- Note over A,M: `fapony stats` reads the frozen ledger back — CLI, because you ask it, not the agent
232
- ```
146
+ optional around that. The **Stop hook** is the only thing fapony *blocks* — once per turn, when a
147
+ commit lands with no new mem row. It never judges what deserves recording. The hints only
148
+ annotate: a big-file read points at `review-seed`, a repeat read of an unchanged file points at
149
+ grep, an edit names the file's importer count before you change its shape.
233
150
 
234
- The Stop hook is the only thing fapony *blocks* — once per turn, when a commit lands with
235
- no new mem row. It never judges what deserves recording; it cannot see whether the work
236
- held up. The hints only annotate and never
237
- block: the **Read** hook adds one factual line when a read is large enough to be cheaper as
238
- `review-seed`, or when the same file is read again in a session and its mtime has not moved
239
- (`FAPONY_NO_REREAD_HINT=1` turns the re-read line off); the **Edit** hook names a file's importer
240
- count, once per session, before you change its shape; OpenCode's **commit** hook nudges after a
241
- `git commit` that left no new mem row. Claude Code receives read/edit *before* the call, OpenCode
242
- *after* it — [What runs where](#what-runs-where) has the full client matrix.
243
-
244
- ### The 3 tools
245
-
246
- | Tool | Tier | Purpose |
247
- |------|------|---------|
248
- | `mem_find` | recall | Search the project's mem log read-only — decisions/bugs/notes matched on the row's `files[]` (text substring for rows written without it), `text`, `kind` (no default filter), `since`. "What was ever decided about this file?" in one call before editing |
249
- | `mem_add` | recall | Append a mem row (decision/bug/note/next/hold) with `files[]` required and rejected when empty — the write half of `mem_find`, so the row is findable when you next touch that file |
250
- | `mem_close` | recall | Close a mem row by id with a tombstone message — a separate tool (not `kind:"close"`) because a close row carries no `files[]`, so sharing `mem_add`'s schema would make required fields depend on another field's value |
151
+ | Tool | Purpose |
152
+ |------|---------|
153
+ | `mem_find` | Search the project's mem log read-only — matched on the row's `files[]` (text substring for older rows), `text`, `kind` (no default filter), `since` |
154
+ | `mem_add` | Append a mem row (decision/bug/note/next/hold) with `files[]` required and rejected when empty — the write half of `mem_find` |
155
+ | `mem_close` | Close a mem row by id with a tombstone message — a separate tool because a close row carries no `files[]` |
251
156
 
252
157
  **A tool earns its schema by being called mid-task without being asked.** Everything you invoke
253
- deliberately is a CLI command instead: the schema is paid as input tokens in every session of
158
+ deliberately is a CLI command instead: a tool schema is paid as input tokens in every session of
254
159
  every client whether or not it is used, while a CLI command costs nothing until it runs. That is
255
- why the handoff/report family is CLI-only, and why `fapony_stats`, `project_health_context`,
256
- `plan_list` and `fapony_usage` left the MCP surface in 2026-09 (`fapony stats` answers the first, `fapony mem
257
- kickoff` the third, `fapony usage-web` the fourth; the second had no caller).
258
- Cutting is not the goal — spending where it pays back is: the three mem tools keep
259
- their schemas because nobody is going to type them at the right moment. `fapony report <run-id>` prints the full report for a frozen-ledger run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
260
-
261
- Full protocol, adapter examples (bash, Python), and safety rules: [docs/mcp-handcheck.md](https://github.com/kire21b/fapony/blob/main/docs/mcp-handcheck.md).
262
-
263
- ### Verdict grades (frozen ledger)
264
-
265
- No new grades are recorded — the tool that filed them left the MCP surface in 2026-09.
266
- The old rows stay readable via `fapony stats` and `fapony report`, and this is the scale
267
- they were filed on. Verification produced a quality grade, not just pass/fail:
268
-
269
- | Grade | Meaning |
270
- |-------|---------|
271
- | `pass-excellent` | Ship-quality, no issues |
272
- | `pass-good` | Minor nits, safe to ship |
273
- | `pass-adequate` | Works, but could be better |
274
- | `pass` | Meets minimum bar |
275
- | `fail` | Needs fixes |
276
- | `uncertain` | Reviewer can't judge — plan may have a problem |
277
-
278
- ### Why measure from the outside
279
-
280
- - **Raw facts are hard to argue with.** Cost, rounds, diff sizes, pass rates — collected from git and session logs, not self-reported. A vendor can dispute a verdict as unfair; they can't dispute their own token count.
281
- - **Agent platforms grading their own homework is a conflict of interest.** fapony is a separate layer that measured any agent the same way, which is what made "model X vs. model Y" or "workflow A vs. workflow B" answerable with real data instead of vibes. That history is still queryable; new accumulation is mem rows, not grades.
282
- - **Verification stays honest about its limits.** The collector runs only commands listed in `.fapony/evidence.json`; commands proposed by the agent outside the allowlist are reported as *proposed — not executed*, never run. And because fapony doesn't control your agent's flow, old verdicts are labeled as one signal — not promised as truth.
160
+ why `stats`, `report`, `usage-web` and friends are CLI-only, and why four tools left the MCP
161
+ surface in 2026-09 — the 3 mem tools keep their schemas because nobody is going to type them at
162
+ the right moment.
283
163
 
284
164
  ## The work side — conveniences, not the contract
285
165
 
286
- Read-only, deterministic, and none of it writes to the ledger. These exist because they were
287
- useful in this project's own dogfooding; use them, use your client's own search, or use
288
- neither. **Nothing here is a precondition for anything in the section above.**
289
-
290
- ### A lookup instead of a file read
291
-
292
- ```mermaid
293
- sequenceDiagram
294
- autonumber
295
- participant A as You + your agent
296
- participant F as fapony CLI (read-only)
297
- participant W as your worktree
298
-
299
- A->>F: review-seed --files src/thing/
300
- F->>W: static scan — exports, importers, untested
301
- W-->>F: facts, no LLM in the middle
302
- F-->>A: the lines worth reading, instead of the whole files
303
- A->>W: build, then commit
304
- Note over A,F: nothing is stored — skip this side entirely and fapony still works
305
- ```
166
+ Read-only, deterministic, none of it writes anything. Skip this side entirely and fapony still
167
+ works. **Nothing here is a precondition for anything above.**
306
168
 
307
- `review-seed --files` takes file names or a directory and answers "what is in here, who
308
- imports it, what is untested" for roughly a thirtieth of the tokens reading those files costs.
309
- That is the whole trick; there is no model in the middle.
169
+ - `fapony review-seed --files src/thing/` — exports, importers, untested, for roughly a thirtieth
170
+ of the tokens reading those files costs. Before touching an unfamiliar file, fire this and Read
171
+ only the line ranges it points at. Directories work too.
172
+ - `fapony digest` — decisions, open bugs, in-flight plans, cost, on one page, from what's already
173
+ on disk.
310
174
 
311
175
  ### Skills
312
176
 
313
- fapony ships seven portable skills, each as `skill/<name>/SKILL.md` — the layout Claude
314
- Code expects, so a client can symlink the directory rather than copy the file:
177
+ fapony ships seven portable skills, each as `skill/<name>/SKILL.md` — the layout Claude Code
178
+ expects, so a client can symlink the directory rather than copy the file:
315
179
 
316
180
  | Skill | Purpose | Trigger |
317
181
  |-------|---------|---------|
@@ -323,98 +187,17 @@ Code expects, so a client can symlink the directory rather than copy the file:
323
187
  | `skill/git-commit-conventional/` | Commit split by concern + conventional message | `/git-commit` |
324
188
  | `skill/git-ship/` | Push branch, open PR with drafted title/body, merge, reset branch onto base | `/ship`, `/pr` |
325
189
 
326
- ### When to call what
327
-
328
- ```mermaid
329
- flowchart TD
330
- I([idea]) --> Q{does it outlive<br/>this session?}
331
- Q -->|"feature, several days"| P["/plan-with-pony<br/>PLAN.md + SPEC.md"]
332
- Q -->|"wire · refactor · fix"| Z["fapony analyze DIR<br/>fapony review-seed --files"]
333
- P --> W[you and your agent build]
334
- Z --> W
335
- W --> C["/git-commit"]
336
- C --> R["/review-pony"]
337
- R -->|findings| W
338
- R -->|clean| S["/git-ship"]
339
- S -->|"there was a PLAN.md"| D["/move-to-done"]
340
- D -.-> H[(mem log + frozen ledger)]
341
- R -.-> H
342
- C -.->|"Stop hook: a commit needs a mem row"| H
343
- H -.->|"pain zones + past model×shape"| Q
344
-
345
- style H fill:#2d333b,stroke:#768390,color:#adbac7
346
- ```
347
-
348
- **The fork at the top is load-bearing.** A plan file is an artifact for work the next session has
349
- to pick up. Wiring, refactors and UI passes finish in one sitting and the PLAN.md gets archived
350
- unread — so `/plan-with-pony` declines those itself and hands over the two seed commands instead.
351
- Both arms meet at the same review and the same ledger.
352
-
353
- **The dotted edges are the whole point.** Mem rows carry `files[]` and standalone text, so
354
- the zones that keep hurting are a query, not a hunch — and the frozen ledger still carries
355
- `regime` on its old rows, so *in this project, which model was worth paying for this shape
356
- of work* stays answerable from history. That is what flows back to the fork — not "this file
357
- broke once", which fapony measured at a 1–9% base rate and demoted.
358
-
359
- | Moment | Call | What fapony gets out of it |
360
- |---|---|---|
361
- | Starting anything | `/plan-with-pony` | decides plan-vs-seed, then reads back how this shape has gone |
362
- | Before editing an unfamiliar file | `fapony review-seed --files` | nothing; it saves you reading the file |
363
- | Before committing | `/git-commit` | nothing; it just keeps commits reviewable |
364
- | Before merging | `/review-pony` | writes a mem row when findings survive (bug/decision + files) |
365
- | Merging | `/git-ship` (`pr` / `land` on a team) | nothing; pure git plumbing |
366
- | After it ships | `/move-to-done` | writes a mem note when the ship taught something, closes the loop |
367
- | Proving a finished run | `fapony report <run-id>` (CLI, not MCP) | git facts + allowlisted evidence, one page |
368
-
369
- **Team flow.** `/git-ship pr` stops once the PR is open and hands you the URL; the reviewer does
370
- their pass; `/git-ship land` merges it after approval. If the default branch requires reviews,
371
- plain `/git-ship` detects that and behaves like `pr` on its own.
372
-
373
- **What this is not.** It doesn't reduce your token bill — an agent that plans against known
374
- failure patterns tends to spend fewer rounds getting there, but fapony measures that, it doesn't
375
- cause it. Use `fapony usage-web` to find out whether it actually happened for you rather than taking
376
- the claim on faith.
377
-
378
- `fapony install --platform claude` (or `opencode`) symlinks these directories into
379
- `~/.claude/skills` rather than copying them, so `fapony update` refreshes every client
380
- at once. ZCode and Codex get the same skills linked into `~/.agents/skills`. A destination
381
- that already exists and isn't a fapony link is reported and left alone — replace it by hand
382
- if you want fapony's version.
383
-
384
190
  `plan-with-pony` is vendor-neutral — the SKILL.md *is* the prompt, so pipe it to any agent:
385
-
386
- ```bash
387
- cat skill/plan-with-pony/SKILL.md | claude -p # Claude Code
388
- cat skill/plan-with-pony/SKILL.md | opencode run # OpenCode
389
- cat skill/plan-with-pony/SKILL.md | <your-agent> # anything that reads stdin
390
- ```
391
-
392
- Example plans produced by it live in [examples/](https://github.com/kire21b/fapony/tree/main/examples).
191
+ `cat skill/plan-with-pony/SKILL.md | claude -p` (or `opencode run`, or anything that reads stdin).
192
+ Example plans it produced: [examples/](https://github.com/kire21b/fapony/tree/main/examples).
393
193
 
394
194
  ### Plans your agent can answer questions about
395
195
 
396
- Plans stay markdown files in your repo — nothing moves into a database. Four optional
397
- frontmatter keys are enough to make a folder of them queryable:
398
-
399
- ```yaml
400
- ---
401
- kind: unit # tracker = a checklist that never finishes
402
- status: blocked # active | blocked | superseded
403
- blocked_by: PLAN-documents.md # a plan, or a sentence
404
- blocks: PLAN-export.md # ordering, stated once instead of buried in prose
405
- ---
406
-
407
- # PLAN — month view in /quick
408
-
409
- ## TL;DR # 15 lines; the only part that changes mid-flight
410
- - **Why:** two menu entries for the same data at different granularity
411
- - [x] chunk 1 — month grid `a1b2c3` 2026-09-13
412
- - [ ] chunk 2 — move overdue out
413
- ```
414
-
415
- Then open the next session with `fapony mem kickoff` — it reads the folder and the mem log and
416
- prints what is next (priority plans, the first unchecked chunk of each, open bugs) without
417
- reading a single 100KB plan body into context:
196
+ Plans stay markdown files in your repo — nothing moves into a database. Four optional frontmatter
197
+ keys (`kind` / `status` / `blocked_by` / `blocks`) make a folder of them queryable; plans with no
198
+ frontmatter still work, because the unchecked checkboxes are enough. Open the next session with
199
+ `fapony mem kickoff` — it prints what is next (priority plans, the first unchecked chunk of each,
200
+ open bugs) without reading a single 100KB plan body into context:
418
201
 
419
202
  ```
420
203
  ## next up
@@ -424,130 +207,101 @@ reading a single 100KB plan body into context:
424
207
  [3] last touched: src/quick/month.tsx, src/lib/money.ts
425
208
  ```
426
209
 
427
- **Plans with no frontmatter still work** — the unchecked checkboxes are enough, so an existing
428
- folder of plans is usable before anyone annotates anything. Two details that keep it honest over
429
- years:
430
-
431
- - The progress tally counts checkboxes in the **first `##` section only**, anchored by position
432
- rather than by the word "TL;DR" — so it works in any language, and a step list deeper in the
433
- file stays detail instead of becoming status.
434
- - **There is no `MASTER.md`.** Every line above is derived from the plan files themselves, so it
435
- cannot drift; a hand-kept master file always does.
436
-
437
- `status` / `blocked_by` / `blocks` / `superseded_by` are read by `fapony mem plan-check`
438
- (dangling refs, blocker shipped but dependent still blocked, waiter cycles, blocked with all
439
- chunks ticked) and by `fapony mem plan-sweep` (blocked view + a `🔓` unblock hint on `--apply`).
440
- Sentence values ("waiting on support email") carry no `PLAN-*.md` token and are never flagged.
441
-
442
- The layout, and why archiving is a plain `git mv`:
443
-
444
- ```
445
- .fapony/plan/PLAN-calendar.md live
446
- .fapony/done/PLAN-calendar.md shipped — same name, same depth, so every relative
447
- link inside the file survives the move untouched
448
- .fapony/spec/SPEC-calendar.md specs are a reference library; they are never archived
449
- ```
450
-
451
- Ship dates live in the plan's own header (`> ✅ **shipped 2026-09-13** (a1b2c3)`), not in the
452
- filename — `grep -h shipped .fapony/done/*.md | sort` answers "what landed when" without paying
453
- to rewrite every inbound link on every ship. Example plans, including an un-annotated one and an
454
- archived one: [examples/](https://github.com/kire21b/fapony/tree/main/examples).
210
+ **There is no `MASTER.md`** — every line above is derived from the plan files themselves, so it
211
+ cannot drift; a hand-kept master file always does. `fapony mem plan-check` verifies ticked chunk
212
+ shas against git history (a ticked box with no sha to check is a claim, not a close) and flags
213
+ dangling `blocked_by` refs; `fapony mem plan-sweep --apply` archives a shipped plan with `git mv`
214
+ into `.fapony/done/` — same name, same depth, so every relative link inside the file survives the
215
+ move. Specs live in `.fapony/spec/` and are never archived.
455
216
 
456
217
  ## CLI
457
218
 
458
219
  ```bash
459
- # Verification & reporting
460
- fapony mcp # MCP server (stdio JSON-RPC — 3 tools)
461
- fapony report <run-id> # verification report for a run
462
- fapony report-web [file] # static HTML report page
463
- fapony usage-scan # scan session logs → cache (incremental, progress bar)
464
- fapony price-scan # fetch model price table → prices.json (cache; query never fetches)
465
- fapony usage-web [port] # live usage comparison dashboard from cache
466
- fapony stats [--mode verdict [--regime code|fix|review|plan|inquiry|test]] # KPIs: pass/stall rate, by-model, by-grade — --mode verdict ranks by quality/tokens instead
467
- fapony digest [--since 7d|YYYY-MM-DD] [--format text|html] [--json] [--out FILE] # single-page summary: decisions, open bugs, in-flight plans, cost, pass/fail — from what's already on disk
468
- fapony plan-seed <name> [--spec] [--scope <path>]... # write PLAN (+SPEC): frontmatter, 8 empty sections, prior-art list, mem-decision context + existing-in-scope; existing plans listed on stdout; SPEC chunks carry signatures, every section capped — the agent fills the judgment
469
- fapony review-seed [--staged|--commit <sha>|--range <a...b>|--files f1,f2,dir|--plan <PLAN.md>] # read-only scope facts for a review (changed files, importers, untested, signatures, plan cross-check)
470
-
471
- # Memory & convention debt
472
- fapony mem add <kind> "<text>" --files f1,f2 [spec.md] # append a mem row (decision/bug/note/next/hold)
473
- fapony mem close <id> "<msg>" # close a bug
220
+ # core: memory + debt
221
+ fapony mem add <kind> "<text>" --files f1,f2 [--key k] [spec.md] # append a mem row (decision/bug/note/next/hold)
222
+ fapony mem close <id> "<msg>" # close a row (a bug stays open without this)
474
223
  fapony mem find ["<text>"] [--kind a,b] [--files f1,f2] [--since <N>d|YYYY-MM-DD] [--limit n] [--open] # search mem log
475
- fapony mem kickoff [<plan.md>] # open a session + a next-up list
224
+ fapony mem kickoff [<plan.md>] [--pick <n>] # open a session + a next-up list
476
225
  fapony mem where # show the resolved mem dir and which step won
477
- fapony mem done | stale # views
478
- fapony debt [--id <convention>] [--where <path>] # ไฟล์ไหนยังไม่ย้ายไป convention ที่ประกาศไว้ (live, read-only)
226
+ fapony mem done | stale | claim | release | synced | plan-sweep | plan-check | rotate
227
+ fapony debt [--id a,b] [--where <path>] # which files haven't migrated to a declared convention (live, read-only)
479
228
  fapony lint-baseline [--cmd ...] [--diff] # separate "already red" from "I made it red"
480
-
481
- # Hooks (wired by `fapony install`, not run by hand)
482
- fapony hook-stop # Stop hook: block turns with commits but no mem row
483
- fapony hook-read-hint # read/re-read annotations
484
- fapony hook-edit-hint # importer count before editing shape
485
- fapony hook-mv-guard # deny raw git mv of plan files into done/
486
- fapony hook-session-start # SessionStart: kickoff into context
229
+ fapony init-mem # delete legacy .memory/ dirs + warn call sites still referencing them
230
+ fapony digest [--since 7d|YYYY-MM-DD] [--format text|html] [--json] [--out FILE] # single-page summary from what's on disk
231
+
232
+ # usage (day one)
233
+ fapony usage-scan # scan session logs → cache (incremental, progress bar)
234
+ fapony price-scan # fetch model price table → prices.json (cache; query never fetches)
235
+ fapony usage-web [port] # usage comparison dashboard from cache
236
+
237
+ # lookup (read-only, never touches state)
238
+ fapony analyze [path] # live repo graph: hubs, orphans, cycles, changed-untested
239
+ fapony review-seed [--staged|--commit <sha>|--range <a...b>|--files f1,f2,dir|--plan <PLAN.md>] # scope facts for a review
240
+ fapony plan-seed <name> [--spec] [--scope <path>[,<path>]]... # write PLAN (+SPEC): frontmatter, capped sections, prior-art list
241
+
242
+ # hooks & MCP (wired by `fapony install`, not run by hand)
243
+ fapony mcp # MCP server (stdio JSON-RPC — 3 tools)
244
+ fapony hook-stop # Stop hook: block turns with commits but no mem row
245
+ fapony hook-read-hint # read/re-read annotations
246
+ fapony hook-edit-hint # importer count before editing shape
247
+ fapony hook-mv-guard # deny raw git mv of plan files into done/
248
+ fapony hook-session-start # SessionStart: kickoff into context
249
+
250
+ # frozen ledger (reads history only — the grading tool left the MCP surface in 2026-09)
251
+ fapony stats [--mode verdict [--regime code|fix|review|plan|inquiry|test]] # KPIs from old graded runs
252
+ fapony report <run-id> # verification report for a run
253
+ fapony report-web [file] # static HTML report page
254
+
255
+ # setup & maintenance
256
+ fapony init <path> # scaffold .fapony/ (plan/spec/memory/evidence)
257
+ fapony install [--all|--platform <name>|--dry-run] # wire MCP + skills into clients
258
+ fapony setup # interactive wizard: config + scaffold in one step
259
+ fapony update # self-update via git pull
260
+ fapony telemetry show|send # opt-in only, default off — see TELEMETRY.md
487
261
  ```
488
262
 
263
+ `fapony report <run-id>` prints the full report for a frozen-ledger run — git facts, handoff
264
+ conformance, allowlisted evidence, the stored verdict, cost — with anything the agent claimed but
265
+ couldn't prove marked as such. Reports are stamped with the producing build's `server_sha`; after
266
+ editing fapony, compare the stamp against `git log -1` before trusting a report from a
267
+ long-lived MCP server. Each allowlisted command gets `timeout_ms` (default 30s), the whole report
268
+ capped at 180s — a command that doesn't fit reports as `timeout`, never as a pass. If your
269
+ `.gitignore` ignores `.fapony/` wholesale, re-include the file: `**/.fapony/*`, then
270
+ `!**/.fapony/evidence.json`.
271
+
489
272
  *When* to call `mem add` is your project's call, not fapony's — write it in your own
490
- `AGENTS.md`/`CLAUDE.md`, not here. A starting point:
273
+ `AGENTS.md`/`CLAUDE.md`. A starting point:
491
274
 
492
275
  ```markdown
493
276
  ## Memory
494
277
  - Found a bug while working (not just user-reported)? Log it before fixing:
495
- `mem_add { kind: "bug", worktree: "<absolute app dir>", files: [...], text: "..." }`
278
+ mem_add { kind: "bug", worktree: "<absolute app dir>", files: [...], text: "..." }
496
279
  - `text` must stand alone — read months later with no chat context: what/where/repro/status.
497
- - Report the row id back in chat.
498
280
  - Don't fold the fix into the same chunk — log first, fix as its own next/chunk if you do.
499
281
  ```
500
282
 
501
- ```bash
502
- # Setup & maintenance
503
- fapony init <path> # scaffold .fapony/ (plan/spec/memory/evidence)
504
- fapony init-mem # delete .memory/ + warn call sites still referencing it
505
- fapony install # detect installed clients, prompt to wire each
506
- fapony install --all # wire all detected clients without prompting
507
- fapony install --platform <name> # force a specific client (bypasses detection)
508
- fapony install --dry-run # show what would happen without writing files
509
- fapony setup # interactive wizard: config + scaffold in one step
510
- fapony update # self-update via git pull
511
- fapony telemetry show|send # opt-in only, default off — see https://github.com/kire21b/fapony/blob/main/TELEMETRY.md
512
- ```
513
-
514
283
  ## Config
515
284
 
516
- `fapony.config.json` lives in the fapony checkout and is gitignored (it's per-machine). Copy [fapony.config.example.json](https://github.com/kire21b/fapony/blob/main/fapony.config.example.json) for a complete working reference; every section is optional with sane defaults. Key fields:
517
-
518
- - `worktrees` — name → absolute path mapping
519
- - `review.maxRounds` — round cap enforced by the gate
520
- - `memory` — shell commands for claim/close/add/kickoff, or `null` to default-wire when a `.fapony/.memory/` dir exists
521
- - `paths` (`planDir`/`doneDir`/`specDir`/`memDir`/`stateDir`) / `safety` — directory layout and the dangerous-command deny-list
522
- - `usageWeb` — optional `{ port, hostname }` for `fapony usage-web` server defaults. Run `fapony usage-scan` first to populate the cache.
523
-
524
- Env overrides: `FAPONY_CONFIG` (config file), `FAPONY_STATE_DIR` (state DB location; default `~/.config/fapony/`), `FAPONY_NO_REREAD_HINT=1` (turn the re-read hint off). Full schema, design decisions, and edge cases live with the code in the repo — this README intentionally doesn't duplicate them.
285
+ `fapony.config.json` lives in the fapony checkout and is gitignored (it's per-machine). Copy
286
+ [fapony.config.example.json](https://github.com/kire21b/fapony/blob/main/fapony.config.example.json)
287
+ for a complete working reference; every section is optional. Key fields: `worktrees`
288
+ (name → path), `memory` (shell commands, or `null` to disable), `paths` / `safety`,
289
+ `usageWeb { port, hostname }`. Env overrides: `FAPONY_CONFIG`, `FAPONY_STATE_DIR` (state DB;
290
+ default `~/.config/fapony/`), `FAPONY_NO_REREAD_HINT=1`.
525
291
 
526
292
  ## Scope
527
293
 
528
- **Supported:**
529
- - MCP server — 3 mem tools via stdio JSON-RPC, works with any MCP client
530
- - Measurement: cross-run KPIs by model/grade/value from the frozen ledger, per-file pain zones from mem rows (`files[]`) + passive usage (tokens, cost)
531
- - Model attribution across clients — resolved from the session log that was live when the old verdict landed, so a frozen row carries a model without the caller having declared one
532
- - Zero setup beyond install: the mem habit ships in the MCP `initialize` response, not in your rules file
533
- - Verification reports (frozen): handoff conformance, 6-grade verdicts, allowlisted evidence collector (`.fapony/evidence.json`); reports stamped with the producing build's `server_sha` — replayable, no new graded runs
534
- - Vendor-neutral executor/reviewer roles — anything that reads stdin
535
- - Memory integration via shell adapter, per project (configurable or default-wired)
536
- - Opt-in telemetry, off by default ([TELEMETRY.md](https://github.com/kire21b/fapony/blob/main/TELEMETRY.md) lists exactly what leaves the machine)
537
- - Bun-only; run state in SQLite via `bun:sqlite` (WAL mode)
538
- - Per-client hooks alongside MCP: Stop hook on Claude Code + Cursor · read/re-read/Edit hints on
539
- Claude Code + OpenCode · commit hint on OpenCode — [What runs where](#what-runs-where)
540
-
541
- **Not supported (yet):**
542
- - PreToolUse hints on Cursor, ZCode or Codex — Cursor has no such hook and the other two expose no
543
- in-process hook surface for read/edit hints (Codex's `apply_patch` sends patch text, not file paths)
544
- - A hosted or shared ledger for a team — `runs.worktree` is the only sharing key today, and it's a
545
- path, not an identity. If you want to try pointing two machines at the same ledger anyway,
546
- `FAPONY_STATE_DIR` can be set to a synced folder (Syncthing, a shared drive) — but SQLite's WAL
547
- mode does not tolerate concurrent writers over most network filesystems (NFS, Dropbox, iCloud
548
- Drive) and can corrupt the db under real contention. Treat this as an experiment you're accepting
549
- the risk on, not a supported path; nothing here is a substitute for a real shared-ledger server.
550
- - Memory migration from `.fapony/.memory/log.jsonl`
294
+ **Supported:** MCP server (3 mem tools, any MCP client) · cross-client usage on one yardstick ·
295
+ per-project mem log + convention debt · per-client hooks ([matrix above](#what-runs-where)) ·
296
+ vendor-neutral skills (anything that reads stdin) · opt-in telemetry, off by default
297
+ ([TELEMETRY.md](https://github.com/kire21b/fapony/blob/main/TELEMETRY.md) lists exactly what
298
+ leaves the machine) · Bun-only; run state in SQLite via `bun:sqlite` (WAL mode).
299
+
300
+ **Not supported (yet):** PreToolUse hints on Cursor, ZCode or Codex — Cursor has no such hook and
301
+ the other two expose no in-process hook surface for read/edit hints. A hosted or shared ledger —
302
+ `FAPONY_STATE_DIR` on a synced folder works as an experiment only; SQLite's WAL mode does not
303
+ tolerate concurrent writers over NFS/Dropbox/iCloud Drive and can corrupt the db under real
304
+ contention.
551
305
 
552
306
  ## License
553
307