fapony 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,26 +6,26 @@
6
6
 
7
7
  [![npm](https://img.shields.io/npm/v/fapony.svg)](https://www.npmjs.com/package/fapony)
8
8
 
9
- **Where did your tokens go?** fapony reads the session logs Claude Code, Codex, OpenCode and ZCode
10
- already write, and puts them all on one yardstick — tokens, cost and time per model, per client,
11
- per workflow. Nothing to instrument, no per-project setup, no waiting for data to accumulate: it
12
- runs on the history already sitting on your disk.
9
+ **See what your coding agents actually cost.** fapony reads the session logs Claude Code, Codex,
10
+ OpenCode and ZCode already write, and puts them all on one yardstick — tokens, cost and time per
11
+ model, per client, per workflow. Nothing to instrument, no per-project setup: it runs on the
12
+ history already sitting on your disk.
13
13
 
14
14
  <p align="center">
15
15
  <img src="images/summary.webp" width="800" alt="fapony usage-web summary cards">
16
16
  </p>
17
17
 
18
18
  ```bash
19
- fapony usage-scan # read the session logs already on your disk
20
- fapony price-scan # fetch the price table (needed once, for cost)
21
- fapony usage-web # every session you already have, all clients, one page
19
+ npm install -g fapony # needs Bun — https://bun.sh
20
+ fapony usage-scan # read the session logs already on your disk
21
+ fapony price-scan # fetch the price table (needed once, for cost)
22
+ fapony usage-web # every session you already have, all clients, one page
22
23
  ```
23
24
 
24
25
  **Cost is the part your client probably isn't logging.** Of the four, only OpenCode writes a real
25
- dollar figure into its session log — Claude Code, ZCode and Codex record `0`. fapony prices those
26
- sessions at published list rates and labels the number `imputed`, so a figure you can compare
27
- across clients exists at all. A model it can't find a rate for stays `unpriced`: nothing is
28
- quietly counted as free.
26
+ dollar figure into its session log — the others record `0`. fapony prices those sessions at
27
+ published list rates and labels the number `imputed`; a model it can't find a rate for stays
28
+ `unpriced` — nothing is quietly counted as free.
29
29
 
30
30
  <details>
31
31
  <summary>full usage-web dashboard preview</summary>
@@ -36,134 +36,88 @@ quietly counted as free.
36
36
 
37
37
  </details>
38
38
 
39
- That is day one. Past that, fapony keeps what coding agents actually did — the frozen
40
- ledger of graded runs (rounds, pass/fail, cost per grade, readable via CLI, no new grades)
41
- plus the live mem log — through 3 MCP tools any agent can call. It checks claims against
42
- git facts: handoff conformance and allowlisted evidence — with everything the agent claimed
43
- but couldn't prove marked as such.
39
+ Raw facts from logs are hard to argue with — a vendor can dispute a verdict as unfair; they can't
40
+ dispute their own token count. That is the whole measurement layer: tokens and cost, nothing
41
+ self-graded.
44
42
 
45
- **The reason to keep it running is the third layer: knowledge accumulation — and the thing it
46
- accumulates is pain.** An agent has no memory of pain across sessions: it writes the 37th
47
- hand-rolled `try/catch` as cheerfully as the first, because every session starts new. Wrappers and
48
- shared libraries get built by *people* who were hurt by the same thing often enough to remember.
49
- That is why a codebase written with agents from day one tends not to grow a shared layer — nobody
50
- in the room remembers. fapony is the part that remembers: mem rows carry
51
- `files[]`, so the zones that keep coming back in re-done work are a query, not a hunch
52
- (the frozen ledger's old graded rows carry them too).
53
- Paired with `fapony debt`, which tracks how far the codebase has actually moved to a convention you
54
- already decided on, that is the loop: notice the repeated cost, name the shared thing, watch the
55
- migration finish. Finding dead code and duplication is *not* part of it — knip and friends already
56
- do that better, and a convention with a `checker` is deliberately left to the checker.
43
+ ## Past day one
57
44
 
58
- **The measurement layer underneath it:** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: what each model costs you, across clients, on one yardstick, in this project (`fapony usage-web`, above) — tokens and cost, nothing self-graded. The frozen ledger still has old verdict rows (`fapony stats`), but the tool that wrote them (`verdict_submit`) is gone from the MCP surface: no new rows, for anyone, ever again, so a count of old rows (`n`) means nothing without the grading that stopped. New accumulation goes to the mem log instead: decisions, bugs and notes with `files[]`, written by the agents doing the work.
45
+ Two more layers, both optional, both compounding:
59
46
 
60
- Three tiers, deliberately: **measurement ships today** and needs no per-project setup — raw facts nobody can call unfair. **Verification is the sharper edge** but stays beta until its evidence layer is hardened; fapony doesn't control your agent's flow, so it never promises "verified" as a headline. **Knowledge accumulation is the compounding one** — it's worthless on run 1 and gets more useful every run after, which is exactly why it's the layer competitors can't clone by copying a feature list.
47
+ **Memory — the mem log.** An agent has no memory of pain across sessions: it writes the 37th
48
+ hand-rolled `try/catch` as cheerfully as the first, because every session starts new. Wrappers
49
+ and shared libraries get built by *people* who were hurt often enough to remember. fapony
50
+ remembers instead: one MCP call per unit of work (`mem_add`) records the decision, bug or note
51
+ with the files it touched, and `mem_find` answers *"what was ever decided about this file?"*
52
+ before the next agent touches it. The log lives in your repo (`.fapony/.memory/`), so it crosses
53
+ machines over git for free.
61
54
 
62
- Adopting it doesn't change your workflow. There is no loop to join and no framework to learn: install the MCP server, point your agent at it, and read the reports. fapony also ships plans, skills and read-only seed commands from its own dogfooding — those are conveniences, kept in their own section below, and deleting all of them costs you nothing the ledger can measure.
55
+ **Convention debt.** `fapony debt` answers the question nothing else does: *we decided this six
56
+ months ago — how far along is the move?* ESLint says this line is wrong; nothing says 11 of 47
57
+ files have migrated. Dead code and duplication it deliberately leaves to knip and friends —
58
+ they already do that better.
63
59
 
64
- ## What fapony is not
60
+ ```mermaid
61
+ flowchart LR
62
+ A[Claude Code] --> F[fapony]
63
+ B[OpenCode] --> F
64
+ C[ZCode] --> F
65
+ D[Codex] --> F
66
+ F --> U[usage — tokens & cost]
67
+ F --> M[mem log — what was decided here]
68
+ F --> D[debt — how far the move has gone]
69
+ ```
65
70
 
66
- Stated up front, because the gap between these two things is where most tooling oversells:
71
+ Adopting it doesn't change your workflow: install it, point your agent at it, read the reports.
72
+ Both layers above are per-project (`fapony init`) and worthless on run 1 — they get more useful
73
+ every run after, which is exactly why they're retention, not the reason to install.
67
74
 
68
- - **It does not run your test suite.** The evidence collector runs an allowlist *you* write in
69
- `.fapony/evidence.json`, and never a command an agent proposes. No allowlist, no evidence.
70
- - **It does not judge your code.** Mem rows *record* decisions, bugs and notes; a human or
71
- a working agent supplies them. fapony is the memory, not the judge. (The frozen
72
- ledger's old grades work the same way — *stored*, never computed.)
73
- - **It checks conformance, not correctness.** What it can verify is that a claim lines up with git
74
- facts and that uncertainty was declared — not that the code works. Those are different
75
- guarantees and fapony only offers the first.
76
- - **Almost nothing blocks.** No CI failure, no gate on your own commands. The one exception is the
77
- Stop hook, once per turn when a commit lands with no new mem row; the read/edit/commit hints only annotate.
78
- Skip the install of all of them and you are back to exactly the workflow you had.
79
- - **Model attribution is inferred, not declared.** A gate in the frozen ledger is attributed to whichever client
80
- session was live in that worktree at that moment. When one model writes the code and another
81
- reviews and files the verdict, the grade lands on the reviewer. Reports label it `inferred`;
82
- read it as such.
83
- - **The knowledge layer is empty on run 1.** It is worth something around run 5 and more every run
84
- after.
85
-
86
- ## Quick start (MCP)
75
+ ## Quick start
87
76
 
88
77
  ```bash
89
78
  # 1. Install (needs Bun — https://bun.sh)
90
- npm install -g fapony # or: bun add -g fapony
79
+ npm install -g fapony
91
80
  # from source instead:
92
81
  # git clone https://github.com/kire21b/fapony.git && cd fapony && bun install && bun link
93
- # note: `bun link` claims the global `fapony` bin by package name, not path — running it
94
- # from a second checkout silently repoints the command there. Re-run it in the one you want.
95
-
96
- # 2. Wire it into your MCP client
97
- fapony install # detects installed clients, asks which to wire
98
- fapony install --all # skip the prompt, wire everything detected
99
- # use --platform <name> to force a specific client (bypasses detection)
100
- # zcode/codex need their config to exist first — open the app once if you never have
101
- # claude/opencode also symlink skill/<name>/ into ~/.claude/skills — an existing
102
- # skill of the same name is reported, never overwritten
103
- # …or add it manually to any MCP client (e.g. Claude Desktop):
104
- # { "mcpServers": { "fapony": { "command": "fapony", "args": ["mcp"] } } }
82
+ # (`bun link` claims the global `fapony` bin by package name, not path — re-run it in the
83
+ # checkout you want to be the one)
105
84
 
106
- # 3. Measure — zero per-project setup
107
- fapony usage-scan # scan the session logs already on disk → cache
108
- fapony price-scan # fetch the OpenRouter price table → ~/.config/fapony/prices.json
109
- fapony usage-web # dashboard; re-run the scans to refresh
85
+ # 2. Measure — zero per-project setup
86
+ fapony usage-scan # scan the session logs already on disk → cache
87
+ fapony price-scan # fetch the OpenRouter price table → ~/.config/fapony/prices.json
88
+ fapony usage-web # dashboard; re-run the scans to refresh
110
89
  # both scans are manual by design — nothing fetches or re-reads session logs behind your back
111
90
 
112
- # 4. Turn on the knowledge layer (per project you want it in)
113
- fapony init /path/to/your-worktree
114
- # creates .fapony/ — .memory/ (the mem log the 3 MCP tools read and write),
115
- # conventions.json for `fapony debt`, plan/spec/done, and evidence.json
116
- # conventions.json + evidence.json are shared rules: commit them
117
- ```
118
-
119
- With `.fapony/evidence.json` in place, any graded run from the frozen ledger can be replayed
120
- as a report. This one is a CLI command, not an MCP tool — the schemas cost every session of
121
- every client and no skill called them (see [The 3 tools](#the-3-tools) below). No new runs
122
- can be created; run ids come from `fapony stats` reading history:
91
+ # 3. Wire your clients
92
+ fapony install # detects installed clients, asks which to wire
93
+ fapony install --all # skip the prompt, wire everything detected
94
+ # claude/opencode also symlink skill/<name>/ into ~/.claude/skills — an existing
95
+ # skill of the same name is reported, never overwritten
123
96
 
124
- ```bash
125
- fapony report <run-id> # run ids come from `fapony stats`
97
+ # 4. Turn on the memory layer (per project you want it in)
98
+ fapony init /path/to/your-worktree
99
+ # creates .fapony/ — .memory/ (the mem log the 3 MCP tools read and write)
100
+ # and conventions.json for `fapony debt` (shared rules: commit them)
126
101
  ```
127
102
 
128
- You get one report: git facts (files, commits, branch), handoff conformance (claims vs. reality), evidence from the allowlisted commands (pass/fail/timeout/unverified), the frozen 6-grade verdict, and cost — with anything the agent claimed but couldn't prove marked as such.
129
-
130
- Sections that have nothing to report say so (`not_run`, `unavailable`) rather than disappearing — a report with no evidence must not read like a report that passed.
131
-
132
- Two details for the allowlist once it is under version control. If your `.gitignore` ignores `.fapony/` wholesale, re-include the file (dir before file): `**/.fapony/*`, then `!**/.fapony/evidence.json`. In a monorepo, give an app its own `apps/<app>/.fapony/evidence.json` — reports whose changed files all sit under that app use it; anything else uses the root one.
103
+ …or add it manually to any MCP client: `{ "mcpServers": { "fapony": { "command": "fapony", "args": ["mcp"] } } }`. Full protocol and adapter examples: [docs/mcp-handcheck.md](https://github.com/kire21b/fapony/blob/main/docs/mcp-handcheck.md).
133
104
 
134
- Two things worth knowing about the report header and budget:
135
-
136
- - **`server_sha`** — every report is stamped with the git SHA of the fapony code that produced it, read once at server start. MCP servers are long-lived: after you edit fapony and don't restart the client, reports keep coming from the old build. Compare the stamp against `git log -1` in the fapony repo; if they differ, reconnect the server before trusting the result.
137
- - **Evidence budget** — each allowlisted command gets `timeout_ms` (default 30s), and the whole report is capped at 180s total. A command that doesn't fit is reported as `timeout`, never as a pass. Time your real suite and set `timeout_ms` accordingly.
138
-
139
- ## How it fits
140
-
141
- ```mermaid
142
- flowchart LR
143
- A[Claude Code] --> F[fapony MCP]
144
- B[OpenCode] --> F
145
- C[ZCode] --> F
146
- D[Codex] --> F
147
- E[Cursor] --> F
148
- F --> G[git facts + session logs]
149
- G --> S[stats / usage]
150
- G --> V[verification report]
151
- G --> M[mem log - what was decided here]
152
- ```
153
-
154
- fapony never drives the agent. It sits on two sides of your work that never touch each
155
- other, and **the rest of this README is organised along that line**: the ledger below is the
156
- product, and everything under "The work side" after it is a convenience you can delete without
157
- losing a single number.
105
+ ## What fapony is not
158
106
 
159
- | | The ledger | The work side |
160
- |---|---|---|
161
- | What it is | 3 MCP tools (mem) + a frozen SQLite ledger (reads history) | plans, skills, read-only seed commands |
162
- | Needs | an MCP client | nothing — or your own tooling instead |
163
- | Writes | one mem row per unit of work, into the project's log | nothing |
164
- | Skip it and | there is no fapony | fapony still answers every question |
107
+ Stated up front, because the gap between these two things is where most tooling oversells:
165
108
 
166
- ### What runs where
109
+ - **It does not run your test suite.** The evidence collector runs an allowlist *you* write in
110
+ `.fapony/evidence.json`, never a command an agent proposes. No allowlist, no evidence.
111
+ - **It does not judge your code.** Mem rows *record* what a human or a working agent supplies.
112
+ fapony is the memory, not the judge.
113
+ - **It checks conformance, not correctness** — that a claim lines up with git facts and that
114
+ uncertainty was declared, not that the code works.
115
+ - **Almost nothing blocks.** The one exception is the Stop hook, once per turn when a commit
116
+ lands with no new mem row; everything else only annotates.
117
+ - **Model attribution is inferred, not declared** — reports label it `inferred`; read it as such.
118
+ - **The knowledge layer is empty on run 1** — worth something around run 5, more every run after.
119
+
120
+ ## What runs where
167
121
 
168
122
  `fapony install` wires five clients (Claude Code, OpenCode, Cursor, ZCode, Codex). MCP is the only
169
123
  piece all of them get — the hooks and in-process hints are per-client, and the read/edit hints
@@ -182,119 +136,46 @@ is `tool.execute.after`. Nothing here is required: skip the hooks and every MCP
182
136
  | Skills symlinked into `~/.agents/skills` | — | — | — | ✅ | ✅ |
183
137
  | `usage-scan` reads this client's session log | ✅ | ✅ | — | ✅ | ✅ |
184
138
 
185
- `—` means not wired, not impossible: Cursor has no PreToolUse hook, and ZCode/Codex expose no
186
- in-process hook surface for read/edit hints yet (Codex's `apply_patch` sends patch text, not
187
- resolved file paths). Codex hooks require trust via `/hooks` before they run — `fapony install`
188
- tells you when. The hints live on hooks rather than MCP on purpose — they must fire mid-turn
189
- without the agent deciding to call anything ([why](#when-to-call-what)).
139
+ `—` means not wired, not impossible. Codex hooks require trust via `/hooks` before they run —
140
+ `fapony install` tells you when. The hints live on hooks rather than MCP on purpose — they must
141
+ fire mid-turn without the agent deciding to call anything.
190
142
 
191
- ## The ledger — this is the product
143
+ ## The ledger — one habit, 3 tools
192
144
 
193
145
  One habit feeds it: record a mem row when a unit of work ends. Everything else on this page is
194
- optional around that. `mem_add` needs no plan file and no skill — any agent that speaks MCP
195
- can call it, and calling it is what turns a pile of session logs into an answer the next
196
- session can find.
197
-
198
- ### One turn, end to end
199
-
200
- ```mermaid
201
- sequenceDiagram
202
- autonumber
203
- participant A as Any MCP client
204
- participant F as fapony MCP
205
- participant M as project mem log (.fapony/.memory)
206
-
207
- Note over A,F: end a turn with a commit and no new mem row → the Stop hook blocks it once
208
- A->>F: mem_add (kind + files + text)
209
- F->>M: one mem row in the project's log, stamped with the model that did it
210
- opt proof, not just a claim — CLI, for runs from the frozen ledger
211
- A->>F: fapony report <run-id>
212
- F-->>A: git facts + evidence from .fapony/evidence.json, stamped with server_sha
213
- end
214
- Note over A,M: `fapony stats` reads the frozen ledger back — CLI, because you ask it, not the agent
215
- ```
216
-
217
- The Stop hook is the only thing fapony *blocks* — once per turn, when a commit lands with
218
- no new mem row. It never judges what deserves recording; it cannot see whether the work
219
- held up. The hints only annotate and never
220
- block: the **Read** hook adds one factual line when a read is large enough to be cheaper as
221
- `review-seed`, or when the same file is read again in a session and its mtime has not moved
222
- (`FAPONY_NO_REREAD_HINT=1` turns the re-read line off); the **Edit** hook names a file's importer
223
- count, once per session, before you change its shape; OpenCode's **commit** hook nudges after a
224
- `git commit` that left no new mem row. Claude Code receives read/edit *before* the call, OpenCode
225
- *after* it — [What runs where](#what-runs-where) has the full client matrix.
226
-
227
- ### The 3 tools
146
+ optional around that. The **Stop hook** is the only thing fapony *blocks* — once per turn, when a
147
+ commit lands with no new mem row. It never judges what deserves recording. The hints only
148
+ annotate: a big-file read points at `review-seed`, a repeat read of an unchanged file points at
149
+ grep, an edit names the file's importer count before you change its shape.
228
150
 
229
- | Tool | Tier | Purpose |
230
- |------|------|---------|
231
- | `mem_find` | recall | Search the project's mem log read-only — decisions/bugs/notes matched on the row's `files[]` (text substring for rows written without it), `text`, `kind` (no default filter), `since`. "What was ever decided about this file?" in one call before editing |
232
- | `mem_add` | recall | Append a mem row (decision/bug/note/next/hold) with `files[]` required and rejected when empty — the write half of `mem_find`, so the row is findable when you next touch that file |
233
- | `mem_close` | recall | Close a mem row by id with a tombstone message — a separate tool (not `kind:"close"`) because a close row carries no `files[]`, so sharing `mem_add`'s schema would make required fields depend on another field's value |
151
+ | Tool | Purpose |
152
+ |------|---------|
153
+ | `mem_find` | Search the project's mem log read-only — matched on the row's `files[]` (text substring for older rows), `text`, `kind` (no default filter), `since` |
154
+ | `mem_add` | Append a mem row (decision/bug/note/next/hold) with `files[]` required and rejected when empty — the write half of `mem_find` |
155
+ | `mem_close` | Close a mem row by id with a tombstone message — a separate tool because a close row carries no `files[]` |
234
156
 
235
157
  **A tool earns its schema by being called mid-task without being asked.** Everything you invoke
236
- deliberately is a CLI command instead: the schema is paid as input tokens in every session of
158
+ deliberately is a CLI command instead: a tool schema is paid as input tokens in every session of
237
159
  every client whether or not it is used, while a CLI command costs nothing until it runs. That is
238
- why the handoff/report family is CLI-only, and why `fapony_stats`, `project_health_context`,
239
- `plan_list` and `fapony_usage` left the MCP surface in 2026-09 (`fapony stats` answers the first, `fapony mem
240
- kickoff` the third, `fapony usage-web` the fourth; the second had no caller).
241
- Cutting is not the goal — spending where it pays back is: the three mem tools keep
242
- their schemas because nobody is going to type them at the right moment. `fapony report <run-id>` prints the full report for a frozen-ledger run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
243
-
244
- Full protocol, adapter examples (bash, Python), and safety rules: [docs/mcp-handcheck.md](https://github.com/kire21b/fapony/blob/main/docs/mcp-handcheck.md).
245
-
246
- ### Verdict grades (frozen ledger)
247
-
248
- No new grades are recorded — the tool that filed them left the MCP surface in 2026-09.
249
- The old rows stay readable via `fapony stats` and `fapony report`, and this is the scale
250
- they were filed on. Verification produced a quality grade, not just pass/fail:
251
-
252
- | Grade | Meaning |
253
- |-------|---------|
254
- | `pass-excellent` | Ship-quality, no issues |
255
- | `pass-good` | Minor nits, safe to ship |
256
- | `pass-adequate` | Works, but could be better |
257
- | `pass` | Meets minimum bar |
258
- | `fail` | Needs fixes |
259
- | `uncertain` | Reviewer can't judge — plan may have a problem |
260
-
261
- ### Why measure from the outside
262
-
263
- - **Raw facts are hard to argue with.** Cost, rounds, diff sizes, pass rates — collected from git and session logs, not self-reported. A vendor can dispute a verdict as unfair; they can't dispute their own token count.
264
- - **Agent platforms grading their own homework is a conflict of interest.** fapony is a separate layer that measured any agent the same way, which is what made "model X vs. model Y" or "workflow A vs. workflow B" answerable with real data instead of vibes. That history is still queryable; new accumulation is mem rows, not grades.
265
- - **Verification stays honest about its limits.** The collector runs only commands listed in `.fapony/evidence.json`; commands proposed by the agent outside the allowlist are reported as *proposed — not executed*, never run. And because fapony doesn't control your agent's flow, old verdicts are labeled as one signal — not promised as truth.
160
+ why `stats`, `report`, `usage-web` and friends are CLI-only, and why four tools left the MCP
161
+ surface in 2026-09 — the 3 mem tools keep their schemas because nobody is going to type them at
162
+ the right moment.
266
163
 
267
164
  ## The work side — conveniences, not the contract
268
165
 
269
- Read-only, deterministic, and none of it writes to the ledger. These exist because they were
270
- useful in this project's own dogfooding; use them, use your client's own search, or use
271
- neither. **Nothing here is a precondition for anything in the section above.**
272
-
273
- ### A lookup instead of a file read
274
-
275
- ```mermaid
276
- sequenceDiagram
277
- autonumber
278
- participant A as You + your agent
279
- participant F as fapony CLI (read-only)
280
- participant W as your worktree
281
-
282
- A->>F: review-seed --files src/thing/
283
- F->>W: static scan — exports, importers, untested
284
- W-->>F: facts, no LLM in the middle
285
- F-->>A: the lines worth reading, instead of the whole files
286
- A->>W: build, then commit
287
- Note over A,F: nothing is stored — skip this side entirely and fapony still works
288
- ```
166
+ Read-only, deterministic, none of it writes anything. Skip this side entirely and fapony still
167
+ works. **Nothing here is a precondition for anything above.**
289
168
 
290
- `review-seed --files` takes file names or a directory and answers "what is in here, who
291
- imports it, what is untested" for roughly a thirtieth of the tokens reading those files costs.
292
- That is the whole trick; there is no model in the middle.
169
+ - `fapony review-seed --files src/thing/` — exports, importers, untested, for roughly a thirtieth
170
+ of the tokens reading those files costs. Before touching an unfamiliar file, fire this and Read
171
+ only the line ranges it points at. Directories work too.
172
+ - `fapony digest` — decisions, open bugs, in-flight plans, cost, on one page, from what's already
173
+ on disk.
293
174
 
294
175
  ### Skills
295
176
 
296
- fapony ships seven portable skills, each as `skill/<name>/SKILL.md` — the layout Claude
297
- Code expects, so a client can symlink the directory rather than copy the file:
177
+ fapony ships seven portable skills, each as `skill/<name>/SKILL.md` — the layout Claude Code
178
+ expects, so a client can symlink the directory rather than copy the file:
298
179
 
299
180
  | Skill | Purpose | Trigger |
300
181
  |-------|---------|---------|
@@ -306,104 +187,17 @@ Code expects, so a client can symlink the directory rather than copy the file:
306
187
  | `skill/git-commit-conventional/` | Commit split by concern + conventional message | `/git-commit` |
307
188
  | `skill/git-ship/` | Push branch, open PR with drafted title/body, merge, reset branch onto base | `/ship`, `/pr` |
308
189
 
309
- ### When to call what
310
-
311
- ```mermaid
312
- flowchart TD
313
- I([idea]) --> Q{does it outlive<br/>this session?}
314
- Q -->|"feature, several days"| P["/plan-with-pony<br/>PLAN.md + SPEC.md"]
315
- Q -->|"wire · refactor · fix"| Z["fapony analyze DIR<br/>fapony review-seed --files"]
316
- P --> W[you and your agent build]
317
- Z --> W
318
- W --> C["/git-commit"]
319
- C --> R["/review-pony"]
320
- R -->|findings| W
321
- R -->|clean| S["/git-ship"]
322
- S -->|"there was a PLAN.md"| D["/move-to-done"]
323
- D -.-> H[(mem log + frozen ledger)]
324
- R -.-> H
325
- C -.->|"Stop hook: a commit needs a mem row"| H
326
- H -.->|"pain zones + past model×shape"| Q
327
-
328
- style H fill:#2d333b,stroke:#768390,color:#adbac7
329
- ```
330
-
331
- **The fork at the top is load-bearing.** A plan file is an artifact for work the next session has
332
- to pick up. Wiring, refactors and UI passes finish in one sitting and the PLAN.md gets archived
333
- unread — so `/plan-with-pony` declines those itself and hands over the two seed commands instead.
334
- Both arms meet at the same review and the same ledger.
335
-
336
- **The dotted edges are the whole point.** Mem rows carry `files[]` and standalone text, so
337
- the zones that keep hurting are a query, not a hunch — and the frozen ledger still carries
338
- `regime` on its old rows, so *in this project, which model was worth paying for this shape
339
- of work* stays answerable from history. That is what flows back to the fork — not "this file
340
- broke once", which fapony measured at a 1–9% base rate and demoted.
341
-
342
- | Moment | Call | What fapony gets out of it |
343
- |---|---|---|
344
- | Starting anything | `/plan-with-pony` | decides plan-vs-seed, then reads back how this shape has gone |
345
- | Before editing an unfamiliar file | `fapony review-seed --files` | nothing; it saves you reading the file |
346
- | Before committing | `/git-commit` | nothing; it just keeps commits reviewable |
347
- | Before merging | `/review-pony` | writes a mem row when findings survive (bug/decision + files) |
348
- | Merging | `/git-ship` (`pr` / `land` on a team) | nothing; pure git plumbing |
349
- | After it ships | `/move-to-done` | writes a mem note when the ship taught something, closes the loop |
350
- | Proving a finished run | `fapony report <run-id>` (CLI, not MCP) | git facts + allowlisted evidence, one page |
351
-
352
- **Team flow.** `/git-ship pr` stops once the PR is open and hands you the URL; the reviewer does
353
- their pass; `/git-ship land` merges it after approval. If the default branch requires reviews,
354
- plain `/git-ship` detects that and behaves like `pr` on its own.
355
-
356
- **What this is not.** It doesn't reduce your token bill — an agent that plans against known
357
- failure patterns tends to spend fewer rounds getting there, but fapony measures that, it doesn't
358
- cause it. Use `fapony usage-web` to find out whether it actually happened for you rather than taking
359
- the claim on faith.
360
-
361
- `fapony install --platform claude` (or `opencode`) symlinks these directories into
362
- `~/.claude/skills` rather than copying them, so `fapony update` refreshes every client
363
- at once. ZCode and Codex get the same skills linked into `~/.agents/skills`. A destination
364
- that already exists and isn't a fapony link is reported and left alone — replace it by hand
365
- if you want fapony's version.
366
-
367
- OpenCode is the one client whose hooks are generated files rather than `fapony hook-*`
368
- commands, so `fapony update` also re-runs the installer in a fresh process to refresh them —
369
- with `--plugins-only`, so a refresh touches fapony's plugin files and never your
370
- `opencode.json`. A fapony-owned plugin is rewritten in place; a file that isn't fapony's is
371
- reported and left alone, same as the skill links.
372
-
373
190
  `plan-with-pony` is vendor-neutral — the SKILL.md *is* the prompt, so pipe it to any agent:
374
-
375
- ```bash
376
- cat skill/plan-with-pony/SKILL.md | claude -p # Claude Code
377
- cat skill/plan-with-pony/SKILL.md | opencode run # OpenCode
378
- cat skill/plan-with-pony/SKILL.md | <your-agent> # anything that reads stdin
379
- ```
380
-
381
- Example plans produced by it live in [examples/](https://github.com/kire21b/fapony/tree/main/examples).
191
+ `cat skill/plan-with-pony/SKILL.md | claude -p` (or `opencode run`, or anything that reads stdin).
192
+ Example plans it produced: [examples/](https://github.com/kire21b/fapony/tree/main/examples).
382
193
 
383
194
  ### Plans your agent can answer questions about
384
195
 
385
- Plans stay markdown files in your repo — nothing moves into a database. Four optional
386
- frontmatter keys are enough to make a folder of them queryable:
387
-
388
- ```yaml
389
- ---
390
- kind: unit # tracker = a checklist that never finishes
391
- status: blocked # active | blocked | superseded
392
- blocked_by: PLAN-documents.md # a plan, or a sentence
393
- blocks: PLAN-export.md # ordering, stated once instead of buried in prose
394
- ---
395
-
396
- # PLAN — month view in /quick
397
-
398
- ## TL;DR # 15 lines; the only part that changes mid-flight
399
- - **Why:** two menu entries for the same data at different granularity
400
- - [x] chunk 1 — month grid `a1b2c3` 2026-09-13
401
- - [ ] chunk 2 — move overdue out
402
- ```
403
-
404
- Then open the next session with `fapony mem kickoff` — it reads the folder and the mem log and
405
- prints what is next (priority plans, the first unchecked chunk of each, open bugs) without
406
- reading a single 100KB plan body into context:
196
+ Plans stay markdown files in your repo — nothing moves into a database. Four optional frontmatter
197
+ keys (`kind` / `status` / `blocked_by` / `blocks`) make a folder of them queryable; plans with no
198
+ frontmatter still work, because the unchecked checkboxes are enough. Open the next session with
199
+ `fapony mem kickoff` — it prints what is next (priority plans, the first unchecked chunk of each,
200
+ open bugs) without reading a single 100KB plan body into context:
407
201
 
408
202
  ```
409
203
  ## next up
@@ -413,130 +207,101 @@ reading a single 100KB plan body into context:
413
207
  [3] last touched: src/quick/month.tsx, src/lib/money.ts
414
208
  ```
415
209
 
416
- **Plans with no frontmatter still work** — the unchecked checkboxes are enough, so an existing
417
- folder of plans is usable before anyone annotates anything. Two details that keep it honest over
418
- years:
419
-
420
- - The progress tally counts checkboxes in the **first `##` section only**, anchored by position
421
- rather than by the word "TL;DR" — so it works in any language, and a step list deeper in the
422
- file stays detail instead of becoming status.
423
- - **There is no `MASTER.md`.** Every line above is derived from the plan files themselves, so it
424
- cannot drift; a hand-kept master file always does.
425
-
426
- `status` / `blocked_by` / `blocks` / `superseded_by` are read by `fapony mem plan-check`
427
- (dangling refs, blocker shipped but dependent still blocked, waiter cycles, blocked with all
428
- chunks ticked) and by `fapony mem plan-sweep` (blocked view + a `🔓` unblock hint on `--apply`).
429
- Sentence values ("waiting on support email") carry no `PLAN-*.md` token and are never flagged.
430
-
431
- The layout, and why archiving is a plain `git mv`:
432
-
433
- ```
434
- .fapony/plan/PLAN-calendar.md live
435
- .fapony/done/PLAN-calendar.md shipped — same name, same depth, so every relative
436
- link inside the file survives the move untouched
437
- .fapony/spec/SPEC-calendar.md specs are a reference library; they are never archived
438
- ```
439
-
440
- Ship dates live in the plan's own header (`> ✅ **shipped 2026-09-13** (a1b2c3)`), not in the
441
- filename — `grep -h shipped .fapony/done/*.md | sort` answers "what landed when" without paying
442
- to rewrite every inbound link on every ship. Example plans, including an un-annotated one and an
443
- archived one: [examples/](https://github.com/kire21b/fapony/tree/main/examples).
210
+ **There is no `MASTER.md`** — every line above is derived from the plan files themselves, so it
211
+ cannot drift; a hand-kept master file always does. `fapony mem plan-check` verifies ticked chunk
212
+ shas against git history (a ticked box with no sha to check is a claim, not a close) and flags
213
+ dangling `blocked_by` refs; `fapony mem plan-sweep --apply` archives a shipped plan with `git mv`
214
+ into `.fapony/done/` — same name, same depth, so every relative link inside the file survives the
215
+ move. Specs live in `.fapony/spec/` and are never archived.
444
216
 
445
217
  ## CLI
446
218
 
447
219
  ```bash
448
- # Verification & reporting
449
- fapony mcp # MCP server (stdio JSON-RPC — 3 tools)
450
- fapony report <run-id> # verification report for a run
451
- fapony report-web [file] # static HTML report page
452
- fapony usage-scan # scan session logs → cache (incremental, progress bar)
453
- fapony price-scan # fetch model price table → prices.json (cache; query never fetches)
454
- fapony usage-web [port] # live usage comparison dashboard from cache
455
- fapony stats [--mode verdict [--regime code|fix|review|plan|inquiry|test]] # KPIs: pass/stall rate, by-model, by-grade — --mode verdict ranks by quality/tokens instead
456
- fapony digest [--since 7d|YYYY-MM-DD] [--format text|html] [--json] [--out FILE] # single-page summary: decisions, open bugs, in-flight plans, cost, pass/fail — from what's already on disk
457
- fapony plan-seed <name> [--spec] [--scope <path>]... # write PLAN (+SPEC): frontmatter, 8 empty sections, prior-art list, mem-decision context + existing-in-scope; existing plans listed on stdout; SPEC chunks carry signatures, every section capped — the agent fills the judgment
458
- fapony review-seed [--staged|--commit <sha>|--range <a...b>|--files f1,f2,dir|--plan <PLAN.md>] # read-only scope facts for a review (changed files, importers, untested, signatures, plan cross-check)
459
-
460
- # Memory & convention debt
461
- fapony mem add <kind> "<text>" --files f1,f2 [spec.md] # append a mem row (decision/bug/note/next/hold)
462
- fapony mem close <id> "<msg>" # close a bug
220
+ # core: memory + debt
221
+ fapony mem add <kind> "<text>" --files f1,f2 [--key k] [spec.md] # append a mem row (decision/bug/note/next/hold)
222
+ fapony mem close <id> "<msg>" # close a row (a bug stays open without this)
463
223
  fapony mem find ["<text>"] [--kind a,b] [--files f1,f2] [--since <N>d|YYYY-MM-DD] [--limit n] [--open] # search mem log
464
- fapony mem kickoff [<plan.md>] # open a session + a next-up list
224
+ fapony mem kickoff [<plan.md>] [--pick <n>] # open a session + a next-up list
465
225
  fapony mem where # show the resolved mem dir and which step won
466
- fapony mem done | stale # views
467
- fapony debt [--id <convention>] [--where <path>] # ไฟล์ไหนยังไม่ย้ายไป convention ที่ประกาศไว้ (live, read-only)
226
+ fapony mem done | stale | claim | release | synced | plan-sweep | plan-check | rotate
227
+ fapony debt [--id a,b] [--where <path>] # which files haven't migrated to a declared convention (live, read-only)
468
228
  fapony lint-baseline [--cmd ...] [--diff] # separate "already red" from "I made it red"
469
-
470
- # Hooks (wired by `fapony install`, not run by hand)
471
- fapony hook-stop # Stop hook: block turns with commits but no mem row
472
- fapony hook-read-hint # read/re-read annotations
473
- fapony hook-edit-hint # importer count before editing shape
474
- fapony hook-mv-guard # deny raw git mv of plan files into done/
475
- fapony hook-session-start # SessionStart: kickoff into context
229
+ fapony init-mem # delete legacy .memory/ dirs + warn call sites still referencing them
230
+ fapony digest [--since 7d|YYYY-MM-DD] [--format text|html] [--json] [--out FILE] # single-page summary from what's on disk
231
+
232
+ # usage (day one)
233
+ fapony usage-scan # scan session logs → cache (incremental, progress bar)
234
+ fapony price-scan # fetch model price table → prices.json (cache; query never fetches)
235
+ fapony usage-web [port] # usage comparison dashboard from cache
236
+
237
+ # lookup (read-only, never touches state)
238
+ fapony analyze [path] # live repo graph: hubs, orphans, cycles, changed-untested
239
+ fapony review-seed [--staged|--commit <sha>|--range <a...b>|--files f1,f2,dir|--plan <PLAN.md>] # scope facts for a review
240
+ fapony plan-seed <name> [--spec] [--scope <path>[,<path>]]... # write PLAN (+SPEC): frontmatter, capped sections, prior-art list
241
+
242
+ # hooks & MCP (wired by `fapony install`, not run by hand)
243
+ fapony mcp # MCP server (stdio JSON-RPC — 3 tools)
244
+ fapony hook-stop # Stop hook: block turns with commits but no mem row
245
+ fapony hook-read-hint # read/re-read annotations
246
+ fapony hook-edit-hint # importer count before editing shape
247
+ fapony hook-mv-guard # deny raw git mv of plan files into done/
248
+ fapony hook-session-start # SessionStart: kickoff into context
249
+
250
+ # frozen ledger (reads history only — the grading tool left the MCP surface in 2026-09)
251
+ fapony stats [--mode verdict [--regime code|fix|review|plan|inquiry|test]] # KPIs from old graded runs
252
+ fapony report <run-id> # verification report for a run
253
+ fapony report-web [file] # static HTML report page
254
+
255
+ # setup & maintenance
256
+ fapony init <path> # scaffold .fapony/ (plan/spec/memory/evidence)
257
+ fapony install [--all|--platform <name>|--dry-run] # wire MCP + skills into clients
258
+ fapony setup # interactive wizard: config + scaffold in one step
259
+ fapony update # self-update via git pull
260
+ fapony telemetry show|send # opt-in only, default off — see TELEMETRY.md
476
261
  ```
477
262
 
263
+ `fapony report <run-id>` prints the full report for a frozen-ledger run — git facts, handoff
264
+ conformance, allowlisted evidence, the stored verdict, cost — with anything the agent claimed but
265
+ couldn't prove marked as such. Reports are stamped with the producing build's `server_sha`; after
266
+ editing fapony, compare the stamp against `git log -1` before trusting a report from a
267
+ long-lived MCP server. Each allowlisted command gets `timeout_ms` (default 30s), the whole report
268
+ capped at 180s — a command that doesn't fit reports as `timeout`, never as a pass. If your
269
+ `.gitignore` ignores `.fapony/` wholesale, re-include the file: `**/.fapony/*`, then
270
+ `!**/.fapony/evidence.json`.
271
+
478
272
  *When* to call `mem add` is your project's call, not fapony's — write it in your own
479
- `AGENTS.md`/`CLAUDE.md`, not here. A starting point:
273
+ `AGENTS.md`/`CLAUDE.md`. A starting point:
480
274
 
481
275
  ```markdown
482
276
  ## Memory
483
277
  - Found a bug while working (not just user-reported)? Log it before fixing:
484
- `mem_add { kind: "bug", worktree: "<absolute app dir>", files: [...], text: "..." }`
278
+ mem_add { kind: "bug", worktree: "<absolute app dir>", files: [...], text: "..." }
485
279
  - `text` must stand alone — read months later with no chat context: what/where/repro/status.
486
- - Report the row id back in chat.
487
280
  - Don't fold the fix into the same chunk — log first, fix as its own next/chunk if you do.
488
281
  ```
489
282
 
490
- ```bash
491
- # Setup & maintenance
492
- fapony init <path> # scaffold .fapony/ (plan/spec/memory/evidence)
493
- fapony init-mem # delete .memory/ + warn call sites still referencing it
494
- fapony install # detect installed clients, prompt to wire each
495
- fapony install --all # wire all detected clients without prompting
496
- fapony install --platform <name> # force a specific client (bypasses detection)
497
- fapony install --dry-run # show what would happen without writing files
498
- fapony setup # interactive wizard: config + scaffold in one step
499
- fapony update # self-update via git pull
500
- fapony telemetry show|send # opt-in only, default off — see https://github.com/kire21b/fapony/blob/main/TELEMETRY.md
501
- ```
502
-
503
283
  ## Config
504
284
 
505
- `fapony.config.json` lives in the fapony checkout and is gitignored (it's per-machine). Copy [fapony.config.example.json](https://github.com/kire21b/fapony/blob/main/fapony.config.example.json) for a complete working reference; every section is optional with sane defaults. Key fields:
506
-
507
- - `worktrees` — name → absolute path mapping
508
- - `review.maxRounds` — round cap enforced by the gate
509
- - `memory` — shell commands for claim/close/add/kickoff, or `null` to default-wire when a `.fapony/.memory/` dir exists
510
- - `paths` (`planDir`/`doneDir`/`specDir`/`memDir`/`stateDir`) / `safety` — directory layout and the dangerous-command deny-list
511
- - `usageWeb` — optional `{ port, hostname }` for `fapony usage-web` server defaults. Run `fapony usage-scan` first to populate the cache.
512
-
513
- Env overrides: `FAPONY_CONFIG` (config file), `FAPONY_STATE_DIR` (state DB location; default `~/.config/fapony/`), `FAPONY_NO_REREAD_HINT=1` (turn the re-read hint off). Full schema, design decisions, and edge cases live with the code in the repo — this README intentionally doesn't duplicate them.
285
+ `fapony.config.json` lives in the fapony checkout and is gitignored (it's per-machine). Copy
286
+ [fapony.config.example.json](https://github.com/kire21b/fapony/blob/main/fapony.config.example.json)
287
+ for a complete working reference; every section is optional. Key fields: `worktrees`
288
+ (name → path), `memory` (shell commands, or `null` to disable), `paths` / `safety`,
289
+ `usageWeb { port, hostname }`. Env overrides: `FAPONY_CONFIG`, `FAPONY_STATE_DIR` (state DB;
290
+ default `~/.config/fapony/`), `FAPONY_NO_REREAD_HINT=1`.
514
291
 
515
292
  ## Scope
516
293
 
517
- **Supported:**
518
- - MCP server — 3 mem tools via stdio JSON-RPC, works with any MCP client
519
- - Measurement: cross-run KPIs by model/grade/value from the frozen ledger, per-file pain zones from mem rows (`files[]`) + passive usage (tokens, cost)
520
- - Model attribution across clients — resolved from the session log that was live when the old verdict landed, so a frozen row carries a model without the caller having declared one
521
- - Zero setup beyond install: the mem habit ships in the MCP `initialize` response, not in your rules file
522
- - Verification reports (frozen): handoff conformance, 6-grade verdicts, allowlisted evidence collector (`.fapony/evidence.json`); reports stamped with the producing build's `server_sha` — replayable, no new graded runs
523
- - Vendor-neutral executor/reviewer roles — anything that reads stdin
524
- - Memory integration via shell adapter, per project (configurable or default-wired)
525
- - Opt-in telemetry, off by default ([TELEMETRY.md](https://github.com/kire21b/fapony/blob/main/TELEMETRY.md) lists exactly what leaves the machine)
526
- - Bun-only; run state in SQLite via `bun:sqlite` (WAL mode)
527
- - Per-client hooks alongside MCP: Stop hook on Claude Code + Cursor · read/re-read/Edit hints on
528
- Claude Code + OpenCode · commit hint on OpenCode — [What runs where](#what-runs-where)
529
-
530
- **Not supported (yet):**
531
- - PreToolUse hints on Cursor, ZCode or Codex — Cursor has no such hook and the other two expose no
532
- in-process hook surface for read/edit hints (Codex's `apply_patch` sends patch text, not file paths)
533
- - A hosted or shared ledger for a team — `runs.worktree` is the only sharing key today, and it's a
534
- path, not an identity. If you want to try pointing two machines at the same ledger anyway,
535
- `FAPONY_STATE_DIR` can be set to a synced folder (Syncthing, a shared drive) — but SQLite's WAL
536
- mode does not tolerate concurrent writers over most network filesystems (NFS, Dropbox, iCloud
537
- Drive) and can corrupt the db under real contention. Treat this as an experiment you're accepting
538
- the risk on, not a supported path; nothing here is a substitute for a real shared-ledger server.
539
- - Memory migration from `.fapony/.memory/log.jsonl`
294
+ **Supported:** MCP server (3 mem tools, any MCP client) · cross-client usage on one yardstick ·
295
+ per-project mem log + convention debt · per-client hooks ([matrix above](#what-runs-where)) ·
296
+ vendor-neutral skills (anything that reads stdin) · opt-in telemetry, off by default
297
+ ([TELEMETRY.md](https://github.com/kire21b/fapony/blob/main/TELEMETRY.md) lists exactly what
298
+ leaves the machine) · Bun-only; run state in SQLite via `bun:sqlite` (WAL mode).
299
+
300
+ **Not supported (yet):** PreToolUse hints on Cursor, ZCode or Codex — Cursor has no such hook and
301
+ the other two expose no in-process hook surface for read/edit hints. A hosted or shared ledger —
302
+ `FAPONY_STATE_DIR` on a synced folder works as an experiment only; SQLite's WAL mode does not
303
+ tolerate concurrent writers over NFS/Dropbox/iCloud Drive and can corrupt the db under real
304
+ contention.
540
305
 
541
306
  ## License
542
307