fapony 0.3.0 → 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -36,15 +36,16 @@ quietly counted as free.
36
36
 
37
37
  </details>
38
38
 
39
- That is day one. Past that, fapony measures what coding agents actually do — rounds, pass/fail,
40
- cost per grade — through 4 MCP tools any agent can call. If you juggle more than one agent, this is
41
- the point: the numbers come from the same yardstick everywhere, so "which model earns its keep on
42
- which kind of task" becomes a data question instead of a vibe. On top of measurement it checks
43
- claims against git facts: handoff conformance, allowlisted evidence, a 6-grade verdict — with
44
- everything the agent claimed but couldn't prove marked as such.
39
+ That is day one. Past that, fapony keeps what coding agents actually did — the frozen
40
+ ledger of graded runs (rounds, pass/fail, cost per grade, readable via CLI, no new grades)
41
+ plus the live mem log — through 3 MCP tools any agent can call. If you juggle more than
42
+ one agent, this is the point: the numbers come from the same yardstick everywhere, so
43
+ "which model earns its keep on which kind of task" becomes a data question instead of a
44
+ vibe. On top of history it checks claims against git facts: handoff conformance and
45
+ allowlisted evidence — with everything the agent claimed but couldn't prove marked as such.
45
46
 
46
- **What that question looks like answered, from one project's own ledger — the top of the `n≥5`
47
- frontier (`fapony stats --mode verdict --regime code`):**
47
+ **What that question looks like answered, from one project's own (frozen — reads history,
48
+ no new grades) ledger — the top of the `n≥5` frontier (`fapony stats --mode verdict --regime code`):**
48
49
 
49
50
  | model | tokens/pass | quality | n |
50
51
  |---|---|---|---|
@@ -63,14 +64,15 @@ accumulates is pain.** An agent has no memory of pain across sessions: it writes
63
64
  hand-rolled `try/catch` as cheerfully as the first, because every session starts new. Wrappers and
64
65
  shared libraries get built by *people* who were hurt by the same thing often enough to remember.
65
66
  That is why a codebase written with agents from day one tends not to grow a shared layer — nobody
66
- in the room remembers. fapony is the part that remembers: graded verdicts and mem rows both carry
67
- `files[]`, so the zones that keep coming back in failed and re-done work are a query, not a hunch.
67
+ in the room remembers. fapony is the part that remembers: mem rows carry
68
+ `files[]`, so the zones that keep coming back in re-done work are a query, not a hunch
69
+ (the frozen ledger's old graded rows carry them too).
68
70
  Paired with `fapony debt`, which tracks how far the codebase has actually moved to a convention you
69
71
  already decided on, that is the loop: notice the repeated cost, name the shared thing, watch the
70
72
  migration finish. Finding dead code and duplication is *not* part of it — knip and friends already
71
73
  do that better, and a convention with a `checker` is deliberately left to the checker.
72
74
 
73
- **The measurement layer underneath it:** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: which model earns its keep on which kind of work **in this project**, at what token cost, graded by whoever reviewed it. Every verdict carries a `regime` (`code` / `fix` / `review` / `plan` / `inquiry` / `test`), and runs split by whether there was a plan at all — so "does planning beat diving in, and for which model" is a table, not an argument.
75
+ **The measurement layer underneath it:** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: which model earns its keep on which kind of work **in this project**, at what token cost. The frozen ledger still answers that from history — every old verdict carries a `regime` (`code` / `fix` / `review` / `plan` / `inquiry` / `test`), and runs split by whether there was a plan at all — so "does planning beat diving in, and for which model" stays a table, not an argument. New accumulation goes to the mem log instead: decisions, bugs and notes with `files[]`, written by the agents doing the work.
74
76
 
75
77
  Three tiers, deliberately: **measurement ships today** and needs no per-project setup — raw facts nobody can call unfair. **Verification is the sharper edge** but stays beta until its evidence layer is hardened; fapony doesn't control your agent's flow, so it never promises "verified" as a headline. **Knowledge accumulation is the compounding one** — it's worthless on run 1 and gets more useful every run after, which is exactly why it's the layer competitors can't clone by copying a feature list.
76
78
 
@@ -82,15 +84,16 @@ Stated up front, because the gap between these two things is where most tooling
82
84
 
83
85
  - **It does not run your test suite.** The evidence collector runs an allowlist *you* write in
84
86
  `.fapony/evidence.json`, and never a command an agent proposes. No allowlist, no evidence.
85
- - **It does not judge your code.** `verdict_submit` *stores* a verdict; a human or a reviewing
86
- agent supplies it. fapony is the ledger, not the judge.
87
+ - **It does not judge your code.** Mem rows *record* decisions, bugs and notes; a human or
88
+ a working agent supplies them. fapony is the memory, not the judge. (The frozen
89
+ ledger's old grades work the same way — *stored*, never computed.)
87
90
  - **It checks conformance, not correctness.** What it can verify is that a claim lines up with git
88
91
  facts and that uncertainty was declared — not that the code works. Those are different
89
92
  guarantees and fapony only offers the first.
90
93
  - **Almost nothing blocks.** No CI failure, no gate on your own commands. The one exception is the
91
- Stop hook, once per turn when a commit ends ungraded; the read/edit/commit hints only annotate.
94
+ Stop hook, once per turn when a commit lands with no new mem row; the read/edit/commit hints only annotate.
92
95
  Skip the install of all of them and you are back to exactly the workflow you had.
93
- - **Model attribution is inferred, not declared.** A gate is attributed to whichever client
96
+ - **Model attribution is inferred, not declared.** A gate in the frozen ledger is attributed to whichever client
94
97
  session was live in that worktree at that moment. When one model writes the code and another
95
98
  reviews and files the verdict, the grade lands on the reviewer. Reports label it `inferred`;
96
99
  read it as such.
@@ -123,21 +126,23 @@ fapony price-scan # fetch the OpenRouter price table →
123
126
  fapony usage-web # dashboard; re-run the scans to refresh
124
127
  # both scans are manual by design — nothing fetches or re-reads session logs behind your back
125
128
 
126
- # 4. Verify (optional, per project) — scaffold the evidence allowlist
129
+ # 4. Turn on the knowledge layer (per project you want it in)
127
130
  fapony init /path/to/your-worktree
128
- # edit .fapony/evidence.json to your real test/typecheck commands, then commit it
131
+ # creates .fapony/ — .memory/ (the mem log the 3 MCP tools read and write),
132
+ # conventions.json for `fapony debt`, plan/spec/done, and evidence.json
133
+ # conventions.json + evidence.json are shared rules: commit them
129
134
  ```
130
135
 
131
- With `.fapony/evidence.json` in place, any graded run can be replayed as a report. This one is
132
- a CLI command, not an MCP tool — the schemas cost every session of every client and no skill
133
- called them (see [The 4 tools](#the-4-tools) below). Grade something first;
134
- `verdict_submit` is what creates the run:
136
+ With `.fapony/evidence.json` in place, any graded run from the frozen ledger can be replayed
137
+ as a report. This one is a CLI command, not an MCP tool — the schemas cost every session of
138
+ every client and no skill called them (see [The 3 tools](#the-3-tools) below). No new runs
139
+ can be created; run ids come from `fapony stats` reading history:
135
140
 
136
141
  ```bash
137
142
  fapony report <run-id> # run ids come from `fapony stats`
138
143
  ```
139
144
 
140
- You get one report: git facts (files, commits, branch), handoff conformance (claims vs. reality), evidence from the allowlisted commands (pass/fail/timeout/unverified), a 6-grade verdict, and cost — with anything the agent claimed but couldn't prove marked as such.
145
+ You get one report: git facts (files, commits, branch), handoff conformance (claims vs. reality), evidence from the allowlisted commands (pass/fail/timeout/unverified), the frozen 6-grade verdict, and cost — with anything the agent claimed but couldn't prove marked as such.
141
146
 
142
147
  Sections that have nothing to report say so (`not_run`, `unavailable`) rather than disappearing — a report with no evidence must not read like a report that passed.
143
148
 
@@ -170,9 +175,9 @@ losing a single number.
170
175
 
171
176
  | | The ledger | The work side |
172
177
  |---|---|---|
173
- | What it is | 4 MCP tools + a SQLite ledger | plans, skills, read-only seed commands |
178
+ | What it is | 3 MCP tools (mem) + a frozen SQLite ledger (reads history) | plans, skills, read-only seed commands |
174
179
  | Needs | an MCP client | nothing — or your own tooling instead |
175
- | Writes | one graded row per unit of work | nothing |
180
+ | Writes | one mem row per unit of work, into the project's log | nothing |
176
181
  | Skip it and | there is no fapony | fapony still answers every question |
177
182
 
178
183
  ### What runs where
@@ -184,12 +189,12 @@ is `tool.execute.after`. Nothing here is required: skip the hooks and every MCP
184
189
 
185
190
  | | Claude Code | OpenCode | Cursor | ZCode | Codex |
186
191
  |---|---|---|---|---|---|
187
- | MCP tools — `mem_find` `mem_add` `mem_close` `verdict_submit` | ✅ | ✅ | ✅ | ✅ | ✅ |
188
- | Stop hook — refuse to end a turn with ungraded commits | ✅ | — | ✅ | — | ✅ after trust |
192
+ | MCP tools — `mem_find` `mem_add` `mem_close` | ✅ | ✅ | ✅ | ✅ | ✅ |
193
+ | Stop hook — refuse to end a turn with commits but no new mem row | ✅ | — | ✅ | — | ✅ after trust |
189
194
  | Read hint — big-file pointer + debt/mem lines | ✅ before | ✅ after | — | — | — |
190
195
  | Re-read hint — unchanged repeat read | ✅ before | ✅ after | — | — | — |
191
196
  | Edit hint — importer count before a shape change | ✅ before | ✅ after | — | — | — |
192
- | Commit hint — `git commit` → ungraded-run nudge | — | ✅ after | — | — | — |
197
+ | Commit hint — `git commit` → record-a-mem-row nudge | — | ✅ after | — | — | — |
193
198
  | Skills symlinked into `~/.claude/skills` | ✅ | ✅ | — | — | — |
194
199
  | Skills symlinked into `~/.agents/skills` | — | — | — | ✅ | ✅ |
195
200
  | `usage-scan` reads this client's session log | ✅ | ✅ | — | ✅ | ✅ |
@@ -202,10 +207,10 @@ without the agent deciding to call anything ([why](#when-to-call-what)).
202
207
 
203
208
  ## The ledger — this is the product
204
209
 
205
- One habit feeds it: grade a unit of work when it ends. Everything else on this page is
206
- optional around that. `verdict_submit` needs no plan file, no skill and no `.fapony/`
207
- directory — any agent that speaks MCP can call it, and calling it is what turns a pile of
208
- session logs into an answer.
210
+ One habit feeds it: record a mem row when a unit of work ends. Everything else on this page is
211
+ optional around that. `mem_add` needs no plan file and no skill — any agent that speaks MCP
212
+ can call it, and calling it is what turns a pile of session logs into an answer the next
213
+ session can find.
209
214
 
210
215
  ### One turn, end to end
211
216
 
@@ -214,32 +219,32 @@ sequenceDiagram
214
219
  autonumber
215
220
  participant A as Any MCP client
216
221
  participant F as fapony MCP
217
- participant L as ~/.config/fapony/state.db
222
+ participant M as project mem log (.fapony/.memory)
218
223
 
219
- Note over A,F: end a turn with a commit and no grade → the Stop hook blocks it once
220
- A->>F: verdict_submit (grade + regime + reason_code + note)
221
- F->>L: one graded unit of work, stamped with the model that did it
222
- opt proof, not just a claim — CLI, once the run exists
224
+ Note over A,F: end a turn with a commit and no new mem row → the Stop hook blocks it once
225
+ A->>F: mem_add (kind + files + text)
226
+ F->>M: one mem row in the project's log, stamped with the model that did it
227
+ opt proof, not just a claim — CLI, for runs from the frozen ledger
223
228
  A->>F: fapony report <run-id>
224
229
  F-->>A: git facts + evidence from .fapony/evidence.json, stamped with server_sha
225
230
  end
226
- Note over A,L: `fapony stats` reads it back — CLI, because you ask it, not the agent
231
+ Note over A,M: `fapony stats` reads the frozen ledger back — CLI, because you ask it, not the agent
227
232
  ```
228
233
 
229
- The Stop hook is the only thing fapony *blocks* — once per turn, when a commit ends ungraded.
230
- It never picks the grade; it cannot see whether the work held up. The hints only annotate and never
234
+ The Stop hook is the only thing fapony *blocks* — once per turn, when a commit lands with
235
+ no new mem row. It never judges what deserves recording; it cannot see whether the work
236
+ held up. The hints only annotate and never
231
237
  block: the **Read** hook adds one factual line when a read is large enough to be cheaper as
232
238
  `review-seed`, or when the same file is read again in a session and its mtime has not moved
233
239
  (`FAPONY_NO_REREAD_HINT=1` turns the re-read line off); the **Edit** hook names a file's importer
234
240
  count, once per session, before you change its shape; OpenCode's **commit** hook nudges after a
235
- `git commit` that left the run ungraded. Claude Code receives read/edit *before* the call, OpenCode
241
+ `git commit` that left no new mem row. Claude Code receives read/edit *before* the call, OpenCode
236
242
  *after* it — [What runs where](#what-runs-where) has the full client matrix.
237
243
 
238
- ### The 4 tools
244
+ ### The 3 tools
239
245
 
240
246
  | Tool | Tier | Purpose |
241
247
  |------|------|---------|
242
- | `verdict_submit` | verify | Store a 6-grade verdict (pass-excellent → uncertain) with a required `regime` — the task shape the grade applies to |
243
248
  | `mem_find` | recall | Search the project's mem log read-only — decisions/bugs/notes matched on the row's `files[]` (text substring for rows written without it), `text`, `kind` (no default filter), `since`. "What was ever decided about this file?" in one call before editing |
244
249
  | `mem_add` | recall | Append a mem row (decision/bug/note/next/hold) with `files[]` required and rejected when empty — the write half of `mem_find`, so the row is findable when you next touch that file |
245
250
  | `mem_close` | recall | Close a mem row by id with a tombstone message — a separate tool (not `kind:"close"`) because a close row carries no `files[]`, so sharing `mem_add`'s schema would make required fields depend on another field's value |
@@ -250,14 +255,16 @@ every client whether or not it is used, while a CLI command costs nothing until
250
255
  why the handoff/report family is CLI-only, and why `fapony_stats`, `project_health_context`,
251
256
  `plan_list` and `fapony_usage` left the MCP surface in 2026-09 (`fapony stats` answers the first, `fapony mem
252
257
  kickoff` the third, `fapony usage-web` the fourth; the second had no caller).
253
- Cutting is not the goal — spending where it pays back is: `mem_find` and `verdict_submit` keep
254
- their schemas because nobody is going to type them at the right moment. `fapony report <run-id>` prints the full report for a run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
258
+ Cutting is not the goal — spending where it pays back is: the three mem tools keep
259
+ their schemas because nobody is going to type them at the right moment. `fapony report <run-id>` prints the full report for a frozen-ledger run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
255
260
 
256
261
  Full protocol, adapter examples (bash, Python), and safety rules: [docs/mcp-handcheck.md](https://github.com/kire21b/fapony/blob/main/docs/mcp-handcheck.md).
257
262
 
258
- ### Verdict grades
263
+ ### Verdict grades (frozen ledger)
259
264
 
260
- Verification produces a quality grade, not just pass/fail:
265
+ No new grades are recorded — the tool that filed them left the MCP surface in 2026-09.
266
+ The old rows stay readable via `fapony stats` and `fapony report`, and this is the scale
267
+ they were filed on. Verification produced a quality grade, not just pass/fail:
261
268
 
262
269
  | Grade | Meaning |
263
270
  |-------|---------|
@@ -271,8 +278,8 @@ Verification produces a quality grade, not just pass/fail:
271
278
  ### Why measure from the outside
272
279
 
273
280
  - **Raw facts are hard to argue with.** Cost, rounds, diff sizes, pass rates — collected from git and session logs, not self-reported. A vendor can dispute a verdict as unfair; they can't dispute their own token count.
274
- - **Agent platforms grading their own homework is a conflict of interest.** fapony is a separate layer that measures any agent the same way, which is what makes "model X vs. model Y" or "workflow A vs. workflow B" answerable with real data instead of vibes.
275
- - **Verification stays honest about its limits.** The collector runs only commands listed in `.fapony/evidence.json`; commands proposed by the agent outside the allowlist are reported as *proposed — not executed*, never run. And because fapony doesn't control your agent's flow, verdicts are labeled as one signal — not promised as truth.
281
+ - **Agent platforms grading their own homework is a conflict of interest.** fapony is a separate layer that measured any agent the same way, which is what made "model X vs. model Y" or "workflow A vs. workflow B" answerable with real data instead of vibes. That history is still queryable; new accumulation is mem rows, not grades.
282
+ - **Verification stays honest about its limits.** The collector runs only commands listed in `.fapony/evidence.json`; commands proposed by the agent outside the allowlist are reported as *proposed — not executed*, never run. And because fapony doesn't control your agent's flow, old verdicts are labeled as one signal — not promised as truth.
276
283
 
277
284
  ## The work side — conveniences, not the contract
278
285
 
@@ -309,7 +316,7 @@ Code expects, so a client can symlink the directory rather than copy the file:
309
316
  | Skill | Purpose | Trigger |
310
317
  |-------|---------|---------|
311
318
  | `skill/plan-with-pony/` | Draft plan + spec from "what's in your head" via conversation | `/plan-with-pony` |
312
- | `skill/review-pony/` | Review as verification, wired to fapony: scope facts before (`review-seed`), verdict after | `/review-pony` |
319
+ | `skill/review-pony/` | Review as verification, wired to fapony: scope facts before (`review-seed`), a mem row after when findings survive | `/review-pony` |
313
320
  | `skill/lookup-before-edit/` | Look up unfamiliar files (`review-seed --files` + mem + debt) before reading/editing them | `/lookup-before-edit` |
314
321
  | `skill/define-convention/` | Turn a not-yet-migrated pattern into a tracked convention (interview + dry-run `debt`) | `/define-convention` |
315
322
  | `skill/move-to-done/` | Archive a shipped PLAN into .fapony/done/ | `/move-to-done` |
@@ -330,10 +337,10 @@ flowchart TD
330
337
  R -->|findings| W
331
338
  R -->|clean| S["/git-ship"]
332
339
  S -->|"there was a PLAN.md"| D["/move-to-done"]
333
- D -.-> H[(fapony ledger)]
340
+ D -.-> H[(mem log + frozen ledger)]
334
341
  R -.-> H
335
- C -.->|"Stop hook: a commit needs a verdict"| H
336
- H -.->|"which model for this shape"| Q
342
+ C -.->|"Stop hook: a commit needs a mem row"| H
343
+ H -.->|"pain zones + past model×shape"| Q
337
344
 
338
345
  style H fill:#2d333b,stroke:#768390,color:#adbac7
339
346
  ```
@@ -343,19 +350,20 @@ to pick up. Wiring, refactors and UI passes finish in one sitting and the PLAN.m
343
350
  unread — so `/plan-with-pony` declines those itself and hands over the two seed commands instead.
344
351
  Both arms meet at the same review and the same ledger.
345
352
 
346
- **The dotted edges are the whole point.** Verdicts carry `regime` and `reason_code`, so the
347
- ledger can answer the one question no single client can: *in this project, which model is worth
348
- paying for this shape of work.* That is what flows back to the fork — not "this file broke once",
349
- which fapony measured at a 1–9% base rate and demoted.
353
+ **The dotted edges are the whole point.** Mem rows carry `files[]` and standalone text, so
354
+ the zones that keep hurting are a query, not a hunch — and the frozen ledger still carries
355
+ `regime` on its old rows, so *in this project, which model was worth paying for this shape
356
+ of work* stays answerable from history. That is what flows back to the fork — not "this file
357
+ broke once", which fapony measured at a 1–9% base rate and demoted.
350
358
 
351
359
  | Moment | Call | What fapony gets out of it |
352
360
  |---|---|---|
353
361
  | Starting anything | `/plan-with-pony` | decides plan-vs-seed, then reads back how this shape has gone |
354
362
  | Before editing an unfamiliar file | `fapony review-seed --files` | nothing; it saves you reading the file |
355
363
  | Before committing | `/git-commit` | nothing; it just keeps commits reviewable |
356
- | Before merging | `/review-pony` | writes a verdict + `reason_code` + `regime` + note |
364
+ | Before merging | `/review-pony` | writes a mem row when findings survive (bug/decision + files) |
357
365
  | Merging | `/git-ship` (`pr` / `land` on a team) | nothing; pure git plumbing |
358
- | After it ships | `/move-to-done` | writes the ship verdict, closes the loop |
366
+ | After it ships | `/move-to-done` | writes a mem note when the ship taught something, closes the loop |
359
367
  | Proving a finished run | `fapony report <run-id>` (CLI, not MCP) | git facts + allowlisted evidence, one page |
360
368
 
361
369
  **Team flow.** `/git-ship pr` stops once the PR is open and hands you the URL; the reviewer does
@@ -448,7 +456,7 @@ archived one: [examples/](https://github.com/kire21b/fapony/tree/main/examples).
448
456
 
449
457
  ```bash
450
458
  # Verification & reporting
451
- fapony mcp # MCP server (stdio JSON-RPC — 4 tools)
459
+ fapony mcp # MCP server (stdio JSON-RPC — 3 tools)
452
460
  fapony report <run-id> # verification report for a run
453
461
  fapony report-web [file] # static HTML report page
454
462
  fapony usage-scan # scan session logs → cache (incremental, progress bar)
@@ -465,10 +473,24 @@ fapony mem close <id> "<msg>" # close a bug
465
473
  fapony mem find "<text>" # substring-search every row
466
474
  fapony mem kickoff [<plan.md>] # open a session + a next-up list
467
475
  fapony mem where # show the resolved mem dir and which step won
468
- fapony mem now | done | stale # views
476
+ fapony mem done | stale # views
469
477
  fapony debt [--id <convention>] [--where <path>] # ไฟล์ไหนยังไม่ย้ายไป convention ที่ประกาศไว้ (live, read-only)
470
478
  fapony lint-baseline [--cmd ...] [--diff] # separate "already red" from "I made it red"
479
+ ```
480
+
481
+ *When* to call `mem add` is your project's call, not fapony's — write it in your own
482
+ `AGENTS.md`/`CLAUDE.md`, not here. A starting point:
471
483
 
484
+ ```markdown
485
+ ## Memory
486
+ - Found a bug while working (not just user-reported)? Log it before fixing:
487
+ `mem_add { kind: "bug", worktree: "<absolute app dir>", files: [...], text: "..." }`
488
+ - `text` must stand alone — read months later with no chat context: what/where/repro/status.
489
+ - Report the row id back in chat.
490
+ - Don't fold the fix into the same chunk — log first, fix as its own next/chunk if you do.
491
+ ```
492
+
493
+ ```bash
472
494
  # Setup & maintenance
473
495
  fapony init <path> # scaffold .fapony/ (plan/spec/memory/evidence)
474
496
  fapony init-mem # delete .memory/ + warn call sites still referencing it
@@ -497,11 +519,11 @@ Env overrides: `FAPONY_CONFIG` (config file), `FAPONY_STATE_DIR` (state DB locat
497
519
  ## Scope
498
520
 
499
521
  **Supported:**
500
- - MCP server — 4 tools via stdio JSON-RPC, works with any MCP client
501
- - Measurement: cross-run KPIs by model/grade/value, per-file risk (graded touches vs. fails) + passive usage (tokens, cost)
502
- - Model attribution across clients — resolved from the session log that was live when the verdict landed, so a verdict carries a model without the caller declaring one
503
- - Zero setup beyond install: the two habits fapony depends on ship in the MCP `initialize` response, not in your rules file
504
- - Verification (beta): handoff conformance, 6-grade verdicts, allowlisted evidence collector (`.fapony/evidence.json`); reports stamped with the producing build's `server_sha`
522
+ - MCP server — 3 mem tools via stdio JSON-RPC, works with any MCP client
523
+ - Measurement: cross-run KPIs by model/grade/value from the frozen ledger, per-file pain zones from mem rows (`files[]`) + passive usage (tokens, cost)
524
+ - Model attribution across clients — resolved from the session log that was live when the old verdict landed, so a frozen row carries a model without the caller having declared one
525
+ - Zero setup beyond install: the mem habit ships in the MCP `initialize` response, not in your rules file
526
+ - Verification reports (frozen): handoff conformance, 6-grade verdicts, allowlisted evidence collector (`.fapony/evidence.json`); reports stamped with the producing build's `server_sha` — replayable, no new graded runs
505
527
  - Vendor-neutral executor/reviewer roles — anything that reads stdin
506
528
  - Memory integration via shell adapter, per project (configurable or default-wired)
507
529
  - Opt-in telemetry, off by default ([TELEMETRY.md](https://github.com/kire21b/fapony/blob/main/TELEMETRY.md) lists exactly what leaves the machine)
package/fapony.ts CHANGED
@@ -20,10 +20,10 @@ import { cmdLintBaseline } from "./src/lint-baseline.js";
20
20
  import { cmdMcp } from "./src/mcp/transport.js";
21
21
  import { cmdMem } from "./src/mem/index.js";
22
22
  import { initStore } from "./src/mem/store.js";
23
- import { cmdPlanSeed } from "./src/plan-seed.js";
24
23
  import { cmdPriceScan } from "./src/price/index.js";
25
24
  import { cmdReport, cmdReportWeb } from "./src/report/index.js";
26
- import { cmdReviewSeed } from "./src/review-seed.js";
25
+ import { cmdPlanSeed } from "./src/seed/plan-seed.js";
26
+ import { cmdReviewSeed } from "./src/seed/review-seed.js";
27
27
  import { cmdSetup } from "./src/setup.js";
28
28
  import { cmdStats } from "./src/stats/index.js";
29
29
  import { cmdTelemetry } from "./src/telemetry.js";
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "fapony",
3
- "version": "0.3.0",
4
- "description": "Measurement layer for coding agents \u2014 measure what agents do, verify what they claim. 4 MCP tools, any agent, no loop required",
3
+ "version": "0.3.3",
4
+ "description": "Token usage across Claude Code, OpenCode, Codex & ZCode on one yardstick — plus a project mem log and convention-debt tracker agents query via 3 MCP tools. No server, your data stays local",
5
5
  "license": "MIT",
6
6
  "author": "delamind (https://github.com/kire21b)",
7
7
  "homepage": "https://github.com/kire21b/fapony#readme",
@@ -24,8 +24,7 @@
24
24
  "fapony.ts",
25
25
  "src/",
26
26
  "templates/",
27
- "skill/",
28
- "images/"
27
+ "skill/"
29
28
  ],
30
29
  "scripts": {
31
30
  "lint": "biome check .",
@@ -33,8 +32,10 @@
33
32
  "test": "bun fapony.ts test",
34
33
  "test:fast": "SKIP_SLOW=1 bun fapony.ts test",
35
34
  "test:one": "bun scripts/test-one.ts",
35
+ "knip": "bunx knip@6 --exclude types,nsTypes || true",
36
36
  "check": "bun run lint && bun run typecheck && bun fapony.ts test",
37
37
  "prepublishOnly": "bash scripts/smoke-publish.sh",
38
+ "release": "git checkout main && git pull --ff-only && npm version patch -m 'release v%s' && git push origin main --follow-tags",
38
39
  "overview": "bun fapony.ts report-web /tmp/fapony-overview.html && open /tmp/fapony-overview.html"
39
40
  },
40
41
  "devDependencies": {
@@ -31,7 +31,7 @@ You are about to move a PLAN that has been shipped to the archive.
31
31
  status: superseded
32
32
  superseded_by: PLAN-bar.md
33
33
  ```
34
- Then skip step 5 — no work shipped, so there is no verdict to record. Never archive a plan as
34
+ Then skip step 5 — no work shipped, so there is nothing to record. Never archive a plan as
35
35
  superseded on your own reading; the user says which plan replaced it.
36
36
 
37
37
  A plan that is merely *waiting* (on a person, a customer, a decision) is **not** dead and does
@@ -60,29 +60,19 @@ You are about to move a PLAN that has been shipped to the archive.
60
60
  chore(plan): archive PLAN-foo.md (shipped <hash>)
61
61
  ```
62
62
 
63
- 5. **Record the verdict** — call the `verdict_submit` MCP tool (fapony) so this ship counts
64
- toward what this project knows about the model that did the work. No `run_id` needed:
65
- - `verdict`: `pass` (adjust if the ship had known rough edges — see VERDICT_GRADES)
66
- - `regime`: **required** — the shape of the work that shipped: `code` for a feature or
67
- refactor, `fix` for a bug fix, `plan` when what shipped was the plan or spec itself
68
- - `reason_code`: **`none` when the ship was clean** — not `other`. `other` means "a real
69
- problem none of the buckets name", so a clean ship filed there shows up in the
70
- recurring-fail-reasons list and crowds out the reasons that mean something. Otherwise
71
- `missing_test` / `scope_mismatch` / `unsafe_command` / `spec_gap` / `incomplete`, and
72
- `other` (with a `note`, which it requires) only when a real finding fits none of them
73
- - `note`: **omit it on a clean ship.** A verdict with no note still counts toward the plan
74
- history future drafts read ("passed round 1 before"), but only notes carry prose forward — so
75
- "clean ship" evicts a note that would have taught the next session something. Write one only when this plan hit something a
76
- reader could not get from the diff: what the symptom looked like, where the cause actually
77
- was, and the rule that follows. Standalone prose — it is read months later with no access
78
- to this conversation.
79
- - `worktree`: **absolute path** to this repo/worktree (`git rev-parse --show-toplevel`) —
80
- every other fapony tool and query scopes by
81
- absolute path too; a bare repo name won't match them
82
- - `plan`: the archived plan's path (post-move, e.g. `.fapony/done/PLAN-foo.md`)
83
- - `files`: repo-relative paths this plan touched (`git diff --name-only <base>..HEAD`) —
84
- the only input to per-file risk history; without it the verdict says something happened
85
- but not where
63
+ 5. **Leave a note when the ship taught something** — `plan-sweep --apply` (step 2)
64
+ already logged the ship itself as a decision row, so a clean ship needs nothing
65
+ more. When the plan hit something a reader could not get from the diff, call the
66
+ `mem_add` MCP tool (fapony) once:
67
+ - `kind`: `note`
68
+ - `text`: what the symptom looked like, where the cause actually was, and the
69
+ rule that follows. Standalone prose — it is read months later with no access
70
+ to this conversation. Write one only then — "clean ship" files nothing, and a
71
+ note that repeats the diff teaches the next session nothing
72
+ - `files`: repo-relative paths this plan touched (`git diff --name-only <base>..HEAD`)
73
+ - `spec`: the archived plan's path (post-move, e.g. `.fapony/done/PLAN-foo.md`)
74
+ - `worktree`: **absolute path** (`git rev-parse --show-toplevel`) — every other
75
+ fapony tool scopes by absolute path too; a bare repo name won't match them
86
76
  Skip only if fapony's MCP tools aren't available in this session — don't block the archive on it.
87
77
 
88
78
  ## Example
@@ -95,17 +85,16 @@ Steps:
95
85
  → moved, links rewritten, decision logged
96
86
  3. spec: untouched, stays in .fapony/spec/
97
87
  4. commit
98
- 5. verdict_submit(verdict="pass", reason_code="none", regime="code", worktree="/Users/you/Project/fapony/wt-fapony", plan=".fapony/done/PLAN-kickoff.md", files=["src/kickoff.ts"])
99
- — clean ship, so no note
88
+ 5. (clean ship — plan-sweep's decision row already recorded it, nothing more to file)
100
89
  ```
101
90
 
102
91
  A ship worth a note looks like this instead:
103
92
 
104
93
  ```
105
- 5. verdict_submit(verdict="pass-adequate", reason_code="spec_gap", regime="fix",
106
- note="sheet scroll reset on open, not close — the restore hook was on the wrong side; the router's own scrollRestoration resets on every navigate(). Check the router option before writing a restore hook.",
107
- worktree="/Users/you/Project/vela", plan=".fapony/done/PLAN-quick-nav.md",
108
- files=["src/routes/expenses/index.tsx"])
94
+ 5. mem_add(kind="note",
95
+ text="sheet scroll reset on open, not close — the restore hook was on the wrong side; the router's own scrollRestoration resets on every navigate(). Check the router option before writing a restore hook.",
96
+ files=["src/routes/expenses/index.tsx"], spec=".fapony/done/PLAN-quick-nav.md",
97
+ worktree="/Users/you/Project/vela")
109
98
  ```
110
99
 
111
100
  ## If fail
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: review-pony
3
- description: Review a plan, PR, diff, or design doc as a verification rather than an opinion — scope first, walk the real path, break it on paper, cite everything. Takes optional effort (low|medium|high|max, widens the walk only, never skips a pass) and --fix (apply CONFIRMED blocker/major findings after the report). Records the verdict to fapony after the report. Trigger on /review-pony and proactively whenever the user asks to review, audit, scrutinize, sanity-check, or get a second opinion on a plan, PR, diff, design doc, or proposed code change.
3
+ description: Review a plan, PR, diff, or design doc as a verification rather than an opinion — scope first, walk the real path, break it on paper, cite everything. Takes optional effort (low|medium|high|max, widens the walk only, never skips a pass) and --fix (apply CONFIRMED blocker/major findings after the report). Records surviving findings to the project's mem log after the report. Trigger on /review-pony and proactively whenever the user asks to review, audit, scrutinize, sanity-check, or get a second opinion on a plan, PR, diff, design doc, or proposed code change.
4
4
  ---
5
5
 
6
6
  # Review Pony
@@ -83,10 +83,10 @@ Write down every place the walk surprises you. Surprises outrank style; chase th
83
83
  - `low` — direct callers only, one hop. No test reading unless the diff touches a test file.
84
84
  - `medium` (default) — as written above: full path, callers, tests on the path.
85
85
  - `high` — also second-degree callers, and read the tests that exercise them, not just the path.
86
- - `max` — also run `get_impact_radius_tool` (or grep if the graph isn't wired) on every changed
87
- file and re-open every `deferred` line from the last review of this scope, if fapony has one.
86
+ - `max` — also grep every changed file for second-degree callers, and re-open
87
+ every `deferred` line from the last review of this scope, if fapony has one.
88
88
 
89
- Whatever level stopped you, say so in the one-line coverage note (rule below) — "walked to 1 hop"
89
+ Whatever level stopped you, say so in the one-line coverage note (see Report) — "walked to 1 hop"
90
90
  is honest, "walked" alone at `low` is not.
91
91
 
92
92
  ## Pass 3 — A finding needs a failing input
@@ -161,56 +161,25 @@ fixed: 1, 2 · skipped: 3 (PLAUSIBLE — could not reach the failing state)
161
161
  ```
162
162
 
163
163
  Fixing changes what actually shipped, not what the review found — re-run pass 4's citation
164
- check on the new state before calling it done, but don't re-run the whole review. Submit the
165
- verdict on what you found, not on the post-fix state (`verdict_submit`'s `note` can say the fix
166
- was applied).
167
-
168
- ## After: record the verdict (fapony)
169
-
170
- Call `verdict_submit` once, after the report is shown. Don't block the report on it, and don't
171
- let it change the report's content. Pass `regime="review"` — required, and a review is what this
172
- was; it is what puts this run in the `regime × model` table.
173
-
174
- | Report verdict | `verdict` |
175
- |---|---|
176
- | ship, 0 findings | `pass-excellent` |
177
- | ship, nit-only findings | `pass-good` |
178
- | fix-then-ship | `pass-adequate` |
179
- | rework / reject | `fail` |
180
- | could not walk enough to have a verdict | `uncertain` |
181
-
182
- `uncertain` is not a softer `fail`. It is the honest answer when the walk never
183
- reached the thing under review — the branch wouldn't build, the path is behind a
184
- service you cannot run, every finding came out `PLAUSIBLE`. Say so in the report
185
- too. Guessing `pass` there is the one outcome that makes the ledger lie.
186
-
187
- `reason_code` — the *lead* (most severe) finding, not a generic bucket:
188
-
189
- - **0 findings, or a clean pass → `none`** — never `other`. `other` means "a real
190
- finding that none of these buckets name", so filing clean passes there puts them
191
- in the recurring-fail-reasons list, where they crowd out the reasons that mean
192
- something. It is the one value in this table that costs other people accuracy.
193
- - missing or weak test coverage on the path you walked → `missing_test`
194
- - change is narrower or wider than the plan / PR description claims → `scope_mismatch`
195
- - a shell/eval/deploy command runs without the guard it needs → `unsafe_command`
196
- - the plan or spec didn't cover a case the walk exposed → `spec_gap`
197
- - the change stops short of what it set out to do → `incomplete`
198
- - a real finding none of the above names → `other`, and then `note` is **required**
199
-
200
- Always attach a one-line `note` — the only field a later review can act on. Say what broke or
201
- was walked, not that a review happened.
202
-
203
- Args: `verdict`, `reason_code`, `note`, `regime`, `worktree` — **absolute path** via
204
- `git rev-parse --show-toplevel`, never a bare name (`runs.worktree` is free text; a bare name
205
- writes where no query reads it and every fapony tool misses the run), `plan` (the PLAN file
206
- path under review, omitted for a bare PR/diff), and `files` — the repo-relative paths you
207
- actually walked. **Always send `files`.** It is the only input to per-file risk history; a
208
- verdict without it tells the next session that something failed but not where. No `run_id` —
209
- fapony reuses the latest still-open run for the same worktree+plan (so round 2+ counts toward
210
- the round cap), creating a row only when none is open. `session_id` (optional) — the client
211
- session id, only if the client exposes it; attribute the model, never block the submit on it.
212
- If `verdict_submit` errors, say so in one line and move on — never re-run a review because
213
- storage failed.
164
+ check on the new state before calling it done, but don't re-run the whole review. Record
165
+ what you found, not the post-fix state (the row's text can say the fix was applied).
166
+
167
+ ## After: record what the next session needs (fapony)
168
+
169
+ If a blocker/major CONFIRMED finding survived, or the verdict is rework/reject,
170
+ call `mem_add` once, after the report is shown. Don't block the report on it,
171
+ and don't let it change the report's content. Clean reviews (ship, nit-only
172
+ findings) record nothing — there is nothing the next session needs to find.
173
+
174
+ kind is `bug` when a finding survived (something is broken), `decision` when
175
+ the verdict turns on scope alone (rework/reject from pass 1 — the review locks
176
+ a direction). text is the report's verdict line, standalone — what broke or was
177
+ decided, not that a review happened. files are the repo-relative paths actually
178
+ walked — required, a row without them is unfindable. worktree is the absolute
179
+ path (`git rev-parse --show-toplevel`), never a bare name. Only record into a
180
+ project that already has a mem log — never create one uninvited.
181
+ If `mem_add` errors, say so in one line and move on — never re-run a review
182
+ because storage failed.
214
183
 
215
184
  ---
216
185
 
@@ -232,10 +201,10 @@ The four passes are the rules. These three are what they fail on in practice:
232
201
  ```
233
202
  1-4. scope holds; walked the new gate branch; ran the evidence command — it exits 0
234
203
  without running the suite (CONFIRMED: `bun test` with no test dir exits 0)
235
- post. verdict_submit(verdict="pass-adequate", reason_code="other", regime="review",
236
- note="evidence entry `bun test` exits 0 while running zero tests — real entry is `bun run test`",
237
- worktree="/Users/you/Project/fapony/wt-fapony",
238
- plan=".fapony/plan/PLAN-verdict-notes.md")
204
+ post. mem_add(kind="bug",
205
+ text="incremental scan replaces cached history with a delta — cache holds 500, truth 1500",
206
+ files=["cache.ts", "claude-code.ts"],
207
+ worktree="/Users/you/Project/fapony/wt-fapony")
239
208
  ```
240
209
 
241
210
  A report in budget — same review that, narrated, ran five paragraphs:
package/src/analyze.ts CHANGED
@@ -147,7 +147,7 @@ export const SCAN_EXTS = new Set([".ts", ".tsx", ".js", ".jsx"]);
147
147
  // have real importers here — scanning them produces false wrapper/orphan
148
148
  // signals (measured: conventions-seed flagged 8 "wrappers" that were all
149
149
  // src/mem/commands/*.ts helpers matched against unrelated identically-
150
- // named calls elsewhere in the repo, e.g. "cmdNow() instead of now(").
150
+ // named calls elsewhere in the repo, e.g. "cmdDone() instead of done(").
151
151
  const SKIP_DIRS = new Set([
152
152
  "node_modules",
153
153
  "dist",
package/src/gate.ts CHANGED
@@ -1,6 +1,6 @@
1
- // gateOnce is the core review-verdict logic — still live, called by the MCP
2
- // verdict_submit tool (src/mcp/tools/verdict.ts). The CLI wrapper (cmdGate)
3
- // was removed with the rest of the execute→review→fix loop (Wave 2).
1
+ // gateOnce is the core review-verdict logic — still live, called by the CLI
2
+ // wrapper (cmdGate) was removed with the rest of the execute→review→fix loop
3
+ // (Wave 2). The MCP verdict_submit tool was removed in PLAN-verdict-to-mem.
4
4
  import {
5
5
  addEvent,
6
6
  getRun,