fapony 0.2.1 → 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/README.md +95 -72
  2. package/fapony.ts +12 -5
  3. package/package.json +5 -4
  4. package/skill/define-convention/SKILL.md +77 -0
  5. package/skill/lookup-before-edit/SKILL.md +48 -0
  6. package/skill/move-to-done/SKILL.md +19 -30
  7. package/skill/review-pony/SKILL.md +38 -61
  8. package/src/analyze.ts +1 -1
  9. package/src/debt/cli.ts +193 -0
  10. package/src/debt/format.ts +107 -0
  11. package/src/debt/index.ts +19 -0
  12. package/src/debt/load.ts +92 -0
  13. package/src/debt/promotion.ts +152 -0
  14. package/src/debt/scan.ts +214 -0
  15. package/src/debt/types.ts +79 -0
  16. package/src/detect.ts +92 -0
  17. package/src/gate.ts +3 -3
  18. package/src/hook.ts +349 -100
  19. package/src/init-mem.ts +57 -71
  20. package/src/init.ts +12 -16
  21. package/src/install/antigravity.ts +112 -0
  22. package/src/install/claude.ts +16 -123
  23. package/src/install/codex.ts +58 -19
  24. package/src/install/detect.ts +17 -7
  25. package/src/install/opencode.ts +131 -6
  26. package/src/install.ts +12 -3
  27. package/src/lint-baseline.ts +2 -2
  28. package/src/mcp/primitives.ts +1 -1
  29. package/src/mcp/tools/index.ts +13 -102
  30. package/src/mcp/tools/mem.ts +71 -0
  31. package/src/mcp/transport.ts +8 -94
  32. package/src/mcp/worktree.ts +1 -1
  33. package/src/mem/commands/read.ts +231 -141
  34. package/src/mem/index.ts +4 -13
  35. package/src/memory.ts +17 -8
  36. package/src/{plan-seed.ts → seed/plan-seed.ts} +14 -32
  37. package/src/seed/primitives.ts +66 -0
  38. package/src/{review-seed.ts → seed/review-seed.ts} +7 -54
  39. package/src/session/helpers.ts +1 -1
  40. package/src/session/registry.ts +3 -6
  41. package/src/setup.ts +1 -1
  42. package/src/stats/data.ts +1 -1
  43. package/src/telemetry.ts +1 -1
  44. package/src/util.ts +61 -0
  45. package/templates/SPEC.md +8 -1
  46. package/images/logo.png +0 -0
  47. package/images/logo.webp +0 -0
  48. package/images/logo@400.webp +0 -0
  49. package/images/sample.webp +0 -0
  50. package/images/summary.webp +0 -0
  51. package/src/debt.ts +0 -811
  52. package/src/math.ts +0 -13
  53. package/src/mcp/tools/usage.ts +0 -211
  54. package/src/mcp/tools/verdict.ts +0 -161
package/README.md CHANGED
@@ -36,15 +36,16 @@ quietly counted as free.
36
36
 
37
37
  </details>
38
38
 
39
- That is day one. Past that, fapony measures what coding agents actually do — rounds, pass/fail,
40
- cost per grade — through 4 MCP tools any agent can call. If you juggle more than one agent, this is
41
- the point: the numbers come from the same yardstick everywhere, so "which model earns its keep on
42
- which kind of task" becomes a data question instead of a vibe. On top of measurement it checks
43
- claims against git facts: handoff conformance, allowlisted evidence, a 6-grade verdict — with
44
- everything the agent claimed but couldn't prove marked as such.
39
+ That is day one. Past that, fapony keeps what coding agents actually did — the frozen
40
+ ledger of graded runs (rounds, pass/fail, cost per grade, readable via CLI, no new grades)
41
+ plus the live mem log — through 3 MCP tools any agent can call. If you juggle more than
42
+ one agent, this is the point: the numbers come from the same yardstick everywhere, so
43
+ "which model earns its keep on which kind of task" becomes a data question instead of a
44
+ vibe. On top of history it checks claims against git facts: handoff conformance and
45
+ allowlisted evidence — with everything the agent claimed but couldn't prove marked as such.
45
46
 
46
- **What that question looks like answered, from one project's own ledger — the top of the `n≥5`
47
- frontier (`fapony stats --mode verdict --regime code`):**
47
+ **What that question looks like answered, from one project's own (frozen — reads history,
48
+ no new grades) ledger — the top of the `n≥5` frontier (`fapony stats --mode verdict --regime code`):**
48
49
 
49
50
  | model | tokens/pass | quality | n |
50
51
  |---|---|---|---|
@@ -63,14 +64,15 @@ accumulates is pain.** An agent has no memory of pain across sessions: it writes
63
64
  hand-rolled `try/catch` as cheerfully as the first, because every session starts new. Wrappers and
64
65
  shared libraries get built by *people* who were hurt by the same thing often enough to remember.
65
66
  That is why a codebase written with agents from day one tends not to grow a shared layer — nobody
66
- in the room remembers. fapony is the part that remembers: graded verdicts and mem rows both carry
67
- `files[]`, so the zones that keep coming back in failed and re-done work are a query, not a hunch.
67
+ in the room remembers. fapony is the part that remembers: mem rows carry
68
+ `files[]`, so the zones that keep coming back in re-done work are a query, not a hunch
69
+ (the frozen ledger's old graded rows carry them too).
68
70
  Paired with `fapony debt`, which tracks how far the codebase has actually moved to a convention you
69
71
  already decided on, that is the loop: notice the repeated cost, name the shared thing, watch the
70
72
  migration finish. Finding dead code and duplication is *not* part of it — knip and friends already
71
73
  do that better, and a convention with a `checker` is deliberately left to the checker.
72
74
 
73
- **The measurement layer underneath it:** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: which model earns its keep on which kind of work **in this project**, at what token cost, graded by whoever reviewed it. Every verdict carries a `regime` (`code` / `fix` / `review` / `plan` / `inquiry` / `test`), and runs split by whether there was a plan at all — so "does planning beat diving in, and for which model" is a table, not an argument.
75
+ **The measurement layer underneath it:** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: which model earns its keep on which kind of work **in this project**, at what token cost. The frozen ledger still answers that from history — every old verdict carries a `regime` (`code` / `fix` / `review` / `plan` / `inquiry` / `test`), and runs split by whether there was a plan at all — so "does planning beat diving in, and for which model" stays a table, not an argument. New accumulation goes to the mem log instead: decisions, bugs and notes with `files[]`, written by the agents doing the work.
74
76
 
75
77
  Three tiers, deliberately: **measurement ships today** and needs no per-project setup — raw facts nobody can call unfair. **Verification is the sharper edge** but stays beta until its evidence layer is hardened; fapony doesn't control your agent's flow, so it never promises "verified" as a headline. **Knowledge accumulation is the compounding one** — it's worthless on run 1 and gets more useful every run after, which is exactly why it's the layer competitors can't clone by copying a feature list.
76
78
 
@@ -82,15 +84,16 @@ Stated up front, because the gap between these two things is where most tooling
82
84
 
83
85
  - **It does not run your test suite.** The evidence collector runs an allowlist *you* write in
84
86
  `.fapony/evidence.json`, and never a command an agent proposes. No allowlist, no evidence.
85
- - **It does not judge your code.** `verdict_submit` *stores* a verdict; a human or a reviewing
86
- agent supplies it. fapony is the ledger, not the judge.
87
+ - **It does not judge your code.** Mem rows *record* decisions, bugs and notes; a human or
88
+ a working agent supplies them. fapony is the memory, not the judge. (The frozen
89
+ ledger's old grades work the same way — *stored*, never computed.)
87
90
  - **It checks conformance, not correctness.** What it can verify is that a claim lines up with git
88
91
  facts and that uncertainty was declared — not that the code works. Those are different
89
92
  guarantees and fapony only offers the first.
90
93
  - **Almost nothing blocks.** No CI failure, no gate on your own commands. The one exception is the
91
- Stop hook, once per turn when a commit ends ungraded; the read/edit/commit hints only annotate.
94
+ Stop hook, once per turn when a commit lands with no new mem row; the read/edit/commit hints only annotate.
92
95
  Skip the install of all of them and you are back to exactly the workflow you had.
93
- - **Model attribution is inferred, not declared.** A gate is attributed to whichever client
96
+ - **Model attribution is inferred, not declared.** A gate in the frozen ledger is attributed to whichever client
94
97
  session was live in that worktree at that moment. When one model writes the code and another
95
98
  reviews and files the verdict, the grade lands on the reviewer. Reports label it `inferred`;
96
99
  read it as such.
@@ -122,23 +125,24 @@ fapony usage-scan # scan the session logs already on dis
122
125
  fapony price-scan # fetch the OpenRouter price table → ~/.config/fapony/prices.json
123
126
  fapony usage-web # dashboard; re-run the scans to refresh
124
127
  # both scans are manual by design — nothing fetches or re-reads session logs behind your back
125
- # ask your agent: "Run fapony_usage — what has it cost me, per model?"
126
128
 
127
- # 4. Verify (optional, per project) — scaffold the evidence allowlist
129
+ # 4. Turn on the knowledge layer (per project you want it in)
128
130
  fapony init /path/to/your-worktree
129
- # edit .fapony/evidence.json to your real test/typecheck commands, then commit it
131
+ # creates .fapony/ — .memory/ (the mem log the 3 MCP tools read and write),
132
+ # conventions.json for `fapony debt`, plan/spec/done, and evidence.json
133
+ # conventions.json + evidence.json are shared rules: commit them
130
134
  ```
131
135
 
132
- With `.fapony/evidence.json` in place, any graded run can be replayed as a report. This one is
133
- a CLI command, not an MCP tool — the schemas cost every session of every client and no skill
134
- called them (see [The 4 tools](#the-4-tools) below). Grade something first;
135
- `verdict_submit` is what creates the run:
136
+ With `.fapony/evidence.json` in place, any graded run from the frozen ledger can be replayed
137
+ as a report. This one is a CLI command, not an MCP tool — the schemas cost every session of
138
+ every client and no skill called them (see [The 3 tools](#the-3-tools) below). No new runs
139
+ can be created; run ids come from `fapony stats` reading history:
136
140
 
137
141
  ```bash
138
142
  fapony report <run-id> # run ids come from `fapony stats`
139
143
  ```
140
144
 
141
- You get one report: git facts (files, commits, branch), handoff conformance (claims vs. reality), evidence from the allowlisted commands (pass/fail/timeout/unverified), a 6-grade verdict, and cost — with anything the agent claimed but couldn't prove marked as such.
145
+ You get one report: git facts (files, commits, branch), handoff conformance (claims vs. reality), evidence from the allowlisted commands (pass/fail/timeout/unverified), the frozen 6-grade verdict, and cost — with anything the agent claimed but couldn't prove marked as such.
142
146
 
143
147
  Sections that have nothing to report say so (`not_run`, `unavailable`) rather than disappearing — a report with no evidence must not read like a report that passed.
144
148
 
@@ -171,9 +175,9 @@ losing a single number.
171
175
 
172
176
  | | The ledger | The work side |
173
177
  |---|---|---|
174
- | What it is | 4 MCP tools + a SQLite ledger | plans, skills, read-only seed commands |
178
+ | What it is | 3 MCP tools (mem) + a frozen SQLite ledger (reads history) | plans, skills, read-only seed commands |
175
179
  | Needs | an MCP client | nothing — or your own tooling instead |
176
- | Writes | one graded row per unit of work | nothing |
180
+ | Writes | one mem row per unit of work, into the project's log | nothing |
177
181
  | Skip it and | there is no fapony | fapony still answers every question |
178
182
 
179
183
  ### What runs where
@@ -185,12 +189,12 @@ is `tool.execute.after`. Nothing here is required: skip the hooks and every MCP
185
189
 
186
190
  | | Claude Code | OpenCode | Cursor | ZCode | Codex |
187
191
  |---|---|---|---|---|---|
188
- | MCP tools — `mem_find` `mem_add` `fapony_usage` `verdict_submit` | ✅ | ✅ | ✅ | ✅ | ✅ |
189
- | Stop hook — refuse to end a turn with ungraded commits | ✅ | — | ✅ | — | ✅ after trust |
192
+ | MCP tools — `mem_find` `mem_add` `mem_close` | ✅ | ✅ | ✅ | ✅ | ✅ |
193
+ | Stop hook — refuse to end a turn with commits but no new mem row | ✅ | — | ✅ | — | ✅ after trust |
190
194
  | Read hint — big-file pointer + debt/mem lines | ✅ before | ✅ after | — | — | — |
191
195
  | Re-read hint — unchanged repeat read | ✅ before | ✅ after | — | — | — |
192
196
  | Edit hint — importer count before a shape change | ✅ before | ✅ after | — | — | — |
193
- | Commit hint — `git commit` → ungraded-run nudge | — | ✅ after | — | — | — |
197
+ | Commit hint — `git commit` → record-a-mem-row nudge | — | ✅ after | — | — | — |
194
198
  | Skills symlinked into `~/.claude/skills` | ✅ | ✅ | — | — | — |
195
199
  | Skills symlinked into `~/.agents/skills` | — | — | — | ✅ | ✅ |
196
200
  | `usage-scan` reads this client's session log | ✅ | ✅ | — | ✅ | ✅ |
@@ -203,10 +207,10 @@ without the agent deciding to call anything ([why](#when-to-call-what)).
203
207
 
204
208
  ## The ledger — this is the product
205
209
 
206
- One habit feeds it: grade a unit of work when it ends. Everything else on this page is
207
- optional around that. `verdict_submit` needs no plan file, no skill and no `.fapony/`
208
- directory — any agent that speaks MCP can call it, and calling it is what turns a pile of
209
- session logs into an answer.
210
+ One habit feeds it: record a mem row when a unit of work ends. Everything else on this page is
211
+ optional around that. `mem_add` needs no plan file and no skill — any agent that speaks MCP
212
+ can call it, and calling it is what turns a pile of session logs into an answer the next
213
+ session can find.
210
214
 
211
215
  ### One turn, end to end
212
216
 
@@ -215,50 +219,52 @@ sequenceDiagram
215
219
  autonumber
216
220
  participant A as Any MCP client
217
221
  participant F as fapony MCP
218
- participant L as ~/.config/fapony/state.db
222
+ participant M as project mem log (.fapony/.memory)
219
223
 
220
- Note over A,F: end a turn with a commit and no grade → the Stop hook blocks it once
221
- A->>F: verdict_submit (grade + regime + reason_code + note)
222
- F->>L: one graded unit of work, stamped with the model that did it
223
- opt proof, not just a claim — CLI, once the run exists
224
+ Note over A,F: end a turn with a commit and no new mem row → the Stop hook blocks it once
225
+ A->>F: mem_add (kind + files + text)
226
+ F->>M: one mem row in the project's log, stamped with the model that did it
227
+ opt proof, not just a claim — CLI, for runs from the frozen ledger
224
228
  A->>F: fapony report <run-id>
225
229
  F-->>A: git facts + evidence from .fapony/evidence.json, stamped with server_sha
226
230
  end
227
- Note over A,L: `fapony stats` reads it back — CLI, because you ask it, not the agent
231
+ Note over A,M: `fapony stats` reads the frozen ledger back — CLI, because you ask it, not the agent
228
232
  ```
229
233
 
230
- The Stop hook is the only thing fapony *blocks* — once per turn, when a commit ends ungraded.
231
- It never picks the grade; it cannot see whether the work held up. The hints only annotate and never
234
+ The Stop hook is the only thing fapony *blocks* — once per turn, when a commit lands with
235
+ no new mem row. It never judges what deserves recording; it cannot see whether the work
236
+ held up. The hints only annotate and never
232
237
  block: the **Read** hook adds one factual line when a read is large enough to be cheaper as
233
238
  `review-seed`, or when the same file is read again in a session and its mtime has not moved
234
239
  (`FAPONY_NO_REREAD_HINT=1` turns the re-read line off); the **Edit** hook names a file's importer
235
240
  count, once per session, before you change its shape; OpenCode's **commit** hook nudges after a
236
- `git commit` that left the run ungraded. Claude Code receives read/edit *before* the call, OpenCode
241
+ `git commit` that left no new mem row. Claude Code receives read/edit *before* the call, OpenCode
237
242
  *after* it — [What runs where](#what-runs-where) has the full client matrix.
238
243
 
239
- ### The 4 tools
244
+ ### The 3 tools
240
245
 
241
246
  | Tool | Tier | Purpose |
242
247
  |------|------|---------|
243
- | `fapony_usage` | measure | Passive usage from OpenCode, ZCode, Claude Code, and Codex sessions (tokens, cost, by-model; `detail:true` adds per-step timing) |
244
- | `verdict_submit` | verify | Store a 6-grade verdict (pass-excellent → uncertain) with a required `regime` — the task shape the grade applies to |
245
248
  | `mem_find` | recall | Search the project's mem log read-only — decisions/bugs/notes matched on the row's `files[]` (text substring for rows written without it), `text`, `kind` (no default filter), `since`. "What was ever decided about this file?" in one call before editing |
246
249
  | `mem_add` | recall | Append a mem row (decision/bug/note/next/hold) with `files[]` required and rejected when empty — the write half of `mem_find`, so the row is findable when you next touch that file |
250
+ | `mem_close` | recall | Close a mem row by id with a tombstone message — a separate tool (not `kind:"close"`) because a close row carries no `files[]`, so sharing `mem_add`'s schema would make required fields depend on another field's value |
247
251
 
248
252
  **A tool earns its schema by being called mid-task without being asked.** Everything you invoke
249
253
  deliberately is a CLI command instead: the schema is paid as input tokens in every session of
250
254
  every client whether or not it is used, while a CLI command costs nothing until it runs. That is
251
- why the handoff/report family is CLI-only, and why `fapony_stats`, `project_health_context` and
252
- `plan_list` left the MCP surface in 2026-09 (`fapony stats` answers the first, `fapony mem
253
- kickoff` the third; the second had no caller).
254
- Cutting is not the goal — spending where it pays back is: `mem_find` and `verdict_submit` keep
255
- their schemas because nobody is going to type them at the right moment. `fapony report <run-id>` prints the full report for a run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
255
+ why the handoff/report family is CLI-only, and why `fapony_stats`, `project_health_context`,
256
+ `plan_list` and `fapony_usage` left the MCP surface in 2026-09 (`fapony stats` answers the first, `fapony mem
257
+ kickoff` the third, `fapony usage-web` the fourth; the second had no caller).
258
+ Cutting is not the goal — spending where it pays back is: the three mem tools keep
259
+ their schemas because nobody is going to type them at the right moment. `fapony report <run-id>` prints the full report for a frozen-ledger run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
256
260
 
257
261
  Full protocol, adapter examples (bash, Python), and safety rules: [docs/mcp-handcheck.md](https://github.com/kire21b/fapony/blob/main/docs/mcp-handcheck.md).
258
262
 
259
- ### Verdict grades
263
+ ### Verdict grades (frozen ledger)
260
264
 
261
- Verification produces a quality grade, not just pass/fail:
265
+ No new grades are recorded — the tool that filed them left the MCP surface in 2026-09.
266
+ The old rows stay readable via `fapony stats` and `fapony report`, and this is the scale
267
+ they were filed on. Verification produced a quality grade, not just pass/fail:
262
268
 
263
269
  | Grade | Meaning |
264
270
  |-------|---------|
@@ -272,8 +278,8 @@ Verification produces a quality grade, not just pass/fail:
272
278
  ### Why measure from the outside
273
279
 
274
280
  - **Raw facts are hard to argue with.** Cost, rounds, diff sizes, pass rates — collected from git and session logs, not self-reported. A vendor can dispute a verdict as unfair; they can't dispute their own token count.
275
- - **Agent platforms grading their own homework is a conflict of interest.** fapony is a separate layer that measures any agent the same way, which is what makes "model X vs. model Y" or "workflow A vs. workflow B" answerable with real data instead of vibes.
276
- - **Verification stays honest about its limits.** The collector runs only commands listed in `.fapony/evidence.json`; commands proposed by the agent outside the allowlist are reported as *proposed — not executed*, never run. And because fapony doesn't control your agent's flow, verdicts are labeled as one signal — not promised as truth.
281
+ - **Agent platforms grading their own homework is a conflict of interest.** fapony is a separate layer that measured any agent the same way, which is what made "model X vs. model Y" or "workflow A vs. workflow B" answerable with real data instead of vibes. That history is still queryable; new accumulation is mem rows, not grades.
282
+ - **Verification stays honest about its limits.** The collector runs only commands listed in `.fapony/evidence.json`; commands proposed by the agent outside the allowlist are reported as *proposed — not executed*, never run. And because fapony doesn't control your agent's flow, old verdicts are labeled as one signal — not promised as truth.
277
283
 
278
284
  ## The work side — conveniences, not the contract
279
285
 
@@ -304,13 +310,15 @@ That is the whole trick; there is no model in the middle.
304
310
 
305
311
  ### Skills
306
312
 
307
- fapony ships five portable skills, each as `skill/<name>/SKILL.md` — the layout Claude
313
+ fapony ships seven portable skills, each as `skill/<name>/SKILL.md` — the layout Claude
308
314
  Code expects, so a client can symlink the directory rather than copy the file:
309
315
 
310
316
  | Skill | Purpose | Trigger |
311
317
  |-------|---------|---------|
312
318
  | `skill/plan-with-pony/` | Draft plan + spec from "what's in your head" via conversation | `/plan-with-pony` |
313
- | `skill/review-pony/` | Review as verification, wired to fapony: scope facts before (`review-seed`), verdict after | `/review-pony` |
319
+ | `skill/review-pony/` | Review as verification, wired to fapony: scope facts before (`review-seed`), a mem row after when findings survive | `/review-pony` |
320
+ | `skill/lookup-before-edit/` | Look up unfamiliar files (`review-seed --files` + mem + debt) before reading/editing them | `/lookup-before-edit` |
321
+ | `skill/define-convention/` | Turn a not-yet-migrated pattern into a tracked convention (interview + dry-run `debt`) | `/define-convention` |
314
322
  | `skill/move-to-done/` | Archive a shipped PLAN into .fapony/done/ | `/move-to-done` |
315
323
  | `skill/git-commit-conventional/` | Commit split by concern + conventional message | `/git-commit` |
316
324
  | `skill/git-ship/` | Push branch, open PR with drafted title/body, merge, reset branch onto base | `/ship`, `/pr` |
@@ -329,10 +337,10 @@ flowchart TD
329
337
  R -->|findings| W
330
338
  R -->|clean| S["/git-ship"]
331
339
  S -->|"there was a PLAN.md"| D["/move-to-done"]
332
- D -.-> H[(fapony ledger)]
340
+ D -.-> H[(mem log + frozen ledger)]
333
341
  R -.-> H
334
- C -.->|"Stop hook: a commit needs a verdict"| H
335
- H -.->|"which model for this shape"| Q
342
+ C -.->|"Stop hook: a commit needs a mem row"| H
343
+ H -.->|"pain zones + past model×shape"| Q
336
344
 
337
345
  style H fill:#2d333b,stroke:#768390,color:#adbac7
338
346
  ```
@@ -342,19 +350,20 @@ to pick up. Wiring, refactors and UI passes finish in one sitting and the PLAN.m
342
350
  unread — so `/plan-with-pony` declines those itself and hands over the two seed commands instead.
343
351
  Both arms meet at the same review and the same ledger.
344
352
 
345
- **The dotted edges are the whole point.** Verdicts carry `regime` and `reason_code`, so the
346
- ledger can answer the one question no single client can: *in this project, which model is worth
347
- paying for this shape of work.* That is what flows back to the fork — not "this file broke once",
348
- which fapony measured at a 1–9% base rate and demoted.
353
+ **The dotted edges are the whole point.** Mem rows carry `files[]` and standalone text, so
354
+ the zones that keep hurting are a query, not a hunch — and the frozen ledger still carries
355
+ `regime` on its old rows, so *in this project, which model was worth paying for this shape
356
+ of work* stays answerable from history. That is what flows back to the fork — not "this file
357
+ broke once", which fapony measured at a 1–9% base rate and demoted.
349
358
 
350
359
  | Moment | Call | What fapony gets out of it |
351
360
  |---|---|---|
352
361
  | Starting anything | `/plan-with-pony` | decides plan-vs-seed, then reads back how this shape has gone |
353
362
  | Before editing an unfamiliar file | `fapony review-seed --files` | nothing; it saves you reading the file |
354
363
  | Before committing | `/git-commit` | nothing; it just keeps commits reviewable |
355
- | Before merging | `/review-pony` | writes a verdict + `reason_code` + `regime` + note |
364
+ | Before merging | `/review-pony` | writes a mem row when findings survive (bug/decision + files) |
356
365
  | Merging | `/git-ship` (`pr` / `land` on a team) | nothing; pure git plumbing |
357
- | After it ships | `/move-to-done` | writes the ship verdict, closes the loop |
366
+ | After it ships | `/move-to-done` | writes a mem note when the ship taught something, closes the loop |
358
367
  | Proving a finished run | `fapony report <run-id>` (CLI, not MCP) | git facts + allowlisted evidence, one page |
359
368
 
360
369
  **Team flow.** `/git-ship pr` stops once the PR is open and hands you the URL; the reviewer does
@@ -363,7 +372,7 @@ plain `/git-ship` detects that and behaves like `pr` on its own.
363
372
 
364
373
  **What this is not.** It doesn't reduce your token bill — an agent that plans against known
365
374
  failure patterns tends to spend fewer rounds getting there, but fapony measures that, it doesn't
366
- cause it. Use `fapony_usage` to find out whether it actually happened for you rather than taking
375
+ cause it. Use `fapony usage-web` to find out whether it actually happened for you rather than taking
367
376
  the claim on faith.
368
377
 
369
378
  `fapony install --platform claude` (or `opencode`) symlinks these directories into
@@ -447,7 +456,7 @@ archived one: [examples/](https://github.com/kire21b/fapony/tree/main/examples).
447
456
 
448
457
  ```bash
449
458
  # Verification & reporting
450
- fapony mcp # MCP server (stdio JSON-RPC — 4 tools)
459
+ fapony mcp # MCP server (stdio JSON-RPC — 3 tools)
451
460
  fapony report <run-id> # verification report for a run
452
461
  fapony report-web [file] # static HTML report page
453
462
  fapony usage-scan # scan session logs → cache (incremental, progress bar)
@@ -464,10 +473,24 @@ fapony mem close <id> "<msg>" # close a bug
464
473
  fapony mem find "<text>" # substring-search every row
465
474
  fapony mem kickoff [<plan.md>] # open a session + a next-up list
466
475
  fapony mem where # show the resolved mem dir and which step won
467
- fapony mem now | done | stale # views
476
+ fapony mem done | stale # views
468
477
  fapony debt [--id <convention>] [--where <path>] # ไฟล์ไหนยังไม่ย้ายไป convention ที่ประกาศไว้ (live, read-only)
469
478
  fapony lint-baseline [--cmd ...] [--diff] # separate "already red" from "I made it red"
479
+ ```
480
+
481
+ *When* to call `mem add` is your project's call, not fapony's — write it in your own
482
+ `AGENTS.md`/`CLAUDE.md`, not here. A starting point:
470
483
 
484
+ ```markdown
485
+ ## Memory
486
+ - Found a bug while working (not just user-reported)? Log it before fixing:
487
+ `mem_add { kind: "bug", worktree: "<absolute app dir>", files: [...], text: "..." }`
488
+ - `text` must stand alone — read months later with no chat context: what/where/repro/status.
489
+ - Report the row id back in chat.
490
+ - Don't fold the fix into the same chunk — log first, fix as its own next/chunk if you do.
491
+ ```
492
+
493
+ ```bash
471
494
  # Setup & maintenance
472
495
  fapony init <path> # scaffold .fapony/ (plan/spec/memory/evidence)
473
496
  fapony init-mem # delete .memory/ + warn call sites still referencing it
@@ -496,11 +519,11 @@ Env overrides: `FAPONY_CONFIG` (config file), `FAPONY_STATE_DIR` (state DB locat
496
519
  ## Scope
497
520
 
498
521
  **Supported:**
499
- - MCP server — 4 tools via stdio JSON-RPC, works with any MCP client
500
- - Measurement: cross-run KPIs by model/grade/value, per-file risk (graded touches vs. fails) + passive usage (tokens, cost)
501
- - Model attribution across clients — resolved from the session log that was live when the verdict landed, so a verdict carries a model without the caller declaring one
502
- - Zero setup beyond install: the two habits fapony depends on ship in the MCP `initialize` response, not in your rules file
503
- - Verification (beta): handoff conformance, 6-grade verdicts, allowlisted evidence collector (`.fapony/evidence.json`); reports stamped with the producing build's `server_sha`
522
+ - MCP server — 3 mem tools via stdio JSON-RPC, works with any MCP client
523
+ - Measurement: cross-run KPIs by model/grade/value from the frozen ledger, per-file pain zones from mem rows (`files[]`) + passive usage (tokens, cost)
524
+ - Model attribution across clients — resolved from the session log that was live when the old verdict landed, so a frozen row carries a model without the caller having declared one
525
+ - Zero setup beyond install: the mem habit ships in the MCP `initialize` response, not in your rules file
526
+ - Verification reports (frozen): handoff conformance, 6-grade verdicts, allowlisted evidence collector (`.fapony/evidence.json`); reports stamped with the producing build's `server_sha` — replayable, no new graded runs
504
527
  - Vendor-neutral executor/reviewer roles — anything that reads stdin
505
528
  - Memory integration via shell adapter, per project (configurable or default-wired)
506
529
  - Opt-in telemetry, off by default ([TELEMETRY.md](https://github.com/kire21b/fapony/blob/main/TELEMETRY.md) lists exactly what leaves the machine)
package/fapony.ts CHANGED
@@ -5,9 +5,14 @@
5
5
 
6
6
  import { existsSync } from "node:fs";
7
7
  import { cmdAnalyze } from "./src/analyze.js";
8
- import { cmdDebt } from "./src/debt.js";
8
+ import { cmdDebt } from "./src/debt/cli.js";
9
9
  import { cmdDigest } from "./src/digest/cli.js";
10
- import { cmdHookEditHint, cmdHookReadHint, cmdHookStop } from "./src/hook.js";
10
+ import {
11
+ cmdHookEditHint,
12
+ cmdHookReadHint,
13
+ cmdHookSessionStart,
14
+ cmdHookStop,
15
+ } from "./src/hook.js";
11
16
  import { cmdInit } from "./src/init.js";
12
17
  import { cmdInitMem } from "./src/init-mem.js";
13
18
  import { cmdInstall } from "./src/install.js";
@@ -15,10 +20,10 @@ import { cmdLintBaseline } from "./src/lint-baseline.js";
15
20
  import { cmdMcp } from "./src/mcp/transport.js";
16
21
  import { cmdMem } from "./src/mem/index.js";
17
22
  import { initStore } from "./src/mem/store.js";
18
- import { cmdPlanSeed } from "./src/plan-seed.js";
19
23
  import { cmdPriceScan } from "./src/price/index.js";
20
24
  import { cmdReport, cmdReportWeb } from "./src/report/index.js";
21
- import { cmdReviewSeed } from "./src/review-seed.js";
25
+ import { cmdPlanSeed } from "./src/seed/plan-seed.js";
26
+ import { cmdReviewSeed } from "./src/seed/review-seed.js";
22
27
  import { cmdSetup } from "./src/setup.js";
23
28
  import { cmdStats } from "./src/stats/index.js";
24
29
  import { cmdTelemetry } from "./src/telemetry.js";
@@ -87,6 +92,8 @@ if (cmd === "analyze") {
87
92
  await cmdHookReadHint();
88
93
  } else if (cmd === "hook-edit-hint") {
89
94
  await cmdHookEditHint();
95
+ } else if (cmd === "hook-session-start") {
96
+ await cmdHookSessionStart();
90
97
  } else if (cmd === "mcp") {
91
98
  cmdMcp();
92
99
  } else if (cmd === "report") {
@@ -108,7 +115,7 @@ if (cmd === "analyze") {
108
115
  } else {
109
116
  console.error(`fapony: unknown command "${cmd ?? ""}"`);
110
117
  console.error(
111
- "usage: fapony <setup|update|stats|telemetry|init|init-mem|mem|install|report|report-web|usage-scan|usage-web|price-scan|analyze|debt|lint-baseline|plan-seed|review-seed|digest|mcp|hook-stop|hook-read-hint|hook-edit-hint|test> [args]",
118
+ "usage: fapony <setup|update|stats|telemetry|init|init-mem|mem|install|report|report-web|usage-scan|usage-web|price-scan|analyze|debt|lint-baseline|plan-seed|review-seed|digest|mcp|hook-stop|hook-read-hint|hook-edit-hint|hook-session-start|test> [args]",
112
119
  );
113
120
  process.exit(1);
114
121
  }
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "fapony",
3
- "version": "0.2.1",
4
- "description": "Measurement layer for coding agents \u2014 measure what agents do, verify what they claim. 4 MCP tools, any agent, no loop required",
3
+ "version": "0.3.3",
4
+ "description": "Token usage across Claude Code, OpenCode, Codex & ZCode on one yardstick — plus a project mem log and convention-debt tracker agents query via 3 MCP tools. No server, your data stays local",
5
5
  "license": "MIT",
6
6
  "author": "delamind (https://github.com/kire21b)",
7
7
  "homepage": "https://github.com/kire21b/fapony#readme",
@@ -24,8 +24,7 @@
24
24
  "fapony.ts",
25
25
  "src/",
26
26
  "templates/",
27
- "skill/",
28
- "images/"
27
+ "skill/"
29
28
  ],
30
29
  "scripts": {
31
30
  "lint": "biome check .",
@@ -33,8 +32,10 @@
33
32
  "test": "bun fapony.ts test",
34
33
  "test:fast": "SKIP_SLOW=1 bun fapony.ts test",
35
34
  "test:one": "bun scripts/test-one.ts",
35
+ "knip": "bunx knip@6 --exclude types,nsTypes || true",
36
36
  "check": "bun run lint && bun run typecheck && bun fapony.ts test",
37
37
  "prepublishOnly": "bash scripts/smoke-publish.sh",
38
+ "release": "git checkout main && git pull --ff-only && npm version patch -m 'release v%s' && git push origin main --follow-tags",
38
39
  "overview": "bun fapony.ts report-web /tmp/fapony-overview.html && open /tmp/fapony-overview.html"
39
40
  },
40
41
  "devDependencies": {
@@ -0,0 +1,77 @@
1
+ ---
2
+ name: define-convention
3
+ description: Turn "files that haven't migrated yet" into a tracked convention — one question, a draft built from real examples, then dry-run fapony debt until the counts hold. Trigger on /define-convention and when the user wants migration tracking or fapony debt shows declared rows with no regex.
4
+ ---
5
+
6
+ # Define Convention — from pain to a counted migration
7
+
8
+ One convention = the pattern to use (`ok`) + the pattern meaning not-yet-migrated
9
+ (`stale`) + scope (`where`) + an optional file condition (`guard`). The output is
10
+ one row in `<worktree>/.fapony/conventions.json` (`{"conventions": [...]}`), in the
11
+ same `.fapony/` dir as the mem log — run `fapony mem where`, go up one level, that
12
+ is where the file lives (app-scoped in a monorepo). No file there yet = create it;
13
+ a file with rows = append only, never rewrite other rows.
14
+
15
+ ## Phase 1 — One question, then the checker question
16
+
17
+ Ask this, and nothing else:
18
+
19
+ > "What should stop appearing, and what does the migrated code look like? Which
20
+ > directory is it in?"
21
+
22
+ Then the iron-rule question (one line, always asked — a convention eslint already
23
+ flags is one fapony must stay silent on):
24
+
25
+ > "Does an eslint rule or script already flag the old pattern? If yes, name it —
26
+ > fapony will record it as `checker` and never report this debt."
27
+
28
+ ## Phase 2 — Real examples before regex
29
+
30
+ Never invent the regex from prose — regex from prose is how entries get dropped.
31
+ `grep` the `where` dir for 2–3 hits of the old pattern and 1–2 of the new one. No
32
+ hits on either side = stop and say so; there is no migration to track yet, only
33
+ an opinion.
34
+
35
+ ## Phase 3 — Draft the row, show it, append it
36
+
37
+ ```json
38
+ {"id": "raw-throw", "rule": "throw failWith, not raw Error",
39
+ "where": "src", "stale": "throw new Error", "ok": "failWith"}
40
+ ```
41
+
42
+ - `id` short, kebab; `rule` one human line; `where` a repo-relative dir that
43
+ exists (`"."` = whole repo).
44
+ - `guard` only when `stale` alone is too broad (the file must ALSO match, e.g.
45
+ `"extends Base"`).
46
+ - `checker` set from Phase 1 = silent by design; verify it with eslint, not fapony.
47
+ - `stale: null` ships a `declared` placeholder — fill it or delete it, never keep it.
48
+ - The shown draft is the confirmation (rule 6c). Append, don't rewrite.
49
+
50
+ ## Phase 4 — Dry-run until the counts hold
51
+
52
+ ```bash
53
+ fapony debt --id <new-id>
54
+ ```
55
+
56
+ Read it literally — every outcome names its fix:
57
+
58
+ - `debt N · moved M` with N > 0 → the convention now counts. Report id, N, moved%
59
+ back in chat.
60
+ - `debt 0 · moved M — clean` → the migration is already done; nothing left to
61
+ track. Report it and change nothing.
62
+ - `debt 0` with no moved → `stale` matches nothing you care about; widen it or
63
+ fix `where`.
64
+ - `⚠ <id>: ... too broad` → narrow `stale`, shrink `where`, or add `guard` (the
65
+ match was capped past 250 files, so the entry was dropped).
66
+ - `⚠ <id>: stale regex broken` / `where: <dir> does not exist` → fix syntax / path.
67
+ - A row prints `declared, no checker, stale not filled in` → `stale` is null; see
68
+ Phase 3.
69
+ - `0 convention(s)` after `--id` → that id does not exist (misspelling — a dropped
70
+ row still prints its `⚠`). Only `no conventions.json in <dir>` means you wrote
71
+ to the wrong `.fapony/` (re-check `mem where`).
72
+
73
+ ## Later
74
+
75
+ A convention whose fix recurs ≥3 times makes `fapony debt` ask "time for a
76
+ checker?" — that promotion answers to a human, never to this skill. `no-checker`
77
+ means "never ask again": say it only deliberately.
@@ -0,0 +1,48 @@
1
+ ---
2
+ name: lookup-before-edit
3
+ description: Look up unfamiliar files before reading or editing them — exports, line numbers, and importers without reading the whole file. Trigger on /lookup-before-edit and proactively whenever you are about to read, edit, or refactor a file you do not already know.
4
+ ---
5
+
6
+ # Lookup Before Edit — scope first, read second
7
+
8
+ You are about to touch files you do not know. Do not `Read` them whole — look them up first.
9
+ A full-file read on a 500-line module costs ~35k tokens of output; a lookup costs under 1k.
10
+
11
+ ## The one command
12
+
13
+ ```bash
14
+ fapony review-seed --files <f1,f2,dir> [--body <sym>] [--callers <sym>]
15
+ ```
16
+
17
+ What it returns: every export with its line number (uncapped), plus the importers — first 12
18
+ per file, the rest as `(+N)`, so the total is still readable. That is your entry map — then
19
+ `Read` only the line ranges you actually need.
20
+
21
+ ## Narrow it
22
+
23
+ - `--body <sym>[,<sym>]` — declaration slice of those exports (truncated at 80 lines each):
24
+ "what does this do" without the file.
25
+ - `--callers <sym>` — symbol→symbol scan across the importers the static graph sees.
26
+ - Directories expand to the source files under them (cap 40, stated when cut). Paths that do
27
+ not exist are dropped with a notice, not counted silently.
28
+
29
+ ## Limits (do not work around them)
30
+
31
+ - **Exports only.** A non-exported function answers "no export named X in scope" — that is
32
+ the correct answer, not a failure. Read the file for internals.
33
+ - **Static only.** Dynamic use is invisible to the graph; an empty caller list means "not
34
+ seen statically", never "unused".
35
+ - **120 lines total.** Past that the tail is cut (`… (+N lines truncated)`) — and the tail is
36
+ the later files' signatures. Many files at once → split the call, don't trust a cut list.
37
+ - No `fapony` CLI or the call errors → read the file normally and carry on. A hint, not a gate.
38
+
39
+ ## History + debt (same paths, two calls)
40
+
41
+ - `mem_find` with `files: [<same paths>]` — "what was ever decided about this file" (MCP, no CLI spawn).
42
+ - `fapony debt --where <dir|file>` — conventions this path still violates; empty until `conventions.json` exists.
43
+
44
+ ## After the lookup
45
+
46
+ 1. Pick line ranges from the export list, `Read` those slices only.
47
+ 2. Before changing a shape, sum shown + `(+N)` importers — that is your blast radius.
48
+ 3. Do not re-read a file whose mtime has not moved; `grep` it instead.