fapony 0.1.3 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/README.md +111 -46
  2. package/fapony.ts +35 -2
  3. package/package.json +3 -2
  4. package/skill/move-to-done/SKILL.md +18 -28
  5. package/skill/plan-with-pony/SKILL.md +8 -3
  6. package/src/analyze.ts +175 -5
  7. package/src/conventions-seed.ts +3 -3
  8. package/src/db/defaults.ts +13 -5
  9. package/src/db/getters.ts +11 -9
  10. package/src/db/load.ts +2 -2
  11. package/src/db/store.ts +0 -75
  12. package/src/db/types.ts +2 -4
  13. package/src/debt.ts +176 -36
  14. package/src/digest/collect.ts +19 -2
  15. package/src/digest/text.ts +25 -0
  16. package/src/hook.ts +665 -24
  17. package/src/init-mem.ts +123 -37
  18. package/src/init.ts +57 -39
  19. package/src/install/claude.ts +44 -8
  20. package/src/install/codex.ts +105 -13
  21. package/src/install/opencode.ts +169 -5
  22. package/src/install.ts +2 -1
  23. package/src/lint-baseline.ts +3 -8
  24. package/src/mcp/evidence.ts +15 -3
  25. package/src/mcp/tools/index.ts +97 -185
  26. package/src/mcp/tools/mem.ts +153 -11
  27. package/src/mcp/transport.ts +5 -13
  28. package/src/mem/commands/read.ts +448 -0
  29. package/src/mem/commands/where.ts +56 -0
  30. package/{templates → src}/mem/commands/write.ts +12 -5
  31. package/src/mem/index.ts +144 -0
  32. package/src/mem/store.ts +350 -0
  33. package/src/memory.ts +238 -40
  34. package/src/plan-seed.ts +26 -6
  35. package/src/review-seed.ts +19 -0
  36. package/src/setup.ts +7 -8
  37. package/src/stats/data.ts +40 -104
  38. package/src/stats/format.ts +8 -9
  39. package/src/stats/index.ts +0 -1
  40. package/templates/PLAN.md +1 -0
  41. package/src/mcp/tools/context.ts +0 -66
  42. package/src/mcp/tools/plans.ts +0 -255
  43. package/src/mcp/tools/stats.ts +0 -96
  44. package/templates/mem/commands/read.ts +0 -194
  45. package/templates/mem/commands/selftest.ts +0 -450
  46. package/templates/mem/mem.ts +0 -68
  47. package/templates/mem/store.ts +0 -285
  48. /package/{templates → src}/mem/commands/plan.ts +0 -0
  49. /package/{templates → src}/mem/commands/rotate.ts +0 -0
  50. /package/{templates → src}/mem/render.ts +0 -0
  51. /package/{templates → src}/mem/selectors.ts +0 -0
package/README.md CHANGED
@@ -37,7 +37,7 @@ quietly counted as free.
37
37
  </details>
38
38
 
39
39
  That is day one. Past that, fapony measures what coding agents actually do — rounds, pass/fail,
40
- cost per grade — through 6 MCP tools any agent can call. If you juggle more than one agent, this is
40
+ cost per grade — through 4 MCP tools any agent can call. If you juggle more than one agent, this is
41
41
  the point: the numbers come from the same yardstick everywhere, so "which model earns its keep on
42
42
  which kind of task" becomes a data question instead of a vibe. On top of measurement it checks
43
43
  claims against git facts: handoff conformance, allowlisted evidence, a 6-grade verdict — with
@@ -58,7 +58,19 @@ on you to hold: work isn't randomly assigned to models, so a gap this size is a
58
58
  controlled trial — you likely route easy tasks to the cheap model already. `n≥5` is fapony's own
59
59
  floor before a model counts toward the frontier at all; below that it's a data point, not a pick.
60
60
 
61
- **The reason to keep it running is the third layer: knowledge accumulation.** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: which model earns its keep on which kind of work **in this project**, at what token cost, graded by whoever reviewed it. Every verdict carries a `regime` (`code` / `fix` / `review` / `plan` / `inquiry` / `test`), and runs split by whether there was a plan at all — so "does planning beat diving in, and for which model" is a table, not an argument.
61
+ **The reason to keep it running is the third layer: knowledge accumulation — and the thing it
62
+ accumulates is pain.** An agent has no memory of pain across sessions: it writes the 37th
63
+ hand-rolled `try/catch` as cheerfully as the first, because every session starts new. Wrappers and
64
+ shared libraries get built by *people* who were hurt by the same thing often enough to remember.
65
+ That is why a codebase written with agents from day one tends not to grow a shared layer — nobody
66
+ in the room remembers. fapony is the part that remembers: graded verdicts and mem rows both carry
67
+ `files[]`, so the zones that keep coming back in failed and re-done work are a query, not a hunch.
68
+ Paired with `fapony debt`, which tracks how far the codebase has actually moved to a convention you
69
+ already decided on, that is the loop: notice the repeated cost, name the shared thing, watch the
70
+ migration finish. Finding dead code and duplication is *not* part of it — knip and friends already
71
+ do that better, and a convention with a `checker` is deliberately left to the checker.
72
+
73
+ **The measurement layer underneath it:** Any single client already logs its own session — timing, tokens, tool calls. What none of them see is *across* runs, clients and task shapes: which model earns its keep on which kind of work **in this project**, at what token cost, graded by whoever reviewed it. Every verdict carries a `regime` (`code` / `fix` / `review` / `plan` / `inquiry` / `test`), and runs split by whether there was a plan at all — so "does planning beat diving in, and for which model" is a table, not an argument.
62
74
 
63
75
  Three tiers, deliberately: **measurement ships today** and needs no per-project setup — raw facts nobody can call unfair. **Verification is the sharper edge** but stays beta until its evidence layer is hardened; fapony doesn't control your agent's flow, so it never promises "verified" as a headline. **Knowledge accumulation is the compounding one** — it's worthless on run 1 and gets more useful every run after, which is exactly why it's the layer competitors can't clone by copying a feature list.
64
76
 
@@ -72,11 +84,12 @@ Stated up front, because the gap between these two things is where most tooling
72
84
  `.fapony/evidence.json`, and never a command an agent proposes. No allowlist, no evidence.
73
85
  - **It does not judge your code.** `verdict_submit` *stores* a verdict; a human or a reviewing
74
86
  agent supplies it. fapony is the ledger, not the judge.
75
- - **`handoff_check` checks conformance, not correctness.** It verifies that what the agent claimed
76
- lines up with git facts and that it declared its uncertainty — not that the code works. Those are
77
- different guarantees and fapony only offers the first.
78
- - **Nothing blocks.** There is no gate, no hook, no CI failure. Forget to call it and you are back
79
- to exactly the workflow you had.
87
+ - **It checks conformance, not correctness.** What it can verify is that a claim lines up with git
88
+ facts and that uncertainty was declared — not that the code works. Those are different
89
+ guarantees and fapony only offers the first.
90
+ - **Almost nothing blocks.** No CI failure, no gate on your own commands. The one exception is the
91
+ Stop hook, once per turn when a commit ends ungraded; the read/edit/commit hints only annotate.
92
+ Skip the install of all of them and you are back to exactly the workflow you had.
80
93
  - **Model attribution is inferred, not declared.** A gate is attributed to whichever client
81
94
  session was live in that worktree at that moment. When one model writes the code and another
82
95
  reviews and files the verdict, the grade lands on the reviewer. Reports label it `inferred`;
@@ -109,7 +122,7 @@ fapony usage-scan # scan the session logs already on dis
109
122
  fapony price-scan # fetch the OpenRouter price table → ~/.config/fapony/prices.json
110
123
  fapony usage-web # dashboard; re-run the scans to refresh
111
124
  # both scans are manual by design — nothing fetches or re-reads session logs behind your back
112
- # ask your agent: "Run fapony_stats and fapony_usage — what has it cost me, per model?"
125
+ # ask your agent: "Run fapony_usage — what has it cost me, per model?"
113
126
 
114
127
  # 4. Verify (optional, per project) — scaffold the evidence allowlist
115
128
  fapony init /path/to/your-worktree
@@ -118,7 +131,7 @@ fapony init /path/to/your-worktree
118
131
 
119
132
  With `.fapony/evidence.json` in place, any graded run can be replayed as a report. This one is
120
133
  a CLI command, not an MCP tool — the schemas cost every session of every client and no skill
121
- called them (see [The 6 tools](#the-6-tools) below). Grade something first;
134
+ called them (see [The 4 tools](#the-4-tools) below). Grade something first;
122
135
  `verdict_submit` is what creates the run:
123
136
 
124
137
  ```bash
@@ -144,10 +157,11 @@ flowchart LR
144
157
  B[OpenCode] --> F
145
158
  C[ZCode] --> F
146
159
  D[Codex] --> F
160
+ E[Cursor] --> F
147
161
  F --> G[git facts + session logs]
148
162
  G --> S[stats / usage]
149
163
  G --> V[verification report]
150
- G --> P[project_health - optional]
164
+ G --> M[mem log - what was decided here]
151
165
  ```
152
166
 
153
167
  fapony never drives the agent. It sits on two sides of your work that never touch each
@@ -157,11 +171,36 @@ losing a single number.
157
171
 
158
172
  | | The ledger | The work side |
159
173
  |---|---|---|
160
- | What it is | 6 MCP tools + a SQLite ledger | plans, skills, read-only seed commands |
174
+ | What it is | 4 MCP tools + a SQLite ledger | plans, skills, read-only seed commands |
161
175
  | Needs | an MCP client | nothing — or your own tooling instead |
162
176
  | Writes | one graded row per unit of work | nothing |
163
177
  | Skip it and | there is no fapony | fapony still answers every question |
164
178
 
179
+ ### What runs where
180
+
181
+ `fapony install` wires five clients (Claude Code, OpenCode, Cursor, ZCode, Codex). MCP is the only
182
+ piece all of them get — the hooks and in-process hints are per-client, and the read/edit hints
183
+ arrive **before** the call on Claude Code but **after** it on OpenCode, whose only annotate channel
184
+ is `tool.execute.after`. Nothing here is required: skip the hooks and every MCP tool still answers.
185
+
186
+ | | Claude Code | OpenCode | Cursor | ZCode | Codex |
187
+ |---|---|---|---|---|---|
188
+ | MCP tools — `mem_find` `mem_add` `fapony_usage` `verdict_submit` | ✅ | ✅ | ✅ | ✅ | ✅ |
189
+ | Stop hook — refuse to end a turn with ungraded commits | ✅ | — | ✅ | — | ✅ after trust |
190
+ | Read hint — big-file pointer + debt/mem lines | ✅ before | ✅ after | — | — | — |
191
+ | Re-read hint — unchanged repeat read | ✅ before | ✅ after | — | — | — |
192
+ | Edit hint — importer count before a shape change | ✅ before | ✅ after | — | — | — |
193
+ | Commit hint — `git commit` → ungraded-run nudge | — | ✅ after | — | — | — |
194
+ | Skills symlinked into `~/.claude/skills` | ✅ | ✅ | — | — | — |
195
+ | Skills symlinked into `~/.agents/skills` | — | — | — | ✅ | ✅ |
196
+ | `usage-scan` reads this client's session log | ✅ | ✅ | — | ✅ | ✅ |
197
+
198
+ `—` means not wired, not impossible: Cursor has no PreToolUse hook, and ZCode/Codex expose no
199
+ in-process hook surface for read/edit hints yet (Codex's `apply_patch` sends patch text, not
200
+ resolved file paths). Codex hooks require trust via `/hooks` before they run — `fapony install`
201
+ tells you when. The hints live on hooks rather than MCP on purpose — they must fire mid-turn
202
+ without the agent deciding to call anything ([why](#when-to-call-what)).
203
+
165
204
  ## The ledger — this is the product
166
205
 
167
206
  One habit feeds it: grade a unit of work when it ends. Everything else on this page is
@@ -185,26 +224,35 @@ sequenceDiagram
185
224
  A->>F: fapony report <run-id>
186
225
  F-->>A: git facts + evidence from .fapony/evidence.json, stamped with server_sha
187
226
  end
188
- A->>F: fapony_stats
189
- F->>L: read across every run, client and project
190
- L-->>A: model x regime x quality — which model to pay for this shape
227
+ Note over A,L: `fapony stats` reads it back — CLI, because you ask it, not the agent
191
228
  ```
192
229
 
193
- The Stop hook is the only thing fapony does *to* you — once per turn, when a commit ends
194
- ungraded. It never picks the grade; it cannot see whether the work held up.
230
+ The Stop hook is the only thing fapony *blocks* — once per turn, when a commit ends ungraded.
231
+ It never picks the grade; it cannot see whether the work held up. The hints only annotate and never
232
+ block: the **Read** hook adds one factual line when a read is large enough to be cheaper as
233
+ `review-seed`, or when the same file is read again in a session and its mtime has not moved
234
+ (`FAPONY_NO_REREAD_HINT=1` turns the re-read line off); the **Edit** hook names a file's importer
235
+ count, once per session, before you change its shape; OpenCode's **commit** hook nudges after a
236
+ `git commit` that left the run ungraded. Claude Code receives read/edit *before* the call, OpenCode
237
+ *after* it — [What runs where](#what-runs-where) has the full client matrix.
195
238
 
196
- ### The 6 tools
239
+ ### The 4 tools
197
240
 
198
241
  | Tool | Tier | Purpose |
199
242
  |------|------|---------|
200
- | `plan_list` | discover | Plan files grouped by state — active / blocked / untouched / superseded / trackers — with a progress tally and each one's run history. Not a raw `ls`; see [Plans your agent can answer questions about](#plans-your-agent-can-answer-questions-about) |
201
- | `fapony_stats` | measure | KPIs across runs: by-model (gates, fail rate, quality, tokens), by-grade, planned vs dove-in, regime x model, per-file risk; `group_by: reason_code\|plan\|file` for top-N slices; `mode: verdict` ranks models by quality vs tokens/pass instead of listing raw counts |
202
243
  | `fapony_usage` | measure | Passive usage from OpenCode, ZCode, Claude Code, and Codex sessions (tokens, cost, by-model; `detail:true` adds per-step timing) |
203
244
  | `verdict_submit` | verify | Store a 6-grade verdict (pass-excellent → uncertain) with a required `regime` — the task shape the grade applies to |
204
- | `project_health_context` | recall | Known-patterns block for the files you are about to touch. Useful when a file does have history; measured across real repos, most do not (1-9% of shipped files come back under a `fix:` within two weeks), so it is optional — never a precondition for editing |
205
- | `mem_find` | recall | Search the project's mem log read-only — decisions/bugs/notes keyed by `files[]`, `text`, `kind` (no default filter), `since`. "What was ever decided about this file?" in one call before editing |
206
-
207
- The handoff/report family is CLI-only — the schemas cost every session of every client and no skill called them. `fapony report <run-id>` prints the full report for a run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
245
+ | `mem_find` | recall | Search the project's mem log read-only — decisions/bugs/notes matched on the row's `files[]` (text substring for rows written without it), `text`, `kind` (no default filter), `since`. "What was ever decided about this file?" in one call before editing |
246
+ | `mem_add` | recall | Append a mem row (decision/bug/note/next/hold) with `files[]` required and rejected when empty — the write half of `mem_find`, so the row is findable when you next touch that file |
247
+
248
+ **A tool earns its schema by being called mid-task without being asked.** Everything you invoke
249
+ deliberately is a CLI command instead: the schema is paid as input tokens in every session of
250
+ every client whether or not it is used, while a CLI command costs nothing until it runs. That is
251
+ why the handoff/report family is CLI-only, and why `fapony_stats`, `project_health_context` and
252
+ `plan_list` left the MCP surface in 2026-09 (`fapony stats` answers the first, `fapony mem
253
+ kickoff` the third; the second had no caller).
254
+ Cutting is not the goal — spending where it pays back is: `mem_find` and `verdict_submit` keep
255
+ their schemas because nobody is going to type them at the right moment. `fapony report <run-id>` prints the full report for a run (facts + handoff conformance + evidence + verdict); `fapony report-web [file]` renders it as a static HTML page (overwrites `file` on every call — safe to reuse the same path). Run `bun run overview` for a one-shot shortcut that writes it to `/tmp/fapony-overview.html` and opens it. `fapony usage-scan` scans session logs and writes a cache file; `fapony usage-web [port]` serves a static HTML dashboard from that cache (no live scanning). Run `fapony usage-scan` periodically to keep data fresh.
208
256
 
209
257
  Full protocol, adapter examples (bash, Python), and safety rules: [docs/mcp-handcheck.md](https://github.com/kire21b/fapony/blob/main/docs/mcp-handcheck.md).
210
258
 
@@ -320,8 +368,9 @@ the claim on faith.
320
368
 
321
369
  `fapony install --platform claude` (or `opencode`) symlinks these directories into
322
370
  `~/.claude/skills` rather than copying them, so `fapony update` refreshes every client
323
- at once. A destination that already exists and isn't a fapony link is reported and left
324
- alone — replace it by hand if you want fapony's version.
371
+ at once. ZCode and Codex get the same skills linked into `~/.agents/skills`. A destination
372
+ that already exists and isn't a fapony link is reported and left alone — replace it by hand
373
+ if you want fapony's version.
325
374
 
326
375
  `plan-with-pony` is vendor-neutral — the SKILL.md *is* the prompt, so pipe it to any agent:
327
376
 
@@ -354,29 +403,31 @@ blocks: PLAN-export.md # ordering, stated once instead of buried in pros
354
403
  - [ ] chunk 2 — move overdue out
355
404
  ```
356
405
 
357
- Then ask your agent *"what's left, and what's blocked?"* — `plan_list` answers from the
358
- frontmatter and from fapony's own run history, without reading a single 100KB plan body into
359
- context (`format: "markdown"`):
406
+ Then open the next session with `fapony mem kickoff` — it reads the folder and the mem log and
407
+ prints what is next (priority plans, the first unchecked chunk of each, open bugs) without
408
+ reading a single 100KB plan body into context:
360
409
 
361
410
  ```
362
- ## active — in order (2)
363
- - [ ] PLAN-calendar — 1/3 · unblocks PLAN-export
364
- - [ ] PLAN-export — never attempted
365
- ## blocked (1)
366
- - [ ] PLAN-attendance — waiting: PLAN-documents.md
367
- ## untouched (14) · trackers (3)
368
- done: 63 archived
411
+ ## next up
412
+ [1] chunk 2 — move overdue out (PLAN-calendar.md)
413
+ [2] bug #mu8t5qve — money drifts in the month grid…
414
+ → fapony mem close mu8t5qve "<msg>"
415
+ [3] last touched: src/quick/month.tsx, src/lib/money.ts
369
416
  ```
370
417
 
371
- **Plans with no frontmatter still work** — they are grouped by run history alone (attempted =
372
- active, never attempted = untouched), so an existing folder of plans is queryable before anyone
373
- annotates anything. Two details that keep it honest over years:
418
+ **Plans with no frontmatter still work** — the unchecked checkboxes are enough, so an existing
419
+ folder of plans is usable before anyone annotates anything. Two details that keep it honest over
420
+ years:
374
421
 
375
422
  - The progress tally counts checkboxes in the **first `##` section only**, anchored by position
376
423
  rather than by the word "TL;DR" — so it works in any language, and a step list deeper in the
377
424
  file stays detail instead of becoming status.
378
- - **There is no `MASTER.md`.** Every line of the list above is derived from frontmatter and
379
- checkboxes, so it cannot drift; a hand-kept master file always does.
425
+ - **There is no `MASTER.md`.** Every line above is derived from the plan files themselves, so it
426
+ cannot drift; a hand-kept master file always does.
427
+
428
+ `status` / `blocked_by` / `blocks` / `superseded_by` are read by people, not by a tool — the one
429
+ that read them, `plan_list`, was removed in 2026-09 once `mem kickoff` answered the same
430
+ question from the CLI, where a schema costs nothing until it runs.
380
431
 
381
432
  The layout, and why archiving is a plain `git mv`:
382
433
 
@@ -396,7 +447,7 @@ archived one: [examples/](https://github.com/kire21b/fapony/tree/main/examples).
396
447
 
397
448
  ```bash
398
449
  # Verification & reporting
399
- fapony mcp # MCP server (stdio JSON-RPC — 6 tools)
450
+ fapony mcp # MCP server (stdio JSON-RPC — 4 tools)
400
451
  fapony report <run-id> # verification report for a run
401
452
  fapony report-web [file] # static HTML report page
402
453
  fapony usage-scan # scan session logs → cache (incremental, progress bar)
@@ -407,9 +458,19 @@ fapony digest [--since 7d|YYYY-MM-DD] [--format text|html] [--json] [--out FILE]
407
458
  fapony plan-seed <name> [--spec] [--scope <path>]... # write PLAN (+SPEC): frontmatter, 8 empty sections, prior-art list, ledger context; SPEC chunks carry signatures, every section capped — the agent fills the judgment
408
459
  fapony review-seed [--staged|--commit <sha>|--range <a...b>|--files f1,f2,dir|--plan <PLAN.md>] # read-only scope facts for a review (changed files, importers, untested, signatures, plan cross-check)
409
460
 
461
+ # Memory & convention debt
462
+ fapony mem add <kind> "<text>" --files f1,f2 [spec.md] # append a mem row (decision/bug/note/next/hold)
463
+ fapony mem close <id> "<msg>" # close a bug
464
+ fapony mem find "<text>" # substring-search every row
465
+ fapony mem kickoff [<plan.md>] # open a session + a next-up list
466
+ fapony mem where # show the resolved mem dir and which step won
467
+ fapony mem now | done | stale # views
468
+ fapony debt [--id <convention>] [--where <path>] # ไฟล์ไหนยังไม่ย้ายไป convention ที่ประกาศไว้ (live, read-only)
469
+ fapony lint-baseline [--cmd ...] [--diff] # separate "already red" from "I made it red"
470
+
410
471
  # Setup & maintenance
411
472
  fapony init <path> # scaffold .fapony/ (plan/spec/memory/evidence)
412
- fapony init-mem [--update] # refresh the memory scaffold from the template
473
+ fapony init-mem # delete .memory/ + warn call sites still referencing it
413
474
  fapony install # detect installed clients, prompt to wire each
414
475
  fapony install --all # wire all detected clients without prompting
415
476
  fapony install --platform <name> # force a specific client (bypasses detection)
@@ -426,16 +487,16 @@ fapony test # self-check
426
487
 
427
488
  - `worktrees` — name → absolute path mapping
428
489
  - `review.maxRounds` — round cap enforced by the gate
429
- - `memory` — shell commands for claim/close/add/kickoff, or `null` to default-wire when `.fapony/.memory/mem.ts` exists
430
- - `paths` (`planDir`/`doneDir`/`specDir`/`memoryEntry`/`stateDir`) / `safety` — directory layout and the dangerous-command deny-list
490
+ - `memory` — shell commands for claim/close/add/kickoff, or `null` to default-wire when a `.fapony/.memory/` dir exists
491
+ - `paths` (`planDir`/`doneDir`/`specDir`/`memDir`/`stateDir`) / `safety` — directory layout and the dangerous-command deny-list
431
492
  - `usageWeb` — optional `{ port, hostname }` for `fapony usage-web` server defaults. Run `fapony usage-scan` first to populate the cache.
432
493
 
433
- Env overrides: `FAPONY_CONFIG` (config file), `FAPONY_STATE_DIR` (state DB location; default `~/.config/fapony/`). Full schema, design decisions, and edge cases live with the code in the repo — this README intentionally doesn't duplicate them.
494
+ Env overrides: `FAPONY_CONFIG` (config file), `FAPONY_STATE_DIR` (state DB location; default `~/.config/fapony/`), `FAPONY_NO_REREAD_HINT=1` (turn the re-read hint off). Full schema, design decisions, and edge cases live with the code in the repo — this README intentionally doesn't duplicate them.
434
495
 
435
496
  ## Scope
436
497
 
437
498
  **Supported:**
438
- - MCP server — 6 tools via stdio JSON-RPC, works with any MCP client
499
+ - MCP server — 4 tools via stdio JSON-RPC, works with any MCP client
439
500
  - Measurement: cross-run KPIs by model/grade/value, per-file risk (graded touches vs. fails) + passive usage (tokens, cost)
440
501
  - Model attribution across clients — resolved from the session log that was live when the verdict landed, so a verdict carries a model without the caller declaring one
441
502
  - Zero setup beyond install: the two habits fapony depends on ship in the MCP `initialize` response, not in your rules file
@@ -444,8 +505,12 @@ Env overrides: `FAPONY_CONFIG` (config file), `FAPONY_STATE_DIR` (state DB locat
444
505
  - Memory integration via shell adapter, per project (configurable or default-wired)
445
506
  - Opt-in telemetry, off by default ([TELEMETRY.md](https://github.com/kire21b/fapony/blob/main/TELEMETRY.md) lists exactly what leaves the machine)
446
507
  - Bun-only; run state in SQLite via `bun:sqlite` (WAL mode)
508
+ - Per-client hooks alongside MCP: Stop hook on Claude Code + Cursor · read/re-read/Edit hints on
509
+ Claude Code + OpenCode · commit hint on OpenCode — [What runs where](#what-runs-where)
447
510
 
448
511
  **Not supported (yet):**
512
+ - PreToolUse hints on Cursor, ZCode or Codex — Cursor has no such hook and the other two expose no
513
+ in-process hook surface for read/edit hints (Codex's `apply_patch` sends patch text, not file paths)
449
514
  - A hosted or shared ledger for a team — `runs.worktree` is the only sharing key today, and it's a
450
515
  path, not an identity. If you want to try pointing two machines at the same ledger anyway,
451
516
  `FAPONY_STATE_DIR` can be set to a synced folder (Syncthing, a shared drive) — but SQLite's WAL
package/fapony.ts CHANGED
@@ -3,15 +3,18 @@
3
3
  // fapony — measure/verify MCP server for coding agents
4
4
  // CLI dispatch: all logic lives in src/
5
5
 
6
+ import { existsSync } from "node:fs";
6
7
  import { cmdAnalyze } from "./src/analyze.js";
7
8
  import { cmdDebt } from "./src/debt.js";
8
9
  import { cmdDigest } from "./src/digest/cli.js";
9
- import { cmdHookReadHint, cmdHookStop } from "./src/hook.js";
10
+ import { cmdHookEditHint, cmdHookReadHint, cmdHookStop } from "./src/hook.js";
10
11
  import { cmdInit } from "./src/init.js";
11
12
  import { cmdInitMem } from "./src/init-mem.js";
12
13
  import { cmdInstall } from "./src/install.js";
13
14
  import { cmdLintBaseline } from "./src/lint-baseline.js";
14
15
  import { cmdMcp } from "./src/mcp/transport.js";
16
+ import { cmdMem } from "./src/mem/index.js";
17
+ import { initStore } from "./src/mem/store.js";
15
18
  import { cmdPlanSeed } from "./src/plan-seed.js";
16
19
  import { cmdPriceScan } from "./src/price/index.js";
17
20
  import { cmdReport, cmdReportWeb } from "./src/report/index.js";
@@ -42,6 +45,34 @@ if (cmd === "analyze") {
42
45
  await cmdTelemetry(a);
43
46
  } else if (cmd === "init-mem") {
44
47
  cmdInitMem(a);
48
+ } else if (cmd === "mem") {
49
+ // `--mem-dir <path>` is global to `mem` and must reach both the writer
50
+ // (initStore) and the resolver behind `mem where` — parse it once and thread
51
+ // it through, never strip it and forget.
52
+ const memDirIdx = a.indexOf("--mem-dir");
53
+ let overrideMemDir: string | undefined;
54
+ let rest = a;
55
+ if (memDirIdx !== -1) {
56
+ overrideMemDir = a[memDirIdx + 1];
57
+ if (!overrideMemDir || overrideMemDir.startsWith("--")) {
58
+ console.error("fapony mem: --mem-dir needs a value");
59
+ process.exit(1);
60
+ }
61
+ if (!existsSync(overrideMemDir)) {
62
+ console.error(
63
+ `fapony mem: --mem-dir path does not exist: ${overrideMemDir}`,
64
+ );
65
+ process.exit(1);
66
+ }
67
+ rest = a.filter((_, i) => i !== memDirIdx && i !== memDirIdx + 1);
68
+ }
69
+ initStore(process.cwd(), overrideMemDir);
70
+ try {
71
+ await cmdMem(rest, overrideMemDir);
72
+ } catch (e) {
73
+ console.error(`fapony mem: ${e instanceof Error ? e.message : String(e)}`);
74
+ process.exit(1);
75
+ }
45
76
  } else if (cmd === "init") {
46
77
  await cmdInit(a);
47
78
  } else if (cmd === "install") {
@@ -54,6 +85,8 @@ if (cmd === "analyze") {
54
85
  await cmdHookStop();
55
86
  } else if (cmd === "hook-read-hint") {
56
87
  await cmdHookReadHint();
88
+ } else if (cmd === "hook-edit-hint") {
89
+ await cmdHookEditHint();
57
90
  } else if (cmd === "mcp") {
58
91
  cmdMcp();
59
92
  } else if (cmd === "report") {
@@ -75,7 +108,7 @@ if (cmd === "analyze") {
75
108
  } else {
76
109
  console.error(`fapony: unknown command "${cmd ?? ""}"`);
77
110
  console.error(
78
- "usage: fapony <setup|update|stats|telemetry|init|init-mem|install|report|report-web|usage-scan|usage-web|price-scan|analyze|debt|lint-baseline|plan-seed|review-seed|digest|mcp|hook-stop|hook-read-hint|test> [args]",
111
+ "usage: fapony <setup|update|stats|telemetry|init|init-mem|mem|install|report|report-web|usage-scan|usage-web|price-scan|analyze|debt|lint-baseline|plan-seed|review-seed|digest|mcp|hook-stop|hook-read-hint|hook-edit-hint|test> [args]",
79
112
  );
80
113
  process.exit(1);
81
114
  }
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "fapony",
3
- "version": "0.1.3",
4
- "description": "Measurement layer for coding agents — measure what agents do, verify what they claim. 6 MCP tools, any agent, no loop required",
3
+ "version": "0.2.1",
4
+ "description": "Measurement layer for coding agents \u2014 measure what agents do, verify what they claim. 4 MCP tools, any agent, no loop required",
5
5
  "license": "MIT",
6
6
  "author": "delamind (https://github.com/kire21b)",
7
7
  "homepage": "https://github.com/kire21b/fapony#readme",
@@ -32,6 +32,7 @@
32
32
  "typecheck": "tsc --noEmit",
33
33
  "test": "bun fapony.ts test",
34
34
  "test:fast": "SKIP_SLOW=1 bun fapony.ts test",
35
+ "test:one": "bun scripts/test-one.ts",
35
36
  "check": "bun run lint && bun run typecheck && bun fapony.ts test",
36
37
  "prepublishOnly": "bash scripts/smoke-publish.sh",
37
38
  "overview": "bun fapony.ts report-web /tmp/fapony-overview.html && open /tmp/fapony-overview.html"
@@ -36,28 +36,19 @@ You are about to move a PLAN that has been shipped to the archive.
36
36
 
37
37
  A plan that is merely *waiting* (on a person, a customer, a decision) is **not** dead and does
38
38
  not move — mark it `status: blocked` + `blocked_by: <what you are waiting for>` and leave it in
39
- `plan/`, where `plan_list` will report it as blocked instead of as backlog.
39
+ `plan/` — the frontmatter is for the next person reading the folder, and the plan stays out
40
+ of `done/`, which is what `plan-sweep` and `kickoff` go by.
40
41
 
41
- 2. **Check inbound links, then `git mv`** — `.fapony/done/` sits *beside* `.fapony/plan/`, at the
42
- same depth, so every relative link *inside* the plan (`../spec/SPEC-x.md`, `../../src/...`)
43
- keeps working untouched. Nothing to normalize. What does change is how *other plans* reach
44
- this one — a sibling reference becomes a `done/` one:
42
+ 2. **Run `plan-sweep --apply`** — this does the `git mv`, rewrites markdown links inside the
43
+ file and inbound links from other plan files, warns about plain-text mentions, and logs
44
+ a decision row — all in one call:
45
45
  ```bash
46
- grep -rln 'PLAN-foo.md' .fapony/plan/ .fapony/spec/ docs/ # who points at it
47
- git mv .fapony/plan/PLAN-foo.md .fapony/done/PLAN-foo.md # same name, same depth
48
- # in .fapony/plan/*.md: (PLAN-foo.md) -> (../done/PLAN-foo.md)
46
+ fapony mem plan-sweep <PLAN-foo.md> --apply
49
47
  ```
50
- Fewer than 5 inbound files → fix them yourself · more → report the list.
51
-
52
- **The filename gets no date prefix.** The ship date is already in the header (step 1), and
53
- duplicating it into the name buys a sortable `ls` at the price of rewriting every inbound link
54
- on every ship, forever. "What shipped on which day" is a question to derive, not to store:
55
- ```bash
56
- grep -h 'shipped' .fapony/done/*.md | sort
57
- ```
58
- If git refuses ("not under version control" — `.fapony/` is gitignored in
59
- this repo), plain `mv` instead; there's nothing to commit for an untracked path, so skip
60
- step 4 in that case.
48
+ It refuses if the file lacks a shipped header or has open mem rows (next/bug/hold/decision/note).
49
+ If git refuses ("not under version control" — `.fapony/` is gitignored in this repo), plain
50
+ `mv` instead; there's nothing to commit for an untracked path, so skip step 4 in that case.
51
+ The filename gets no date prefix — the ship date is already in the header (step 1).
61
52
 
62
53
  3. **Leave the spec where it is** — `.fapony/spec/` is a reference library, not a queue. A spec
63
54
  answers "how does this work", which is asked long after the plan that ordered it shipped, and
@@ -80,15 +71,14 @@ You are about to move a PLAN that has been shipped to the archive.
80
71
  `missing_test` / `scope_mismatch` / `unsafe_command` / `spec_gap` / `incomplete`, and
81
72
  `other` (with a `note`, which it requires) only when a real finding fits none of them
82
73
  - `note`: **omit it on a clean ship.** A verdict with no note still counts toward the plan
83
- history future drafts read ("passed round 1 before"), but only notes reach the three
84
- free-text slots `project_health_context` shows — so "clean ship" evicts a note that would
85
- have taught the next session something. Write one only when this plan hit something a
74
+ history future drafts read ("passed round 1 before"), but only notes carry prose forward — so
75
+ "clean ship" evicts a note that would have taught the next session something. Write one only when this plan hit something a
86
76
  reader could not get from the diff: what the symptom looked like, where the cause actually
87
77
  was, and the rule that follows. Standalone prose — it is read months later with no access
88
78
  to this conversation.
89
79
  - `worktree`: **absolute path** to this repo/worktree (`git rev-parse --show-toplevel`) —
90
- every other fapony tool (`fapony_usage`, `fapony_stats`, `project_health_context`)
91
- scopes by absolute path too; a bare repo name won't match those queries
80
+ every other fapony tool and query scopes by
81
+ absolute path too; a bare repo name won't match them
92
82
  - `plan`: the archived plan's path (post-move, e.g. `.fapony/done/PLAN-foo.md`)
93
83
  - `files`: repo-relative paths this plan touched (`git diff --name-only <base>..HEAD`) —
94
84
  the only input to per-file risk history; without it the verdict says something happened
@@ -101,8 +91,8 @@ You are about to move a PLAN that has been shipped to the archive.
101
91
  Input: .fapony/plan/PLAN-kickoff.md, no shipped header yet
102
92
  Steps:
103
93
  1. stamp header: > ✅ **shipped 2026-09-13** (a1b2c3)
104
- 2. inbound: README.md, .fapony/plan/PLAN-loop.md → (PLAN-kickoff.md) becomes (../done/PLAN-kickoff.md)
105
- git mv .fapony/plan/PLAN-kickoff.md .fapony/done/PLAN-kickoff.md
94
+ 2. fapony mem plan-sweep .fapony/plan/PLAN-kickoff.md --apply
95
+ → moved, links rewritten, decision logged
106
96
  3. spec: untouched, stays in .fapony/spec/
107
97
  4. commit
108
98
  5. verdict_submit(verdict="pass", reason_code="none", regime="code", worktree="/Users/you/Project/fapony/wt-fapony", plan=".fapony/done/PLAN-kickoff.md", files=["src/kickoff.ts"])
@@ -122,5 +112,5 @@ A ship worth a note looks like this instead:
122
112
 
123
113
  - No git repo / no commits (can't derive a shipped hash) → tell user: "Add header > ✅ **shipped** (<hash>) first"
124
114
  - Stamped the header yourself → always say which hash you used
125
- - Link normalize fails → report which paths normalized wrong
126
- - Too many inbound links → report full list, don't fix yourself
115
+ - plan-sweep refuses (open mem rows) → close them or use `MEM_FORCE=1`
116
+ - Too many inbound links → plan-sweep reports them; too many to fix → report the list
@@ -106,6 +106,11 @@ So the seed buys you structure; the draft budget goes on judgment:
106
106
  - **Run the Phase −1 commands for facts** when the idea needs them, and put the numbers in the
107
107
  section they answer — a number you measured beats a number the seed guessed at.
108
108
  - **Signatures live in the SPEC chunks only.** Never paste them into plan §7 — link to the spec.
109
+ - **Does this zone already owe a convention?** If the feature touches a directory, run
110
+ `fapony debt` once and read only the entries whose files overlap it. An open migration
111
+ ("36 files still throw raw errors") is a constraint for §4, not a side quest — a plan that
112
+ adds the 37th is how the debt got there. Nothing overlaps, or no `conventions.json`? Say
113
+ nothing and move on.
109
114
  - If the CLI is missing, skip silently and draft from scratch (Phase 2 as written) — never block
110
115
  on a missing tool.
111
116
 
@@ -140,8 +145,8 @@ normal — writing to the default there scatters plans into a directory nobody r
140
145
  by hand, this check is yours.)
141
146
 
142
147
  **Editing a plan someone is executing right now is a different job from drafting one.** Ask the
143
- dev, or call `plan_list` — it joins plan files against run history, so a plan with an open run is
144
- one an agent is working from this minute. When that is the case:
148
+ dev, or run `fapony mem kickoff` — it reads the same plan files and names the first unchecked
149
+ chunk, so a plan already in flight is the one you are about to edit under someone. When that is the case:
145
150
 
146
151
  - **Anything you add is an instruction, not a note.** A measured fact parked under "don't do"
147
152
  still reads as a to-do to an agent mid-execution — the numbers are what make it tempting.
@@ -180,7 +185,7 @@ only place that ordering stays true.
180
185
 
181
186
  **The TL;DR is 15 lines, hard cap, and is the only part that changes while the work is in flight**
182
187
  (tick a box, stamp a short sha). Everything below it is the agreement. A TL;DR allowed to grow
183
- becomes a second copy of the plan, and then neither copy can be trusted. `plan_list` tallies the
188
+ becomes a second copy of the plan, and then neither copy can be trusted. `fapony mem kickoff` reads the
184
189
  checkboxes in the **first `##` section only**, so section 6 stays detail rather than status.
185
190
 
186
191
  Section 6 — every step must be verifiable. Section 8 — must link back to anything it came from.