fv-skills-baif 2.3.1 → 2.3.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (81) hide show
  1. package/CHANGELOG.md +51 -0
  2. package/README.md +27 -7
  3. package/agents/fvs-crypto-thinker.md +31 -18
  4. package/bin/install.js +179 -131
  5. package/commands/fvs/aeneas-extract.md +21 -23
  6. package/commands/fvs/configure.md +159 -0
  7. package/commands/fvs/crypto-eval.md +52 -28
  8. package/commands/fvs/crypto-execute.md +14 -28
  9. package/commands/fvs/crypto-followup.md +21 -13
  10. package/commands/fvs/crypto-plan.md +32 -17
  11. package/commands/fvs/crypto-review.md +54 -20
  12. package/commands/fvs/fc-plan.md +10 -5
  13. package/commands/fvs/help.md +23 -12
  14. package/commands/fvs/lean-formalise.md +10 -15
  15. package/commands/fvs/lean-refactor.md +10 -15
  16. package/commands/fvs/lean-spec-review.md +26 -5
  17. package/commands/fvs/lean-specify.md +13 -15
  18. package/commands/fvs/lean-verify.md +10 -15
  19. package/commands/fvs/manage.md +3 -2
  20. package/commands/fvs/map-code.md +10 -19
  21. package/commands/fvs/natural-language.md +12 -2
  22. package/commands/fvs/reapply-patches.md +48 -38
  23. package/commands/fvs/sync-aeneas-verif.md +12 -13
  24. package/commands/fvs/trust-audit.md +11 -6
  25. package/fv-skills/VERSION +1 -1
  26. package/fv-skills/references/crypto-plan-review.md +29 -8
  27. package/fv-skills/references/fc-spec-review.md +18 -6
  28. package/fv-skills/references/model-profiles.md +252 -169
  29. package/fv-skills/references/review-diagnostics.md +31 -0
  30. package/fv-skills/references/review-grounding.md +59 -0
  31. package/fv-skills/references/review-policy.md +53 -0
  32. package/fv-skills/templates/config.json +28 -4
  33. package/fv-skills/workflows/aeneas-extract.md +13 -5
  34. package/fv-skills/workflows/crypto-eval.md +34 -15
  35. package/fv-skills/workflows/crypto-execute.md +6 -3
  36. package/fv-skills/workflows/crypto-followup.md +4 -2
  37. package/fv-skills/workflows/crypto-plan.md +7 -2
  38. package/fv-skills/workflows/crypto-review.md +39 -13
  39. package/fv-skills/workflows/fc-plan.md +4 -5
  40. package/fv-skills/workflows/lean-formalise.md +4 -13
  41. package/fv-skills/workflows/lean-refactor.md +7 -13
  42. package/fv-skills/workflows/lean-spec-review.md +77 -40
  43. package/fv-skills/workflows/lean-specify.md +7 -13
  44. package/fv-skills/workflows/lean-verify.md +4 -13
  45. package/fv-skills/workflows/map-code.md +4 -14
  46. package/fv-skills/workflows/sync-aeneas-verif.md +8 -5
  47. package/fv-skills/workflows/trust-audit.md +13 -4
  48. package/package.json +11 -3
  49. package/pi/skills/fvs-aeneas/SKILL.md +23 -0
  50. package/pi/skills/fvs-aeneas-extract/SKILL.md +224 -0
  51. package/pi/skills/fvs-checkpoint/SKILL.md +154 -0
  52. package/pi/skills/fvs-configure/SKILL.md +160 -0
  53. package/pi/skills/fvs-context/SKILL.md +22 -0
  54. package/pi/skills/fvs-crypto-eval/SKILL.md +204 -0
  55. package/pi/skills/fvs-crypto-execute/SKILL.md +218 -0
  56. package/pi/skills/fvs-crypto-followup/SKILL.md +272 -0
  57. package/pi/skills/fvs-crypto-plan/SKILL.md +322 -0
  58. package/pi/skills/fvs-crypto-review/SKILL.md +162 -0
  59. package/pi/skills/fvs-fc/SKILL.md +31 -0
  60. package/pi/skills/fvs-fc-plan/SKILL.md +289 -0
  61. package/pi/skills/fvs-formalise/SKILL.md +37 -0
  62. package/pi/skills/fvs-help/SKILL.md +457 -0
  63. package/pi/skills/fvs-kb-setup/SKILL.md +322 -0
  64. package/pi/skills/fvs-lean-formalise/SKILL.md +397 -0
  65. package/pi/skills/fvs-lean-refactor/SKILL.md +293 -0
  66. package/pi/skills/fvs-lean-spec-review/SKILL.md +61 -0
  67. package/pi/skills/fvs-lean-specify/SKILL.md +424 -0
  68. package/pi/skills/fvs-lean-verify/SKILL.md +461 -0
  69. package/pi/skills/fvs-manage/SKILL.md +29 -0
  70. package/pi/skills/fvs-map-code/SKILL.md +349 -0
  71. package/pi/skills/fvs-natural-language/SKILL.md +220 -0
  72. package/pi/skills/fvs-pause-work/SKILL.md +154 -0
  73. package/pi/skills/fvs-reapply-patches/SKILL.md +27 -0
  74. package/pi/skills/fvs-resume-work/SKILL.md +96 -0
  75. package/pi/skills/fvs-sync-aeneas-verif/SKILL.md +235 -0
  76. package/pi/skills/fvs-trust-audit/SKILL.md +234 -0
  77. package/pi/skills/fvs-update/SKILL.md +24 -0
  78. package/scripts/build-plugin.cjs +129 -6
  79. package/scripts/fvs-codex-think.mjs +156 -64
  80. package/scripts/fvs-review-grounding.mjs +101 -0
  81. package/scripts/fvs-spec-review.mjs +384 -58
package/CHANGELOG.md CHANGED
@@ -4,6 +4,57 @@ All notable changes to FVS (Formal Verification Skills) will be documented in th
4
4
 
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/).
6
6
 
7
+ ## [Unreleased]
8
+
9
+ ## [2.3.4] - 2026-09-20
10
+
11
+ ### Fixed
12
+
13
+ - Crypto evaluation now trusts Lean's kernel for checked proof terms and concentrates adversarial
14
+ review on statement/source conformance and trust-boundary evidence. It reuses a current executor
15
+ build log or runs one bounded fallback without retry, rather than repeating proof derivations,
16
+ executor gates, or external numeric computations.
17
+
18
+ ## [2.3.3] - 2026-09-20
19
+
20
+ ### Added
21
+
22
+ - FVS is now a native Pi package with 29 generated skills, provider-qualified model handling,
23
+ `/skill:fvs-*` invocation guidance, and direct `pi install npm:fv-skills-baif` support (#62).
24
+ - `/fvs:configure` and every delegated workflow now use stage-scoped, runtime-aware model profiles.
25
+ Quality selects authority, execution, and scout tiers from the live runtime catalog; commands show
26
+ one confirmable model/effort manifest and fail closed when a provider or dispatch cannot honor it
27
+ (#64).
28
+
29
+ ### Fixed
30
+
31
+ - Fresh local Codex installs create their configuration directory before patch discovery. Local
32
+ patch metadata now separates immutable history from active pending patches, preserves real edits
33
+ without treating unchanged generated TOML as custom, and retires resolved patches atomically
34
+ (#60).
35
+ - FC and crypto reviews require explicit catalog-resolved model/effort selections, preserve raw and
36
+ normalized responses separately, retain writable diagnostic scratch without exposing source
37
+ writes, and enforce consistent verdict, severity, authority, and thread-count contracts (#60).
38
+
39
+ ## [2.3.2] - 2026-09-10
40
+
41
+ ### Fixed
42
+
43
+ - Updates preserve added and modified local FVS files in complete, versioned patch bundles,
44
+ recover legacy unlisted backups, and track generated Codex agent/hook files. Failed backup
45
+ validation stops installation before replacement (#55).
46
+ - Claude crypto and FC reviewers can run bounded diagnostic probes in native-sandboxed scratch
47
+ space with explicit generated Lake output paths. Failed launches and invalid responses retain
48
+ local evidence. Sources/plans remain protected; model/effort selection is unchanged (#58).
49
+
50
+ ### Added
51
+
52
+ - Crypto and FC reviews receive bounded source/reuse inventories with verbatim signature spans
53
+ and freshness checks. FC grounds behavior in implementation source and checks project/mathlib
54
+ helper reuse. Review contracts cover scope economy, CONTENT/PROCESS findings, mathematical
55
+ coverage, and explicit author dispositions. Complete finding validation and mechanical-only
56
+ formatting repair preserve review substance; APPROVE-WITH-EDITS remains terminal (#56).
57
+
7
58
  ## [2.3.1] - 2026-09-09
8
59
 
9
60
  ### Added
package/README.md CHANGED
@@ -13,7 +13,7 @@
13
13
  npx fv-skills-baif
14
14
  ```
15
15
 
16
- **Works on Mac, Windows, and Linux. Supports Claude Code, Codex, OpenCode, and Gemini CLI.**
16
+ **Works on Mac, Windows, and Linux. Supports Pi, Claude Code, Codex, OpenCode, and Gemini CLI.**
17
17
 
18
18
  <br>
19
19
 
@@ -42,6 +42,18 @@ Framework-specific commands (currently Lean) handle the actual specification and
42
42
 
43
43
  ## Getting Started
44
44
 
45
+ ### Pi package
46
+
47
+ Install FVS directly from npm as a Pi package:
48
+
49
+ ```bash
50
+ pi install npm:fv-skills-baif
51
+ ```
52
+
53
+ Start a new session, then run `/skill:fvs-help`. Bundle routers such as `/skill:fvs-fc` and
54
+ `/skill:fvs-formalise`, plus member skills such as `/skill:fvs-crypto-plan`, are available directly.
55
+ Update an unpinned install with `pi update npm:fv-skills-baif`.
56
+
45
57
  ### Plugin marketplace (Claude Code and Codex)
46
58
 
47
59
  The Beneficial AI Foundation maintains one catalog for FVS and future BAIF plugins. Add the catalog
@@ -74,7 +86,7 @@ The BAIF Git catalog is a versioned distribution source that can list multiple i
74
86
  released plugins. It is separate from OpenAI's universal public Plugins Directory, which has its
75
87
  own per-plugin submission process.
76
88
 
77
- ### npm installer (all runtimes)
89
+ ### npm installer (Claude Code, Codex, OpenCode, and Gemini)
78
90
 
79
91
  ```bash
80
92
  npx fv-skills-baif
@@ -172,11 +184,10 @@ Commands are grouped into five bundles. Each bundle has a **router** command tha
172
184
  rejects ordinary Lean identifiers with three or more namespace dots, steering generated code
173
185
  toward scoped namespaces, `open`, and local names.
174
186
 
175
- After `lean-specify`, an interactive review menu asks reviewer, then model, then effort; it never
176
- auto-selects a choice. It offers the other runtime first, a fresh reviewer in the current runtime,
177
- or another provider. Choose GPT Sol or Astra for Codex, Fable for Claude, or a cheaper/custom model;
178
- effort defaults to `max` and can be lowered. Other providers use an exported source packet and
179
- imported response. One-run `Skip review` records `Unreviewed (user skipped)` and does not begin
187
+ After `lean-specify`, an interactive review menu resolves reviewer, exact catalog model, and
188
+ model-supported effort; it never auto-selects a choice. It offers the other runtime first, a fresh
189
+ reviewer in the current runtime, or another provider. Other providers use an exported source packet
190
+ and imported response. One-run `Skip review` records `Unreviewed (user skipped)` and does not begin
180
191
  proof work. Reviews and source hashes live under `.formalising/spec-reviews/`.
181
192
 
182
193
  The reviewer stays read-only and `review.md` stays immutable. The `lean-specify` authoring seat
@@ -233,10 +244,19 @@ secrets, raw transcripts, ephemeral error dumps, unsupported guesses, or inferre
233
244
  | Command | Description |
234
245
  |---------|-------------|
235
246
  | `/fvs:help` | Show available FVS commands and usage guide |
247
+ | `/fvs:configure` | Configure runtime-aware subagent models, effort levels, and review defaults |
236
248
  | `/fvs:update` | Update FVS through the current installation channel |
237
249
  | `/fvs:reapply-patches` | Preserve customizations across FVS updates (patches for npm installs; fork guidance for plugin installs) |
238
250
  | `/fvs:kb-setup` | Set up NotebookLM knowledge base integration (venv, auth, config) |
239
251
 
252
+ `/fvs:configure` stores concrete model IDs under the runtime and exact stage that reported them.
253
+ The quality profile is role-aware: authority artifacts use the strongest detected model with max
254
+ reasoning, execution/proof filling uses the executor model with xhigh, and research/eval/audit work
255
+ uses a smaller model with high. Claude and Codex family preferences never cross runtimes; Pi uses
256
+ provider-qualified IDs and keeps its active provider for ordinary work. Each interactive command
257
+ shows one model/effort selection manifest with one-run, save, notes, and cancel paths. Missing models
258
+ or unsupported efforts require a user choice; unresolved noninteractive runs fail before dispatch.
259
+
240
260
  ---
241
261
 
242
262
  ## How It Works
@@ -6,15 +6,14 @@ color: purple
6
6
  ---
7
7
 
8
8
  <role>
9
- You are the FVS crypto formalisation thinker. You are the high-effort author of the loop: you
10
- re-derive everything independently, from the branch state and the paper-grounded sources, and you
11
- return your reasoning as text. You are NOT the executor -- a separate `fvs-executor`-style agent in
12
- the current runtime runs the plans you author. You author; they execute.
9
+ You are the FVS crypto formalisation thinker. You are the high-effort author of the loop: in plan and
10
+ follow-up modes you derive bounded work independently from the branch state and paper-grounded
11
+ sources, then return your reasoning as text. You are NOT the executor -- a separate
12
+ `fvs-executor`-style agent in the current runtime runs the plans you author. You author; they execute.
13
13
 
14
- Planning is ALWAYS high reasoning effort -- you never produce a sketch and call it a plan. The eval
15
- stage is ALWAYS adversarial: you take the posture of a reviewer who is actively trying to REFUTE the
16
- spec, the proof, and the stated assumptions, not one who is looking for a reason to wave them
17
- through. A plan or proof survives only by surviving your attempt to break it.
14
+ Planning is ALWAYS high reasoning effort -- you never produce a sketch and call it a plan. Eval mode
15
+ is adversarial about landed statements, modeling assumptions, source fidelity, and trust boundaries;
16
+ it follows the bounded kernel-trusting contract below instead of reproducing checked proof work.
18
17
 
19
18
  You are read-only with respect to the deliverable: you do NOT write or modify any project file. You
20
19
  RETURN the bounded plan / the adversarial eval / the follow-up as text, and the orchestrating
@@ -66,6 +65,10 @@ runtime's executor with no thinker in the loop. State EVERY field explicitly:
66
65
  run is expected to produce or update.
67
66
 
68
67
  End with `## PLAN COMPLETE`.
68
+ Include `## Reuse audit`: map proposed declarations to existing project and pinned dependency
69
+ APIs with exact signatures/citations, or documented searches finding no analog. Justify forks;
70
+ name each helper's consumer and remove unused parameters, trivial wrappers, and deferred work
71
+ from this iteration. Apply the same audit to follow-up plans.
69
72
  </mode>
70
73
 
71
74
  <mode name="eval">
@@ -73,18 +76,28 @@ End with `## PLAN COMPLETE`.
73
76
  **Input:** the executor's run output, the touched files, the plan it was run against, the KB sources.
74
77
  **Output (returned as text):** an adversarial review ending in exactly ONE decision verb.
75
78
 
76
- This stage is ALWAYS adversarial. Re-derive independently; do not echo the executor's reasoning.
77
- Actively try to REFUTE: does the spec actually capture the paper's claim? Does the proof close the
78
- goal it claims, or does it lean on an unstated assumption? Is every `sorry` a named obligation with
79
- the correct statement, or is it papering over a real gap? Name the exact input, caller, or modeling
80
- assumption that would make the argument FALSE.
79
+ This stage TRUSTS THE LEAN KERNEL for kernel-checked proof terms and stays adversarial about statements
80
+ and trust boundaries. Check the landed definitions, theorem signatures, constants, and API shape
81
+ against the paper/standard and the approved plan. Run cheap scans over touched Lean files and import
82
+ changes for reserved names, forbidden imports, `sorry`, unexpected `axiom`, `native_decide`, and
83
+ `set_option`. Classify every hit in context; unexplained or disallowed hits prevent `ACCEPT`.
81
84
 
82
- A `sorry` is acceptable ONLY as an intentional, named obligation carrying the correct statement --
83
- never judged by count, never waved through because "the build is green".
85
+ Reuse a successful current executor `build.log`. If it is missing, failed, or does not cover the
86
+ landed files, run at most one fallback `LEAN_NUM_THREADS="${LEAN_NUM_THREADS:-4}" nice -n 19 lake build`;
87
+ do not retry. A failed fallback is `BLOCKED`. Never re-elaborate individual files, replay the
88
+ executor's per-gate builds, search for proof terms, or use an outside script to recompute certified
89
+ numbers. Compare landed constants directly with explicit plan/source values; missing derivation
90
+ evidence is `FOLLOWUP`, not permission to recompute it.
91
+
92
+ A `sorry` is an intentional, named unmet obligation carrying the correct statement. It is never
93
+ judged by count or waved through because the build is green, and it does not inherit the
94
+ kernel-complete status of checked proof terms. Name the exact input, caller, statement, or modeling
95
+ assumption that would make the landed claim false.
84
96
 
85
97
  End with EXACTLY ONE of these decision verbs, on its own:
86
98
 
87
- - **ACCEPT** -- the spec/proof survives the adversarial pass; the obligations are honest.
99
+ - **ACCEPT** -- statement/source conformance, classified trust-boundary evidence, and required green
100
+ build evidence all pass; named unmet obligations are honestly recorded.
88
101
  - **FOLLOWUP** -- the work is sound but incomplete; a bounded follow-up plan is warranted.
89
102
  - **HUMAN_RULING** -- a modeling decision is required that you must NOT make yourself (see followup).
90
103
  - **BLOCKED** -- the work cannot proceed (e.g. the build will not compile, a prerequisite is absent).
@@ -141,7 +154,7 @@ Adversarial eval:
141
154
 
142
155
  **Stage:** eval
143
156
  **Decision:** ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
144
- **Refutation attempted:** {the strongest counter you raised}
157
+ **Statement challenge:** {the strongest source/spec counterexample you tested}
145
158
  ```
146
159
 
147
160
  On HALT / failure:
@@ -156,7 +169,7 @@ On HALT / failure:
156
169
 
157
170
  <success_criteria>
158
171
  - [ ] In `plan`/`followup` mode, authored a bounded, runtime-neutral plan stating branch/state, exact target files+theorems, immutable public statements, old->new API map (if a port), allowed-`sorry` policy, stop conditions, `LEAN_NUM_THREADS="${LEAN_NUM_THREADS:-4}" nice -n 19 lake build` verification, and expected artifact updates
159
- - [ ] In `eval` mode, took an adversarial posture (tried to refute), judged each `sorry` as a named obligation not by count, and ended in exactly one of ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
172
+ - [ ] In `eval` mode, trusted kernel-checked proof terms, challenged statement/source conformance and trust boundaries, reused current build evidence or ran one guarded fallback without retry, and ended in exactly one of ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
160
173
  - [ ] On `HUMAN_RULING`, HALTed and asked for the modeling decision -- never fabricated a plan
161
174
  - [ ] Author-by-return: no project file written or modified; no `gh` auto-open; Lean-via-Aeneas pipeline only; no bare `lake build`
162
175
  - [ ] Result returned with the ## PLAN COMPLETE / ## EVAL COMPLETE / ## ERROR header