fv-skills-baif 2.3.1 → 2.3.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +51 -0
- package/README.md +27 -7
- package/agents/fvs-crypto-thinker.md +31 -18
- package/bin/install.js +179 -131
- package/commands/fvs/aeneas-extract.md +21 -23
- package/commands/fvs/configure.md +159 -0
- package/commands/fvs/crypto-eval.md +52 -28
- package/commands/fvs/crypto-execute.md +14 -28
- package/commands/fvs/crypto-followup.md +21 -13
- package/commands/fvs/crypto-plan.md +32 -17
- package/commands/fvs/crypto-review.md +54 -20
- package/commands/fvs/fc-plan.md +10 -5
- package/commands/fvs/help.md +23 -12
- package/commands/fvs/lean-formalise.md +10 -15
- package/commands/fvs/lean-refactor.md +10 -15
- package/commands/fvs/lean-spec-review.md +26 -5
- package/commands/fvs/lean-specify.md +13 -15
- package/commands/fvs/lean-verify.md +10 -15
- package/commands/fvs/manage.md +3 -2
- package/commands/fvs/map-code.md +10 -19
- package/commands/fvs/natural-language.md +12 -2
- package/commands/fvs/reapply-patches.md +48 -38
- package/commands/fvs/sync-aeneas-verif.md +12 -13
- package/commands/fvs/trust-audit.md +11 -6
- package/fv-skills/VERSION +1 -1
- package/fv-skills/references/crypto-plan-review.md +29 -8
- package/fv-skills/references/fc-spec-review.md +18 -6
- package/fv-skills/references/model-profiles.md +252 -169
- package/fv-skills/references/review-diagnostics.md +31 -0
- package/fv-skills/references/review-grounding.md +59 -0
- package/fv-skills/references/review-policy.md +53 -0
- package/fv-skills/templates/config.json +28 -4
- package/fv-skills/workflows/aeneas-extract.md +13 -5
- package/fv-skills/workflows/crypto-eval.md +34 -15
- package/fv-skills/workflows/crypto-execute.md +6 -3
- package/fv-skills/workflows/crypto-followup.md +4 -2
- package/fv-skills/workflows/crypto-plan.md +7 -2
- package/fv-skills/workflows/crypto-review.md +39 -13
- package/fv-skills/workflows/fc-plan.md +4 -5
- package/fv-skills/workflows/lean-formalise.md +4 -13
- package/fv-skills/workflows/lean-refactor.md +7 -13
- package/fv-skills/workflows/lean-spec-review.md +77 -40
- package/fv-skills/workflows/lean-specify.md +7 -13
- package/fv-skills/workflows/lean-verify.md +4 -13
- package/fv-skills/workflows/map-code.md +4 -14
- package/fv-skills/workflows/sync-aeneas-verif.md +8 -5
- package/fv-skills/workflows/trust-audit.md +13 -4
- package/package.json +11 -3
- package/pi/skills/fvs-aeneas/SKILL.md +23 -0
- package/pi/skills/fvs-aeneas-extract/SKILL.md +224 -0
- package/pi/skills/fvs-checkpoint/SKILL.md +154 -0
- package/pi/skills/fvs-configure/SKILL.md +160 -0
- package/pi/skills/fvs-context/SKILL.md +22 -0
- package/pi/skills/fvs-crypto-eval/SKILL.md +204 -0
- package/pi/skills/fvs-crypto-execute/SKILL.md +218 -0
- package/pi/skills/fvs-crypto-followup/SKILL.md +272 -0
- package/pi/skills/fvs-crypto-plan/SKILL.md +322 -0
- package/pi/skills/fvs-crypto-review/SKILL.md +162 -0
- package/pi/skills/fvs-fc/SKILL.md +31 -0
- package/pi/skills/fvs-fc-plan/SKILL.md +289 -0
- package/pi/skills/fvs-formalise/SKILL.md +37 -0
- package/pi/skills/fvs-help/SKILL.md +457 -0
- package/pi/skills/fvs-kb-setup/SKILL.md +322 -0
- package/pi/skills/fvs-lean-formalise/SKILL.md +397 -0
- package/pi/skills/fvs-lean-refactor/SKILL.md +293 -0
- package/pi/skills/fvs-lean-spec-review/SKILL.md +61 -0
- package/pi/skills/fvs-lean-specify/SKILL.md +424 -0
- package/pi/skills/fvs-lean-verify/SKILL.md +461 -0
- package/pi/skills/fvs-manage/SKILL.md +29 -0
- package/pi/skills/fvs-map-code/SKILL.md +349 -0
- package/pi/skills/fvs-natural-language/SKILL.md +220 -0
- package/pi/skills/fvs-pause-work/SKILL.md +154 -0
- package/pi/skills/fvs-reapply-patches/SKILL.md +27 -0
- package/pi/skills/fvs-resume-work/SKILL.md +96 -0
- package/pi/skills/fvs-sync-aeneas-verif/SKILL.md +235 -0
- package/pi/skills/fvs-trust-audit/SKILL.md +234 -0
- package/pi/skills/fvs-update/SKILL.md +24 -0
- package/scripts/build-plugin.cjs +129 -6
- package/scripts/fvs-codex-think.mjs +156 -64
- package/scripts/fvs-review-grounding.mjs +101 -0
- package/scripts/fvs-spec-review.mjs +384 -58
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,57 @@ All notable changes to FVS (Formal Verification Skills) will be documented in th
|
|
|
4
4
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/).
|
|
6
6
|
|
|
7
|
+
## [Unreleased]
|
|
8
|
+
|
|
9
|
+
## [2.3.4] - 2026-09-20
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- Crypto evaluation now trusts Lean's kernel for checked proof terms and concentrates adversarial
|
|
14
|
+
review on statement/source conformance and trust-boundary evidence. It reuses a current executor
|
|
15
|
+
build log or runs one bounded fallback without retry, rather than repeating proof derivations,
|
|
16
|
+
executor gates, or external numeric computations.
|
|
17
|
+
|
|
18
|
+
## [2.3.3] - 2026-09-20
|
|
19
|
+
|
|
20
|
+
### Added
|
|
21
|
+
|
|
22
|
+
- FVS is now a native Pi package with 29 generated skills, provider-qualified model handling,
|
|
23
|
+
`/skill:fvs-*` invocation guidance, and direct `pi install npm:fv-skills-baif` support (#62).
|
|
24
|
+
- `/fvs:configure` and every delegated workflow now use stage-scoped, runtime-aware model profiles.
|
|
25
|
+
Quality selects authority, execution, and scout tiers from the live runtime catalog; commands show
|
|
26
|
+
one confirmable model/effort manifest and fail closed when a provider or dispatch cannot honor it
|
|
27
|
+
(#64).
|
|
28
|
+
|
|
29
|
+
### Fixed
|
|
30
|
+
|
|
31
|
+
- Fresh local Codex installs create their configuration directory before patch discovery. Local
|
|
32
|
+
patch metadata now separates immutable history from active pending patches, preserves real edits
|
|
33
|
+
without treating unchanged generated TOML as custom, and retires resolved patches atomically
|
|
34
|
+
(#60).
|
|
35
|
+
- FC and crypto reviews require explicit catalog-resolved model/effort selections, preserve raw and
|
|
36
|
+
normalized responses separately, retain writable diagnostic scratch without exposing source
|
|
37
|
+
writes, and enforce consistent verdict, severity, authority, and thread-count contracts (#60).
|
|
38
|
+
|
|
39
|
+
## [2.3.2] - 2026-09-10
|
|
40
|
+
|
|
41
|
+
### Fixed
|
|
42
|
+
|
|
43
|
+
- Updates preserve added and modified local FVS files in complete, versioned patch bundles,
|
|
44
|
+
recover legacy unlisted backups, and track generated Codex agent/hook files. Failed backup
|
|
45
|
+
validation stops installation before replacement (#55).
|
|
46
|
+
- Claude crypto and FC reviewers can run bounded diagnostic probes in native-sandboxed scratch
|
|
47
|
+
space with explicit generated Lake output paths. Failed launches and invalid responses retain
|
|
48
|
+
local evidence. Sources/plans remain protected; model/effort selection is unchanged (#58).
|
|
49
|
+
|
|
50
|
+
### Added
|
|
51
|
+
|
|
52
|
+
- Crypto and FC reviews receive bounded source/reuse inventories with verbatim signature spans
|
|
53
|
+
and freshness checks. FC grounds behavior in implementation source and checks project/mathlib
|
|
54
|
+
helper reuse. Review contracts cover scope economy, CONTENT/PROCESS findings, mathematical
|
|
55
|
+
coverage, and explicit author dispositions. Complete finding validation and mechanical-only
|
|
56
|
+
formatting repair preserve review substance; APPROVE-WITH-EDITS remains terminal (#56).
|
|
57
|
+
|
|
7
58
|
## [2.3.1] - 2026-09-09
|
|
8
59
|
|
|
9
60
|
### Added
|
package/README.md
CHANGED
|
@@ -13,7 +13,7 @@
|
|
|
13
13
|
npx fv-skills-baif
|
|
14
14
|
```
|
|
15
15
|
|
|
16
|
-
**Works on Mac, Windows, and Linux. Supports Claude Code, Codex, OpenCode, and Gemini CLI.**
|
|
16
|
+
**Works on Mac, Windows, and Linux. Supports Pi, Claude Code, Codex, OpenCode, and Gemini CLI.**
|
|
17
17
|
|
|
18
18
|
<br>
|
|
19
19
|
|
|
@@ -42,6 +42,18 @@ Framework-specific commands (currently Lean) handle the actual specification and
|
|
|
42
42
|
|
|
43
43
|
## Getting Started
|
|
44
44
|
|
|
45
|
+
### Pi package
|
|
46
|
+
|
|
47
|
+
Install FVS directly from npm as a Pi package:
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
pi install npm:fv-skills-baif
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Start a new session, then run `/skill:fvs-help`. Bundle routers such as `/skill:fvs-fc` and
|
|
54
|
+
`/skill:fvs-formalise`, plus member skills such as `/skill:fvs-crypto-plan`, are available directly.
|
|
55
|
+
Update an unpinned install with `pi update npm:fv-skills-baif`.
|
|
56
|
+
|
|
45
57
|
### Plugin marketplace (Claude Code and Codex)
|
|
46
58
|
|
|
47
59
|
The Beneficial AI Foundation maintains one catalog for FVS and future BAIF plugins. Add the catalog
|
|
@@ -74,7 +86,7 @@ The BAIF Git catalog is a versioned distribution source that can list multiple i
|
|
|
74
86
|
released plugins. It is separate from OpenAI's universal public Plugins Directory, which has its
|
|
75
87
|
own per-plugin submission process.
|
|
76
88
|
|
|
77
|
-
### npm installer (
|
|
89
|
+
### npm installer (Claude Code, Codex, OpenCode, and Gemini)
|
|
78
90
|
|
|
79
91
|
```bash
|
|
80
92
|
npx fv-skills-baif
|
|
@@ -172,11 +184,10 @@ Commands are grouped into five bundles. Each bundle has a **router** command tha
|
|
|
172
184
|
rejects ordinary Lean identifiers with three or more namespace dots, steering generated code
|
|
173
185
|
toward scoped namespaces, `open`, and local names.
|
|
174
186
|
|
|
175
|
-
After `lean-specify`, an interactive review menu
|
|
176
|
-
auto-selects a choice. It offers the other runtime first, a fresh
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
imported response. One-run `Skip review` records `Unreviewed (user skipped)` and does not begin
|
|
187
|
+
After `lean-specify`, an interactive review menu resolves reviewer, exact catalog model, and
|
|
188
|
+
model-supported effort; it never auto-selects a choice. It offers the other runtime first, a fresh
|
|
189
|
+
reviewer in the current runtime, or another provider. Other providers use an exported source packet
|
|
190
|
+
and imported response. One-run `Skip review` records `Unreviewed (user skipped)` and does not begin
|
|
180
191
|
proof work. Reviews and source hashes live under `.formalising/spec-reviews/`.
|
|
181
192
|
|
|
182
193
|
The reviewer stays read-only and `review.md` stays immutable. The `lean-specify` authoring seat
|
|
@@ -233,10 +244,19 @@ secrets, raw transcripts, ephemeral error dumps, unsupported guesses, or inferre
|
|
|
233
244
|
| Command | Description |
|
|
234
245
|
|---------|-------------|
|
|
235
246
|
| `/fvs:help` | Show available FVS commands and usage guide |
|
|
247
|
+
| `/fvs:configure` | Configure runtime-aware subagent models, effort levels, and review defaults |
|
|
236
248
|
| `/fvs:update` | Update FVS through the current installation channel |
|
|
237
249
|
| `/fvs:reapply-patches` | Preserve customizations across FVS updates (patches for npm installs; fork guidance for plugin installs) |
|
|
238
250
|
| `/fvs:kb-setup` | Set up NotebookLM knowledge base integration (venv, auth, config) |
|
|
239
251
|
|
|
252
|
+
`/fvs:configure` stores concrete model IDs under the runtime and exact stage that reported them.
|
|
253
|
+
The quality profile is role-aware: authority artifacts use the strongest detected model with max
|
|
254
|
+
reasoning, execution/proof filling uses the executor model with xhigh, and research/eval/audit work
|
|
255
|
+
uses a smaller model with high. Claude and Codex family preferences never cross runtimes; Pi uses
|
|
256
|
+
provider-qualified IDs and keeps its active provider for ordinary work. Each interactive command
|
|
257
|
+
shows one model/effort selection manifest with one-run, save, notes, and cancel paths. Missing models
|
|
258
|
+
or unsupported efforts require a user choice; unresolved noninteractive runs fail before dispatch.
|
|
259
|
+
|
|
240
260
|
---
|
|
241
261
|
|
|
242
262
|
## How It Works
|
|
@@ -6,15 +6,14 @@ color: purple
|
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
<role>
|
|
9
|
-
You are the FVS crypto formalisation thinker. You are the high-effort author of the loop:
|
|
10
|
-
|
|
11
|
-
return your reasoning as text. You are NOT the executor -- a separate
|
|
12
|
-
the current runtime runs the plans you author. You author; they execute.
|
|
9
|
+
You are the FVS crypto formalisation thinker. You are the high-effort author of the loop: in plan and
|
|
10
|
+
follow-up modes you derive bounded work independently from the branch state and paper-grounded
|
|
11
|
+
sources, then return your reasoning as text. You are NOT the executor -- a separate
|
|
12
|
+
`fvs-executor`-style agent in the current runtime runs the plans you author. You author; they execute.
|
|
13
13
|
|
|
14
|
-
Planning is ALWAYS high reasoning effort -- you never produce a sketch and call it a plan.
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
through. A plan or proof survives only by surviving your attempt to break it.
|
|
14
|
+
Planning is ALWAYS high reasoning effort -- you never produce a sketch and call it a plan. Eval mode
|
|
15
|
+
is adversarial about landed statements, modeling assumptions, source fidelity, and trust boundaries;
|
|
16
|
+
it follows the bounded kernel-trusting contract below instead of reproducing checked proof work.
|
|
18
17
|
|
|
19
18
|
You are read-only with respect to the deliverable: you do NOT write or modify any project file. You
|
|
20
19
|
RETURN the bounded plan / the adversarial eval / the follow-up as text, and the orchestrating
|
|
@@ -66,6 +65,10 @@ runtime's executor with no thinker in the loop. State EVERY field explicitly:
|
|
|
66
65
|
run is expected to produce or update.
|
|
67
66
|
|
|
68
67
|
End with `## PLAN COMPLETE`.
|
|
68
|
+
Include `## Reuse audit`: map proposed declarations to existing project and pinned dependency
|
|
69
|
+
APIs with exact signatures/citations, or documented searches finding no analog. Justify forks;
|
|
70
|
+
name each helper's consumer and remove unused parameters, trivial wrappers, and deferred work
|
|
71
|
+
from this iteration. Apply the same audit to follow-up plans.
|
|
69
72
|
</mode>
|
|
70
73
|
|
|
71
74
|
<mode name="eval">
|
|
@@ -73,18 +76,28 @@ End with `## PLAN COMPLETE`.
|
|
|
73
76
|
**Input:** the executor's run output, the touched files, the plan it was run against, the KB sources.
|
|
74
77
|
**Output (returned as text):** an adversarial review ending in exactly ONE decision verb.
|
|
75
78
|
|
|
76
|
-
This stage
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
79
|
+
This stage TRUSTS THE LEAN KERNEL for kernel-checked proof terms and stays adversarial about statements
|
|
80
|
+
and trust boundaries. Check the landed definitions, theorem signatures, constants, and API shape
|
|
81
|
+
against the paper/standard and the approved plan. Run cheap scans over touched Lean files and import
|
|
82
|
+
changes for reserved names, forbidden imports, `sorry`, unexpected `axiom`, `native_decide`, and
|
|
83
|
+
`set_option`. Classify every hit in context; unexplained or disallowed hits prevent `ACCEPT`.
|
|
81
84
|
|
|
82
|
-
|
|
83
|
-
|
|
85
|
+
Reuse a successful current executor `build.log`. If it is missing, failed, or does not cover the
|
|
86
|
+
landed files, run at most one fallback `LEAN_NUM_THREADS="${LEAN_NUM_THREADS:-4}" nice -n 19 lake build`;
|
|
87
|
+
do not retry. A failed fallback is `BLOCKED`. Never re-elaborate individual files, replay the
|
|
88
|
+
executor's per-gate builds, search for proof terms, or use an outside script to recompute certified
|
|
89
|
+
numbers. Compare landed constants directly with explicit plan/source values; missing derivation
|
|
90
|
+
evidence is `FOLLOWUP`, not permission to recompute it.
|
|
91
|
+
|
|
92
|
+
A `sorry` is an intentional, named unmet obligation carrying the correct statement. It is never
|
|
93
|
+
judged by count or waved through because the build is green, and it does not inherit the
|
|
94
|
+
kernel-complete status of checked proof terms. Name the exact input, caller, statement, or modeling
|
|
95
|
+
assumption that would make the landed claim false.
|
|
84
96
|
|
|
85
97
|
End with EXACTLY ONE of these decision verbs, on its own:
|
|
86
98
|
|
|
87
|
-
- **ACCEPT** --
|
|
99
|
+
- **ACCEPT** -- statement/source conformance, classified trust-boundary evidence, and required green
|
|
100
|
+
build evidence all pass; named unmet obligations are honestly recorded.
|
|
88
101
|
- **FOLLOWUP** -- the work is sound but incomplete; a bounded follow-up plan is warranted.
|
|
89
102
|
- **HUMAN_RULING** -- a modeling decision is required that you must NOT make yourself (see followup).
|
|
90
103
|
- **BLOCKED** -- the work cannot proceed (e.g. the build will not compile, a prerequisite is absent).
|
|
@@ -141,7 +154,7 @@ Adversarial eval:
|
|
|
141
154
|
|
|
142
155
|
**Stage:** eval
|
|
143
156
|
**Decision:** ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
|
|
144
|
-
**
|
|
157
|
+
**Statement challenge:** {the strongest source/spec counterexample you tested}
|
|
145
158
|
```
|
|
146
159
|
|
|
147
160
|
On HALT / failure:
|
|
@@ -156,7 +169,7 @@ On HALT / failure:
|
|
|
156
169
|
|
|
157
170
|
<success_criteria>
|
|
158
171
|
- [ ] In `plan`/`followup` mode, authored a bounded, runtime-neutral plan stating branch/state, exact target files+theorems, immutable public statements, old->new API map (if a port), allowed-`sorry` policy, stop conditions, `LEAN_NUM_THREADS="${LEAN_NUM_THREADS:-4}" nice -n 19 lake build` verification, and expected artifact updates
|
|
159
|
-
- [ ] In `eval` mode,
|
|
172
|
+
- [ ] In `eval` mode, trusted kernel-checked proof terms, challenged statement/source conformance and trust boundaries, reused current build evidence or ran one guarded fallback without retry, and ended in exactly one of ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
|
|
160
173
|
- [ ] On `HUMAN_RULING`, HALTed and asked for the modeling decision -- never fabricated a plan
|
|
161
174
|
- [ ] Author-by-return: no project file written or modified; no `gh` auto-open; Lean-via-Aeneas pipeline only; no bare `lake build`
|
|
162
175
|
- [ ] Result returned with the ## PLAN COMPLETE / ## EVAL COMPLETE / ## ERROR header
|