fv-skills-baif 2.3.2 → 2.3.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (78) hide show
  1. package/CHANGELOG.md +30 -0
  2. package/README.md +27 -7
  3. package/agents/fvs-crypto-thinker.md +27 -18
  4. package/bin/install.js +64 -35
  5. package/commands/fvs/aeneas-extract.md +21 -23
  6. package/commands/fvs/configure.md +159 -0
  7. package/commands/fvs/crypto-eval.md +52 -28
  8. package/commands/fvs/crypto-execute.md +14 -28
  9. package/commands/fvs/crypto-followup.md +21 -13
  10. package/commands/fvs/crypto-plan.md +26 -16
  11. package/commands/fvs/crypto-review.md +44 -23
  12. package/commands/fvs/fc-plan.md +10 -5
  13. package/commands/fvs/help.md +23 -12
  14. package/commands/fvs/lean-formalise.md +10 -15
  15. package/commands/fvs/lean-refactor.md +10 -15
  16. package/commands/fvs/lean-spec-review.md +26 -5
  17. package/commands/fvs/lean-specify.md +10 -15
  18. package/commands/fvs/lean-verify.md +10 -15
  19. package/commands/fvs/manage.md +3 -2
  20. package/commands/fvs/map-code.md +10 -19
  21. package/commands/fvs/natural-language.md +12 -2
  22. package/commands/fvs/reapply-patches.md +28 -21
  23. package/commands/fvs/sync-aeneas-verif.md +12 -13
  24. package/commands/fvs/trust-audit.md +11 -6
  25. package/fv-skills/VERSION +1 -1
  26. package/fv-skills/references/crypto-plan-review.md +10 -3
  27. package/fv-skills/references/fc-spec-review.md +4 -2
  28. package/fv-skills/references/model-profiles.md +252 -169
  29. package/fv-skills/references/review-diagnostics.md +10 -8
  30. package/fv-skills/templates/config.json +28 -4
  31. package/fv-skills/workflows/aeneas-extract.md +13 -5
  32. package/fv-skills/workflows/crypto-eval.md +34 -15
  33. package/fv-skills/workflows/crypto-execute.md +6 -3
  34. package/fv-skills/workflows/crypto-followup.md +4 -2
  35. package/fv-skills/workflows/crypto-plan.md +4 -2
  36. package/fv-skills/workflows/crypto-review.md +34 -14
  37. package/fv-skills/workflows/fc-plan.md +4 -5
  38. package/fv-skills/workflows/lean-formalise.md +4 -13
  39. package/fv-skills/workflows/lean-refactor.md +7 -13
  40. package/fv-skills/workflows/lean-spec-review.md +66 -42
  41. package/fv-skills/workflows/lean-specify.md +4 -13
  42. package/fv-skills/workflows/lean-verify.md +4 -13
  43. package/fv-skills/workflows/map-code.md +4 -14
  44. package/fv-skills/workflows/sync-aeneas-verif.md +8 -5
  45. package/fv-skills/workflows/trust-audit.md +13 -4
  46. package/package.json +11 -3
  47. package/pi/skills/fvs-aeneas/SKILL.md +23 -0
  48. package/pi/skills/fvs-aeneas-extract/SKILL.md +224 -0
  49. package/pi/skills/fvs-checkpoint/SKILL.md +154 -0
  50. package/pi/skills/fvs-configure/SKILL.md +160 -0
  51. package/pi/skills/fvs-context/SKILL.md +22 -0
  52. package/pi/skills/fvs-crypto-eval/SKILL.md +204 -0
  53. package/pi/skills/fvs-crypto-execute/SKILL.md +218 -0
  54. package/pi/skills/fvs-crypto-followup/SKILL.md +272 -0
  55. package/pi/skills/fvs-crypto-plan/SKILL.md +322 -0
  56. package/pi/skills/fvs-crypto-review/SKILL.md +162 -0
  57. package/pi/skills/fvs-fc/SKILL.md +31 -0
  58. package/pi/skills/fvs-fc-plan/SKILL.md +289 -0
  59. package/pi/skills/fvs-formalise/SKILL.md +37 -0
  60. package/pi/skills/fvs-help/SKILL.md +457 -0
  61. package/pi/skills/fvs-kb-setup/SKILL.md +322 -0
  62. package/pi/skills/fvs-lean-formalise/SKILL.md +397 -0
  63. package/pi/skills/fvs-lean-refactor/SKILL.md +293 -0
  64. package/pi/skills/fvs-lean-spec-review/SKILL.md +61 -0
  65. package/pi/skills/fvs-lean-specify/SKILL.md +424 -0
  66. package/pi/skills/fvs-lean-verify/SKILL.md +461 -0
  67. package/pi/skills/fvs-manage/SKILL.md +29 -0
  68. package/pi/skills/fvs-map-code/SKILL.md +349 -0
  69. package/pi/skills/fvs-natural-language/SKILL.md +220 -0
  70. package/pi/skills/fvs-pause-work/SKILL.md +154 -0
  71. package/pi/skills/fvs-reapply-patches/SKILL.md +27 -0
  72. package/pi/skills/fvs-resume-work/SKILL.md +96 -0
  73. package/pi/skills/fvs-sync-aeneas-verif/SKILL.md +235 -0
  74. package/pi/skills/fvs-trust-audit/SKILL.md +234 -0
  75. package/pi/skills/fvs-update/SKILL.md +24 -0
  76. package/scripts/build-plugin.cjs +128 -6
  77. package/scripts/fvs-codex-think.mjs +128 -54
  78. package/scripts/fvs-spec-review.mjs +135 -39
package/CHANGELOG.md CHANGED
@@ -6,6 +6,36 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [2.3.4] - 2026-09-20
10
+
11
+ ### Fixed
12
+
13
+ - Crypto evaluation now trusts Lean's kernel for checked proof terms and concentrates adversarial
14
+ review on statement/source conformance and trust-boundary evidence. It reuses a current executor
15
+ build log or runs one bounded fallback without retry, rather than repeating proof derivations,
16
+ executor gates, or external numeric computations.
17
+
18
+ ## [2.3.3] - 2026-09-20
19
+
20
+ ### Added
21
+
22
+ - FVS is now a native Pi package with 29 generated skills, provider-qualified model handling,
23
+ `/skill:fvs-*` invocation guidance, and direct `pi install npm:fv-skills-baif` support (#62).
24
+ - `/fvs:configure` and every delegated workflow now use stage-scoped, runtime-aware model profiles.
25
+ Quality selects authority, execution, and scout tiers from the live runtime catalog; commands show
26
+ one confirmable model/effort manifest and fail closed when a provider or dispatch cannot honor it
27
+ (#64).
28
+
29
+ ### Fixed
30
+
31
+ - Fresh local Codex installs create their configuration directory before patch discovery. Local
32
+ patch metadata now separates immutable history from active pending patches, preserves real edits
33
+ without treating unchanged generated TOML as custom, and retires resolved patches atomically
34
+ (#60).
35
+ - FC and crypto reviews require explicit catalog-resolved model/effort selections, preserve raw and
36
+ normalized responses separately, retain writable diagnostic scratch without exposing source
37
+ writes, and enforce consistent verdict, severity, authority, and thread-count contracts (#60).
38
+
9
39
  ## [2.3.2] - 2026-09-10
10
40
 
11
41
  ### Fixed
package/README.md CHANGED
@@ -13,7 +13,7 @@
13
13
  npx fv-skills-baif
14
14
  ```
15
15
 
16
- **Works on Mac, Windows, and Linux. Supports Claude Code, Codex, OpenCode, and Gemini CLI.**
16
+ **Works on Mac, Windows, and Linux. Supports Pi, Claude Code, Codex, OpenCode, and Gemini CLI.**
17
17
 
18
18
  <br>
19
19
 
@@ -42,6 +42,18 @@ Framework-specific commands (currently Lean) handle the actual specification and
42
42
 
43
43
  ## Getting Started
44
44
 
45
+ ### Pi package
46
+
47
+ Install FVS directly from npm as a Pi package:
48
+
49
+ ```bash
50
+ pi install npm:fv-skills-baif
51
+ ```
52
+
53
+ Start a new session, then run `/skill:fvs-help`. Bundle routers such as `/skill:fvs-fc` and
54
+ `/skill:fvs-formalise`, plus member skills such as `/skill:fvs-crypto-plan`, are available directly.
55
+ Update an unpinned install with `pi update npm:fv-skills-baif`.
56
+
45
57
  ### Plugin marketplace (Claude Code and Codex)
46
58
 
47
59
  The Beneficial AI Foundation maintains one catalog for FVS and future BAIF plugins. Add the catalog
@@ -74,7 +86,7 @@ The BAIF Git catalog is a versioned distribution source that can list multiple i
74
86
  released plugins. It is separate from OpenAI's universal public Plugins Directory, which has its
75
87
  own per-plugin submission process.
76
88
 
77
- ### npm installer (all runtimes)
89
+ ### npm installer (Claude Code, Codex, OpenCode, and Gemini)
78
90
 
79
91
  ```bash
80
92
  npx fv-skills-baif
@@ -172,11 +184,10 @@ Commands are grouped into five bundles. Each bundle has a **router** command tha
172
184
  rejects ordinary Lean identifiers with three or more namespace dots, steering generated code
173
185
  toward scoped namespaces, `open`, and local names.
174
186
 
175
- After `lean-specify`, an interactive review menu asks reviewer, then model, then effort; it never
176
- auto-selects a choice. It offers the other runtime first, a fresh reviewer in the current runtime,
177
- or another provider. Choose GPT Sol or Astra for Codex, Fable for Claude, or a cheaper/custom model;
178
- effort defaults to `max` and can be lowered. Other providers use an exported source packet and
179
- imported response. One-run `Skip review` records `Unreviewed (user skipped)` and does not begin
187
+ After `lean-specify`, an interactive review menu resolves reviewer, exact catalog model, and
188
+ model-supported effort; it never auto-selects a choice. It offers the other runtime first, a fresh
189
+ reviewer in the current runtime, or another provider. Other providers use an exported source packet
190
+ and imported response. One-run `Skip review` records `Unreviewed (user skipped)` and does not begin
180
191
  proof work. Reviews and source hashes live under `.formalising/spec-reviews/`.
181
192
 
182
193
  The reviewer stays read-only and `review.md` stays immutable. The `lean-specify` authoring seat
@@ -233,10 +244,19 @@ secrets, raw transcripts, ephemeral error dumps, unsupported guesses, or inferre
233
244
  | Command | Description |
234
245
  |---------|-------------|
235
246
  | `/fvs:help` | Show available FVS commands and usage guide |
247
+ | `/fvs:configure` | Configure runtime-aware subagent models, effort levels, and review defaults |
236
248
  | `/fvs:update` | Update FVS through the current installation channel |
237
249
  | `/fvs:reapply-patches` | Preserve customizations across FVS updates (patches for npm installs; fork guidance for plugin installs) |
238
250
  | `/fvs:kb-setup` | Set up NotebookLM knowledge base integration (venv, auth, config) |
239
251
 
252
+ `/fvs:configure` stores concrete model IDs under the runtime and exact stage that reported them.
253
+ The quality profile is role-aware: authority artifacts use the strongest detected model with max
254
+ reasoning, execution/proof filling uses the executor model with xhigh, and research/eval/audit work
255
+ uses a smaller model with high. Claude and Codex family preferences never cross runtimes; Pi uses
256
+ provider-qualified IDs and keeps its active provider for ordinary work. Each interactive command
257
+ shows one model/effort selection manifest with one-run, save, notes, and cancel paths. Missing models
258
+ or unsupported efforts require a user choice; unresolved noninteractive runs fail before dispatch.
259
+
240
260
  ---
241
261
 
242
262
  ## How It Works
@@ -6,15 +6,14 @@ color: purple
6
6
  ---
7
7
 
8
8
  <role>
9
- You are the FVS crypto formalisation thinker. You are the high-effort author of the loop: you
10
- re-derive everything independently, from the branch state and the paper-grounded sources, and you
11
- return your reasoning as text. You are NOT the executor -- a separate `fvs-executor`-style agent in
12
- the current runtime runs the plans you author. You author; they execute.
9
+ You are the FVS crypto formalisation thinker. You are the high-effort author of the loop: in plan and
10
+ follow-up modes you derive bounded work independently from the branch state and paper-grounded
11
+ sources, then return your reasoning as text. You are NOT the executor -- a separate
12
+ `fvs-executor`-style agent in the current runtime runs the plans you author. You author; they execute.
13
13
 
14
- Planning is ALWAYS high reasoning effort -- you never produce a sketch and call it a plan. The eval
15
- stage is ALWAYS adversarial: you take the posture of a reviewer who is actively trying to REFUTE the
16
- spec, the proof, and the stated assumptions, not one who is looking for a reason to wave them
17
- through. A plan or proof survives only by surviving your attempt to break it.
14
+ Planning is ALWAYS high reasoning effort -- you never produce a sketch and call it a plan. Eval mode
15
+ is adversarial about landed statements, modeling assumptions, source fidelity, and trust boundaries;
16
+ it follows the bounded kernel-trusting contract below instead of reproducing checked proof work.
18
17
 
19
18
  You are read-only with respect to the deliverable: you do NOT write or modify any project file. You
20
19
  RETURN the bounded plan / the adversarial eval / the follow-up as text, and the orchestrating
@@ -77,18 +76,28 @@ from this iteration. Apply the same audit to follow-up plans.
77
76
  **Input:** the executor's run output, the touched files, the plan it was run against, the KB sources.
78
77
  **Output (returned as text):** an adversarial review ending in exactly ONE decision verb.
79
78
 
80
- This stage is ALWAYS adversarial. Re-derive independently; do not echo the executor's reasoning.
81
- Actively try to REFUTE: does the spec actually capture the paper's claim? Does the proof close the
82
- goal it claims, or does it lean on an unstated assumption? Is every `sorry` a named obligation with
83
- the correct statement, or is it papering over a real gap? Name the exact input, caller, or modeling
84
- assumption that would make the argument FALSE.
79
+ This stage TRUSTS THE LEAN KERNEL for kernel-checked proof terms and stays adversarial about statements
80
+ and trust boundaries. Check the landed definitions, theorem signatures, constants, and API shape
81
+ against the paper/standard and the approved plan. Run cheap scans over touched Lean files and import
82
+ changes for reserved names, forbidden imports, `sorry`, unexpected `axiom`, `native_decide`, and
83
+ `set_option`. Classify every hit in context; unexplained or disallowed hits prevent `ACCEPT`.
85
84
 
86
- A `sorry` is acceptable ONLY as an intentional, named obligation carrying the correct statement --
87
- never judged by count, never waved through because "the build is green".
85
+ Reuse a successful current executor `build.log`. If it is missing, failed, or does not cover the
86
+ landed files, run at most one fallback `LEAN_NUM_THREADS="${LEAN_NUM_THREADS:-4}" nice -n 19 lake build`;
87
+ do not retry. A failed fallback is `BLOCKED`. Never re-elaborate individual files, replay the
88
+ executor's per-gate builds, search for proof terms, or use an outside script to recompute certified
89
+ numbers. Compare landed constants directly with explicit plan/source values; missing derivation
90
+ evidence is `FOLLOWUP`, not permission to recompute it.
91
+
92
+ A `sorry` is an intentional, named unmet obligation carrying the correct statement. It is never
93
+ judged by count or waved through because the build is green, and it does not inherit the
94
+ kernel-complete status of checked proof terms. Name the exact input, caller, statement, or modeling
95
+ assumption that would make the landed claim false.
88
96
 
89
97
  End with EXACTLY ONE of these decision verbs, on its own:
90
98
 
91
- - **ACCEPT** -- the spec/proof survives the adversarial pass; the obligations are honest.
99
+ - **ACCEPT** -- statement/source conformance, classified trust-boundary evidence, and required green
100
+ build evidence all pass; named unmet obligations are honestly recorded.
92
101
  - **FOLLOWUP** -- the work is sound but incomplete; a bounded follow-up plan is warranted.
93
102
  - **HUMAN_RULING** -- a modeling decision is required that you must NOT make yourself (see followup).
94
103
  - **BLOCKED** -- the work cannot proceed (e.g. the build will not compile, a prerequisite is absent).
@@ -145,7 +154,7 @@ Adversarial eval:
145
154
 
146
155
  **Stage:** eval
147
156
  **Decision:** ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
148
- **Refutation attempted:** {the strongest counter you raised}
157
+ **Statement challenge:** {the strongest source/spec counterexample you tested}
149
158
  ```
150
159
 
151
160
  On HALT / failure:
@@ -160,7 +169,7 @@ On HALT / failure:
160
169
 
161
170
  <success_criteria>
162
171
  - [ ] In `plan`/`followup` mode, authored a bounded, runtime-neutral plan stating branch/state, exact target files+theorems, immutable public statements, old->new API map (if a port), allowed-`sorry` policy, stop conditions, `LEAN_NUM_THREADS="${LEAN_NUM_THREADS:-4}" nice -n 19 lake build` verification, and expected artifact updates
163
- - [ ] In `eval` mode, took an adversarial posture (tried to refute), judged each `sorry` as a named obligation not by count, and ended in exactly one of ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
172
+ - [ ] In `eval` mode, trusted kernel-checked proof terms, challenged statement/source conformance and trust boundaries, reused current build evidence or ran one guarded fallback without retry, and ended in exactly one of ACCEPT | FOLLOWUP | HUMAN_RULING | BLOCKED
164
173
  - [ ] On `HUMAN_RULING`, HALTed and asked for the modeling decision -- never fabricated a plan
165
174
  - [ ] Author-by-return: no project file written or modified; no `gh` auto-open; Lean-via-Aeneas pipeline only; no bare `lake build`
166
175
  - [ ] Result returned with the ## PLAN COMPLETE / ## EVAL COMPLETE / ## ERROR header
package/bin/install.js CHANGED
@@ -70,37 +70,30 @@ const CODEX_AGENT_SANDBOX = {
70
70
  'fvs-axiom-auditor': 'read-only',
71
71
  };
72
72
 
73
- // Codex agents inherit the user's selected Codex model — the converter never
74
- // pins a `model`, it sets only the reasoning-effort budget. The user owns the
75
- // model choice and may change it mid-session; FVS controls only how hard each
76
- // agent thinks. Precision-critical roles (spec authoring, proof attempt and
77
- // refactor, explanation) run at the highest effort; the research/mapping
78
- // support role runs one tier lower. Codex accepts minimal|low|medium|high|xhigh;
79
- // any unmapped agent defaults to high at the lookup site.
73
+ // Codex agent TOMLs cannot express FVS's stage-aware quality matrix: shared
74
+ // agents serve authority, work, and scout stages. These are safe installation
75
+ // fallbacks only. The command-level selection manifest is authoritative and
76
+ // passes a stage effort dynamically when spawn_agent supports it; otherwise it
77
+ // must surface the limitation instead of claiming the requested stage setting.
78
+ // Codex accepts minimal|low|medium|high|xhigh in agent TOML; any unmapped agent
79
+ // defaults to high at the lookup site.
80
80
  const FVS_CODEX_AGENT_EFFORT = {
81
81
  'fvs-executor': 'xhigh',
82
82
  'fvs-lean-refactorer': 'xhigh',
83
- 'fvs-explainer': 'xhigh',
83
+ 'fvs-explainer': 'high',
84
84
  'fvs-researcher': 'high',
85
- // Extraction loop + doc-sync workers. The two roles whose output is the most
86
- // reasoning-sensitive — the independent equivalence assessor (judging whether a
87
- // meaning-bearing change preserves behaviour) and the bisector (minimizing a
88
- // novel blocker to a faithful failing example) — run at xhigh; the remaining
89
- // workers run at high.
90
- 'fvs-equivalence-assessor': 'xhigh',
85
+ // Extraction fallbacks follow the confirmed work/scout split.
86
+ 'fvs-equivalence-assessor': 'high',
91
87
  'fvs-extract-bisector': 'xhigh',
92
88
  'fvs-extract-classifier': 'high',
93
- 'fvs-extract-applier': 'high',
89
+ 'fvs-extract-applier': 'xhigh',
94
90
  'fvs-draft-investigator': 'high',
95
91
  'fvs-doc-syncer': 'high',
96
- // Crypto loop + trust audit. The thinker authors bounded plans and runs the
97
- // always-adversarial eval -- the most reasoning-sensitive role in the loop --
98
- // so it runs at xhigh (the dual-runtime Codex thinker must think at least this
99
- // hard); the read-only auditor introspects axioms at high.
92
+ // The thinker spans authority plan/follow-up and scout eval stages, so xhigh
93
+ // is the strongest safe static fallback; the manifest supplies max/high when
94
+ // dynamic effort is available. The executor is a work stage.
100
95
  'fvs-crypto-thinker': 'xhigh',
101
- // The crypto executor is the dial-down implementation stage (vs the thinker's
102
- // xhigh authoring) -- it executes a fully-specified plan, so it runs at high.
103
- 'fvs-crypto-executor': 'high',
96
+ 'fvs-crypto-executor': 'xhigh',
104
97
  'fvs-axiom-auditor': 'high',
105
98
  };
106
99
 
@@ -762,7 +755,7 @@ function getCodexSkillAdapterHeader(skillName, options = {}) {
762
755
  : '';
763
756
  const typedModelNote = pluginName
764
757
  ? 'The marketplace plugin does not install Codex agent TOML. Use this mapping only when the exact FVS agent type is registered independently; otherwise use the bundled-agent workaround.'
765
- : 'FVS bakes each agent\'s reasoning effort into its `.toml` at install time and the model is inherited from the user\'s Codex configuration.';
758
+ : 'Installed agent TOML supplies only a fallback effort; it does not prove that the command-level stage selection was applied.';
766
759
  const fallbackSteps = pluginName
767
760
  ? `1. Read \`\${CLAUDE_PLUGIN_ROOT}/agents/<agent-name>.md\` and extract its instructions. If the token is still literal, resolve the path from this SKILL.md as described above.\n2. Spawn a generic/default agent and inject those instructions as a role preamble before the task prompt.\n3. Label results clearly as \"generic-agent workaround\" so the user knows typed guarantees are not in effect.\n4. Where typed dispatch is mandatory for correctness, fail closed and report the schema limitation rather than silently degrading.`
768
761
  : `1. Resolve your active Codex config root (the directory containing your \`config.toml\`), then read \`agents/<agent-name>.toml\` relative to that root to extract the agent's instructions.\n2. Inject those instructions as a role-preamble into a generic \`spawn_agent(message=...)\` call.\n3. Label results clearly as \"generic-agent workaround\" so the user knows typed guarantees are not in effect.\n4. Where typed dispatch is mandatory for correctness, fail closed and report the schema limitation rather than silently degrading.`;
@@ -807,9 +800,15 @@ FVS workflows use \`Task(...)\` (Claude Code syntax). Translate to Codex collabo
807
800
 
808
801
  Before spawning, inspect the \`spawn_agent\` tool's visible parameter schema to determine which form is active.${typedDispatchQualification}
809
802
 
803
+ Selection-capability gate (before manifest confirmation):
804
+ - Compare the requested model with the exact active/inherited Codex model. Because \`spawn_agent\` has no inline model field, any different requested model is unresolved. Rebuild the manifest around the actual active model and ask explicitly, choose a capable external runner, or fail before dispatch in noninteractive mode.
805
+ - When \`reasoning_effort\` is absent from the schema, compare the requested effort with the installed/runtime default. A mismatch is unresolved and follows the same rebuild/ask-or-fail rule.
806
+ - Never confirm a requested model or effort and then omit it with a warning. ${typedModelNote}
807
+
810
808
  Typed mapping (agent_type-capable schema only):
811
809
  - \`Task(subagent_type="X", prompt="Y")\` -> \`spawn_agent(agent_type="X", message="Y")\`
812
- - \`Task(model="...")\` -> omit. \`spawn_agent\` has no inline \`model\` parameter. ${typedModelNote}
810
+ - \`Task(model="...")\` -> omit only after the selection-capability gate proves it equals the active/inherited model.
811
+ - \`Task(reasoning_effort="...")\` -> \`spawn_agent(reasoning_effort="...")\` when that field is present. If absent, dispatch only after the gate proves the installed/runtime default equals the confirmed effort.
813
812
  - \`fork_context: false\` by default -- FVS agents load their own context via \`<files_to_read>\` blocks.
814
813
 
815
814
  Generic-agent workaround (schema with NO agent_type field):
@@ -2639,11 +2638,14 @@ const MANIFEST_NAME = 'fvs-file-manifest.json';
2639
2638
  /**
2640
2639
  * Compute SHA256 hash of file contents
2641
2640
  */
2642
- function fileHash(filePath) {
2643
- const content = fs.readFileSync(filePath);
2641
+ function contentHash(content) {
2644
2642
  return crypto.createHash('sha256').update(content).digest('hex');
2645
2643
  }
2646
2644
 
2645
+ function fileHash(filePath) {
2646
+ return contentHash(fs.readFileSync(filePath));
2647
+ }
2648
+
2647
2649
  /**
2648
2650
  * Recursively collect all files in dir with their hashes
2649
2651
  */
@@ -2710,12 +2712,25 @@ function writeManifest(configDir, runtime = 'claude') {
2710
2712
  return manifest;
2711
2713
  }
2712
2714
 
2715
+ function matchesGeneratedCodexAgentToml(configDir, rel, digest) {
2716
+ if (!/^agents\/fvs-.*\.toml$/.test(rel)) return false;
2717
+ const markdownPath = path.join(configDir, rel.replace(/\.toml$/, '.md'));
2718
+ if (!fs.existsSync(markdownPath)) return false;
2719
+ const converted = fs.readFileSync(markdownPath, 'utf8');
2720
+ const { frontmatter, body } = extractFrontmatterAndBody(converted);
2721
+ if (!frontmatter || !/<codex_agent_role>/.test(body)) return false;
2722
+ const originalBody = body.replace(/^\s*<codex_agent_role>[\s\S]*?<\/codex_agent_role>\s*/, '');
2723
+ const reconstructed = `---\n${frontmatter}\n---\n\n${originalBody}`;
2724
+ const agentName = path.basename(rel, '.toml');
2725
+ return contentHash(generateCodexAgentToml(agentName, reconstructed)) === digest;
2726
+ }
2713
2727
 
2714
2728
  /**
2715
2729
  * Detect user-modified FVS files by comparing against install manifest.
2716
2730
  * Backs up modified files to fvs-local-patches/ for reapply after update.
2717
2731
  */
2718
2732
  function saveLocalPatches(configDir, runtime = 'claude') {
2733
+ fs.mkdirSync(configDir, { recursive: true });
2719
2734
  configDir = fs.realpathSync(configDir);
2720
2735
  const manifestPath = path.join(configDir, MANIFEST_NAME);
2721
2736
  // An unreadable baseline must stop the update before any destructive copy.
@@ -2725,7 +2740,12 @@ function saveLocalPatches(configDir, runtime = 'claude') {
2725
2740
  throw new Error('Invalid FVS manifest; repair it before updating. Existing files were preserved.');
2726
2741
  }
2727
2742
  const current = ownedFileHashes(configDir, runtime);
2728
- const modified = Object.keys(current).filter(rel => current[rel] !== manifest.files[rel]);
2743
+ const modified = Object.keys(current).filter(rel => {
2744
+ if (current[rel] === manifest.files[rel]) return false;
2745
+ if (!Object.prototype.hasOwnProperty.call(manifest.files, rel) && runtime === 'codex' &&
2746
+ matchesGeneratedCodexAgentToml(configDir, rel, current[rel])) return false;
2747
+ return true;
2748
+ });
2729
2749
  const patchesDir = path.join(configDir, PATCHES_DIR_NAME);
2730
2750
  if (fs.lstatSync(patchesDir, { throwIfNoEntry: false })?.isSymbolicLink()) {
2731
2751
  throw new Error('Local patch directory is a symlink; refusing to update');
@@ -2736,15 +2756,22 @@ function saveLocalPatches(configDir, runtime = 'claude') {
2736
2756
  const addBundle = (directory, meta) => {
2737
2757
  const hashes = generateManifest(directory);
2738
2758
  delete hashes['backup-meta.json'];
2739
- for (const rel of meta?.files ?? []) {
2759
+ const listed = new Set(meta?.files ?? []);
2760
+ const pending = new Set(meta?.pending ?? meta?.files ?? Object.keys(hashes));
2761
+ for (const rel of listed) {
2740
2762
  if (!Object.prototype.hasOwnProperty.call(hashes, rel)) throw new Error(`Missing backed-up patch: ${rel}; refusing to update`);
2741
2763
  if (meta.hashes?.[rel] && meta.hashes[rel] !== hashes[rel]) {
2742
2764
  throw new Error(`Backed-up patch changed: ${rel}; inspect it before updating`);
2743
2765
  }
2744
2766
  }
2745
2767
  for (const rel of Object.keys(hashes)) {
2746
- if (meta && !meta.files?.includes(rel)) console.warn(` Unlisted local patch recovered: ${rel}`);
2747
- sources.set(rel, { file: path.join(directory, rel), kind: meta?.kinds?.[rel] ?? 'legacy' });
2768
+ if (meta && !listed.has(rel)) {
2769
+ console.warn(` Unlisted local patch recovered: ${rel}`);
2770
+ pending.add(rel);
2771
+ }
2772
+ if (pending.has(rel)) {
2773
+ sources.set(rel, { file: path.join(directory, rel), kind: meta?.kinds?.[rel] ?? 'legacy' });
2774
+ }
2748
2775
  }
2749
2776
  };
2750
2777
  if (prior?.bundle) {
@@ -2777,8 +2804,9 @@ function saveLocalPatches(configDir, runtime = 'claude') {
2777
2804
  fs.mkdirSync(bundles, { recursive: true });
2778
2805
  const staging = fs.mkdtempSync(path.join(bundles, '.pending-'));
2779
2806
  const bundle = `bundles/bundle-${path.basename(staging).slice(9)}`;
2780
- const meta = { version: 2, bundle, backed_up_at: new Date().toISOString(),
2781
- from_version: manifest.version, files: [...sources.keys()].sort(), hashes: {}, kinds: {} };
2807
+ const meta = { version: 3, bundle, backed_up_at: new Date().toISOString(),
2808
+ from_version: manifest.version, files: [...sources.keys()].sort(), pending: [], hashes: {}, kinds: {} };
2809
+ meta.pending = [...meta.files];
2782
2810
  for (const rel of meta.files) {
2783
2811
  const dest = path.join(staging, rel);
2784
2812
  fs.mkdirSync(path.dirname(dest), { recursive: true });
@@ -2808,7 +2836,8 @@ function reportLocalPatches(configDir, runtime = 'claude') {
2808
2836
  let meta;
2809
2837
  try { meta = JSON.parse(fs.readFileSync(metaPath, 'utf8')); } catch { return []; }
2810
2838
 
2811
- if (meta.files && meta.files.length > 0) {
2839
+ const pending = meta.pending ?? meta.files ?? [];
2840
+ if (pending.length > 0) {
2812
2841
  const reapplyCommand = runtime === 'opencode'
2813
2842
  ? '/fvs-reapply-patches'
2814
2843
  : runtime === 'codex'
@@ -2816,7 +2845,7 @@ function reportLocalPatches(configDir, runtime = 'claude') {
2816
2845
  : '/fvs:reapply-patches';
2817
2846
  console.log('');
2818
2847
  console.log(' ' + yellow + 'Local patches detected' + reset + ' (from v' + meta.from_version + '):');
2819
- for (const f of meta.files) {
2848
+ for (const f of pending) {
2820
2849
  console.log(' ' + orange + f + reset);
2821
2850
  }
2822
2851
  console.log('');
@@ -2825,7 +2854,7 @@ function reportLocalPatches(configDir, runtime = 'claude') {
2825
2854
  console.log(' Or manually compare and merge the files.');
2826
2855
  console.log('');
2827
2856
  }
2828
- return meta.files || [];
2857
+ return pending;
2829
2858
  }
2830
2859
 
2831
2860
  /**
@@ -47,27 +47,22 @@ idempotent and skips already-clean blockers. There is no separate status/resume
47
47
 
48
48
  <process>
49
49
 
50
- ## Step 1: Read config and resolve subagent models
50
+ ## Step 1: Read config and resolve subagent models + effort
51
51
 
52
- Read the project config and resolve the model for each subagent dispatch using the
53
- model-profiles dispatch sequence (config `model_overrides` first, then the profile table,
54
- then `inherit` for unknown agents):
52
+ Read the complete project config and apply the canonical resolution contract in
53
+ `model-profiles.md`. Declare these stage keys rather than deriving tiers from agent names:
55
54
 
56
- ```bash
57
- CONFIG=$(cat .formalising/fvs-config.json 2>/dev/null)
58
- # profile = config.model_profile || "balanced"
59
- # for each agent: model = model_overrides[agent] ?? PROFILE_TABLE[agent][profile]
60
- ```
61
-
62
- Resolve and store:
63
- - `$CLASSIFIER_MODEL` for `fvs-extract-classifier`
64
- - `$APPLIER_MODEL` for `fvs-extract-applier`
65
- - `$BISECTOR_MODEL` for `fvs-extract-bisector`
66
- - `$ASSESSOR_MODEL` for `fvs-equivalence-assessor`
67
- - `$DRAFT_MODEL` for `fvs-draft-investigator`
55
+ - `fvs-extract-classifier` -> `extract_classify`
56
+ - `fvs-extract-applier` -> `extract_apply`
57
+ - `fvs-extract-bisector` -> `extract_bisect`
58
+ - `fvs-equivalence-assessor` -> `extract_assess`
59
+ - `fvs-draft-investigator` -> `extract_investigate`
68
60
 
69
- On Codex (which does not support dynamic model selection) the `model=` parameter is silently
70
- ignored; the same dispatches work unchanged.
61
+ Resolve model and effort independently for every distinct stage. Before the first dispatch, show
62
+ one command-level selection manifest and obtain confirmation. `Adjust once` changes only this run;
63
+ `Save override` writes the exact runtime+stage entry. Notes rebuild the manifest and require
64
+ reconfirmation. Missing preferred models or unsupported efforts prompt interactively and fail with
65
+ exact remediation in noninteractive mode. Pass only validated native model/effort fields.
71
66
 
72
67
  ## Step 2: PRE-FLIGHT (pin audit + clone resolution)
73
68
 
@@ -117,13 +112,13 @@ fix escalates immediately).
117
112
  - **EXTRACT:** run extraction and build under `set -o pipefail` + `LEAN_NUM_THREADS="${LEAN_NUM_THREADS:-4}" nice -n 19 lake build`;
118
113
  read the tool's real exit status via `${PIPESTATUS[0]}`, never the tail of a piped log.
119
114
  Clean -> success oracle (Step 5). Failure -> classify.
120
- - **CLASSIFY:** `Task(subagent_type="fvs-extract-classifier", model="$CLASSIFIER_MODEL", ...)`
115
+ - **CLASSIFY:** `Task(subagent_type="fvs-extract-classifier", model="$CLASSIFIER_MODEL", reasoning_effort="$CLASSIFIER_EFFORT", ...)`
121
116
  -> `{ layer, symbol, signature, match }`.
122
117
  - **DISPATCH by category + coverage_impact:**
123
- - Category-A -> `Task(subagent_type="fvs-extract-applier", model="$APPLIER_MODEL", ...)`.
118
+ - Category-A -> `Task(subagent_type="fvs-extract-applier", model="$APPLIER_MODEL", reasoning_effort="$APPLIER_EFFORT", ...)`.
124
119
  The coverage-escalation guard: A-opacity on a function inside required coverage is NOT
125
120
  auto-applied -- it becomes Category-B and routes to the gate.
126
- - NOVEL -> `Task(subagent_type="fvs-extract-bisector", model="$BISECTOR_MODEL", ...)` to
121
+ - NOVEL -> `Task(subagent_type="fvs-extract-bisector", model="$BISECTOR_MODEL", reasoning_effort="$BISECTOR_EFFORT", ...)` to
127
122
  minimize to an MFE, write a schema-conformant candidate, and PROPOSE a fix.
128
123
  - Forced Category-B -> the GATE (Step 4).
129
124
  - Escalate conditions (attempt-cap hit, same-signature recurrence, forced
@@ -146,6 +141,7 @@ a fixing subagent:
146
141
 
147
142
  ```
148
143
  Task(subagent_type="fvs-equivalence-assessor", model="$ASSESSOR_MODEL",
144
+ reasoning_effort="$ASSESSOR_EFFORT", // when supported; otherwise apply the capability gate
149
145
  description="Draft equivalence-gate packet sections 2-6",
150
146
  prompt="...minimized diff (section 1) + the blocker context...")
151
147
  ```
@@ -182,6 +178,7 @@ ONLY on the user's acceptance -- on acceptance dispatch the draft-investigator:
182
178
 
183
179
  ```
184
180
  Task(subagent_type="fvs-draft-investigator", model="$DRAFT_MODEL",
181
+ reasoning_effort="$DRAFT_EFFORT", // when supported; otherwise apply the capability gate
185
182
  description="Mine precedent + draft an evidence-cited HTML+MD report",
186
183
  prompt="...the MFE + the escalation context...")
187
184
  ```
@@ -213,8 +210,9 @@ On Codex, every interactive HALT in this command -- the pin-audit warn-and-confi
213
210
  the Category-B gate presentation (Step 4), and the escalation draft offer (Step 6) -- degrades
214
211
  to a plain-text question and WAITS for the user. It is fail-closed: it never auto-picks a
215
212
  default, never self-ratifies a gate packet, and never writes an upstream artifact. The
216
- `Task(...)` dispatches survive intact (the `model=` parameter is silently ignored on Codex,
217
- per model-profiles runtime handling).
213
+ Before any Codex dispatch, apply the model-profile capability gate: confirm only the actual
214
+ active/inherited model and applicable effort, or fail before dispatch. Never silently ignore a
215
+ confirmed field.
218
216
  </codex_skill_adapter>
219
217
 
220
218
  <success_criteria>
@@ -0,0 +1,159 @@
1
+ ---
2
+ name: fvs:configure
3
+ description: Configure runtime-aware FVS stage models, effort levels, and review defaults
4
+ argument-hint: ""
5
+ allowed-tools:
6
+ - Read
7
+ - Write
8
+ - Edit
9
+ - Bash
10
+ - AskUserQuestion
11
+ ---
12
+
13
+ <objective>
14
+ Edit `.formalising/fvs-config.json` through bounded choice menus. Store exact runtime model IDs only
15
+ when the runtime reports them, keep quality role-aware by stage, preserve one-run user control, and
16
+ preserve unrelated project configuration.
17
+ </objective>
18
+
19
+ <execution_context>
20
+ @~/.claude/fv-skills/references/model-profiles.md
21
+ @~/.claude/fv-skills/templates/config.json
22
+ </execution_context>
23
+
24
+ <process>
25
+
26
+ ## 1. Load without clobbering
27
+
28
+ Read `.formalising/fvs-config.json` when it exists; otherwise start from the shipped template. Reject
29
+ malformed JSON and stop before writing. Preserve unknown keys and every unrelated nested value.
30
+ Never replace the whole file with a partial settings object.
31
+
32
+ Add missing current-schema containers without deleting old settings: `stage_overrides`,
33
+ `model_overrides`, `effort_overrides`, `spec_review`, and `crypto_review`.
34
+
35
+ ## 2. Detect runtime, provider, and catalog
36
+
37
+ Prefer explicit host identity over directory guesses:
38
+
39
+ 1. `PI_CODING_AGENT=true` or `AI_AGENT=pi` means Pi.
40
+ 2. A host adapter or `AI_AGENT` value naming Claude, Codex, OpenCode, or Gemini wins next.
41
+ 3. If identity is still unknown, ask which runtime is being configured.
42
+
43
+ Read the runtime's current model catalog or picker. On Pi, `pi --list-models` is the CLI fallback.
44
+ Keep Pi IDs provider-qualified and identify its active provider. Never invent or persist a guessed
45
+ slug. Read the canonical quality stage and runtime matrices from `model-profiles.md`; do not copy or
46
+ weaken them here.
47
+
48
+ For ordinary Pi stages, stay within the active provider. Review routing is configured separately and
49
+ may use an authenticated opposite-runtime CLI. Never treat a Claude Code subscription as a Pi
50
+ Anthropic credential.
51
+
52
+ ## 3. Main settings menu
53
+
54
+ Use `AskUserQuestion` and repeat until the user saves or cancels. Retain the question UI's notes or
55
+ custom-answer path. Present at most these four choices:
56
+
57
+ - **Profile** — choose `quality`, `balanced`, or `budget`.
58
+ - **Stage overrides** — configure one exact canonical stage for the active runtime.
59
+ - **Advanced defaults** — open agent-compatibility or review-default settings.
60
+ - **Save/exit** — preview, save, discard, or continue.
61
+
62
+ If structured questions are unavailable, print the same numbered choices plus a notes/custom line,
63
+ then stop and wait. Never select on the user's behalf.
64
+
65
+ ## 4. Configure a stage override
66
+
67
+ Read stage keys and tiers from the sole canonical table in `model-profiles.md`. First choose a tier,
68
+ then a stage in that tier. Do not infer a stage from an agent name.
69
+
70
+ Show the quality recommendation for the detected runtime/provider as an exact catalog match. The
71
+ model menu offers:
72
+
73
+ 1. the exact detected quality model when available;
74
+ 2. `inherit`;
75
+ 3. up to one other detected runtime/provider model;
76
+ 4. catalog/custom selection.
77
+
78
+ If the preferred model is absent, show the detected catalog and require a choice; do not silently
79
+ choose a lower family. Custom input must exactly match a catalog entry. On Pi it must include the
80
+ provider. For ordinary Pi stages, reject entries whose catalog provider differs from the active
81
+ provider; review stages use their separate opposite-provider policy.
82
+
83
+ For effort, offer every value the selected catalog model reports as supported, with the stage
84
+ tier's quality effort first. Never advertise a value the model cannot accept. Save either or both fields under:
85
+
86
+ ```
87
+ stage_overrides[runtime][stage] = { model?, effort? }
88
+ ```
89
+
90
+ Removing both fields removes the stage entry. A saved stage override affects only that runtime and
91
+ stage.
92
+
93
+ ## 5. Configure advanced defaults
94
+
95
+ Use a submenu with at most three choices: **Agent compatibility**, **Review defaults**, **Back**.
96
+
97
+ ### Agent compatibility
98
+
99
+ Existing broad overrides remain available for projects that need one setting across several
100
+ stages. Choose one shipped agent and store exact values under
101
+ `model_overrides[runtime][agent]` / `effort_overrides[runtime][agent]`. Explain that a stage override
102
+ wins and is preferred when a shared agent performs different roles.
103
+
104
+ Shipped groups:
105
+
106
+ - FC: `fvs-researcher`, `fvs-executor`, `fvs-lean-refactorer`, `fvs-explainer`.
107
+ - Crypto: `fvs-crypto-thinker`, `fvs-crypto-executor`.
108
+ - Extraction: `fvs-extract-classifier`, `fvs-extract-applier`, `fvs-extract-bisector`,
109
+ `fvs-equivalence-assessor`, `fvs-draft-investigator`, `fvs-doc-syncer`.
110
+ - Audit: `fvs-axiom-auditor`.
111
+
112
+ ### Review defaults
113
+
114
+ Choose `spec_review`, `crypto_review`, or both. Null reviewer/model/effort values keep role-aware
115
+ profile routing and the per-command confirmation active. For each explicit saved review setting:
116
+
117
+ 1. choose reviewer runtime `codex`, `claude`, `pi`, or `other` (`pi` only when the active host is Pi and the selected provider is authenticated);
118
+ 2. choose an exact model from that reviewer's catalog;
119
+ 3. choose a supported effort;
120
+ 4. keep `automatic` unchanged unless explicitly edited.
121
+
122
+ Quality recommends the authenticated opposite provider/runtime and its authority-tier model+effort.
123
+ On Pi, OpenAI review of Claude-authored work prefers a fresh provider-qualified Pi seat only when
124
+ the child facility can honor and report the exact per-child model/effort, then Codex CLI. Review of
125
+ OpenAI-authored work uses authenticated Claude Code CLI when available. If the opposite runner is
126
+ unavailable, do not silently save a same-runtime fallback.
127
+
128
+ ## 6. Preview and save
129
+
130
+ Show all changed keys and a selection preview. Ask for confirmation with `Save`, `Continue editing`,
131
+ `Discard`, and `Cancel`. Notes rebuild the preview and require another confirmation.
132
+
133
+ Create `.formalising/` if needed, serialize with two-space indentation and a trailing newline,
134
+ validate the complete JSON in a temporary file, then atomically replace
135
+ `.formalising/fvs-config.json`. On validation or write failure, leave the old file untouched.
136
+
137
+ Report runtime/provider, profile, changed stage/agent/review settings, and precedence:
138
+
139
+ 1. explicit one-run selection;
140
+ 2. saved review value for review stages;
141
+ 3. runtime+stage override;
142
+ 4. runtime+agent compatibility override;
143
+ 5. profile/catalog resolution.
144
+
145
+ Also report that missing preferred models, ordinary-Pi provider mismatches, unsupported efforts,
146
+ and unavailable per-child controls prompt interactively, while noninteractive unresolved choices
147
+ fail before dispatch with exact remediation.
148
+
149
+ </process>
150
+
151
+ <success_criteria>
152
+ - [ ] Runtime/provider identified explicitly; every concrete model came from its current catalog.
153
+ - [ ] Ordinary Pi stage overrides use the active provider; review overrides follow review routing.
154
+ - [ ] Quality recommendations use the canonical authority/work/scout matrices.
155
+ - [ ] Stage overrides are stored under the exact runtime+stage key and win over agent overrides.
156
+ - [ ] Choice menus expose notes and reconfirm after notes change the preview.
157
+ - [ ] Review defaults may remain null so opposite-runtime profile routing stays active.
158
+ - [ ] Unknown config keys were preserved and malformed JSON was never overwritten.
159
+ </success_criteria>