@chrono-meta/fh-gate 1.4.51 → 1.4.52

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CATALOG.md CHANGED
@@ -8,6 +8,12 @@ AI reads this file first when searching past work. Open individual files for det
8
8
 
9
9
  <!-- Add entries in reverse date order (newest at top) -->
10
10
 
11
+ ### 2026-07-10 | forge-harness | #sonnet-floor, #doctrine, #loop-engineering, #tier-census, #cross-family, #pre-commit-gate, #dispatch-first
12
+ **File:** knowledge/shared/harness-core/sonnet_floor_doctrine.md · knowledge/shared/harness-core/loop_engineering.md
13
+ Encoded the operator-declared **Sonnet-Floor Doctrine** (base ops 100% Sonnet-runnable; tier-gated capability = defect; escalation = dispatch, never substrate; depth ladder = effort→dispatch→anchored-Sonnet) as a canonical axiom node, plus **loop_engineering.md** (5-question design-time discipline + FH loop inventory MECH/PROSE census + evidence-threshold hardening backlog). Cross-family evolution pass: codex gpt-5.5 xhigh repo census (T1 tier refs / T2 loop legs / T3 contradictions) + agy Gemini 3.1 Pro breadth (pattern-level only, zero citations imported — phantom-risk URLs). All 6 identified availability-gates fixed: pre-commit Axis-2 gains a **sonnet-floor lane** (anchor-required, R-tier auto-queue, 8/8 regression fixtures in scripts/test_marker_floor_lanes.sh), self_evolution weekly dead-end recast dispatch-first, Mode D notice re-directed (keep Sonnet + dispatch primary), canary opus-judge → Sonnet-governor+anchor, verify-bidirectional "never stay at sonnet" fixed, 9 SKILL.md `model: opus` hard pins retired (session-inherit). Trust-floors tightened to run-first/ask-last (full Sonnet autonomy; gates stay). Sonnet blind sims: 2 dispatched, 1 salience miss caught (loop-stub enumeration) → hardened → re-sim PASS.
14
+ - Decision: Sonnet = the optimization target, measured spine = H1 (harness benefit largest on weaker tiers); Opus/Fable-only capability is now a named defect class with a census discipline.
15
+ - Open: quarterly/substrate loop rows are governor self-assessment (R-tier external census pending); sonnet-floor markers queue via below_floor_scan.sh R-tier lane.
16
+
11
17
  ### 2026-07-07 | forge-harness | #sister-asset, #cross-audit, #revfactory, #harness-100, #agent-composer, #benchmarking, #linkedin, #source-verification, #diffusion-llm
12
18
  **File:** tracks/_audit/session_2026_07_07_revfactory-harness.md
13
19
  Sister-asset cross-audit of `revfactory/harness` + `revfactory/harness-100` (AX TF lead-recommended benchmarking target) vs FH — their axis = one-shot team-architecture generation + a 200-harness ready library (breadth/quick-start); FH's axis = dynamic composition (`agent-composer`) + governance (4-axis gate, irreversibility floors, continuity), which their pipeline lacks entirely. Plus LinkedIn source-verification for the operator's 2 queued insight links: diffusion-LLM paradigm-shift paper confirmed accurate (ICML 2026 Outstanding Paper Award, JustGRPO, arXiv:2601.15165 via official ICML blog); "Claude Code loop-engineering" post confirmed to be about `k021/claude-code-skills` — the *same* sister-asset Gemini already analyzed, not new content.
package/CLAUDE.md CHANGED
@@ -38,6 +38,7 @@ Four foundational assets for hub operations. **Mandatory pre-reference** before
38
38
  | `knowledge/shared/harness-core/hub_compounding_loop.md` | Feedback automation | Weekly/monthly/quarterly cycles. Axis-6 Compounding automation |
39
39
  | `knowledge/shared/dialogue/ai_dialogue_playbook.md` | Dialogue principles (should) | Session start, token efficiency, rule hierarchy, amplifier/coach dual mode |
40
40
  | `knowledge/shared/dialogue/claude_code_runtime_flow.md` | Runtime behavior (does) | Chronological flow during a session · sub-agent delegation flowchart |
41
+ | `knowledge/shared/harness-core/sonnet_floor_doctrine.md` | Canonical invariant | **Sonnet-Floor**: base ops 100% Sonnet-runnable · tier-gated capability = defect · escalation = dispatch (consent-gated), never substrate. Loop companion: `loop_engineering.md` |
41
42
 
42
43
  ## Voice / Tone — Soft Charisma (delivery layer only)
43
44
 
@@ -209,11 +210,14 @@ a session *following prose instructions* (salience-dependent — rules, onboardi
209
210
  trigger behavior), or is it mechanically enforced (hooks, scripts — tier-independent, normal 4-axis
210
211
  path, exempt)? For salience-dependent changes, verify with a **blind simulation in an isolated Agent**
211
212
  (no main-session reasoning inherited — isolation is the FH mechanism that keeps the sim honest) with
212
- `model:` pinned to the tier the change must survive on. Application strength scales with context:
213
+ `model:` pinned to the tier the change must survive on **default sim tier = Sonnet** (the base
214
+ floor every FH behavior must survive on, `sonnet_floor_doctrine.md`). Application strength scales
215
+ with context:
213
216
  - **Mode D (FH self-dev) — near-mandatory**: any salience-dependent FH asset change runs the sim
214
- before Done. Mandatory without exception when the change fixes a behavioral miss *observed* on a
215
- specific tier — sim at that same tier (the verification tier must match the failure tier; fixing on
216
- a stronger model and verifying by review alone leaves "does it fire on the weaker tier?" unanswered).
217
+ before Done, at Sonnet by default. Mandatory without exception when the change fixes a behavioral
218
+ miss *observed* on a specific tier — sim at that same tier, even below Sonnet (the verification
219
+ tier must match the failure tier; fixing on a stronger model and verifying by review alone leaves
220
+ "does it fire on the weaker tier?" unanswered).
217
221
  - **Field harness assets (templates/ propagated via Full-Harness Mode) — conditional**: sim at the
218
222
  default field tier (Sonnet) when the behavior is load-bearing (gates, onboarding, destructive/publish
219
223
  paths); skip with a one-line note for low-stakes prose.
@@ -221,8 +225,9 @@ path, exempt)? For salience-dependent changes, verify with a **blind simulation
221
225
  (hook logic, scripts, file moves — tier-independent by construction).
222
226
 
223
227
  **Autonomy floor**: the skip/run *judgment* on conditional cases is itself depth-sensitive — trust it
224
- only at opus-tier or above. A below-floor orchestrator does not silently skip: it runs the sim or asks
225
- the operator (one line), mirroring §Floor governance.
228
+ only at opus-tier or above. A below-floor orchestrator does not silently skip and does not stall:
229
+ its default is to RUN the sim (the conservative branch needs no trust); it asks the operator only when
230
+ no runnable path exists (run-first, ask-last — sonnet_floor_doctrine.md §Autonomy at Sonnet).
226
231
 
227
232
  Record sim results in the Axes 2–3 marker + sub-agent invocation log.
228
233
 
@@ -241,8 +246,9 @@ measurement is trusted.
241
246
  **Floor-tier canary (optional pre-screen — token-free, *below* the Sonnet sim)**: a local model ≤ Sonnet
242
247
  can blind-pre-screen a salience-dependent edit *before* the Sonnet dispatch is spent. **Canary, NOT gate**:
243
248
  a PASS adds cheap floor confidence and you still run the Sonnet sim; a FAIL never blocks alone. The terminal
244
- verdict stays with the frontier (Sonnet sim + opus judge) **no judge-only path**, no weak-local-judge
245
- regression of the judge-robustness principle (mechanical anchor over judge-only verdict).
249
+ verdict stays with the **Sonnet-or-higher governor bound to a mechanical anchor** (an opus judge is the
250
+ dispatch-recommended strengthener, not a requirement `sonnet_floor_doctrine.md`) **no judge-only path**,
251
+ no weak-local-judge regression of the judge-robustness principle (mechanical anchor over judge-only verdict).
246
252
 
247
253
  > **Detail**: See `knowledge/shared/harness-core/claude_md_gate_details.md §Floor-Tier-Canary` — the local
248
254
  > model/panel options, the blind-probe procedure, dogfood evidence, and the FAIL-triage (real salience gap
@@ -272,7 +278,8 @@ the target-tier sim all shared — the decorrelation value made concrete.
272
278
 
273
279
  When FH self-dev begins (an FH asset is about to change), check the **session model** and surface **one
274
280
  line**, then proceed — never block, **never switch the model** (human override inviolable): opus-tier+ →
275
- no notice · below-opus → recommend `/model opus`+ · unknown static fallback recommend. Once per session;
281
+ no notice · below-opus → **dispatch-first recommend** (keep Sonnet + route depth turns to sidecar/opus
282
+ dispatch; `/model opus` pin = secondary — `sonnet_floor_doctrine.md`) · unknown → static fallback recommend. Once per session;
276
283
  field-project (non-FH-asset) sessions never see it. Whether a session actually *escalates* (not just this
277
284
  advisory) is governed separately by `capability_escalation_consent.md`.
278
285
 
@@ -338,7 +345,8 @@ same as the FH cross-family complement. **In autonomous loops** (innovator loop-
338
345
  `/goal` · cluster orchestration): this gate is **part of the delegated pipeline**, not an
339
346
  afterthought — a load-bearing field change produced autonomously runs the lint → cross-family →
340
347
  converge loop *before* it is Done. Autonomy floor (§Floor governance): the skip/run judgment is
341
- trusted only at opus-tier+; below-floor runs the review or asks, never silently skips.
348
+ trusted only at opus-tier+; below-floor RUNS the review by default (run-first, ask-last asks only
349
+ when no runnable path exists), never silently skips (sonnet_floor_doctrine.md §Autonomy at Sonnet).
342
350
 
343
351
  > **Detail** (discretion principle · 4-face signature · gate mechanics · n=7 qasp evidence):
344
352
  > `knowledge/shared/harness-core/field_verdict_crossfamily_gate.md`.
@@ -361,6 +369,7 @@ dozen skills to invoke. Every fix is HITL — the diagnostic **proposes**, never
361
369
  | **Token / salience** | salience-split candidates (`/context-doctor` · `/salience-splitter` targets) | oversized always-loaded SKILL.md / CLAUDE.md — trim candidates |
362
370
  | **Structure** | `/harness-doctor` (L1–L4) | orphaned/redundant/decorative units, missing Done-When, ≥70% overlap |
363
371
  | **Verdict/gate degrade** | `scripts/degrade_direction_scan.sh` | a field verdict/gate helper that degrades toward permissive (advisory pre-screen) |
372
+ | **Loop-readiness** (황민호 loop-eng 5-question lens, 2026-07-10 — detail home: `loop_engineering.md`, incl. the FH loop inventory + design-time discipline) | *Loop-runtime axis — net-new vs Structure* (harness-doctor scans static form; this scans whether the path closes a loop). **Mechanical grep**: `/goal-quench`·`/loop` wiring present · check-class token declared. **Judged**: is the persisted state (card/handoff/memory) actually reloaded · is the declared check-class anchored, not judged-only · does the path halt. Done-When *presence* → see Structure row (no double-grep). **Adversarial pair** (for the judged sub-checks — decorrelated, behavior-vs-checklist): a target-tier blind sim that *runs* the path and observes whether it halts + persists, rather than re-checklisting it (the harness litmus shares this lens's axis, so it is a co-lens, not the adversary). | an agent path that *runs but doesn't loop*: no completion criterion (Done-When absent), judged-only validation with no anchor, no halt/budget guard (runaway/cost), or no state carried to the next run — the 5 questions (initiate · complete · validate · halt · persist) with 0 answers |
364
373
 
365
374
  **Output**: one ranked list, `M` (must-fix) / `S` (should-fix) / `R` (recommended) — same tiering as
366
375
  harness-doctor — each item stating *lens · file:line · one-line fix*. **Then HITL**: the operator approves
@@ -771,7 +780,8 @@ Closing phrase detected ("wrap up", "done", "good work", "end session", etc.)
771
780
  > when executing that close step.
772
781
 
773
782
  **Card-last guard**: ①–④-c (incl. ①-b open-PR sweep, ④-c handoff lifecycle) must ALL complete before
774
- ⑤ runs. Any new information produced during ①–④ (new commits from a merged self-PR, model changes,
783
+ ⑤ runs. **Mechanical floor**: `bash scripts/session_close_check.sh` before
784
+ exit 1 (card-last violated / required close artifact missing) blocks the push step until fixed. Any new information produced during ①–④ (new commits from a merged self-PR, model changes,
775
785
  new findings, a carry item flipped to DONE) feeds INTO ⑤ — card is never written mid-sequence and
776
786
  then left open for more work to accumulate after it.
777
787
 
@@ -123,3 +123,10 @@ to offload exec cost off the paid API entirely. The protocol turns hard-won econ
123
123
  protocol never flips the model itself.
124
124
  - **Sonnet floor is a floor, not a ceiling** — declined-floor-up still escalates *within* Sonnet's
125
125
  harness depth (sub-agents, good structure); it does not cap capability, only cost/trust surface.
126
+
127
+ ---
128
+
129
+ > **Canonical axiom cross-ref (2026-07-10)**: the "Sonnet floor is first-class, not degraded" stance
130
+ > this protocol operationalizes is now named — `sonnet_floor_doctrine.md` (base ops 100% Sonnet;
131
+ > tier-gated capability = defect; escalation = dispatch, consent-gated **here**). This file remains
132
+ > the consent mechanics home; the doctrine node does not restate them.
@@ -69,10 +69,13 @@ demand a strict YES/NO + one-line reason, judge whether the rule fired (mechanis
69
69
  directions — a claim checkable against that skill — re-validating that day's salience-binding fix at a
70
70
  sub-Sonnet tier).
71
71
 
72
- **FAIL-triage**: a FAIL never blocks alone — the opus orchestrator triages it as a *real salience gap* (fix
73
- the rule) vs a *floor-model quirk* (small-model loop/hallucination, per the public "Local AI is not Opus"
74
- finding + the cheap-oracle ceiling a small model adds nothing where one grep already settles the check).
75
- The terminal verdict stays with the frontier (Sonnet sim + opus judge) no judge-only path, no
72
+ **FAIL-triage**: a FAIL never blocks alone — the orchestrator (whatever tier is driving; the triage
73
+ judgment is *trusted* at opus+ and run-or-ask below, per §Floor governance) triages it as a *real
74
+ salience gap* (fix the rule) vs a *floor-model quirk* (small-model loop/hallucination, per the public
75
+ "Local AI is not Opus" finding + the cheap-oracle ceiling a small model adds nothing where one grep
76
+ already settles the check). The terminal verdict stays with the **Sonnet-or-higher governor bound to a
77
+ mechanical anchor** (Sonnet sim verdict + the anchor evidence; an opus judge is the *dispatch-recommended*
78
+ strengthener, not a requirement — Sonnet-Floor Doctrine 2026-07-10) — no judge-only path, no
76
79
  weak-local-judge regression of the judge-robustness principle (mechanical anchor over judge-only verdict).
77
80
  The cross-family-panel upgrade spec lives in the private companion store's `handoff/` design note.
78
81
 
@@ -190,11 +193,15 @@ to be modified), check the **session model** (self-identity; if the runtime with
190
193
  unknown) and surface **one line** — then proceed, never block:
191
194
 
192
195
  - Model known and opus-tier or above → no notice (already optimal).
193
- - Model known and below opus-tier → *"이 작업은 FH 자체개발(Mode D)입니다 가용 최강 모델 핀을
194
- 권장합니다 (`/model opus` 이상; 측정 근거: README §Model setup). 그대로 진행해도 floored
195
- 디스패치가 깊이 턴을 커버하지만, 세션-레벨 설계 깊이는 핀이 좌우합니다."*
196
+ - Model known and below opus-tier → **dispatch-first** (Sonnet-Floor Doctrine 2026-07-10the
197
+ primary recommendation keeps the Sonnet substrate and routes depth to dispatch; a session pin is
198
+ the *secondary* option): *"이 작업은 FH 자체개발(Mode D)입니다 — Sonnet 그대로 진행하면서 깊이
199
+ 턴(적대검증·설계리뷰)은 사이드카/opus 디스패치로 커버하는 걸 권장합니다(동의 게이트:
200
+ capability_escalation_consent). 세션 전체가 설계-깊이 중심이면 차선으로 `/model opus` 핀도
201
+ 가능합니다."*
196
202
  - Model unknown (runtime withholds identity) → static fallback: *"FH 자체개발 작업입니다 — 세션
197
- 모델이 opus 이상이 아니라면 전환을 권장합니다 (`/model opus`+)."*
203
+ 모델이 opus 미만이면 깊이 턴을 디스패치로 커버하세요(권장); 설계-깊이 세션이면 `/model opus`
204
+ 핀이 차선입니다."*
198
205
 
199
206
  **Guards**: once per session · advisory only — **never switch the session model** (human override is
200
207
  inviolable; a pin is not a cap — tier-floor resolution §Floor governance) · field-project operation
@@ -23,7 +23,7 @@ research-heavy task can pull it.
23
23
  | Rung | Capability | When it applies | Tier note |
24
24
  |---|---|---|---|
25
25
  | **1. Agentic research skill** | An autonomous multi-step researcher **present in the live session skill list** — in Claude Code that is `octo:research` (Claude Octopus, multi-AI synthesis) when installed. A native `/deep-research` was **not registered in CC in this install** (measured 2026-06-14: Skill `deep-research` → "Unknown skill"; the Claude **app** surfaces it highlighted, this CC build does not). Treat it as **app-side / operator-invoked unless a future CC build surfaces it** — re-detect from the live skill list, don't assume CC *cannot* have it (capability is install- and version-dependent — `[[feedback_verify_before_downgrade]]`) | Best when agent-fireable: runs its own search→read→synthesize loop | Self-contained; an external multi-AI path (Octopus → Gemini/Codex) bills **outside** CC's budget |
26
- | **2. Claude multi-source synthesis** | `WebSearch` + `WebFetch` tools, synthesized in-context | The always-available floor for any Claude session — no extra install | **Tier-sensitive**: synthesis depth tracks the session model. Routine survey = Sonnet default; deep analysis / contested findings = pin Opus (tier-floor, `multi_model_sidecar_strategy.md §Tier-floor resolution`) |
26
+ | **2. Claude multi-source synthesis** | `WebSearch` + `WebFetch` tools, synthesized in-context | The always-available floor for any Claude session — no extra install | **Tier-sensitive**: synthesis depth tracks the session model. Routine survey = Sonnet default; deep analysis / contested findings = dispatch-first (route the deep-read to an opus/sidecar agent, consent-gated; session pin secondary — sonnet_floor_doctrine.md; tier-floor mechanics: `multi_model_sidecar_strategy.md §Tier-floor resolution`) |
27
27
  | **3. `frontier-digest`** | The narrow specialization — HN + arxiv trend scan with FH-context synthesis | Use **only** when the research *is* AI/harness trend-scanning, not general topic research | FH-native; has its own WebSearch fallback |
28
28
 
29
29
  **Resolution rule**: detect research-heavy intent → check the live skill list → take rung 1 if a
@@ -175,6 +175,10 @@ priority: high|medium|low
175
175
 
176
176
  **forge-harness is not meant to use more tokens** — standard tier delivers meaningful improvements while minimizing token usage.
177
177
 
178
+ > Terminology guard: the S/M/L/XL **execution tier is a token-depth budget, NOT a model tier** — it is
179
+ > orthogonal to the Sonnet-floor / model-floor axis (`sonnet_floor_doctrine.md`); an XL run on Sonnet and
180
+ > an S run on Opus are both legal combinations.
181
+
178
182
  ```yaml
179
183
  EXECUTION_TIER: standard # light / standard / full / max
180
184
  ```
@@ -0,0 +1,80 @@
1
+ # Loop Engineering — the 5-question discipline and FH's loop inventory
2
+
3
+ > Companion node to the CLAUDE.md §Field-Harness Diagnostic **Loop-readiness lens** (its detail
4
+ > home) and to `sonnet_floor_doctrine.md` (PROSE loop legs are exactly where Sonnet-tier misses
5
+ > live — one spine, two lenses). Origin: 황민호 loop-eng 5-question review (2026-07-10, C-tier
6
+ > sister ledger — dedup hit on coverage, the *lens* absorbed) + Loop Engineering sister
7
+ > (`[[project_loop_engineering_sister]]`, Prompt→Context→Harness→Loop layering).
8
+
9
+ ## The 5 questions — from diagnostic to design-time discipline
10
+
11
+ A path that *runs* is not a path that *loops*. Before authoring any autonomous path (skill step
12
+ chain, routine, close sequence), answer all five **at design time** — the diagnostic lens then only
13
+ re-checks what authoring already declared:
14
+
15
+ | # | Question | FH's mechanical form |
16
+ |---|---|---|
17
+ | 1 | **Initiate** — what starts it, mechanically or by utterance? | trigger phrases (≥3, gate-checked) · hooks (SessionStart/Stop/pre-commit) · cadence rules |
18
+ | 2 | **Complete** — is there a Done-When? | Done-When with declared check class (mandatory-pass / measured / judged) — already a skill-gate item |
19
+ | 3 | **Validate** — is the check anchored, not judge-only? | mechanical anchor over judge verdict (`[[feedback_judge_robustness_mechanical_anchor]]`); judged conditions name their adversarial pairing |
20
+ | 4 | **Halt** — budget guard, convergence detection, runaway stop? | token-budget-gate · convergence-loop N-round cap · goal-quench thresholds · the 1000-agent backstop |
21
+ | 5 | **Persist** — does state reach the next run? | session card · handoff STATUS stamps · edit-manifest predict-verify · memory |
22
+
23
+ **Design-time rule**: a new autonomous path whose author cannot answer one of the five has found a
24
+ defect *before shipping it* — cheaper than the diagnostic finding it later. The field-asset
25
+ scaffold (auto_project_mapping §6) carries halt + persist stubs so field skills answer #4/#5 by
26
+ construction, the same by-construction pattern as the gate-compliant skeleton.
27
+
28
+ ## FH loop inventory — legs, enforcement class, measured gaps
29
+
30
+ Census 2026-07-10, all rows cross-family source-verified (codex gpt-5.5 — xhigh for micro/session/
31
+ weekly, high for quarterly/substrate; the substrate row's original governor self-assessment was
32
+ 4/5 REFUTED by the external pass — the census discipline earning its keep). MECH = hook /
33
+ script / exit-code (tier-independent); PROSE = salience-dependent (Sonnet-floor risk surface).
34
+
35
+ | Loop | initiate | complete | validate | halt | persist |
36
+ |---|---|---|---|---|---|
37
+ | **Micro** — goal-quench | MECH (`.active` state) | mixed (Done-When explicit; `/goal` invocation manual) | MECH (`.pending` + pipeline-conductor gate) | **PROSE** (mid-run thresholds instructional) | MECH (calibration record) |
38
+ | **Session** — close chain ①–⑥ | PROSE (closing-phrase trigger) | **mixed** (card-last + step coverage now verified by `scripts/session_close_check.sh` — exit 1 on violation; the *performing* stays with the session, the *catching* is MECH, 2026-07-10) | mixed (git/PR inputs mech, synthesis remembered) | PROSE (no budget/stop guard) | **mixed** (companion-store sync script + SessionStart STATUS map + close-check ⑤ invariant, 2026-07-10) |
39
+ | **Weekly** — harvest-loop / audit cycle | PROSE (proposal at session start) | PROSE (self-reported Done-When) | mixed (`below_floor_scan.sh` exit code; rest hand-gathered) | mixed (critic retry cap 1; no global budget) | PROSE (audit file by hand) |
40
+ | **Quarterly** — maturity roadmap | PROSE (~90d cadence, no auto-detection) | PROSE (phase gates, checked manually) | PROSE (basis-path obligations, no anchor bundle) | **PROSE** (transition deferral / Phase-regression guards exist — hub_maturity_roadmap §6.1 — prose, NOT n/a) | PROSE (roadmap doc, no canonical state file) |
41
+ | **Substrate** — self-adaptation mission | **mixed→MECH improving** (routine schedules + context-entry proposals; the "substrate-version jump" detector now EXISTS — `scripts/substrate_jump_detector.sh`, SessionStart-wired, silent-unless-jump; was a phantom until 2026-07-10) | PROSE (routine terminal states are prompt-following) | mixed (weekly change runs the 4-axis gate — marker form MECH at commit; removal approval prose/HITL) | PROSE (one-proposal-per-week · deferred-draft-PR fallback · stop-after-PR — guards exist, all prose, NOT n/a) | mixed (GitHub issue comments + draft PRs are the durable spine, plus gate markers/edit-manifest/fh_signals — not "memory") |
42
+
43
+ **Reading the map**: the 4-axis auto-gate is FH's only all-MECH loop (initiate=hook detect,
44
+ complete=all-axes-or-block, halt=missing-marker-fails, persist=marker+manifest) — and it is also
45
+ FH's most trusted loop. That correlation is the doctrine: **trust tracks mechanization, not model
46
+ tier.** The PROSE-densest loops (session close, weekly) are where the measured misses actually
47
+ occurred (card staleness 2026-07-10; audit cadence slips).
48
+
49
+ ## Case study — persist-leg mechanization (2026-07-10)
50
+
51
+ Company sessions push results to the companion store but never run the local close chain, so the
52
+ card's ⑤ update was the only reconcile point — prose, and mtime-blind: a status stamp landing
53
+ *before* a card rewrite became permanently invisible to the "newer than card" list. Fix: the
54
+ SessionStart hook now emits an **mtime-independent STATUS map** (all DONE/SUPERSEDED/RESOLVED
55
+ stamps, every session) with an explicit cross-check imperative. One measured miss → one mechanized
56
+ leg — the standing pattern (`sonnet_floor_doctrine.md §Why`).
57
+
58
+ ## Hardening backlog — evidence-threshold, NOT built speculatively
59
+
60
+ The census surfaces candidate hardenings (close-chain checklist script, weekly-audit scaffold
61
+ script, harvest-loop evidence bundle, goal-quench checkpoint files). Per the build discipline
62
+ (`[[feedback_evidence_threshold_build_discipline]]`), each is built **only when its miss is
63
+ measured** (a real slip attributable to that PROSE leg), mirroring how the SessionStart hook and
64
+ STATUS map each shipped on a production miss, not a guess. Recording the map here *is* the
65
+ instrument: the next slip finds its leg pre-diagnosed.
66
+
67
+ | Backlog item | Fires when (measured trigger) |
68
+ |---|---|
69
+ | ~~Close-chain ordered-checklist script~~ | **BUILT 2026-07-10** (`scripts/session_close_check.sh`) — operator strengthen-instruction; the miss class (card staleness) was already measured, only the build trigger was overridden (recorded, not silent) |
70
+ | Weekly-audit scaffold + data-gather script | a weekly audit missed or hand-gathered wrong window data |
71
+ | harvest-loop Step 0-b/0-c evidence check | a harvest run misses completed items despite `fh_completed_*` existing |
72
+ | goal-quench mid-run checkpoint files (70/85/95%) | a /goal run blows through a threshold unnoticed |
73
+ | ~~Substrate-jump detector~~ | **BUILT 2026-07-10** (`scripts/substrate_jump_detector.sh`, SessionStart-wired) — same operator instruction; structure-enforcing class (out-of-context drift), permanent per the durable-mechanization criterion |
74
+ | Quarterly maturity checker (`quarterly_maturity_check.sh` — criterion status + §6.1 BLOCKED emit) | a quarterly re-diagnosis is missed >90d or a phase transition skips the simplification checklist |
75
+
76
+ ## Done When (for a new/changed autonomous path)
77
+
78
+ - All 5 questions answered at design time, each leg labeled MECH or PROSE *(mandatory-pass)*.
79
+ - Any judged validate-leg names its adversarial pairing *(mandatory-pass — inherits the skill gate)*.
80
+ - PROSE legs on load-bearing paths carry a Sonnet blind-sim verdict *(measured — doctrine §ladder step 2)*.
@@ -396,7 +396,11 @@ because it has opus?". A floor is satisfied by the chosen engine's **strongest f
396
396
  **Human override is inviolable — and a pin is not a cap**: if the operator pins a session default
397
397
  (stronger or weaker), FH follows it for **session turns**; floors govern FH's **own sub-agent
398
398
  dispatches** and a session pin does not lower them — that separation *is* the Sonnet-main +
399
- Opus-dispatch doctrine (pinned-sonnet sessions still dispatch floored agents at opus).
399
+ Opus-dispatch doctrine (pinned-sonnet sessions still dispatch floored agents at opus). Canonical
400
+ axiom + defect-class + prescription ladder: `sonnet_floor_doctrine.md` (2026-07-10) — this section
401
+ remains the operating mechanics (F1/F2, floor governance) under that axiom; SKILL.md hard `model:`
402
+ pins were retired the same day (session-inherit + dispatch recommendation), agent-side dispatch
403
+ floors unchanged.
400
404
 
401
405
  **Field depth-escalation (the upward complement)**: floors push *dispatches* up automatically, but
402
406
  main-thread depth on a field session has no floor — so the templates bundle carries a Field
@@ -32,7 +32,7 @@ Three platform facts shape the design below:
32
32
  → post the digest as a COMMENT on the standing Issue "🛰️ Frontier Digest Log"
33
33
  → no repo commit, no PR (cheap, ephemeral-safe)
34
34
 
35
- [Weekly Routine] (schedule: weekly, model: OPUS required, see §4 floor note)
35
+ [Weekly Routine] (schedule: weekly, model: opus preferred / sonnet first-class, see §4 floor note)
36
36
  ① read the last 7 days of comments on the "🛰️ Frontier Digest Log" issue
37
37
  ② persona-innovator Mode F — gap + external-frontier scan against current FH assets
38
38
  ③ pick AT MOST ONE concrete, scoped improvement candidate (or none — "no proposal this week" is valid)
@@ -79,17 +79,22 @@ bash templates/regression_guard.sh --pr "$(git branch --show-current)" # Axis
79
79
  # Axis 4 edit-manifest RECORD → append today's predicted-impact line to tracks/_meta/edit_manifest.yaml
80
80
  ```
81
81
 
82
- **Floor note — why the weekly routine MUST be opus (load-bearing).** The pre-commit hook cross-checks
83
- the marker's `axis2-model` against `floor-status`: a `sonnet`/`haiku` model claiming `at-floor` or
84
- `above-floor` is **rejected**, and `below-floor` is rejected too **unless** a `below-floor-ack:` line
85
- quotes a live operator utterance. A routine runs with **no human online**, so there is no utterance to
86
- quote meaning a Sonnet weekly run **cannot legally pass the gate and cannot commit**. That dead-end
87
- is itself a hazard: under "produce a draft PR" pressure a session may improvise an unsanctioned escape
88
- (fabricate an ack, mislabel the floor, `--no-verify`). **Pin the weekly routine to opus** so the
89
- honest marker is `floor-status: at-floor` and the gate passes unattended. (If opus is ever unavailable
90
- in routines, the sanctioned fallback is: do **not** commit — attach the proposed patch to a draft PR
91
- opened via the GitHub tools as a diff in the PR body / a patch file, bypassing the local hook, and
92
- label it `gate: deferredopus re-run needed` for the operator. Never `--no-verify`.)
82
+ **Floor note — dispatch-first, Sonnet first-class (re-semanticized 2026-07-10, Sonnet-Floor
83
+ Doctrine).** The pre-commit hook cross-checks the marker's `axis2-model` against `floor-status`.
84
+ The routine's honest, unattended paths in preference order:
85
+ 1. **Dispatch the Axis-2 audit** (cross-family sidecar or an opus sub-agent, consent permitting) —
86
+ marker is `floor-status: at-floor` with the dispatched engine recorded; strongest and preferred.
87
+ 2. **Opus inline** (when the routine session itself runs at opus) `at-floor`, as before.
88
+ 3. **Sonnet inline** first-class, no operator utterance needed: `floor-status: sonnet-floor` +
89
+ an `axis2-anchor:` line naming the mechanical evidence that grounds the judged verdict (a
90
+ regression test, scan output, probe count). The marker auto-enters the weekly re-validation
91
+ queue (`below_floor_scan.sh`, R-tier advisory). The old dead-end "a Sonnet run cannot legally
92
+ commit, so it improvises an escape" is gone *because* the sanctioned lane exists; the anchor
93
+ requirement is what keeps the lane from being a free pass.
94
+ Sub-Sonnet tiers remain `below-floor` + operator ack (a routine with no human online genuinely
95
+ cannot pass there — that residual is intended). If no anchor can be produced at Sonnet either, the
96
+ fallback stays: do **not** commit — attach the patch to a draft PR opened via the GitHub tools,
97
+ label it `gate: deferred — floor re-run needed`. Never `--no-verify`.
93
98
 
94
99
  **Marker carve-out (the one sanctioned `tracks/**` write).** The two gate files
95
100
  (`tracks/_meta/.axes_23_passed_*.marker` and `tracks/_meta/edit_manifest.yaml`) live under gitignored
@@ -0,0 +1,125 @@
1
+ # Sonnet-Floor Doctrine — the harness's optimization target
2
+
3
+ > **Canonical axiom node** (operator-declared 2026-07-10). Short by design: this file names the
4
+ > invariant, its defect class, and the prescription ladder. The operating mechanics live in their
5
+ > existing homes — floor resolution & dispatch: `multi_model_sidecar_strategy.md §Tier-floor`
6
+ > (F1/F2, "Sonnet-main + Opus-dispatch"); escalation consent: `capability_escalation_consent.md`;
7
+ > mechanical enforcement: `templates/.git-hooks/pre-commit` (Axis-2 floor fields) +
8
+ > `scripts/below_floor_scan.sh`. Do not restate their details here; do not restate this axiom there.
9
+
10
+ ## The invariant
11
+
12
+ **FH's base operation must run 100% at Sonnet-tier.** Every gate, onboarding path, diagnostic,
13
+ close-chain step, and skill must fire and complete on Sonnet 5. A capability that is only
14
+ discoverable, or only fires, on Opus/Fable-tier is a **harness defect** — the same severity class
15
+ as a phantom reference. A harness exists to hold quality high *on weaker models*; if it needs the
16
+ strongest model to work at all, it has failed as a harness.
17
+
18
+ **Escalation is dispatch, never substrate.** Depth beyond Sonnet's ceiling is reached by
19
+ *recommending* a dispatch — an Opus/Fable same-family sub-agent, or a cross-family sidecar
20
+ (codex / agy) — consent-gated per `capability_escalation_consent.md`. The session substrate stays
21
+ whatever the operator chose. A Sonnet-only environment is a **first-class mode**: run everything at
22
+ Sonnet, extract the harness's maximum, and name residuals honestly (below-floor / sonnet-floor
23
+ markers) — never silently drop a capability.
24
+
25
+ ## Why this is the optimization target (measured, not aspirational)
26
+
27
+ - **H1 (2026-07-05)**: the anchor-emit harness reduced borderline verdict flips **more on weaker
28
+ models** — Flash −18.5pp vs Pro −11.1pp (within-model deltas, 3-measurement convergence). The
29
+ harness's value peaks exactly where the model is weakest; optimizing FH for the strong tier
30
+ optimizes it where it matters least.
31
+ - **Every confirmed Sonnet-tier miss in FH history was closed by mechanization or salience
32
+ hardening, never by requiring a stronger model**: the task-first companion-load miss (2026-07-05
33
+ → SessionStart hook), the tone salience gap (2026-07-08 → recorded, prompt-layer), the
34
+ card-reconcile blind spot (2026-07-10 → mtime-independent STATUS map in the hook). The doctrine
35
+ is a name for what the fix pattern already was.
36
+
37
+ ## The defect class: tier-gated capability
38
+
39
+ When auditing (harness-doctor, weekly audit, or a dedicated census), enumerate candidates with
40
+ `bash scripts/tier_census_grep.sh <files>` (word-boundary patterns + N/A-sense hints — mechanized
41
+ 2026-07-10 after a probe's naive grep false-positived on "fron**tier**"), then classify every hit:
42
+
43
+ | Class | Shape | Verdict |
44
+ |---|---|---|
45
+ | **Trust-floor** | a *judgment* (skip/run, compose/rank) is trusted only at opus+; below-floor = run the check anyway or ask | **Compatible** — Sonnet still runs everything; degrade direction is run-or-ask, never skip |
46
+ | **Availability-gate** | a capability is *absent, blocked, or dead-ended* below a tier (hard `model:` pin, "cannot pass", opus-only judge path) | **Defect** — fix via the ladder below |
47
+ | **Advisory** | recommends a tier, never blocks (Mode D Model Notice, depth-escalation notices) | Compatible — but the recommendation direction must be **dispatch-first** (keep Sonnet + dispatch the depth), with a session pin as the secondary option |
48
+
49
+ ## Prescription ladder (for a confirmed tier-gated capability)
50
+
51
+ 1. **Mechanize** — move the behavior to a hook / script / exit code. Tier-independent by
52
+ construction; the strongest fix. (SessionStart load, STATUS map, pre-commit gate.)
53
+ 2. **Salience-harden** — split, imperative pointers, turn-0 injection; then verify with a
54
+ **Sonnet blind sim** (the target-tier sim gate's default tier *is* Sonnet for this reason).
55
+ 3. **Reclassify as dispatch** — if the capability is irreducibly judgment-heavy (adversarial
56
+ depth, wide design synthesis), it becomes a *dispatch surface*: at Sonnet the harness surfaces
57
+ a one-line escalation recommendation (sidecar or floor-up, consent-gated) and proceeds at the
58
+ floor with a named residual. Silent absence is never an option.
59
+
60
+ ## Floor semantics under the doctrine
61
+
62
+ - **Sonnet = the base floor.** Meeting it is `at-floor` for base operation. Judged-class verdicts
63
+ produced at Sonnet on depth-critical roles remain **provisional** (`sonnet-floor` marker,
64
+ auto-queued for the weekly audit's re-run-or-write-off pass) — first-class ≠ free of residuals.
65
+ - **Sub-Sonnet (Haiku, local canaries)** stays `below-floor`: canary/producer roles only, explicit
66
+ ack required on gate surfaces. The doctrine raises no ceiling and lowers no guard rails there.
67
+ - **Depth ladder on a depth turn at Sonnet** (cheapest rung first — effort is depth, model is
68
+ ceiling, `[[feedback_workflow_stage_effort_routing]]`): ① **raise reasoning effort on the same
69
+ substrate** (Sonnet medium → high — free, no consent needed, no boundary crossed) → ② dispatch
70
+ an audit/research sidecar (consent permitting — cross-family preferred for decorrelation) →
71
+ ③ proceed at Sonnet-high with mechanical anchors + named residual. A hard model requirement is
72
+ never a rung.
73
+
74
+ ## Autonomy at Sonnet — run-first, ask-last (full-potential clause, 2026-07-10)
75
+
76
+ The intended FH surface — including its **full autonomous potential** (goal-quench max runs,
77
+ harvest-loop full mode, overnight loops, cluster orchestration) — must be *executable* at
78
+ Sonnet medium-high effort, not merely available-if-a-human-answers. Two rules make that safe:
79
+
80
+ - **Trust-floor degrade order is RUN → ASK, never ask-first**: where a judgment is trusted at
81
+ opus+ ("skip/run", "compose/rank"), a Sonnet session's default is to **run the full check /
82
+ present the full result** — the conservative branch that needs no trust. Asking is reserved for
83
+ the case where no mechanical or anchored path exists at all (a pure-judged fork with no anchor).
84
+ A Sonnet loop that stalls on "ask" when running-the-check was available has mis-degraded.
85
+ - **The defense is the gate layer, not the model tier**: FH's mechanical floors — pre-commit
86
+ 4-axis, pre-push Destructive-Op, prepublish scan, consent protocol, HITL irreversibility floors —
87
+ are tier-independent hooks. They hold *regardless of who is driving*, which is precisely what
88
+ makes Sonnet full-autonomy safe: **autonomy removes the prompt, never the gate** (the same
89
+ clause the Autopilot's full-autonomy mode already carries). Irreversible-surface HITL floors are
90
+ surface-class rules and do not scale down with tier — a Sonnet loop gets the same hard walls,
91
+ not softer ones.
92
+
93
+ ## What survives model evolution — the durable-mechanization criterion (operator insight, 2026-07-10)
94
+
95
+ Sidecar dispatch is the *cheap* way to chase LLM evolution (swap the engine, keep the harness), and
96
+ internal mechanization could chase capability gaps forever — so which mechanization is worth
97
+ building? Split by **what the mechanization compensates for**:
98
+
99
+ | Class | Compensates for | Fate as models improve | Examples |
100
+ |---|---|---|---|
101
+ | **Capability-compensating** | the model being *weak* — reasoning depth, salience, attention discipline | **evaporates** — scaffolding to shed (`[[feedback_frontier_substrate_self_adaptation]]`); build only on measured misses, keep cheap to delete | salience splits · turn-0 imperatives · word-boundary grep discipline (partially — see note) |
102
+ | **Structure-enforcing** | what a *perfect* model still cannot see or is still incentivized to fumble: information outside the context boundary (cross-machine state, version drift), ordering invariants across ephemeral contexts, ship-pressure optimism, irreversible surfaces | **permanent** — model evolution never fixes "the card lives on another machine" or "the runner controls what the hook sees" | STATUS map (machine boundary) · card-last check (ordering invariant) · substrate-jump detector (out-of-context drift) · fail-closed gates · consent floors |
103
+
104
+ **The test question when proposing mechanization: "would an infinitely strong model still miss
105
+ this?"** Yes → structure-enforcing, build it, it compounds. No → capability-compensating, prefer
106
+ dispatch first, mechanize only on a measured miss, and tag it shed-eligible (the substrate loop's
107
+ shed/advance pass is its consumer).
108
+
109
+ *Note on determinism*: some capability-class tools survive anyway because they are **cheaper and
110
+ deterministic** (a grep never has an attention lapse and costs nothing) — determinism is a second
111
+ survival axis, orthogonal to capability. A deterministic check that replaces a per-session judged
112
+ step keeps paying even when the model no longer needs the help.
113
+
114
+ ## Done When (for any change citing this doctrine)
115
+
116
+ - No availability-gate remains in the touched surface *(check class: measured — tier-reference
117
+ census grep, classify per the table)*.
118
+ - Salience-dependent changes pass a Sonnet blind sim *(measured — sim verdict recorded in the
119
+ Axes 2–3 marker)*.
120
+ - Depth needs express as dispatch recommendations, not requirements *(judged, pair: adversarial
121
+ review asks "where does this silently require opus?")*.
122
+
123
+ Cross-refs: `[[feedback_tier_invariant_over_treadmill]]` · `[[feedback_harness_aerodynamics_perceived_perf]]`
124
+ · `[[feedback_fh_rides_on_cc_harness]]` · `[[feedback_h1_two_tier_closure]]` · `loop_engineering.md`
125
+ (PROSE legs are where Sonnet-tier misses live — the two lenses share one spine).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chrono-meta/fh-gate",
3
- "version": "1.4.51",
3
+ "version": "1.4.52",
4
4
  "description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
5
5
  "license": "MIT",
6
6
  "keywords": [
@@ -3,7 +3,7 @@ name: deliberation
3
3
  description: Multi-perspective synthesis structure — Innovator (propose) → Devil-Advocate (challenge) → Mediator (synthesize) 3-layer execution. Outputs conditional verdicts without binary win/loss. Activates on "deliberation", "battle this out", "weigh the pros and cons", "review from multiple angles", "which side is right?". Optional deep-insight persona jurors for domain-specific views. Designed for design decisions, skill proposals, and architectural choices.
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Bash", "Agent"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  origin: fh-meta
8
8
  ---
9
9
 
@@ -3,7 +3,7 @@ name: agent-composer
3
3
  description: Reads the current work context and plans the optimal agent dispatch. Clarifies direction with 1-2 questions when unclear; infers and proceeds immediately when execution path is unclear. Runs an automatic recording gate after each Wave completes. Triggered by "compose agents", "which agent should I use?", "run in parallel", or "agent-composer".
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Bash", "Glob", "Grep"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  ---
8
8
 
9
9
  # agent-composer — Agent Composition Layer
@@ -3,7 +3,7 @@ name: apex-review
3
3
  description: Reviews a technical proposal from the perspective of organizational decision-makers (CTO, technical lead, QA lead, conference reviewers, etc.) and generates an HTML presentation deck. Outputs approval gate results per persona and connects to sim-conductor for improvement suggestions.
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Write", "Bash", "Agent"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  ---
8
8
 
9
9
  # apex-review — Decision-Maker Review Layer
@@ -3,7 +3,7 @@ name: auto-decorrelation
3
3
  description: Recruits cross-family verifier sidecars (codex, agy, local 4090 over Tailscale) for adversarial verification of load-bearing changes, maximizing model-family diversity against the orchestrator. Mechanically discovers the available sidecar panel, recruits at least one cross-family verifier when present, and degrades gracefully when none are. The governor (Claude) keeps the terminal verdict; sidecar findings must be source-grounded before acceptance. Opt-in via one-time consent, stored in the UAP; fires only on load-bearing changes. Triggered by "recruit a cross-family check", "decorrelate this verification", "use the idle sidecars to verify", "auto-decorrelation".
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Bash", "Grep", "Glob"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  ---
8
8
 
9
9
  # auto-decorrelation — Cross-Family Verifier Recruitment
@@ -94,7 +94,7 @@ When context is near the limit and you want to *preserve state* rather than rese
94
94
 
95
95
  | Current task | Recommended | Command |
96
96
  |---|---|---|
97
- | Complex design decisions · architecture review | Opus | `/model opus` |
97
+ | Complex design decisions · architecture review | Opus — dispatch-first: package into an opus/sidecar agent dispatch (consent-gated); session pin secondary | dispatch · or `/model opus` |
98
98
  | Code writing · file editing · refactoring | Sonnet (default) | — |
99
99
  | Simple file lookup · short Q&A | Haiku | `/model haiku` |
100
100
 
@@ -3,7 +3,7 @@ name: harvest-loop
3
3
  description: A self-evolution pipeline that runs automatically after field sessions end. field-harvest (pattern extraction) → contention-layer (collision signals) → [Agent(subagent_type="challenger") + persona-innovator parallel] → synthesizer (challenger/innovator collision harvest) → Critic isolated Agent (SAGE automated critique) → harness-doctor (health check) → verify-bidirectional (consistency validation) → curator (skill lifecycle management) — 8 steps. Session learnings are automatically absorbed back into the FH ecosystem so the harness evolves on its own. In the main development environment, runs automatically at session end. For external FH users, proposes execution first. Triggered by "session harvest", "learning absorption", "fh evolution", or "harvest-loop". (The phrase "run the pipeline" is ceded to pipeline-conductor to avoid a trigger collision — for end-to-end verification sweeps use pipeline-conductor.)
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Write", "Bash", "Grep", "Glob", "Agent"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  ---
8
8
 
9
9
  # harvest-loop — Field Session → FH Self-Evolution Pipeline
@@ -3,7 +3,7 @@ name: install-wizard
3
3
  description: Run when setting up a new project for the first time or onboarding after installing FH (first setup, initial configuration, onboarding start, configure project, help me set up). Performs environment detection → gap diagnosis → item-by-item suggestions → user approval → execution → acceleration baseline setup in sequence. Use --dry-run to output diagnosis report only (bg dispatch compatible).
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Write", "Bash", "Glob", "Grep", "Edit"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  category: Composability Gate
8
8
  ---
9
9
 
@@ -3,7 +3,7 @@ name: meta-prompt-builder
3
3
  description: Generates structured prompts to send to each agent in an agent dispatch plan. Triggered by "write the instructions", "what do I say to the agent?", "write the prompt for me", "meta-prompt-builder". Bridges agent-composer (which agents) and prompt content (what to say). Uses Goal/Context/Constraints/Done When structure.
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Bash", "Glob", "Grep"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  ---
8
8
 
9
9
  # meta-prompt-builder — Prompt Delegation Skill
@@ -3,7 +3,7 @@ name: sim-conductor
3
3
  description: Autonomously runs external user reaction simulations, internal audits, ideation scans, artifact validation, and quality reviews. Profiles the target artifact first, then derives task-appropriate personas, dispatches them as parallel agents, classifies findings into M/S/R tiers, and completes the pipeline through to commit automatically.
4
4
  user-invocable: true
5
5
  allowed-tools: ["Read", "Write", "Bash", "Grep", "Glob", "Agent"]
6
- model: opus
6
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
7
7
  ---
8
8
 
9
9
  # sim-conductor — Meta-Simulation Automation Orchestrator
@@ -13,7 +13,7 @@ description: >-
13
13
  "steel quench", "deep pre-completion inspection", "did it really pass?".
14
14
  user-invocable: true
15
15
  allowed-tools: ["Read", "Write", "Edit", "Bash", "Grep", "Glob", "WebSearch", "Agent"]
16
- model: opus
16
+ model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
17
17
  ---
18
18
 
19
19
  # steel-quench — All-Angle Verification Meta-Skill
@@ -10,7 +10,7 @@ complexity_routing:
10
10
  escalate_when:
11
11
  - full_revalidation
12
12
  - high_stakes
13
- - fail_verdict # AI recommendation was wrong → baseline overwrite is high-stakes, never stay at sonnet
13
+ - fail_verdict # AI recommendation was wrong → baseline overwrite is high-stakes: at Sonnet, bind the overwrite to a mechanical anchor (diff review + source re-check) and RECOMMEND an opus/sidecar dispatch (consent-gated) — Sonnet+anchor is a legitimate path (sonnet_floor_doctrine.md), silent judged-only overwrite is not
14
14
  ---
15
15
 
16
16
  # verify-bidirectional — Bidirectional Self-Validation Automation