@chrono-meta/fh-gate 1.4.51 → 1.4.52
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CATALOG.md +6 -0
- package/CLAUDE.md +21 -11
- package/knowledge/shared/harness-core/capability_escalation_consent.md +7 -0
- package/knowledge/shared/harness-core/claude_md_gate_details.md +15 -8
- package/knowledge/shared/harness-core/deep_research_capability_ladder.md +1 -1
- package/knowledge/shared/harness-core/fh_detail_protocols.md +4 -0
- package/knowledge/shared/harness-core/loop_engineering.md +80 -0
- package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +5 -1
- package/knowledge/shared/harness-core/self_evolution_routine.md +17 -12
- package/knowledge/shared/harness-core/sonnet_floor_doctrine.md +125 -0
- package/package.json +1 -1
- package/plugins/fh-commons/skills/deliberation/SKILL.md +1 -1
- package/plugins/fh-meta/skills/agent-composer/SKILL.md +1 -1
- package/plugins/fh-meta/skills/apex-review/SKILL.md +1 -1
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +1 -1
- package/plugins/fh-meta/skills/context-doctor/SKILL.md +1 -1
- package/plugins/fh-meta/skills/harvest-loop/SKILL.md +1 -1
- package/plugins/fh-meta/skills/install-wizard/SKILL.md +1 -1
- package/plugins/fh-meta/skills/meta-prompt-builder/SKILL.md +1 -1
- package/plugins/fh-meta/skills/sim-conductor/SKILL.md +1 -1
- package/plugins/fh-meta/skills/steel-quench/SKILL.md +1 -1
- package/plugins/fh-meta/skills/verify-bidirectional/SKILL.md +1 -1
package/CATALOG.md
CHANGED
|
@@ -8,6 +8,12 @@ AI reads this file first when searching past work. Open individual files for det
|
|
|
8
8
|
|
|
9
9
|
<!-- Add entries in reverse date order (newest at top) -->
|
|
10
10
|
|
|
11
|
+
### 2026-07-10 | forge-harness | #sonnet-floor, #doctrine, #loop-engineering, #tier-census, #cross-family, #pre-commit-gate, #dispatch-first
|
|
12
|
+
**File:** knowledge/shared/harness-core/sonnet_floor_doctrine.md · knowledge/shared/harness-core/loop_engineering.md
|
|
13
|
+
Encoded the operator-declared **Sonnet-Floor Doctrine** (base ops 100% Sonnet-runnable; tier-gated capability = defect; escalation = dispatch, never substrate; depth ladder = effort→dispatch→anchored-Sonnet) as a canonical axiom node, plus **loop_engineering.md** (5-question design-time discipline + FH loop inventory MECH/PROSE census + evidence-threshold hardening backlog). Cross-family evolution pass: codex gpt-5.5 xhigh repo census (T1 tier refs / T2 loop legs / T3 contradictions) + agy Gemini 3.1 Pro breadth (pattern-level only, zero citations imported — phantom-risk URLs). All 6 identified availability-gates fixed: pre-commit Axis-2 gains a **sonnet-floor lane** (anchor-required, R-tier auto-queue, 8/8 regression fixtures in scripts/test_marker_floor_lanes.sh), self_evolution weekly dead-end recast dispatch-first, Mode D notice re-directed (keep Sonnet + dispatch primary), canary opus-judge → Sonnet-governor+anchor, verify-bidirectional "never stay at sonnet" fixed, 9 SKILL.md `model: opus` hard pins retired (session-inherit). Trust-floors tightened to run-first/ask-last (full Sonnet autonomy; gates stay). Sonnet blind sims: 2 dispatched, 1 salience miss caught (loop-stub enumeration) → hardened → re-sim PASS.
|
|
14
|
+
- Decision: Sonnet = the optimization target, measured spine = H1 (harness benefit largest on weaker tiers); Opus/Fable-only capability is now a named defect class with a census discipline.
|
|
15
|
+
- Open: quarterly/substrate loop rows are governor self-assessment (R-tier external census pending); sonnet-floor markers queue via below_floor_scan.sh R-tier lane.
|
|
16
|
+
|
|
11
17
|
### 2026-07-07 | forge-harness | #sister-asset, #cross-audit, #revfactory, #harness-100, #agent-composer, #benchmarking, #linkedin, #source-verification, #diffusion-llm
|
|
12
18
|
**File:** tracks/_audit/session_2026_07_07_revfactory-harness.md
|
|
13
19
|
Sister-asset cross-audit of `revfactory/harness` + `revfactory/harness-100` (AX TF lead-recommended benchmarking target) vs FH — their axis = one-shot team-architecture generation + a 200-harness ready library (breadth/quick-start); FH's axis = dynamic composition (`agent-composer`) + governance (4-axis gate, irreversibility floors, continuity), which their pipeline lacks entirely. Plus LinkedIn source-verification for the operator's 2 queued insight links: diffusion-LLM paradigm-shift paper confirmed accurate (ICML 2026 Outstanding Paper Award, JustGRPO, arXiv:2601.15165 via official ICML blog); "Claude Code loop-engineering" post confirmed to be about `k021/claude-code-skills` — the *same* sister-asset Gemini already analyzed, not new content.
|
package/CLAUDE.md
CHANGED
|
@@ -38,6 +38,7 @@ Four foundational assets for hub operations. **Mandatory pre-reference** before
|
|
|
38
38
|
| `knowledge/shared/harness-core/hub_compounding_loop.md` | Feedback automation | Weekly/monthly/quarterly cycles. Axis-6 Compounding automation |
|
|
39
39
|
| `knowledge/shared/dialogue/ai_dialogue_playbook.md` | Dialogue principles (should) | Session start, token efficiency, rule hierarchy, amplifier/coach dual mode |
|
|
40
40
|
| `knowledge/shared/dialogue/claude_code_runtime_flow.md` | Runtime behavior (does) | Chronological flow during a session · sub-agent delegation flowchart |
|
|
41
|
+
| `knowledge/shared/harness-core/sonnet_floor_doctrine.md` | Canonical invariant | **Sonnet-Floor**: base ops 100% Sonnet-runnable · tier-gated capability = defect · escalation = dispatch (consent-gated), never substrate. Loop companion: `loop_engineering.md` |
|
|
41
42
|
|
|
42
43
|
## Voice / Tone — Soft Charisma (delivery layer only)
|
|
43
44
|
|
|
@@ -209,11 +210,14 @@ a session *following prose instructions* (salience-dependent — rules, onboardi
|
|
|
209
210
|
trigger behavior), or is it mechanically enforced (hooks, scripts — tier-independent, normal 4-axis
|
|
210
211
|
path, exempt)? For salience-dependent changes, verify with a **blind simulation in an isolated Agent**
|
|
211
212
|
(no main-session reasoning inherited — isolation is the FH mechanism that keeps the sim honest) with
|
|
212
|
-
`model:` pinned to the tier the change must survive on
|
|
213
|
+
`model:` pinned to the tier the change must survive on — **default sim tier = Sonnet** (the base
|
|
214
|
+
floor every FH behavior must survive on, `sonnet_floor_doctrine.md`). Application strength scales
|
|
215
|
+
with context:
|
|
213
216
|
- **Mode D (FH self-dev) — near-mandatory**: any salience-dependent FH asset change runs the sim
|
|
214
|
-
before Done. Mandatory without exception when the change fixes a behavioral
|
|
215
|
-
specific tier — sim at that same tier
|
|
216
|
-
a stronger model and verifying by review alone leaves
|
|
217
|
+
before Done, at Sonnet by default. Mandatory without exception when the change fixes a behavioral
|
|
218
|
+
miss *observed* on a specific tier — sim at that same tier, even below Sonnet (the verification
|
|
219
|
+
tier must match the failure tier; fixing on a stronger model and verifying by review alone leaves
|
|
220
|
+
"does it fire on the weaker tier?" unanswered).
|
|
217
221
|
- **Field harness assets (templates/ propagated via Full-Harness Mode) — conditional**: sim at the
|
|
218
222
|
default field tier (Sonnet) when the behavior is load-bearing (gates, onboarding, destructive/publish
|
|
219
223
|
paths); skip with a one-line note for low-stakes prose.
|
|
@@ -221,8 +225,9 @@ path, exempt)? For salience-dependent changes, verify with a **blind simulation
|
|
|
221
225
|
(hook logic, scripts, file moves — tier-independent by construction).
|
|
222
226
|
|
|
223
227
|
**Autonomy floor**: the skip/run *judgment* on conditional cases is itself depth-sensitive — trust it
|
|
224
|
-
only at opus-tier or above. A below-floor orchestrator does not silently skip
|
|
225
|
-
the
|
|
228
|
+
only at opus-tier or above. A below-floor orchestrator does not silently skip — and does not stall:
|
|
229
|
+
its default is to RUN the sim (the conservative branch needs no trust); it asks the operator only when
|
|
230
|
+
no runnable path exists (run-first, ask-last — sonnet_floor_doctrine.md §Autonomy at Sonnet).
|
|
226
231
|
|
|
227
232
|
Record sim results in the Axes 2–3 marker + sub-agent invocation log.
|
|
228
233
|
|
|
@@ -241,8 +246,9 @@ measurement is trusted.
|
|
|
241
246
|
**Floor-tier canary (optional pre-screen — token-free, *below* the Sonnet sim)**: a local model ≤ Sonnet
|
|
242
247
|
can blind-pre-screen a salience-dependent edit *before* the Sonnet dispatch is spent. **Canary, NOT gate**:
|
|
243
248
|
a PASS adds cheap floor confidence and you still run the Sonnet sim; a FAIL never blocks alone. The terminal
|
|
244
|
-
verdict stays with the
|
|
245
|
-
|
|
249
|
+
verdict stays with the **Sonnet-or-higher governor bound to a mechanical anchor** (an opus judge is the
|
|
250
|
+
dispatch-recommended strengthener, not a requirement — `sonnet_floor_doctrine.md`) — **no judge-only path**,
|
|
251
|
+
no weak-local-judge regression of the judge-robustness principle (mechanical anchor over judge-only verdict).
|
|
246
252
|
|
|
247
253
|
> **Detail**: See `knowledge/shared/harness-core/claude_md_gate_details.md §Floor-Tier-Canary` — the local
|
|
248
254
|
> model/panel options, the blind-probe procedure, dogfood evidence, and the FAIL-triage (real salience gap
|
|
@@ -272,7 +278,8 @@ the target-tier sim all shared — the decorrelation value made concrete.
|
|
|
272
278
|
|
|
273
279
|
When FH self-dev begins (an FH asset is about to change), check the **session model** and surface **one
|
|
274
280
|
line**, then proceed — never block, **never switch the model** (human override inviolable): opus-tier+ →
|
|
275
|
-
no notice · below-opus → recommend
|
|
281
|
+
no notice · below-opus → **dispatch-first recommend** (keep Sonnet + route depth turns to sidecar/opus
|
|
282
|
+
dispatch; `/model opus` pin = secondary — `sonnet_floor_doctrine.md`) · unknown → static fallback recommend. Once per session;
|
|
276
283
|
field-project (non-FH-asset) sessions never see it. Whether a session actually *escalates* (not just this
|
|
277
284
|
advisory) is governed separately by `capability_escalation_consent.md`.
|
|
278
285
|
|
|
@@ -338,7 +345,8 @@ same as the FH cross-family complement. **In autonomous loops** (innovator loop-
|
|
|
338
345
|
`/goal` · cluster orchestration): this gate is **part of the delegated pipeline**, not an
|
|
339
346
|
afterthought — a load-bearing field change produced autonomously runs the lint → cross-family →
|
|
340
347
|
converge loop *before* it is Done. Autonomy floor (§Floor governance): the skip/run judgment is
|
|
341
|
-
trusted only at opus-tier+; below-floor
|
|
348
|
+
trusted only at opus-tier+; below-floor RUNS the review by default (run-first, ask-last — asks only
|
|
349
|
+
when no runnable path exists), never silently skips (sonnet_floor_doctrine.md §Autonomy at Sonnet).
|
|
342
350
|
|
|
343
351
|
> **Detail** (discretion principle · 4-face signature · gate mechanics · n=7 qasp evidence):
|
|
344
352
|
> `knowledge/shared/harness-core/field_verdict_crossfamily_gate.md`.
|
|
@@ -361,6 +369,7 @@ dozen skills to invoke. Every fix is HITL — the diagnostic **proposes**, never
|
|
|
361
369
|
| **Token / salience** | salience-split candidates (`/context-doctor` · `/salience-splitter` targets) | oversized always-loaded SKILL.md / CLAUDE.md — trim candidates |
|
|
362
370
|
| **Structure** | `/harness-doctor` (L1–L4) | orphaned/redundant/decorative units, missing Done-When, ≥70% overlap |
|
|
363
371
|
| **Verdict/gate degrade** | `scripts/degrade_direction_scan.sh` | a field verdict/gate helper that degrades toward permissive (advisory pre-screen) |
|
|
372
|
+
| **Loop-readiness** (황민호 loop-eng 5-question lens, 2026-07-10 — detail home: `loop_engineering.md`, incl. the FH loop inventory + design-time discipline) | *Loop-runtime axis — net-new vs Structure* (harness-doctor scans static form; this scans whether the path closes a loop). **Mechanical grep**: `/goal-quench`·`/loop` wiring present · check-class token declared. **Judged**: is the persisted state (card/handoff/memory) actually reloaded · is the declared check-class anchored, not judged-only · does the path halt. Done-When *presence* → see Structure row (no double-grep). **Adversarial pair** (for the judged sub-checks — decorrelated, behavior-vs-checklist): a target-tier blind sim that *runs* the path and observes whether it halts + persists, rather than re-checklisting it (the harness litmus shares this lens's axis, so it is a co-lens, not the adversary). | an agent path that *runs but doesn't loop*: no completion criterion (Done-When absent), judged-only validation with no anchor, no halt/budget guard (runaway/cost), or no state carried to the next run — the 5 questions (initiate · complete · validate · halt · persist) with 0 answers |
|
|
364
373
|
|
|
365
374
|
**Output**: one ranked list, `M` (must-fix) / `S` (should-fix) / `R` (recommended) — same tiering as
|
|
366
375
|
harness-doctor — each item stating *lens · file:line · one-line fix*. **Then HITL**: the operator approves
|
|
@@ -771,7 +780,8 @@ Closing phrase detected ("wrap up", "done", "good work", "end session", etc.)
|
|
|
771
780
|
> when executing that close step.
|
|
772
781
|
|
|
773
782
|
**Card-last guard**: ①–④-c (incl. ①-b open-PR sweep, ④-c handoff lifecycle) must ALL complete before
|
|
774
|
-
⑤ runs.
|
|
783
|
+
⑤ runs. **Mechanical floor**: `bash scripts/session_close_check.sh` before ⑥ —
|
|
784
|
+
exit 1 (card-last violated / required close artifact missing) blocks the push step until fixed. Any new information produced during ①–④ (new commits from a merged self-PR, model changes,
|
|
775
785
|
new findings, a carry item flipped to DONE) feeds INTO ⑤ — card is never written mid-sequence and
|
|
776
786
|
then left open for more work to accumulate after it.
|
|
777
787
|
|
|
@@ -123,3 +123,10 @@ to offload exec cost off the paid API entirely. The protocol turns hard-won econ
|
|
|
123
123
|
protocol never flips the model itself.
|
|
124
124
|
- **Sonnet floor is a floor, not a ceiling** — declined-floor-up still escalates *within* Sonnet's
|
|
125
125
|
harness depth (sub-agents, good structure); it does not cap capability, only cost/trust surface.
|
|
126
|
+
|
|
127
|
+
---
|
|
128
|
+
|
|
129
|
+
> **Canonical axiom cross-ref (2026-07-10)**: the "Sonnet floor is first-class, not degraded" stance
|
|
130
|
+
> this protocol operationalizes is now named — `sonnet_floor_doctrine.md` (base ops 100% Sonnet;
|
|
131
|
+
> tier-gated capability = defect; escalation = dispatch, consent-gated **here**). This file remains
|
|
132
|
+
> the consent mechanics home; the doctrine node does not restate them.
|
|
@@ -69,10 +69,13 @@ demand a strict YES/NO + one-line reason, judge whether the rule fired (mechanis
|
|
|
69
69
|
directions — a claim checkable against that skill — re-validating that day's salience-binding fix at a
|
|
70
70
|
sub-Sonnet tier).
|
|
71
71
|
|
|
72
|
-
**FAIL-triage**: a FAIL never blocks alone — the
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
72
|
+
**FAIL-triage**: a FAIL never blocks alone — the orchestrator (whatever tier is driving; the triage
|
|
73
|
+
judgment is *trusted* at opus+ and run-or-ask below, per §Floor governance) triages it as a *real
|
|
74
|
+
salience gap* (fix the rule) vs a *floor-model quirk* (small-model loop/hallucination, per the public
|
|
75
|
+
"Local AI is not Opus" finding + the cheap-oracle ceiling — a small model adds nothing where one grep
|
|
76
|
+
already settles the check). The terminal verdict stays with the **Sonnet-or-higher governor bound to a
|
|
77
|
+
mechanical anchor** (Sonnet sim verdict + the anchor evidence; an opus judge is the *dispatch-recommended*
|
|
78
|
+
strengthener, not a requirement — Sonnet-Floor Doctrine 2026-07-10) — no judge-only path, no
|
|
76
79
|
weak-local-judge regression of the judge-robustness principle (mechanical anchor over judge-only verdict).
|
|
77
80
|
The cross-family-panel upgrade spec lives in the private companion store's `handoff/` design note.
|
|
78
81
|
|
|
@@ -190,11 +193,15 @@ to be modified), check the **session model** (self-identity; if the runtime with
|
|
|
190
193
|
unknown) and surface **one line** — then proceed, never block:
|
|
191
194
|
|
|
192
195
|
- Model known and opus-tier or above → no notice (already optimal).
|
|
193
|
-
- Model known and below opus-tier →
|
|
194
|
-
|
|
195
|
-
|
|
196
|
+
- Model known and below opus-tier → **dispatch-first** (Sonnet-Floor Doctrine 2026-07-10 — the
|
|
197
|
+
primary recommendation keeps the Sonnet substrate and routes depth to dispatch; a session pin is
|
|
198
|
+
the *secondary* option): *"이 작업은 FH 자체개발(Mode D)입니다 — Sonnet 그대로 진행하면서 깊이
|
|
199
|
+
턴(적대검증·설계리뷰)은 사이드카/opus 디스패치로 커버하는 걸 권장합니다(동의 게이트:
|
|
200
|
+
capability_escalation_consent). 세션 전체가 설계-깊이 중심이면 차선으로 `/model opus` 핀도
|
|
201
|
+
가능합니다."*
|
|
196
202
|
- Model unknown (runtime withholds identity) → static fallback: *"FH 자체개발 작업입니다 — 세션
|
|
197
|
-
모델이 opus
|
|
203
|
+
모델이 opus 미만이면 깊이 턴을 디스패치로 커버하세요(권장); 설계-깊이 세션이면 `/model opus`
|
|
204
|
+
핀이 차선입니다."*
|
|
198
205
|
|
|
199
206
|
**Guards**: once per session · advisory only — **never switch the session model** (human override is
|
|
200
207
|
inviolable; a pin is not a cap — tier-floor resolution §Floor governance) · field-project operation
|
|
@@ -23,7 +23,7 @@ research-heavy task can pull it.
|
|
|
23
23
|
| Rung | Capability | When it applies | Tier note |
|
|
24
24
|
|---|---|---|---|
|
|
25
25
|
| **1. Agentic research skill** | An autonomous multi-step researcher **present in the live session skill list** — in Claude Code that is `octo:research` (Claude Octopus, multi-AI synthesis) when installed. A native `/deep-research` was **not registered in CC in this install** (measured 2026-06-14: Skill `deep-research` → "Unknown skill"; the Claude **app** surfaces it highlighted, this CC build does not). Treat it as **app-side / operator-invoked unless a future CC build surfaces it** — re-detect from the live skill list, don't assume CC *cannot* have it (capability is install- and version-dependent — `[[feedback_verify_before_downgrade]]`) | Best when agent-fireable: runs its own search→read→synthesize loop | Self-contained; an external multi-AI path (Octopus → Gemini/Codex) bills **outside** CC's budget |
|
|
26
|
-
| **2. Claude multi-source synthesis** | `WebSearch` + `WebFetch` tools, synthesized in-context | The always-available floor for any Claude session — no extra install | **Tier-sensitive**: synthesis depth tracks the session model. Routine survey = Sonnet default; deep analysis / contested findings = pin
|
|
26
|
+
| **2. Claude multi-source synthesis** | `WebSearch` + `WebFetch` tools, synthesized in-context | The always-available floor for any Claude session — no extra install | **Tier-sensitive**: synthesis depth tracks the session model. Routine survey = Sonnet default; deep analysis / contested findings = dispatch-first (route the deep-read to an opus/sidecar agent, consent-gated; session pin secondary — sonnet_floor_doctrine.md; tier-floor mechanics: `multi_model_sidecar_strategy.md §Tier-floor resolution`) |
|
|
27
27
|
| **3. `frontier-digest`** | The narrow specialization — HN + arxiv trend scan with FH-context synthesis | Use **only** when the research *is* AI/harness trend-scanning, not general topic research | FH-native; has its own WebSearch fallback |
|
|
28
28
|
|
|
29
29
|
**Resolution rule**: detect research-heavy intent → check the live skill list → take rung 1 if a
|
|
@@ -175,6 +175,10 @@ priority: high|medium|low
|
|
|
175
175
|
|
|
176
176
|
**forge-harness is not meant to use more tokens** — standard tier delivers meaningful improvements while minimizing token usage.
|
|
177
177
|
|
|
178
|
+
> Terminology guard: the S/M/L/XL **execution tier is a token-depth budget, NOT a model tier** — it is
|
|
179
|
+
> orthogonal to the Sonnet-floor / model-floor axis (`sonnet_floor_doctrine.md`); an XL run on Sonnet and
|
|
180
|
+
> an S run on Opus are both legal combinations.
|
|
181
|
+
|
|
178
182
|
```yaml
|
|
179
183
|
EXECUTION_TIER: standard # light / standard / full / max
|
|
180
184
|
```
|
|
@@ -0,0 +1,80 @@
|
|
|
1
|
+
# Loop Engineering — the 5-question discipline and FH's loop inventory
|
|
2
|
+
|
|
3
|
+
> Companion node to the CLAUDE.md §Field-Harness Diagnostic **Loop-readiness lens** (its detail
|
|
4
|
+
> home) and to `sonnet_floor_doctrine.md` (PROSE loop legs are exactly where Sonnet-tier misses
|
|
5
|
+
> live — one spine, two lenses). Origin: 황민호 loop-eng 5-question review (2026-07-10, C-tier
|
|
6
|
+
> sister ledger — dedup hit on coverage, the *lens* absorbed) + Loop Engineering sister
|
|
7
|
+
> (`[[project_loop_engineering_sister]]`, Prompt→Context→Harness→Loop layering).
|
|
8
|
+
|
|
9
|
+
## The 5 questions — from diagnostic to design-time discipline
|
|
10
|
+
|
|
11
|
+
A path that *runs* is not a path that *loops*. Before authoring any autonomous path (skill step
|
|
12
|
+
chain, routine, close sequence), answer all five **at design time** — the diagnostic lens then only
|
|
13
|
+
re-checks what authoring already declared:
|
|
14
|
+
|
|
15
|
+
| # | Question | FH's mechanical form |
|
|
16
|
+
|---|---|---|
|
|
17
|
+
| 1 | **Initiate** — what starts it, mechanically or by utterance? | trigger phrases (≥3, gate-checked) · hooks (SessionStart/Stop/pre-commit) · cadence rules |
|
|
18
|
+
| 2 | **Complete** — is there a Done-When? | Done-When with declared check class (mandatory-pass / measured / judged) — already a skill-gate item |
|
|
19
|
+
| 3 | **Validate** — is the check anchored, not judge-only? | mechanical anchor over judge verdict (`[[feedback_judge_robustness_mechanical_anchor]]`); judged conditions name their adversarial pairing |
|
|
20
|
+
| 4 | **Halt** — budget guard, convergence detection, runaway stop? | token-budget-gate · convergence-loop N-round cap · goal-quench thresholds · the 1000-agent backstop |
|
|
21
|
+
| 5 | **Persist** — does state reach the next run? | session card · handoff STATUS stamps · edit-manifest predict-verify · memory |
|
|
22
|
+
|
|
23
|
+
**Design-time rule**: a new autonomous path whose author cannot answer one of the five has found a
|
|
24
|
+
defect *before shipping it* — cheaper than the diagnostic finding it later. The field-asset
|
|
25
|
+
scaffold (auto_project_mapping §6) carries halt + persist stubs so field skills answer #4/#5 by
|
|
26
|
+
construction, the same by-construction pattern as the gate-compliant skeleton.
|
|
27
|
+
|
|
28
|
+
## FH loop inventory — legs, enforcement class, measured gaps
|
|
29
|
+
|
|
30
|
+
Census 2026-07-10, all rows cross-family source-verified (codex gpt-5.5 — xhigh for micro/session/
|
|
31
|
+
weekly, high for quarterly/substrate; the substrate row's original governor self-assessment was
|
|
32
|
+
4/5 REFUTED by the external pass — the census discipline earning its keep). MECH = hook /
|
|
33
|
+
script / exit-code (tier-independent); PROSE = salience-dependent (Sonnet-floor risk surface).
|
|
34
|
+
|
|
35
|
+
| Loop | initiate | complete | validate | halt | persist |
|
|
36
|
+
|---|---|---|---|---|---|
|
|
37
|
+
| **Micro** — goal-quench | MECH (`.active` state) | mixed (Done-When explicit; `/goal` invocation manual) | MECH (`.pending` + pipeline-conductor gate) | **PROSE** (mid-run thresholds instructional) | MECH (calibration record) |
|
|
38
|
+
| **Session** — close chain ①–⑥ | PROSE (closing-phrase trigger) | **mixed** (card-last + step coverage now verified by `scripts/session_close_check.sh` — exit 1 on violation; the *performing* stays with the session, the *catching* is MECH, 2026-07-10) | mixed (git/PR inputs mech, synthesis remembered) | PROSE (no budget/stop guard) | **mixed** (companion-store sync script + SessionStart STATUS map + close-check ⑤ invariant, 2026-07-10) |
|
|
39
|
+
| **Weekly** — harvest-loop / audit cycle | PROSE (proposal at session start) | PROSE (self-reported Done-When) | mixed (`below_floor_scan.sh` exit code; rest hand-gathered) | mixed (critic retry cap 1; no global budget) | PROSE (audit file by hand) |
|
|
40
|
+
| **Quarterly** — maturity roadmap | PROSE (~90d cadence, no auto-detection) | PROSE (phase gates, checked manually) | PROSE (basis-path obligations, no anchor bundle) | **PROSE** (transition deferral / Phase-regression guards exist — hub_maturity_roadmap §6.1 — prose, NOT n/a) | PROSE (roadmap doc, no canonical state file) |
|
|
41
|
+
| **Substrate** — self-adaptation mission | **mixed→MECH improving** (routine schedules + context-entry proposals; the "substrate-version jump" detector now EXISTS — `scripts/substrate_jump_detector.sh`, SessionStart-wired, silent-unless-jump; was a phantom until 2026-07-10) | PROSE (routine terminal states are prompt-following) | mixed (weekly change runs the 4-axis gate — marker form MECH at commit; removal approval prose/HITL) | PROSE (one-proposal-per-week · deferred-draft-PR fallback · stop-after-PR — guards exist, all prose, NOT n/a) | mixed (GitHub issue comments + draft PRs are the durable spine, plus gate markers/edit-manifest/fh_signals — not "memory") |
|
|
42
|
+
|
|
43
|
+
**Reading the map**: the 4-axis auto-gate is FH's only all-MECH loop (initiate=hook detect,
|
|
44
|
+
complete=all-axes-or-block, halt=missing-marker-fails, persist=marker+manifest) — and it is also
|
|
45
|
+
FH's most trusted loop. That correlation is the doctrine: **trust tracks mechanization, not model
|
|
46
|
+
tier.** The PROSE-densest loops (session close, weekly) are where the measured misses actually
|
|
47
|
+
occurred (card staleness 2026-07-10; audit cadence slips).
|
|
48
|
+
|
|
49
|
+
## Case study — persist-leg mechanization (2026-07-10)
|
|
50
|
+
|
|
51
|
+
Company sessions push results to the companion store but never run the local close chain, so the
|
|
52
|
+
card's ⑤ update was the only reconcile point — prose, and mtime-blind: a status stamp landing
|
|
53
|
+
*before* a card rewrite became permanently invisible to the "newer than card" list. Fix: the
|
|
54
|
+
SessionStart hook now emits an **mtime-independent STATUS map** (all DONE/SUPERSEDED/RESOLVED
|
|
55
|
+
stamps, every session) with an explicit cross-check imperative. One measured miss → one mechanized
|
|
56
|
+
leg — the standing pattern (`sonnet_floor_doctrine.md §Why`).
|
|
57
|
+
|
|
58
|
+
## Hardening backlog — evidence-threshold, NOT built speculatively
|
|
59
|
+
|
|
60
|
+
The census surfaces candidate hardenings (close-chain checklist script, weekly-audit scaffold
|
|
61
|
+
script, harvest-loop evidence bundle, goal-quench checkpoint files). Per the build discipline
|
|
62
|
+
(`[[feedback_evidence_threshold_build_discipline]]`), each is built **only when its miss is
|
|
63
|
+
measured** (a real slip attributable to that PROSE leg), mirroring how the SessionStart hook and
|
|
64
|
+
STATUS map each shipped on a production miss, not a guess. Recording the map here *is* the
|
|
65
|
+
instrument: the next slip finds its leg pre-diagnosed.
|
|
66
|
+
|
|
67
|
+
| Backlog item | Fires when (measured trigger) |
|
|
68
|
+
|---|---|
|
|
69
|
+
| ~~Close-chain ordered-checklist script~~ | **BUILT 2026-07-10** (`scripts/session_close_check.sh`) — operator strengthen-instruction; the miss class (card staleness) was already measured, only the build trigger was overridden (recorded, not silent) |
|
|
70
|
+
| Weekly-audit scaffold + data-gather script | a weekly audit missed or hand-gathered wrong window data |
|
|
71
|
+
| harvest-loop Step 0-b/0-c evidence check | a harvest run misses completed items despite `fh_completed_*` existing |
|
|
72
|
+
| goal-quench mid-run checkpoint files (70/85/95%) | a /goal run blows through a threshold unnoticed |
|
|
73
|
+
| ~~Substrate-jump detector~~ | **BUILT 2026-07-10** (`scripts/substrate_jump_detector.sh`, SessionStart-wired) — same operator instruction; structure-enforcing class (out-of-context drift), permanent per the durable-mechanization criterion |
|
|
74
|
+
| Quarterly maturity checker (`quarterly_maturity_check.sh` — criterion status + §6.1 BLOCKED emit) | a quarterly re-diagnosis is missed >90d or a phase transition skips the simplification checklist |
|
|
75
|
+
|
|
76
|
+
## Done When (for a new/changed autonomous path)
|
|
77
|
+
|
|
78
|
+
- All 5 questions answered at design time, each leg labeled MECH or PROSE *(mandatory-pass)*.
|
|
79
|
+
- Any judged validate-leg names its adversarial pairing *(mandatory-pass — inherits the skill gate)*.
|
|
80
|
+
- PROSE legs on load-bearing paths carry a Sonnet blind-sim verdict *(measured — doctrine §ladder step 2)*.
|
|
@@ -396,7 +396,11 @@ because it has opus?". A floor is satisfied by the chosen engine's **strongest f
|
|
|
396
396
|
**Human override is inviolable — and a pin is not a cap**: if the operator pins a session default
|
|
397
397
|
(stronger or weaker), FH follows it for **session turns**; floors govern FH's **own sub-agent
|
|
398
398
|
dispatches** and a session pin does not lower them — that separation *is* the Sonnet-main +
|
|
399
|
-
Opus-dispatch doctrine (pinned-sonnet sessions still dispatch floored agents at opus).
|
|
399
|
+
Opus-dispatch doctrine (pinned-sonnet sessions still dispatch floored agents at opus). Canonical
|
|
400
|
+
axiom + defect-class + prescription ladder: `sonnet_floor_doctrine.md` (2026-07-10) — this section
|
|
401
|
+
remains the operating mechanics (F1/F2, floor governance) under that axiom; SKILL.md hard `model:`
|
|
402
|
+
pins were retired the same day (session-inherit + dispatch recommendation), agent-side dispatch
|
|
403
|
+
floors unchanged.
|
|
400
404
|
|
|
401
405
|
**Field depth-escalation (the upward complement)**: floors push *dispatches* up automatically, but
|
|
402
406
|
main-thread depth on a field session has no floor — so the templates bundle carries a Field
|
|
@@ -32,7 +32,7 @@ Three platform facts shape the design below:
|
|
|
32
32
|
→ post the digest as a COMMENT on the standing Issue "🛰️ Frontier Digest Log"
|
|
33
33
|
→ no repo commit, no PR (cheap, ephemeral-safe)
|
|
34
34
|
|
|
35
|
-
[Weekly Routine] (schedule: weekly, model:
|
|
35
|
+
[Weekly Routine] (schedule: weekly, model: opus preferred / sonnet first-class, see §4 floor note)
|
|
36
36
|
① read the last 7 days of comments on the "🛰️ Frontier Digest Log" issue
|
|
37
37
|
② persona-innovator Mode F — gap + external-frontier scan against current FH assets
|
|
38
38
|
③ pick AT MOST ONE concrete, scoped improvement candidate (or none — "no proposal this week" is valid)
|
|
@@ -79,17 +79,22 @@ bash templates/regression_guard.sh --pr "$(git branch --show-current)" # Axis
|
|
|
79
79
|
# Axis 4 edit-manifest RECORD → append today's predicted-impact line to tracks/_meta/edit_manifest.yaml
|
|
80
80
|
```
|
|
81
81
|
|
|
82
|
-
**Floor note —
|
|
83
|
-
the marker's `axis2-model` against `floor-status
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
82
|
+
**Floor note — dispatch-first, Sonnet first-class (re-semanticized 2026-07-10, Sonnet-Floor
|
|
83
|
+
Doctrine).** The pre-commit hook cross-checks the marker's `axis2-model` against `floor-status`.
|
|
84
|
+
The routine's honest, unattended paths in preference order:
|
|
85
|
+
1. **Dispatch the Axis-2 audit** (cross-family sidecar or an opus sub-agent, consent permitting) —
|
|
86
|
+
marker is `floor-status: at-floor` with the dispatched engine recorded; strongest and preferred.
|
|
87
|
+
2. **Opus inline** (when the routine session itself runs at opus) — `at-floor`, as before.
|
|
88
|
+
3. **Sonnet inline** — first-class, no operator utterance needed: `floor-status: sonnet-floor` +
|
|
89
|
+
an `axis2-anchor:` line naming the mechanical evidence that grounds the judged verdict (a
|
|
90
|
+
regression test, scan output, probe count). The marker auto-enters the weekly re-validation
|
|
91
|
+
queue (`below_floor_scan.sh`, R-tier advisory). The old dead-end — "a Sonnet run cannot legally
|
|
92
|
+
commit, so it improvises an escape" — is gone *because* the sanctioned lane exists; the anchor
|
|
93
|
+
requirement is what keeps the lane from being a free pass.
|
|
94
|
+
Sub-Sonnet tiers remain `below-floor` + operator ack (a routine with no human online genuinely
|
|
95
|
+
cannot pass there — that residual is intended). If no anchor can be produced at Sonnet either, the
|
|
96
|
+
fallback stays: do **not** commit — attach the patch to a draft PR opened via the GitHub tools,
|
|
97
|
+
label it `gate: deferred — floor re-run needed`. Never `--no-verify`.
|
|
93
98
|
|
|
94
99
|
**Marker carve-out (the one sanctioned `tracks/**` write).** The two gate files
|
|
95
100
|
(`tracks/_meta/.axes_23_passed_*.marker` and `tracks/_meta/edit_manifest.yaml`) live under gitignored
|
|
@@ -0,0 +1,125 @@
|
|
|
1
|
+
# Sonnet-Floor Doctrine — the harness's optimization target
|
|
2
|
+
|
|
3
|
+
> **Canonical axiom node** (operator-declared 2026-07-10). Short by design: this file names the
|
|
4
|
+
> invariant, its defect class, and the prescription ladder. The operating mechanics live in their
|
|
5
|
+
> existing homes — floor resolution & dispatch: `multi_model_sidecar_strategy.md §Tier-floor`
|
|
6
|
+
> (F1/F2, "Sonnet-main + Opus-dispatch"); escalation consent: `capability_escalation_consent.md`;
|
|
7
|
+
> mechanical enforcement: `templates/.git-hooks/pre-commit` (Axis-2 floor fields) +
|
|
8
|
+
> `scripts/below_floor_scan.sh`. Do not restate their details here; do not restate this axiom there.
|
|
9
|
+
|
|
10
|
+
## The invariant
|
|
11
|
+
|
|
12
|
+
**FH's base operation must run 100% at Sonnet-tier.** Every gate, onboarding path, diagnostic,
|
|
13
|
+
close-chain step, and skill must fire and complete on Sonnet 5. A capability that is only
|
|
14
|
+
discoverable, or only fires, on Opus/Fable-tier is a **harness defect** — the same severity class
|
|
15
|
+
as a phantom reference. A harness exists to hold quality high *on weaker models*; if it needs the
|
|
16
|
+
strongest model to work at all, it has failed as a harness.
|
|
17
|
+
|
|
18
|
+
**Escalation is dispatch, never substrate.** Depth beyond Sonnet's ceiling is reached by
|
|
19
|
+
*recommending* a dispatch — an Opus/Fable same-family sub-agent, or a cross-family sidecar
|
|
20
|
+
(codex / agy) — consent-gated per `capability_escalation_consent.md`. The session substrate stays
|
|
21
|
+
whatever the operator chose. A Sonnet-only environment is a **first-class mode**: run everything at
|
|
22
|
+
Sonnet, extract the harness's maximum, and name residuals honestly (below-floor / sonnet-floor
|
|
23
|
+
markers) — never silently drop a capability.
|
|
24
|
+
|
|
25
|
+
## Why this is the optimization target (measured, not aspirational)
|
|
26
|
+
|
|
27
|
+
- **H1 (2026-07-05)**: the anchor-emit harness reduced borderline verdict flips **more on weaker
|
|
28
|
+
models** — Flash −18.5pp vs Pro −11.1pp (within-model deltas, 3-measurement convergence). The
|
|
29
|
+
harness's value peaks exactly where the model is weakest; optimizing FH for the strong tier
|
|
30
|
+
optimizes it where it matters least.
|
|
31
|
+
- **Every confirmed Sonnet-tier miss in FH history was closed by mechanization or salience
|
|
32
|
+
hardening, never by requiring a stronger model**: the task-first companion-load miss (2026-07-05
|
|
33
|
+
→ SessionStart hook), the tone salience gap (2026-07-08 → recorded, prompt-layer), the
|
|
34
|
+
card-reconcile blind spot (2026-07-10 → mtime-independent STATUS map in the hook). The doctrine
|
|
35
|
+
is a name for what the fix pattern already was.
|
|
36
|
+
|
|
37
|
+
## The defect class: tier-gated capability
|
|
38
|
+
|
|
39
|
+
When auditing (harness-doctor, weekly audit, or a dedicated census), enumerate candidates with
|
|
40
|
+
`bash scripts/tier_census_grep.sh <files>` (word-boundary patterns + N/A-sense hints — mechanized
|
|
41
|
+
2026-07-10 after a probe's naive grep false-positived on "fron**tier**"), then classify every hit:
|
|
42
|
+
|
|
43
|
+
| Class | Shape | Verdict |
|
|
44
|
+
|---|---|---|
|
|
45
|
+
| **Trust-floor** | a *judgment* (skip/run, compose/rank) is trusted only at opus+; below-floor = run the check anyway or ask | **Compatible** — Sonnet still runs everything; degrade direction is run-or-ask, never skip |
|
|
46
|
+
| **Availability-gate** | a capability is *absent, blocked, or dead-ended* below a tier (hard `model:` pin, "cannot pass", opus-only judge path) | **Defect** — fix via the ladder below |
|
|
47
|
+
| **Advisory** | recommends a tier, never blocks (Mode D Model Notice, depth-escalation notices) | Compatible — but the recommendation direction must be **dispatch-first** (keep Sonnet + dispatch the depth), with a session pin as the secondary option |
|
|
48
|
+
|
|
49
|
+
## Prescription ladder (for a confirmed tier-gated capability)
|
|
50
|
+
|
|
51
|
+
1. **Mechanize** — move the behavior to a hook / script / exit code. Tier-independent by
|
|
52
|
+
construction; the strongest fix. (SessionStart load, STATUS map, pre-commit gate.)
|
|
53
|
+
2. **Salience-harden** — split, imperative pointers, turn-0 injection; then verify with a
|
|
54
|
+
**Sonnet blind sim** (the target-tier sim gate's default tier *is* Sonnet for this reason).
|
|
55
|
+
3. **Reclassify as dispatch** — if the capability is irreducibly judgment-heavy (adversarial
|
|
56
|
+
depth, wide design synthesis), it becomes a *dispatch surface*: at Sonnet the harness surfaces
|
|
57
|
+
a one-line escalation recommendation (sidecar or floor-up, consent-gated) and proceeds at the
|
|
58
|
+
floor with a named residual. Silent absence is never an option.
|
|
59
|
+
|
|
60
|
+
## Floor semantics under the doctrine
|
|
61
|
+
|
|
62
|
+
- **Sonnet = the base floor.** Meeting it is `at-floor` for base operation. Judged-class verdicts
|
|
63
|
+
produced at Sonnet on depth-critical roles remain **provisional** (`sonnet-floor` marker,
|
|
64
|
+
auto-queued for the weekly audit's re-run-or-write-off pass) — first-class ≠ free of residuals.
|
|
65
|
+
- **Sub-Sonnet (Haiku, local canaries)** stays `below-floor`: canary/producer roles only, explicit
|
|
66
|
+
ack required on gate surfaces. The doctrine raises no ceiling and lowers no guard rails there.
|
|
67
|
+
- **Depth ladder on a depth turn at Sonnet** (cheapest rung first — effort is depth, model is
|
|
68
|
+
ceiling, `[[feedback_workflow_stage_effort_routing]]`): ① **raise reasoning effort on the same
|
|
69
|
+
substrate** (Sonnet medium → high — free, no consent needed, no boundary crossed) → ② dispatch
|
|
70
|
+
an audit/research sidecar (consent permitting — cross-family preferred for decorrelation) →
|
|
71
|
+
③ proceed at Sonnet-high with mechanical anchors + named residual. A hard model requirement is
|
|
72
|
+
never a rung.
|
|
73
|
+
|
|
74
|
+
## Autonomy at Sonnet — run-first, ask-last (full-potential clause, 2026-07-10)
|
|
75
|
+
|
|
76
|
+
The intended FH surface — including its **full autonomous potential** (goal-quench max runs,
|
|
77
|
+
harvest-loop full mode, overnight loops, cluster orchestration) — must be *executable* at
|
|
78
|
+
Sonnet medium-high effort, not merely available-if-a-human-answers. Two rules make that safe:
|
|
79
|
+
|
|
80
|
+
- **Trust-floor degrade order is RUN → ASK, never ask-first**: where a judgment is trusted at
|
|
81
|
+
opus+ ("skip/run", "compose/rank"), a Sonnet session's default is to **run the full check /
|
|
82
|
+
present the full result** — the conservative branch that needs no trust. Asking is reserved for
|
|
83
|
+
the case where no mechanical or anchored path exists at all (a pure-judged fork with no anchor).
|
|
84
|
+
A Sonnet loop that stalls on "ask" when running-the-check was available has mis-degraded.
|
|
85
|
+
- **The defense is the gate layer, not the model tier**: FH's mechanical floors — pre-commit
|
|
86
|
+
4-axis, pre-push Destructive-Op, prepublish scan, consent protocol, HITL irreversibility floors —
|
|
87
|
+
are tier-independent hooks. They hold *regardless of who is driving*, which is precisely what
|
|
88
|
+
makes Sonnet full-autonomy safe: **autonomy removes the prompt, never the gate** (the same
|
|
89
|
+
clause the Autopilot's full-autonomy mode already carries). Irreversible-surface HITL floors are
|
|
90
|
+
surface-class rules and do not scale down with tier — a Sonnet loop gets the same hard walls,
|
|
91
|
+
not softer ones.
|
|
92
|
+
|
|
93
|
+
## What survives model evolution — the durable-mechanization criterion (operator insight, 2026-07-10)
|
|
94
|
+
|
|
95
|
+
Sidecar dispatch is the *cheap* way to chase LLM evolution (swap the engine, keep the harness), and
|
|
96
|
+
internal mechanization could chase capability gaps forever — so which mechanization is worth
|
|
97
|
+
building? Split by **what the mechanization compensates for**:
|
|
98
|
+
|
|
99
|
+
| Class | Compensates for | Fate as models improve | Examples |
|
|
100
|
+
|---|---|---|---|
|
|
101
|
+
| **Capability-compensating** | the model being *weak* — reasoning depth, salience, attention discipline | **evaporates** — scaffolding to shed (`[[feedback_frontier_substrate_self_adaptation]]`); build only on measured misses, keep cheap to delete | salience splits · turn-0 imperatives · word-boundary grep discipline (partially — see note) |
|
|
102
|
+
| **Structure-enforcing** | what a *perfect* model still cannot see or is still incentivized to fumble: information outside the context boundary (cross-machine state, version drift), ordering invariants across ephemeral contexts, ship-pressure optimism, irreversible surfaces | **permanent** — model evolution never fixes "the card lives on another machine" or "the runner controls what the hook sees" | STATUS map (machine boundary) · card-last check (ordering invariant) · substrate-jump detector (out-of-context drift) · fail-closed gates · consent floors |
|
|
103
|
+
|
|
104
|
+
**The test question when proposing mechanization: "would an infinitely strong model still miss
|
|
105
|
+
this?"** Yes → structure-enforcing, build it, it compounds. No → capability-compensating, prefer
|
|
106
|
+
dispatch first, mechanize only on a measured miss, and tag it shed-eligible (the substrate loop's
|
|
107
|
+
shed/advance pass is its consumer).
|
|
108
|
+
|
|
109
|
+
*Note on determinism*: some capability-class tools survive anyway because they are **cheaper and
|
|
110
|
+
deterministic** (a grep never has an attention lapse and costs nothing) — determinism is a second
|
|
111
|
+
survival axis, orthogonal to capability. A deterministic check that replaces a per-session judged
|
|
112
|
+
step keeps paying even when the model no longer needs the help.
|
|
113
|
+
|
|
114
|
+
## Done When (for any change citing this doctrine)
|
|
115
|
+
|
|
116
|
+
- No availability-gate remains in the touched surface *(check class: measured — tier-reference
|
|
117
|
+
census grep, classify per the table)*.
|
|
118
|
+
- Salience-dependent changes pass a Sonnet blind sim *(measured — sim verdict recorded in the
|
|
119
|
+
Axes 2–3 marker)*.
|
|
120
|
+
- Depth needs express as dispatch recommendations, not requirements *(judged, pair: adversarial
|
|
121
|
+
review asks "where does this silently require opus?")*.
|
|
122
|
+
|
|
123
|
+
Cross-refs: `[[feedback_tier_invariant_over_treadmill]]` · `[[feedback_harness_aerodynamics_perceived_perf]]`
|
|
124
|
+
· `[[feedback_fh_rides_on_cc_harness]]` · `[[feedback_h1_two_tier_closure]]` · `loop_engineering.md`
|
|
125
|
+
(PROSE legs are where Sonnet-tier misses live — the two lenses share one spine).
|
package/package.json
CHANGED
|
@@ -3,7 +3,7 @@ name: deliberation
|
|
|
3
3
|
description: Multi-perspective synthesis structure — Innovator (propose) → Devil-Advocate (challenge) → Mediator (synthesize) 3-layer execution. Outputs conditional verdicts without binary win/loss. Activates on "deliberation", "battle this out", "weigh the pros and cons", "review from multiple angles", "which side is right?". Optional deep-insight persona jurors for domain-specific views. Designed for design decisions, skill proposals, and architectural choices.
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Bash", "Agent"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
origin: fh-meta
|
|
8
8
|
---
|
|
9
9
|
|
|
@@ -3,7 +3,7 @@ name: agent-composer
|
|
|
3
3
|
description: Reads the current work context and plans the optimal agent dispatch. Clarifies direction with 1-2 questions when unclear; infers and proceeds immediately when execution path is unclear. Runs an automatic recording gate after each Wave completes. Triggered by "compose agents", "which agent should I use?", "run in parallel", or "agent-composer".
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Bash", "Glob", "Grep"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
# agent-composer — Agent Composition Layer
|
|
@@ -3,7 +3,7 @@ name: apex-review
|
|
|
3
3
|
description: Reviews a technical proposal from the perspective of organizational decision-makers (CTO, technical lead, QA lead, conference reviewers, etc.) and generates an HTML presentation deck. Outputs approval gate results per persona and connects to sim-conductor for improvement suggestions.
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Write", "Bash", "Agent"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
# apex-review — Decision-Maker Review Layer
|
|
@@ -3,7 +3,7 @@ name: auto-decorrelation
|
|
|
3
3
|
description: Recruits cross-family verifier sidecars (codex, agy, local 4090 over Tailscale) for adversarial verification of load-bearing changes, maximizing model-family diversity against the orchestrator. Mechanically discovers the available sidecar panel, recruits at least one cross-family verifier when present, and degrades gracefully when none are. The governor (Claude) keeps the terminal verdict; sidecar findings must be source-grounded before acceptance. Opt-in via one-time consent, stored in the UAP; fires only on load-bearing changes. Triggered by "recruit a cross-family check", "decorrelate this verification", "use the idle sidecars to verify", "auto-decorrelation".
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Bash", "Grep", "Glob"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
# auto-decorrelation — Cross-Family Verifier Recruitment
|
|
@@ -94,7 +94,7 @@ When context is near the limit and you want to *preserve state* rather than rese
|
|
|
94
94
|
|
|
95
95
|
| Current task | Recommended | Command |
|
|
96
96
|
|---|---|---|
|
|
97
|
-
| Complex design decisions · architecture review | Opus | `/model opus` |
|
|
97
|
+
| Complex design decisions · architecture review | Opus — dispatch-first: package into an opus/sidecar agent dispatch (consent-gated); session pin secondary | dispatch · or `/model opus` |
|
|
98
98
|
| Code writing · file editing · refactoring | Sonnet (default) | — |
|
|
99
99
|
| Simple file lookup · short Q&A | Haiku | `/model haiku` |
|
|
100
100
|
|
|
@@ -3,7 +3,7 @@ name: harvest-loop
|
|
|
3
3
|
description: A self-evolution pipeline that runs automatically after field sessions end. field-harvest (pattern extraction) → contention-layer (collision signals) → [Agent(subagent_type="challenger") + persona-innovator parallel] → synthesizer (challenger/innovator collision harvest) → Critic isolated Agent (SAGE automated critique) → harness-doctor (health check) → verify-bidirectional (consistency validation) → curator (skill lifecycle management) — 8 steps. Session learnings are automatically absorbed back into the FH ecosystem so the harness evolves on its own. In the main development environment, runs automatically at session end. For external FH users, proposes execution first. Triggered by "session harvest", "learning absorption", "fh evolution", or "harvest-loop". (The phrase "run the pipeline" is ceded to pipeline-conductor to avoid a trigger collision — for end-to-end verification sweeps use pipeline-conductor.)
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Write", "Bash", "Grep", "Glob", "Agent"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
# harvest-loop — Field Session → FH Self-Evolution Pipeline
|
|
@@ -3,7 +3,7 @@ name: install-wizard
|
|
|
3
3
|
description: Run when setting up a new project for the first time or onboarding after installing FH (first setup, initial configuration, onboarding start, configure project, help me set up). Performs environment detection → gap diagnosis → item-by-item suggestions → user approval → execution → acceleration baseline setup in sequence. Use --dry-run to output diagnosis report only (bg dispatch compatible).
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Write", "Bash", "Glob", "Grep", "Edit"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
category: Composability Gate
|
|
8
8
|
---
|
|
9
9
|
|
|
@@ -3,7 +3,7 @@ name: meta-prompt-builder
|
|
|
3
3
|
description: Generates structured prompts to send to each agent in an agent dispatch plan. Triggered by "write the instructions", "what do I say to the agent?", "write the prompt for me", "meta-prompt-builder". Bridges agent-composer (which agents) and prompt content (what to say). Uses Goal/Context/Constraints/Done When structure.
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Bash", "Glob", "Grep"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
# meta-prompt-builder — Prompt Delegation Skill
|
|
@@ -3,7 +3,7 @@ name: sim-conductor
|
|
|
3
3
|
description: Autonomously runs external user reaction simulations, internal audits, ideation scans, artifact validation, and quality reviews. Profiles the target artifact first, then derives task-appropriate personas, dispatches them as parallel agents, classifies findings into M/S/R tiers, and completes the pipeline through to commit automatically.
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Read", "Write", "Bash", "Grep", "Glob", "Agent"]
|
|
6
|
-
model: opus
|
|
6
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
# sim-conductor — Meta-Simulation Automation Orchestrator
|
|
@@ -13,7 +13,7 @@ description: >-
|
|
|
13
13
|
"steel quench", "deep pre-completion inspection", "did it really pass?".
|
|
14
14
|
user-invocable: true
|
|
15
15
|
allowed-tools: ["Read", "Write", "Edit", "Bash", "Grep", "Glob", "WebSearch", "Agent"]
|
|
16
|
-
model: opus
|
|
16
|
+
model-note: session-inherit — Sonnet base is first-class (sonnet_floor_doctrine.md); depth-critical judged steps route to dispatch (opus agent / cross-family sidecar, consent-gated), never a substrate requirement
|
|
17
17
|
---
|
|
18
18
|
|
|
19
19
|
# steel-quench — All-Angle Verification Meta-Skill
|
|
@@ -10,7 +10,7 @@ complexity_routing:
|
|
|
10
10
|
escalate_when:
|
|
11
11
|
- full_revalidation
|
|
12
12
|
- high_stakes
|
|
13
|
-
- fail_verdict # AI recommendation was wrong → baseline overwrite is high-stakes,
|
|
13
|
+
- fail_verdict # AI recommendation was wrong → baseline overwrite is high-stakes: at Sonnet, bind the overwrite to a mechanical anchor (diff review + source re-check) and RECOMMEND an opus/sidecar dispatch (consent-gated) — Sonnet+anchor is a legitimate path (sonnet_floor_doctrine.md), silent judged-only overwrite is not
|
|
14
14
|
---
|
|
15
15
|
|
|
16
16
|
# verify-bidirectional — Bidirectional Self-Validation Automation
|