@chrono-meta/fh-gate 1.4.49 → 1.4.51

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (25) hide show
  1. package/CATALOG.md +8 -2
  2. package/CHEATSHEET.md +1 -1
  3. package/CLAUDE.md +86 -15
  4. package/README.md +2 -2
  5. package/knowledge/shared/harness-core/capability_escalation_consent.md +125 -0
  6. package/knowledge/shared/harness-core/claude_md_gate_details.md +23 -0
  7. package/knowledge/shared/harness-core/fh_detail_protocols.md +23 -4
  8. package/knowledge/shared/harness-core/measurement-integrity-checklist.md +23 -3
  9. package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +29 -0
  10. package/package.json +1 -1
  11. package/plugins/fh-meta/skills/agent-composer/SKILL.md +7 -0
  12. package/plugins/fh-meta/skills/agent-composer/SKILL_detail.md +22 -0
  13. package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +10 -1
  14. package/plugins/fh-meta/skills/context-doctor/SKILL.md +3 -3
  15. package/plugins/fh-meta/skills/context-doctor/SKILL_detail.md +2 -2
  16. package/plugins/fh-meta/skills/goal-quench/SKILL.md +17 -10
  17. package/plugins/fh-meta/skills/harness-doctor/SKILL.md +2 -2
  18. package/plugins/fh-meta/skills/harvest-loop/SKILL.md +1 -1
  19. package/plugins/fh-meta/skills/install-wizard/SKILL.md +3 -1
  20. package/plugins/fh-meta/skills/phantom-quench/SKILL.md +25 -0
  21. package/plugins/fh-meta/skills/phantom-quench/SKILL_detail.md +57 -0
  22. package/plugins/fh-meta/skills/public-surface-audit/SKILL.md +36 -110
  23. package/plugins/fh-meta/skills/public-surface-audit/SKILL_detail.md +150 -0
  24. package/plugins/fh-meta/skills/{skill-splitter → salience-splitter}/SKILL.md +19 -9
  25. package/plugins/fh-meta/skills/{skill-splitter → salience-splitter}/SKILL_detail.md +6 -6
package/CATALOG.md CHANGED
@@ -8,6 +8,12 @@ AI reads this file first when searching past work. Open individual files for det
8
8
 
9
9
  <!-- Add entries in reverse date order (newest at top) -->
10
10
 
11
+ ### 2026-07-07 | forge-harness | #sister-asset, #cross-audit, #revfactory, #harness-100, #agent-composer, #benchmarking, #linkedin, #source-verification, #diffusion-llm
12
+ **File:** tracks/_audit/session_2026_07_07_revfactory-harness.md
13
+ Sister-asset cross-audit of `revfactory/harness` + `revfactory/harness-100` (AX TF lead-recommended benchmarking target) vs FH — their axis = one-shot team-architecture generation + a 200-harness ready library (breadth/quick-start); FH's axis = dynamic composition (`agent-composer`) + governance (4-axis gate, irreversibility floors, continuity), which their pipeline lacks entirely. Plus LinkedIn source-verification for the operator's 2 queued insight links: diffusion-LLM paradigm-shift paper confirmed accurate (ICML 2026 Outstanding Paper Award, JustGRPO, arXiv:2601.15165 via official ICML blog); "Claude Code loop-engineering" post confirmed to be about `k021/claude-code-skills` — the *same* sister-asset Gemini already analyzed, not new content.
14
+ - Decision: no functional import beyond a C-tier team-pattern naming label for `agent-composer` output; FH's governance moat holds (revfactory doesn't compete on that axis). Both LinkedIn links closed.
15
+ - Open: adopt 6-pattern naming in agent-composer (operator HITL); harness-100-style pre-built library stays gated behind the existing 3+-recurrence trigger.
16
+
11
17
  ### 2026-06-27 | forge-harness | #sister-asset, #cross-audit, #hermes-agent, #nous-research, #self-improving-agent, #skills, #memory, #messaging-gateway
12
18
  **File:** tracks/_contrib/session_2026_06_27_hermes-agent-nous-self-improving-cross-audit.md (committed via _contrib consent lane — authored in an ephemeral cloud session where tracks/_audit/ is gitignored/non-durable)
13
19
  Sister-asset cross-audit of **Hermes Agent (Nous Research)** vs FH, triggered by a LinkedIn post (esperer) distilling Hermes' official *Tips & Best Practices* (post = faithful doc summary, not original methodology; same summary circulates on Threads). ~90% of Hermes' best-practice surface is already present in FH (persistent memory · auto-skill-from-repetition · skill self-improvement · context economy · delegation · model selection — all grounded to `plugins/*/skills/`), and on the **self-improvement + governance** axis FH is *ahead*: Hermes *advises* "review auto-generated skills," FH *mechanically enforces* it (pre-commit 4-axis gate + steel/phantom-quench + HITL). Key honest finding — most apparent "gaps" dissolve: cron/daemon is a **deliberate FH boundary** (`self_evolution_routine.md` §8 "recommendation surface, not a daemon"), external-memory-providers **already audited** (companion-store pluggable, 2026-06-11). Only genuine absence = **messaging gateway** (Telegram/Slack daily-driver), which is a *delivery channel*, not methodology.
@@ -177,7 +183,7 @@ CC built-ins utilization imports (operator-approved; video claims verified 9/13
177
183
  ### 2026-06-10 | forge-harness | #ingest-gate, #contradiction-scan, #crossref-lint, #llm-wiki, #karpathy
178
184
  **File:** .claude/rules/sync_push_protocols.md (+ harness-doctor SKILL.md, probes.md)
179
185
  Karpathy LLM-Wiki sister-audit imports (operator-approved; convergence case n=5, citable primary source): I1 — contradiction scan as Sync step 3 (ingest gate, judged + verify-bidirectional pair): new knowledge grepped against existing claims before indexing, conflicts flagged in both files, old-claim removal is HITL. I2 — harness-doctor L4 knowledge cross-ref lint: no CATALOG entry = S-tier index orphan, no inbound ref = R-tier orphan page. Probes G-SYNC-01/G-LINT-01 added (30 total).
180
- - Decision: scale escape (W1) deliberately NOT built — watch-item with trigger (CATALOG hundreds of entries / repeated search misses); operator-preferred first remedy = skill-splitter-style CATALOG split-mapping, RAG hybrid only after that.
186
+ - Decision: scale escape (W1) deliberately NOT built — watch-item with trigger (CATALOG hundreds of entries / repeated search misses); operator-preferred first remedy = salience-splitter-style CATALOG split-mapping, RAG hybrid only after that.
181
187
  - Open: npm republish (harness-doctor SKILL.md shipped) — folded into the open 1.4.8 handoff.
182
188
 
183
189
  ### 2026-06-10 | forge-harness | #golden-probes, #offline-eval, #doc-code-coupling, #anthropic-4layer
@@ -282,7 +288,7 @@ Post-merge micro R-tier cleanup: corrected two prompt-regression probe expectati
282
288
 
283
289
  ### 2026-06-03 | forge-harness | #goal-quench, #skill-evolution, #mode-ladder, #sidecar-routing
284
290
  **File:** plugins/fh-meta/skills/goal-quench/SKILL.md
285
- Evolved goal-quench into a fluid core→pro→max mode ladder. Core: token-budget-gate + pipeline-conductor --quick; pro: +context-doctor +agent-composer; max: +plugin-recommender +cross-ecosystem-synergy-detection. Phase-1 budget verdict auto-recommends mode. Ran full-harness dogfood sweep (33 skills): fixed phantom refs, dead blocks, stale agent forks (4 deleted), trigger collision, and 3 skill-splitter splits.
291
+ Evolved goal-quench into a fluid core→pro→max mode ladder. Core: token-budget-gate + pipeline-conductor --quick; pro: +context-doctor +agent-composer; max: +plugin-recommender +cross-ecosystem-synergy-detection. Phase-1 budget verdict auto-recommends mode. Ran full-harness dogfood sweep (33 skills): fixed phantom refs, dead blocks, stale agent forks (4 deleted), trigger collision, and 3 salience-splitter splits.
286
292
  - Decision: RED tier reframed as max-mode decomposition on-ramp, not hard block
287
293
 
288
294
  ### 2026-06-02 | _audit | sister-asset, token-efficiency, compression, headroom
package/CHEATSHEET.md CHANGED
@@ -511,7 +511,7 @@ Claude agents feature
511
511
  |---|---|---|
512
512
  | `install-wizard` | First-install onboarding (zshrc, sentinels, the FH self-gate) | "first-time setup", "run the install wizard" |
513
513
  | `hub-cc-pr-reviewer` | Reads a PR diff → 8-matrix baseline-consistency check → review comment + merge call | "review this PR", "check this diff" |
514
- | `skill-splitter` | Splits an over-large SKILL.md that "does everything" into scoped files | "this skill is bloated", "SKILL.md too large", "split this skill" |
514
+ | `salience-splitter` | Splits an over-large SKILL.md that "does everything" into scoped files | "this skill is bloated", "SKILL.md too large", "split this skill" |
515
515
 
516
516
  ### Agents (sub-agents, dispatched — not slash commands)
517
517
 
package/CLAUDE.md CHANGED
@@ -142,7 +142,7 @@ Compose session-card candidates **into door ③ (field) and the 🔧 door (FH-de
142
142
 
143
143
  **Identity marker**: every greeting response (Step ②) opens with 🐿️ then an identity-revealing welcome line **on the same line** (a space after 🐿️; exact count not significant — the renderer collapses it — the invariant is *same-line*, not 🐿️ alone) — new / exploratory = "Welcome to FH." · returning = "Welcome back to FH." · operator (FH-dev state) = "The FH operator — good to see you." It is embedded in all skeletons above (do not strip it when composing doors); the exploratory branch template (`fh_detail_protocols.md` Step 2) uses the "Welcome to FH." line.
144
144
 
145
- **Guards**: explicit task-entry utterance → skip onboarding · once per session · code/debug requests → start working directly · project routing is a suggestion, mention at most once
145
+ **Guards**: explicit task-entry utterance → skip onboarding **menu** (the door skeleton / greeting) — but this **never skips the Mode D companion-store freshness load** (pull + INDEX read + card-vs-commit reconcile); that is a data-load, not the menu, and it fires even when the first message is a task (measured miss 2026-07-05: task-first entry skipped the companion-store pull → stale memory → wrong recommendations; now hook-backed via `scripts/fh_session_load.sh`, see `modes_and_value.md §Session-start freshness`) · once per session · code/debug requests → start working directly · project routing is a suggestion, mention at most once
146
146
  **Metadata-is-not-intent guard**: the trigger is the user's **typed message only**. Session metadata — branch name (auto-derived from the first message, e.g. `claude/korean-greeting-*`), repo name, file paths — is **never** a task spec and never suppresses or redirects the greeting trigger. A bare greeting fires onboarding even when the branch name looks like a feature request; if the only "task" signal lives in metadata and not in what the user typed, treat the message as a greeting and run the greeting branch + door skeleton above.
147
147
 
148
148
  ## New Skill Creation Pre-Commit Gate
@@ -270,20 +270,15 @@ the target-tier sim all shared — the decorrelation value made concrete.
270
270
 
271
271
  ### Mode D Model Notice (fires once, at the same trigger as this gate)
272
272
 
273
- The moment FH self-development work begins (= the gate's own activation trigger: an FH asset is about
274
- to be modified), check the **session model** (self-identity; if the runtime withholds it, treat as
275
- unknown) and surface **one line** then proceed, never block:
273
+ When FH self-dev begins (an FH asset is about to change), check the **session model** and surface **one
274
+ line**, then proceed — never block, **never switch the model** (human override inviolable): opus-tier+
275
+ no notice · below-opus recommend `/model opus`+ · unknown → static fallback recommend. Once per session;
276
+ field-project (non-FH-asset) sessions never see it. Whether a session actually *escalates* (not just this
277
+ advisory) is governed separately by `capability_escalation_consent.md`.
276
278
 
277
- - Model known and opus-tier or above → no notice (already optimal).
278
- - Model known and below opus-tier *"이 작업은 FH 자체개발(Mode D)입니다가용 최강 모델 핀을
279
- 권장합니다 (`/model opus` 이상; 측정 근거: README §Model setup). 그대로 진행해도 floored
280
- 디스패치가 깊이 턴을 커버하지만, 세션-레벨 설계 깊이는 핀이 좌우합니다."*
281
- - Model unknown (runtime withholds identity) → static fallback: *"FH 자체개발 작업입니다 — 세션
282
- 모델이 opus 이상이 아니라면 핀 전환을 권장합니다 (`/model opus`+)."*
283
-
284
- **Guards**: once per session · advisory only — **never switch the session model** (human override is
285
- inviolable; a pin is not a cap — tier-floor resolution §Floor governance) · field-project operation
286
- sessions (no FH asset modification) never see this notice — the Sonnet default stays friction-free.
279
+ > **Detail**: See `knowledge/shared/harness-core/claude_md_gate_details.md §Mode-D-Model-Notice` the
280
+ > exact 3-branch wording (한글), the full guards, and the capability-escalation-consent cross-refread
281
+ > when surfacing the notice.
287
282
 
288
283
  ## Field-Harness Load-Bearing Change Gate (cross-family, pre-merge)
289
284
 
@@ -348,6 +343,80 @@ trusted only at opus-tier+; below-floor runs the review or asks, never silently
348
343
  > **Detail** (discretion principle · 4-face signature · gate mechanics · n=7 qasp evidence):
349
344
  > `knowledge/shared/harness-core/field_verdict_crossfamily_gate.md`.
350
345
 
346
+ ## Field-Harness Diagnostic — "진단해줘 / 개선해줘" on a mapped project (compose → rank → HITL)
347
+
348
+ The gate above fires on a **specific field code change**. This is its **on-demand pull sibling**: when
349
+ the operator, working in a mapped project, asks to *diagnose* or *improve* the harness itself ("진단해줘",
350
+ "개선해줘", "check this project"), don't hand-pick one skill — **compose the checks FH already has into a
351
+ single ranked diagnostic list and get per-item approval.** The value is that the operator asks once and
352
+ the harness surfaces *everything* worth fixing, ranked, instead of the operator having to know which of a
353
+ dozen skills to invoke. Every fix is HITL — the diagnostic **proposes**, never auto-edits.
354
+
355
+ **Composition (no-reinvention — every row is an existing check; the diagnostic only *routes and ranks*):**
356
+
357
+ | Lens | Existing check | Catches (real examples from 2026-07-08) |
358
+ |---|---|---|
359
+ | **Confidentiality / leak** | `/public-surface-audit` (incl. Step 3c ignore-verification) | a hardcoded internal API host literal in a SKILL body; a `local_*_context.md` that is **tracked** when it should be gitignored (the gitignore-mistake class) |
360
+ | **Split integrity** | `/phantom-quench` **Step 2.7** (bidirectional) | orphan detail sections + phantom pointers in a SKILL.md ↔ SKILL_detail.md pair |
361
+ | **Token / salience** | salience-split candidates (`/context-doctor` · `/salience-splitter` targets) | oversized always-loaded SKILL.md / CLAUDE.md — trim candidates |
362
+ | **Structure** | `/harness-doctor` (L1–L4) | orphaned/redundant/decorative units, missing Done-When, ≥70% overlap |
363
+ | **Verdict/gate degrade** | `scripts/degrade_direction_scan.sh` | a field verdict/gate helper that degrades toward permissive (advisory pre-screen) |
364
+
365
+ **Output**: one ranked list, `M` (must-fix) / `S` (should-fix) / `R` (recommended) — same tiering as
366
+ harness-doctor — each item stating *lens · file:line · one-line fix*. **Then HITL**: the operator approves
367
+ per item (or a batch); an approved fix routes to the owning skill's normal path (and, if it is itself a
368
+ load-bearing field change, through the Load-Bearing Change Gate above). **Nothing is auto-fixed** — the
369
+ diagnostic's job is the *intelligent list*, the human's job is the *go*.
370
+
371
+ **Guards**: (a) fires on a **project-level** "진단/개선" ask, not a single-file edit request (those go
372
+ straight to the relevant skill); (b) **once per ask** — not a per-turn nag; (c) **company residency** —
373
+ run leak/confidentiality lenses locally, sanitize before any cross-family dispatch, and *surface*
374
+ company-sensitive findings (tracked company hosts, git-history rewrites) for operator decision rather
375
+ than auto-fixing them (dogfood 2026-07-08: the `local_pmh_context.md` tracked-company-hosts finding was
376
+ surfaced, not auto-untracked — history rewrite is the operator's call); (d) **autonomy floor** — the
377
+ compose/rank judgment is trusted at opus-tier+; below-floor, run the individual checks and present raw
378
+ rather than silently skipping a lens. Scale to the ask: a quick "뭐 고칠 거 있어?" runs the cheap
379
+ mechanical lenses (leak · split · token); "제대로 진단해줘" runs all five + harness-doctor depth.
380
+
381
+ ## Onboarding / Acceleration Autopilot — "새 프로젝트 · 하네스 작성 · 가속화" (discover → compose → rank → install-HITL)
382
+
383
+ The **install-direction twin of the Field-Harness Diagnostic**: same `compose → rank → HITL` engine, but
384
+ it decides *what to install/wire* instead of *what to fix*. When the operator enters an onboarding /
385
+ acceleration door (returning-menu ①②③: "새 프로젝트", "하네스 작성/작성해줘", "이 프로젝트 가속화",
386
+ "harness-ify", "accelerate this project"), don't hand-run one skill — **auto-discover the local state,
387
+ let the innovator center a recommend cascade, produce a ranked install plan, and gate every install.**
388
+
389
+ **Flow:**
390
+
391
+ 1. **Phase 0 — State Audit + branch (auto-discovery)**: read the target's existing `.claude/agents|skills`,
392
+ `CLAUDE.md`, mapped `tracks/`, **locally-connected sibling repos** (the env-delta SessionStart hook already
393
+ emits "N unmapped sibling repos"), and the `LOCAL_SKILL_REGISTRY` + stack/language. Then **branch**:
394
+ *new-build* (no prior harness) · *extend-existing* (harness present → found→extend, never fork) ·
395
+ *maintain* (mature harness → route to the Field-Harness Diagnostic instead). This audit-and-branch pre-step
396
+ is imported from the revfactory/harness Phase-0 State Audit (sister-audit 2026-07-07) — it tightens FH's
397
+ found→extend reflex and is the "이미 로컬에 연결돼 있으면 자동 탐색" mechanism.
398
+ 2. **Innovator-centered recommend**: `persona-innovator` centers the cascade (Mode I on acceleration / Mode F
399
+ on FH-dev), composing `plugin-recommender` (Tier 0 platform → Tier 1 official → Tier 2/3) +
400
+ `cross-ecosystem-synergy-detection` (locally-connected skills worth wiring) + inferred technical level
401
+ (conversation-cue read, also imported from revfactory) to shape *what* and *how much*.
402
+ 3. **Ranked install plan**: one list, `M`/`S`/`R`, each item = *what · why · source (Tier 0 built-in / Tier 1
403
+ official / local sibling / FH scaffold) · exact install command*. No-reinvention: an official/built-in that
404
+ covers the need ranks above a net-new scaffold.
405
+ 4. **Install — HITL, non-overwriting**: per-item approval; **never clobber an existing `.claude/`** (propose
406
+ merge/skip if present — this is FH's edge over revfactory's post-plan auto-write and harness-100's raw
407
+ `cp`). Any generated/installed FH asset runs the **4-axis gate**; a field scaffold runs
408
+ `asset-placement-gate` + `steel-quench`. **"끝까지 해줘 / 자율로 완주" → full-autonomy**: run the whole
409
+ plan under the `/goal-quench` budget+quality gate (token cost accepted by the operator), still
410
+ non-overwriting and still gated per asset — autonomy removes the per-item *prompt*, never the *gate*.
411
+
412
+ **Guards**: (a) **non-overwriting is inviolable** — the one thing both revfactory surfaces get wrong; FH
413
+ proposes merge, never clobbers; (b) **no-reinvention** — Tier 0/1 first, scaffold only what adds governance;
414
+ (c) **company residency** — discovery of a company sibling repo surfaces it, does not auto-map/leak it;
415
+ (d) **autonomy floor** — the discover/rank judgment is trusted at opus-tier+; below-floor, present the raw
416
+ recommend and ask; (e) **once per door-entry**, not a per-turn nag. This is the door ③ (accelerate) engine
417
+ and the new-project/harness-write path made autonomous — the operator asks once and the harness discovers,
418
+ ranks, and (on request) installs everything worth wiring.
419
+
351
420
  ## Irreversibility Gates — Surface-Class Degrade Invariant (shared spine of the two gates below)
352
421
 
353
422
  The two gates that follow (Pre-Publish, Destructive-Op) guard **irreversible surfaces**. The floor they
@@ -537,7 +606,7 @@ Proposal format: `"I see [X]. Want me to run /[skill] to [one-line description]?
537
606
  | "add this MCP server", "mount this MCP", "mcp.json에 추가", "connect this tool server" (external-MCP mount intent — **proactive**, fire *before* first tool call; mount intent only — a failing/erroring mounted server is `/mcp-circuit-breaker`'s row above) | `templates/.claude/rules/mcp_tool_gating.md` (name-keyed ask/allow table — never trust server annotations or names; fill §3 at mount time) |
538
607
  | "token budget", "how expensive", "estimate tokens", "will this cost a lot" | `/token-budget-gate` |
539
608
  | "did my rule change break anything", "regression check", "test harness changes" | `/prompt-regression` |
540
- | "SKILL.md too large", "split this skill", "skill is bloated", "skill file too long" | `/skill-splitter` |
609
+ | "SKILL.md too large", "split this skill", "skill is bloated", "skill file too long" | `/salience-splitter` |
541
610
  | "review for the team", "CTO review", "decision-maker", "share with leadership", "approval deck" | `/apex-review` |
542
611
  | "run full pipeline", "verify everything", "end-to-end sweep", "chain all verifications" | `/pipeline-conductor` |
543
612
  | "help me write a prompt", "build a prompt", "improve this prompt", "prompt template" | `/meta-prompt-builder` |
@@ -546,6 +615,8 @@ Proposal format: `"I see [X]. Want me to run /[skill] to [one-line description]?
546
615
  | "memory feels bloated", "clean up memory", "memory too large", "memory hygiene" | `/memory-hygiene` |
547
616
  | "ready to PR", "about to push", "merge this", "PR 올려줘", FH asset changed in session | 4-axis auto-gate (see above — runs automatically, no proposal needed) |
548
617
  | **field verdict/gate/safety/irreversible code changed** in a mapped project (function returning a verdict enum / gate exit code / safety-invariant · publish/delete/history path) — **proactive, before merge** | **Field-Harness Load-Bearing Change Gate** (see above → degrade-lint → cross-family review → converge; same rigor as FH assets, applied to field code) |
618
+ | **"진단해줘", "개선해줘", "diagnose this", "improve this harness", "check this project", "audit this project"** — said while working **in a mapped project** (not a single-file ask) | **Field-Harness Diagnostic** (see §Field-Harness Diagnostic below → compose existing checks into one ranked M/S/R list → HITL approval per item, nothing auto-fixed) |
619
+ | **"새 프로젝트", "하네스 작성해줘", "이 프로젝트 가속화", "harness-ify this", "accelerate this project"** — an onboarding/acceleration door (returning-menu ①②③) | **Onboarding / Acceleration Autopilot** (see §Onboarding / Acceleration Autopilot below → Phase 0 auto-discover + branch → innovator-centered recommend → ranked install plan → HITL per item, non-overwriting; "끝까지 자율로" → full-autonomy under /goal-quench gate) |
549
620
 
550
621
  **Guard**: Do not propose a skill that is already running. One signal = one-line proposal (no pressure). Before proposing, consult the UAP (§Operational Adaptation Loop): a skill the user has rejected 3+ times is **suppressed**, not re-proposed.
551
622
  For per-skill utterance patterns, see the relevant `SKILL.md §Trigger Phrases` section.
package/README.md CHANGED
@@ -214,7 +214,7 @@ two more signatures keep it running: `harvest-loop` (each session's lessons beco
214
214
  | `token-budget-gate` *(fh-commons)* | Pre-task token cost estimate | "How expensive is this?" |
215
215
  | `mcp-circuit-breaker` *(fh-commons)* | MCP tool failure pattern detection | "MCP keeps failing" |
216
216
  | `quench-challenger` *(fh-commons)* | Adversarial pressure-test agent | "Challenge this with a devil" |
217
- | *(+ additional assets)* | marketplace-gate · contention-layer · edit-manifest · fact-checker · goal-quench · hub-persona-auditor · install-doctor · memory-hygiene · persona-innovator · prompt-regression · public-surface-audit · skill-splitter | |
217
+ | *(+ additional assets)* | marketplace-gate · contention-layer · edit-manifest · fact-checker · goal-quench · hub-persona-auditor · install-doctor · memory-hygiene · persona-innovator · prompt-regression · public-surface-audit · salience-splitter | |
218
218
 
219
219
  | Active count | Diagnosis |
220
220
  |:---:|---|
@@ -233,7 +233,7 @@ two more signatures keep it running: `harvest-loop` (each session's lessons beco
233
233
  | Gate / Guard | `token-budget-gate` · `asset-placement-gate` · `marketplace-gate` |
234
234
  | Discovery | `plugin-recommender` · `cross-ecosystem-synergy-detection` · `frontier-digest` · `verify-bidirectional` |
235
235
  | Content / Simulation | `sim-conductor` · `apex-review` · `meta-prompt-builder` · `deep-clarify` |
236
- | Setup | `install-wizard` · `hub-cc-pr-reviewer` · `skill-splitter` |
236
+ | Setup | `install-wizard` · `hub-cc-pr-reviewer` · `salience-splitter` |
237
237
 
238
238
  > **Full phrasebook** — every skill + agent with its one-line definition and the plain-language phrase
239
239
  > that triggers it: [`CHEATSHEET.md` §12](CHEATSHEET.md#12-skills--agents--what-each-does-and-what-to-say).
@@ -0,0 +1,125 @@
1
+ # Capability-Escalation Consent Protocol
2
+
3
+ > **Principle**: any escalation that raises **cost or trust surface** — recruiting a **cross-family
4
+ > sidecar** (external model families see the work) or a **model-tier floor-up** (Sonnet → Opus, higher
5
+ > $/token) — is **consent-gated and negotiated up front**, never sprung silently. A user who was never
6
+ > asked, then finds a paid/external escalation happened, feels **blindsided (뒤통수)**. The harness
7
+ > must run **intelligently on Claude Code alone at the Sonnet floor** for anyone who declined, using
8
+ > **sub-agents** in place of cross-family sidecars — as a *first-class mode*, not a degraded one.
9
+
10
+ This is a **규약 (convention)**, not a per-call decision: the same answer holds for the whole session
11
+ once settled, and is **remembered across sessions** (UAP) so it is asked at most once.
12
+
13
+ ---
14
+
15
+ ## Two escalation axes (same consent shape)
16
+
17
+ | Axis | Escalation | Why it needs consent | Floor when declined |
18
+ |---|---|---|---|
19
+ | **(a) Cross-family sidecar** | recruit codex / agy / gemini / local-4090 as a verifier | work leaves the Claude boundary (external family sees it) + external billing | **Tier-3 CC-only sub-agent** (isolation-decorrelation, honest same-family note) |
20
+ | **(b) Model-tier floor-up** | Sonnet → Opus for a depth-heavy turn | higher $/token (Opus ≈ 3–5× Sonnet); the operator's real cost lever | **stay at Sonnet** (the established minimum-recommended floor) |
21
+
22
+ **Sonnet = minimum-recommended model is already established** — so the declined-floor is safe: the
23
+ harness is *designed* to run well at Sonnet, not crippled by it (`[[feedback_harness_aerodynamics_perceived_perf]]`).
24
+
25
+ ---
26
+
27
+ ## When consent is settled — two entry points
28
+
29
+ ### 1. Onboarding negotiation (install-wizard) — the no-surprise path
30
+
31
+ install-wizard **explicitly negotiates both axes** at setup, as named items:
32
+ - *"Allow cross-family sidecars (codex/agy/local) for adversarial verification of load-bearing changes? They add external-family decorrelation but external billing applies."* → records `sidecar_consent`.
33
+ - *"Allow automatic Sonnet→Opus floor-up on depth-heavy turns? Opus is ~3–5× the cost; declining keeps you at the Sonnet floor and asks per-occasion instead."* → records `floorup_consent`.
34
+
35
+ Settling this at onboarding is the **whole point**: the escalation is then *expected*, never a surprise
36
+ on the bill or the egress log.
37
+
38
+ ### 2. Runtime ask-once (for anyone who skipped onboarding setup) — the graceful path
39
+
40
+ A user who skipped explicit setup is **not** auto-escalated. Instead, the **first time** an escalation
41
+ is actually needed:
42
+ - **Ask once**, at the moment of need, framed with the cost/trust reason:
43
+ *"This turn needs Opus depth (higher cost) — proceed at Opus, or stay at Sonnet?"* /
44
+ *"This load-bearing change is best verified cross-family (external billing) — recruit a sidecar, or verify with CC sub-agents only?"*
45
+ - **Accept** → record consent (UAP), proceed, no re-ask.
46
+ - **Decline** → record decline (UAP), **mark the escalation "recommended only"** going forward (surface
47
+ it as a one-line recommendation when relevant, **never a re-nag**), and **proceed at the floor**
48
+ (Sonnet / Tier-3 sub-agent).
49
+
50
+ The ask fires **once per axis**; the answer is remembered. A declined axis becomes a standing floor, not
51
+ a per-turn question.
52
+
53
+ ---
54
+
55
+ ## The declined mode is first-class, not "degraded"
56
+
57
+ When an axis is declined (or no sidecar is reachable), the fallback is **the user's chosen normal
58
+ operating mode**, and must be framed that way:
59
+
60
+ - **Cross-family declined → Tier-3 CC-only sub-agent verification.** Multiple **isolated** Claude
61
+ sub-agents adversarially verify (isolation-decorrelation), with an **honest same-family note**
62
+ ("no cross-family diversity — verification is isolation-decorrelated only"). This is intelligent
63
+ operation, **not** "reduced value / degraded" language. (`[[feedback_judge_robustness_mechanical_anchor]]`
64
+ still holds: the governor keeps the terminal verdict + a mechanical anchor.)
65
+ - **Floor-up declined → stay at Sonnet**, run the turn with good harness structure. No apology framing.
66
+
67
+ **Reframe rule**: the Sidecar Resolution Protocol's Tier-3 line and auto-decorrelation's degrade ladder
68
+ must not read "no diversity; reduced value" for a *declined* user — that pathologizes their choice.
69
+ Distinguish **declined** (chosen floor — first-class) from **unavailable-but-wanted** (genuine
70
+ degrade-with-note). Only the latter carries the degrade framing.
71
+
72
+ ### Reconciliation with the corp fail-closed invariant
73
+
74
+ The pmh corp degrade-invariant (`local_pmh_context.md` auto-decorrelation Step 6) says a load-bearing
75
+ change with **no reachable cross-family panel → NOT-CONVERGED / ask operator**. That is the
76
+ **unavailable-but-expected** case (consent given, panel down). The **declined** case is different: the
77
+ user opted out, so Tier-3 sub-agent verification **is** the standing bar (with the honest same-family
78
+ note), and load-bearing changes proceed under it — not blocked. Membership: `floorup_consent`/
79
+ `sidecar_consent == declined` in the UAP routes to the first-class floor; absence of a *wanted* panel
80
+ routes to fail-closed.
81
+
82
+ ---
83
+
84
+ ## UAP persistence (behavioral pref — never domain content)
85
+
86
+ `operational_adaptation.md` UAP gains two fields:
87
+ - `sidecar_consent: accepted | declined | unset`
88
+ - `floorup_consent: accepted | declined | unset`
89
+
90
+ **READ** (session start / at point of need): `declined` → route to floor, surface as recommendation
91
+ only, no re-nag. `accepted` → escalation available (still per-need, no auto-run). `unset` → ask-once
92
+ on first need (runtime path above).
93
+ **WRITE**: on onboarding settle, or on the first runtime accept/decline.
94
+
95
+ Ephemeral/cloud sessions (UAP wiped) → operate from the **Sonnet floor + CC-only** default (the safe
96
+ floor), do not fabricate consent.
97
+
98
+ ---
99
+
100
+ ## Cost lens (why this IS the cost-governance mechanism)
101
+
102
+ The floor-up consent prompt is **the operator's cost lever made explicit** (measured: Sonnet-economized
103
+ ≈ 34× subscription; Opus-full 40–60×, `cost_report_2026-07-07`). Default Sonnet floor + escalate-Opus-
104
+ **only-where-depth-pays**-with-consent is precisely "amortize Opus on the turns that earn it." Pair with
105
+ a **free internal-model execution lane** where the environment has one (`[[feedback_multimodel_3lane_architecture]]`)
106
+ to offload exec cost off the paid API entirely. The protocol turns hard-won economizing into a mechanical default.
107
+
108
+ ---
109
+
110
+ ## Done When
111
+
112
+ - **Both axes negotiated at onboarding** (install-wizard names sidecar + floor-up as consent items). *[mandatory-pass — grep install-wizard for both items]*
113
+ - **Runtime ask-once wired** for `unset` consent at first need; decline → recommend-only + floor, no re-nag. *[judged — pair: Sonnet target-tier blind sim of a skipped-onboarding session hitting first floor-up]*
114
+ - **UAP fields present + READ applied** (`sidecar_consent`, `floorup_consent`). *[mandatory-pass — grep operational_adaptation.md]*
115
+ - **Declined-mode reframed first-class** (no "degraded/reduced-value" language on a declined user's Tier-3 path). *[judged — pair: challenger reads auto-decorrelation + Sidecar Resolution for pathologizing language]*
116
+ - **Corp fail-closed reconciled** (declined ≠ unavailable-but-wanted). *[judged — pair: the target-tier sim above exercises both branches]*
117
+
118
+ ## Guards
119
+
120
+ - **Human override inviolable** — a consent is a default, never a cap; the operator can always force a
121
+ tier/sidecar for a given turn (`[[feedback_verify_before_downgrade]]` floor governance).
122
+ - **Never auto-switch the session model** — the floor-up ASK proposes; the human acts (`/model`). The
123
+ protocol never flips the model itself.
124
+ - **Sonnet floor is a floor, not a ceiling** — declined-floor-up still escalates *within* Sonnet's
125
+ harness depth (sub-agents, good structure); it does not cap capability, only cost/trust surface.
@@ -182,3 +182,26 @@ freshness + each operator's local session-start binding.
182
182
  **Salience-dependent** — prose, not hook-enforced; on a weaker tier may silently not fire. Backstops: ⑤'s
183
183
  removal obligation + the reader-side result-file read. A hook-enforced writer-side is a future hardening
184
184
  candidate, not built today (keep the surface thin).
185
+
186
+ ## §Mode-D-Model-Notice
187
+
188
+ The moment FH self-development work begins (= the gate's own activation trigger: an FH asset is about
189
+ to be modified), check the **session model** (self-identity; if the runtime withholds it, treat as
190
+ unknown) and surface **one line** — then proceed, never block:
191
+
192
+ - Model known and opus-tier or above → no notice (already optimal).
193
+ - Model known and below opus-tier → *"이 작업은 FH 자체개발(Mode D)입니다 — 가용 최강 모델 핀을
194
+ 권장합니다 (`/model opus` 이상; 측정 근거: README §Model setup). 그대로 진행해도 floored
195
+ 디스패치가 깊이 턴을 커버하지만, 세션-레벨 설계 깊이는 핀이 좌우합니다."*
196
+ - Model unknown (runtime withholds identity) → static fallback: *"FH 자체개발 작업입니다 — 세션
197
+ 모델이 opus 이상이 아니라면 핀 전환을 권장합니다 (`/model opus`+)."*
198
+
199
+ **Guards**: once per session · advisory only — **never switch the session model** (human override is
200
+ inviolable; a pin is not a cap — tier-floor resolution §Floor governance) · field-project operation
201
+ sessions (no FH asset modification) never see this notice — the Sonnet default stays friction-free.
202
+
203
+ > **Related — capability-escalation consent**: whether a session actually *escalates* to a stronger
204
+ > model or a cross-family sidecar (not just this advisory notice) is governed separately by
205
+ > `knowledge/shared/harness-core/capability_escalation_consent.md` — the negotiated-consent protocol
206
+ > (UAP `sidecar_consent`/`floorup_consent`) that decides ask-once vs. no-surprise floor-up/sidecar use.
207
+ > This notice is the passive advisory; that doc is the active escalation gate.
@@ -37,12 +37,31 @@ ls ../ | grep -iE '(forge-harness|meta-harness|-harness|-hub)'
37
37
  ls .claude/registry/LOCAL_SKILL_REGISTRY.md 2>/dev/null
38
38
  ```
39
39
  - File exists and modified within 7 days → load into session
40
- - Missing or older than 7 days → regenerate:
40
+ - Missing or older than 7 days → regenerate.
41
+
42
+ **No hardcoded root — derive the install location (users install FH anywhere).** The projects root is
43
+ the *parent of the FH repo*, discovered at runtime, never a literal `~/projects` / `~/PycharmProjects`
44
+ (a hardcoded root silently returns 0 on any machine whose layout differs — the 2026-07-05 dead-path
45
+ `fail-open` bug: `find ~/projects` on a `~/PycharmProjects` machine → 0 catches → the registry is
46
+ overwritten empty and cross-project summon goes dark):
41
47
  ```bash
42
- find ~/projects -path "*/.claude/skills/*/SKILL.md" \
43
- -not -path "*/forge-harness/*" 2>/dev/null
48
+ HUB="${CLAUDE_PROJECT_DIR:-$(pwd)}" # FH 레포 위치 (설치 위치 무관, 감지)
49
+ ROOT="$(cd "$HUB/.." 2>/dev/null && pwd)" # 형제 프로젝트가 사는 부모 = 프로젝트 루트
50
+ # 두 레이아웃 모두 포착: .claude/skills/*/SKILL.md AND 루트-레벨 */SKILL.md (예: gstack).
51
+ # vendored(.venv·site-packages·node_modules·.git) 제외 — 없으면 playwright/streamlit 스킬까지 삼킴.
52
+ FOUND="$(find "$ROOT" -name SKILL.md \
53
+ -not -path "*/.venv/*" -not -path "*/site-packages/*" \
54
+ -not -path "*/node_modules/*" -not -path "*/.git/*" \
55
+ -not -path "$HUB/*" 2>/dev/null)" # exclude FH's own subtree by DERIVED path, not a name-literal (works when FH is cloned under any dir name)
44
56
  ```
45
- Group by project update `.claude/registry/LOCAL_SKILL_REGISTRY.md`. Propose cross-project skills when request maps to registry. Scan once per session.
57
+ Then **fail-closed** (irreversible-ish: a silent empty overwrite blinds the bus): if `$FOUND` is empty
58
+ **and** the existing registry has >0 entries, do **not** overwrite — flag `⚠️ scan returned 0 (root=$ROOT);
59
+ kept existing registry` and skip the rewrite. Only rewrite when the scan is non-empty (or the registry
60
+ was absent). Group by project (parent dir name). Record per skill: name · path · description · trigger
61
+ phrases · `requires_cwd` · `direct-executable` · `origin(FH|project|external)`+trust. **Non-FH skills are
62
+ propose-only (ask-tier), never auto-run** — a cross-project skill body is an injection surface. Propose
63
+ cross-project skills when a request maps to the registry. Scan once per session. (Detection belongs at
64
+ install too — `/install-wizard` records HUB/ROOT so the runtime never guesses; see install-wizard.)
46
65
 
47
66
  ### Step 2 — Active Proposal
48
67
 
@@ -1,7 +1,7 @@
1
1
  # Measurement-Integrity Checklist — cross-model measurement pre-flight
2
2
 
3
3
  > A cross-model measurement is only trustworthy if its **instrument** is verified first.
4
- > Measurement integrity is a *precondition*, not a result. Three observed failure modes, each with a
4
+ > Measurement integrity is a *precondition*, not a result. Four observed failure modes, each with a
5
5
  > concrete countermeasure. Consult this before any FH measurement that compares models (sims, sidecar
6
6
  > comparisons, capability-equalizer runs, the-bible model panels, A6-class experiments).
7
7
 
@@ -16,6 +16,7 @@ becomes a gate other skills invoke, revisit the weight.
16
16
  | 1 | **Silent model fallback** — passing a model *slug* silently resolved to a weaker model (e.g. an `agy` slug fell back to Flash) instead of the intended one. The run *looks* like the named model but isn't. | **Pin the display name, not the slug** (e.g. `"Gemini 3.1 Pro (High)"`, not a bare slug). Confirm the resolved identity, don't assume the slug binds. |
17
17
  | 2 | **Non-deterministic borderline verdicts** — contested/borderline cases flip across runs (observed: haiku 4/4 flip; flagship models flip too — flipping is **not** a tier signal). A single draw is noise, not a measurement. | **reps ≥ 3 on any borderline/contested verdict.** A single run on a contested case is inadmissible. Report the flip pattern (STABLE vs FLIP), not just the modal verdict. |
18
18
  | 3 | **Generic self-identity probe** — a probe any model passes ("are you working? → OK") proves nothing about *which* model answered. | **Use a discriminating probe** — one that two different models answer *differently*. A generic-pass probe is invalid. The probe is a **pattern, not a fixed string**: a probe that discriminates Opus 4.8 from Sonnet 4.6 today may both-pass a future model generation, so **re-validate the probe each model generation** (same staleness class `memory-hygiene` exists to catch). |
19
+ | 4 | **Serving-path / quantization variance** — the *same* display-name model served over two different backends (different quantization/infra) is a **different instrument** and yields materially different measurements. Observed: one GLM-5.2 model family gave effect-size delta **+0.21** when served via an internal NVFP4-quantized deployment vs **+0.08** via an OpenRouter relay — same model name, ~2.6× different effect (n=864, reps≥3). A correctly-pinned display name (item #1) is **necessary but not sufficient**. | **Pin *and record* the serving path** — backend host + quantization, not just the display name. Two runs are comparable only if the serving path matches; a name match across different infra is an implicit apples-to-oranges. When you cannot hold it fixed, **report the serving path as a measured variable**, not a constant. |
19
20
 
20
21
  ## Why these are entangled (and why they matter beyond their own scope)
21
22
 
@@ -28,10 +29,17 @@ single-draw artifact. Item #3 (discriminating probe) **embodies** the judge-robu
28
29
  mechanical-anchor principle — don't trust self-reported identity, prove it discriminatingly
29
30
  ([[feedback_judge_robustness_mechanical_anchor]]).
30
31
 
32
+ Item #4 (serving-path variance) **sharpens** item #1 into a two-part identity: #1 catches the *wrong
33
+ model* (a slug that fell back); #4 catches the *right model on the wrong instrument* (a correct name
34
+ served over a different quantization/backend). The verified identity a measurement records is therefore
35
+ **name + serving path**, not name alone — a family-decorrelation claim (cross-family sidecar) is only
36
+ sound once the serving path of each family is itself pinned, else "different family" silently smuggles
37
+ "different infra" ([[reference_measurement_serving_path_variance]]).
38
+
31
39
  ## Done When
32
40
 
33
- - The checklist enumerates all three failure modes, each with its countermeasure.
34
- *Check class: mandatory-pass (binary — three items present, each with a countermeasure).*
41
+ - The checklist enumerates all four failure modes, each with its countermeasure.
42
+ *Check class: mandatory-pass (binary — four items present, each with a countermeasure).*
35
43
  - The probe item specifies a **discriminating** test and rejects generic probes.
36
44
  *Check class: judged, pair: a probe that two different models both pass must FAIL this check; a
37
45
  discriminating one must distinguish them.*
@@ -57,3 +65,15 @@ mechanical log.
57
65
  ambiguity). Sister findings: [[feedback_correlated_blindspot_union_over_majority]] (reps≥3 prerequisite),
58
66
  [[feedback_judge_robustness_mechanical_anchor]] (discriminating-probe = mechanical anchor),
59
67
  [[reference_agy_model_catalog]] (display-name pin — agy slug fallback documented there).
68
+
69
+ **#4 added** (2026-07-05): serving-path variance surfaced in a cross-family verdict-invariance run
70
+ (n=864, borderline fixtures × 2 conditions × K=6 paraphrase × reps≥3). An identical GLM-5.2 model name
71
+ served over an internal NVFP4-quantized deployment vs an OpenRouter relay gave +0.21 vs +0.08
72
+ effect-size delta — quantifying that "same model name ⇒ same measurement" is false. Provenance +
73
+ generalizable finding: [[reference_measurement_serving_path_variance]].
74
+
75
+ **External corroboration** (2026-07): the local-LLM community independently reports the same hazard —
76
+ practitioners conflate "running model X" with running a *pruned/quantized derivative* of X (aggressive
77
+ low-bit quantization + expert pruning measurably degrade long-context quality while the model *name* is
78
+ unchanged). This is a general measurement pitfall, not FH-specific: a leaderboard or replication that
79
+ pins only the display name silently compares different instruments across serving paths.
@@ -167,6 +167,35 @@ buff and degrades that model's realized intelligence** — not just Claude's. Th
167
167
  is autocomplete/QA only, which is exactly its demoted role. (Derived 2026-07-03, operator + cross-vendor
168
168
  Gemini concurrence; extends the governor=native-CC point to every vendor.)
169
169
 
170
+ ### Batch-judging corollary — the native harness is for interactive/agentic work, not batch scoring
171
+
172
+ The vendor-native harness gives a model its highest capability for **interactive, agentic** tasks
173
+ (repo-grounded audit, multi-step design, tool-use) — but that *same* agentic loop is a **liability for
174
+ high-volume batch judging**: a deterministic verdict emitted over N fixtures, where there is nothing for a
175
+ tool-use loop to do. Measured 2026-07-04 (H1 verdict-invariance run): the native `codex exec` spins a full
176
+ agentic session per judge (hooks + reasoning ≈ an order of magnitude more tokens than a bare completion),
177
+ and native `agy -p` (once its headless permission-wait is cleared) returns *agentic prose* — a "Summary of
178
+ Work" — rather than a parseable last-line verdict. Both **complete**, but at a cost/parse profile wrong for
179
+ batch.
180
+
181
+ So the dispatch splits by *shape of the task*, not just by family:
182
+
183
+ - **Batch cross-family judging** (steel-quench Step 0.6 verdict-invariance, auto-decorrelation over many
184
+ items, any fixed-fixture flip count) → **clean completion APIs** (OpenRouter, model pinned by
185
+ *display-name* + `served`-field silent-route check) **+ free local** (a 4090 ollama endpoint). Clean,
186
+ cheap, parseable, per-call pinnable. *Caveat*: a local thinking model needs a large enough output budget
187
+ or it truncates inside `<think>` and emits an empty verdict — a config axis, not a capacity limit.
188
+ - **Interactive / agentic verification** (repo-grounded catching, the divergence audits where cross-family
189
+ disagreement *localizes* a bug) → the **native CLIs** (`codex`, `agy`/Antigravity), where the harness
190
+ earns its overhead.
191
+
192
+ This is **not** a contradiction of the harness-depth thesis — it *is* it. The harness lifts capability
193
+ exactly where judgment + tools + iteration matter; for a one-shot self-contained verdict the loop has no
194
+ work, so its depth becomes pure cost. Pick the naked API for batch scoring, the native harness for agentic
195
+ audit. (Derived 2026-07-04, operator + H1 measurement; the batch-side dual of the vendor-native thesis
196
+ above. The native-harness Gemini path via `agy` is headless-usable again once tool-permission auto-proceed
197
+ is set — see [[reference_agy_model_catalog]] for the pin/permission mechanics.)
198
+
170
199
  **Maintenance-Cost Rule** — a compatibility layer is cheap as a *thin entrypoint*, expensive when it
171
200
  *duplicates canonical knowledge*. The test:
172
201
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chrono-meta/fh-gate",
3
- "version": "1.4.49",
3
+ "version": "1.4.51",
4
4
  "description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
5
5
  "license": "MIT",
6
6
  "keywords": [
@@ -157,6 +157,13 @@ agent-composer — Composition Plan
157
157
  Execute? (Y: run all / E: edit then run / N: cancel)
158
158
  ```
159
159
 
160
+ Optionally name the composition's shape in the plan header (`Composition pattern: {name}`) so the
161
+ operator recognizes the structure at a glance — a recognition aid, not a routing input.
162
+
163
+ > **Detail**: See `SKILL_detail.md §Composition-Pattern-Labels` — the six revfactory-vocabulary labels
164
+ > (Pipeline / Fan-out-in / Expert Pool / Producer-Reviewer / Supervisor / Hierarchical) mapped to the
165
+ > existing FH constructs they name — read when labeling a plan header.
166
+
160
167
  ---
161
168
 
162
169
  ## Step 2.5 — Model Routing Decision (complexity_routing)
@@ -462,3 +462,25 @@ agent-composer also acts as Curator — surveys existing agents/skills/assets an
462
462
  - **External positioning**: Independent convergence with hermes-agent (Nous Research) curator.py pattern
463
463
 
464
464
  > Architecture basis: Anthropic [Harness Design for Long-Running Apps](https://www.anthropic.com/engineering/harness-design-long-running-apps) — single agent ($9, fails) vs. multi-agent harness ($200, perfect). Cost gap justified by quality gap.
465
+
466
+ ---
467
+
468
+ ## §Composition-Pattern-Labels — revfactory Team-Pattern Vocabulary (naming only)
469
+
470
+ Optional recognition labels for the `Composition pattern: {name}` line in the Step 2 plan header.
471
+ Borrowed from the revfactory/harness team-pattern vocabulary (sister-asset cross-audit
472
+ `tracks/_audit/session_2026_07_07_revfactory-harness.md`, 2026-07-07). **Each label names an FH
473
+ construct that already exists** — this is a vocabulary import, not a new dispatch mechanism
474
+ (no-reinvention). Use a label only when it fits cleanly; a bespoke composition needs no forced label.
475
+
476
+ | Pattern label | = existing FH construct |
477
+ |---|---|
478
+ | **Pipeline** | sequential Waves, each consuming the prior's fan-in (Wave 0→1→2) |
479
+ | **Fan-out-in** | parallel Wave 1 split → Step 4 fan-in integration |
480
+ | **Expert Pool** | Step 0.2 capability-fit routing to specialist agents |
481
+ | **Producer-Reviewer** | a generating agent + an adversarial reviewer (`challenger` / Critic) |
482
+ | **Supervisor** | one orchestrator/governor delegates then integrates (the default here) |
483
+ | **Hierarchical** | nested supervisors — cluster orchestration (memory `project_fh_cluster_orchestration`) |
484
+
485
+ This is a recognition aid, not a routing input — the actual plan still comes from Steps 0.2–2. If a
486
+ composition matches none of the six, omit the label rather than stretching one to fit.
@@ -100,11 +100,20 @@ Local 4090 = **canary tier** (evidence-of, never terminal verdict).
100
100
 
101
101
  ## Step 6 — Degrade ladder (the intelligent scale-down)
102
102
 
103
+ **Consent branch first — declined ≠ degraded** (`[[capability_escalation_consent]]`): if the UAP has
104
+ `sidecar_consent: declined`, do **not** probe/recruit — route straight to **Tier-3 CC-only sub-agent
105
+ verification** (multiple isolated Claude sub-agents, isolation-decorrelation) as a **first-class chosen
106
+ mode**, with an honest *same-family* note but **no "reduced value / degraded" framing** — the user chose
107
+ this floor. `unset` → ask-once at first load-bearing need (accept → proceed; decline → record + this
108
+ branch). Only proceed to the discovery ladder below when consent is `accepted`.
109
+
103
110
  1. frontier cross-family CLI present → recruit it (decorrelated, at-floor) — best.
104
111
  2. only local 4090 present → canary pre-screen + in-session opus governor (canary, not full decorrelation).
105
112
  3. nothing present → in-session same-family + **honest below-floor/same-family note** (residual named).
106
113
 
107
- Env non-determinism (CLI presence varies) → **silent degrade, never hard-fail**.
114
+ Env non-determinism (CLI presence varies) → **silent degrade, never hard-fail**. Distinguish this
115
+ **unavailable-but-wanted** case (consent given, panel down → degrade-with-note; for a *load-bearing corp*
116
+ surface, fail-closed per `local_pmh_context.md`) from the **declined** case above (chosen floor, first-class).
108
117
 
109
118
  ## Step 7 — Output
110
119
 
@@ -117,7 +117,7 @@ Run the audit bash (§Step-Bash) and apply thresholds:
117
117
  | CLAUDE.md | Exceeds 300 lines | Section-by-section compression / move completed sections to archive |
118
118
  | MEMORY.md | Exceeds 180 lines | Check entry count + move `✅ CLOSED` items to archive section |
119
119
  | memory/*.md single file | Exceeds 30K (300 lines) | Suggest splitting accumulated history into separate files |
120
- | SKILL.md (any) | > 300 lines AND no SKILL_detail.md | Propose `/skill-splitter` — governance-semantic split (not compression); compression removes content, splitting routes it on-demand |
120
+ | SKILL.md (any) | > 300 lines AND no SKILL_detail.md | Propose `/salience-splitter` — governance-semantic split (not compression); compression removes content, splitting routes it on-demand |
121
121
 
122
122
  **Frequency**: When explicitly called with `/context-doctor` or auto-invoked at session start when MEMORY.md is detected at 180+ lines.
123
123
 
@@ -259,7 +259,7 @@ context-doctor (token/context) · harness-doctor (structure) · sim-conductor (s
259
259
  |---|---|
260
260
  | Want to also check structure after resolving token waste | `/harness-doctor` |
261
261
  | Want to validate prescription results from external user perspective | `/sim-conductor Area A` |
262
- | SKILL.md diagnosed as over-loaded (> 300 lines, no SKILL_detail.md) | `/skill-splitter` — governance-semantic split |
262
+ | SKILL.md diagnosed as over-loaded (> 300 lines, no SKILL_detail.md) | `/salience-splitter` — governance-semantic split |
263
263
  | All three skills mentioned simultaneously | Three-Doctor Loop circuit activated — diagnosis→prescription→re-diagnosis cycle
264
264
 
265
265
  ## Done When
@@ -278,4 +278,4 @@ context-doctor (token/context) · harness-doctor (structure) · sim-conductor (s
278
278
  **→ Three-Doctor Loop chain (auto-propose after diagnosis):**
279
279
  - Prescription modifies SKILL.md / rules / CLAUDE.md → **propose `/harness-doctor`** re-check after fix (structural integrity)
280
280
  - Prescription addresses user-facing context (onboarding, README, install guides) → **propose `/sim-conductor Area A`** (external user impact validation)
281
- - SKILL.md detected as over-loaded → **auto-propose `/skill-splitter`**: `"I see [skill-name] SKILL.md is [N] lines with no SKILL_detail.md. Want me to run /skill-splitter to do a governance-semantic split?"`
281
+ - SKILL.md detected as over-loaded → **auto-propose `/salience-splitter`**: `"I see [skill-name] SKILL.md is [N] lines with no SKILL_detail.md. Want me to run /salience-splitter to do a governance-semantic split?"`
@@ -73,11 +73,11 @@ wc -l memory/MEMORY.md 2>/dev/null
73
73
  # memory/*.md files exceeding 30K
74
74
  find memory -name "*.md" -size +30k 2>/dev/null | xargs wc -l | sort -rn | head -10
75
75
 
76
- # SKILL.md files > 300 lines with no SKILL_detail.md (skill-splitter candidates)
76
+ # SKILL.md files > 300 lines with no SKILL_detail.md (salience-splitter candidates)
77
77
  find plugins -name "SKILL.md" 2>/dev/null | while read f; do
78
78
  lines=$(wc -l < "$f")
79
79
  detail=$(dirname "$f")/SKILL_detail.md
80
- [ "$lines" -gt 300 ] && [ ! -f "$detail" ] && echo "[skill-splitter candidate] $f ($lines lines)"
80
+ [ "$lines" -gt 300 ] && [ ! -f "$detail" ] && echo "[salience-splitter candidate] $f ($lines lines)"
81
81
  done
82
82
  ```
83
83