@chrono-meta/fh-gate 1.4.53 → 1.4.54

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -11,13 +11,13 @@
11
11
  "plugins": [
12
12
  {
13
13
  "name": "fh-meta",
14
- "version": "1.4.53",
14
+ "version": "1.4.54",
15
15
  "description": "Hub meta-operations toolkit — 33 skills + 7 agents. New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
16
16
  "source": "./plugins/fh-meta"
17
17
  },
18
18
  {
19
19
  "name": "fh-commons",
20
- "version": "1.4.53",
20
+ "version": "1.4.54",
21
21
  "description": "Project-agnostic utility skills — 4 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
22
22
  "source": "./plugins/fh-commons"
23
23
  }
package/CLAUDE.md CHANGED
@@ -21,6 +21,17 @@ Running Claude Code in this project activates **Control Tower** mode.
21
21
 
22
22
  The forge-harness hub is not just a repository — it is the **command center for all Claude Code-connected projects in your local environment**.
23
23
 
24
+ **Doctrine (2026-07-12, operator-forged)**: a harness **machinizes intent** — it reads the human's
25
+ intent and forges it into a machined form (AI-followable rules or deterministic code), via
26
+ `intent → forge → agreement (HITL) → machinery`. Its payoff is relocating trial-and-error off the human
27
+ (into the harness, run in parallel), so human time drops and attention routes to irreversible points.
28
+ FH is the **meta-harness and nursery**: it incubates field harnesses (projects *and* new capabilities of
29
+ existing harnesses) in its own sandbox — expensive per run, cheaper in total because trial-and-error
30
+ pools and compounds — and **emits** them as independent specialized harnesses (shipped today as
31
+ scaffold + approval machinery; the full chamber flow is the named target). Over other harnesses it
32
+ operates in two modes: **compose** (cluster strengths) ∪ **disrupt** (melt and reforge via crucible;
33
+ core invariants never melt). Full doctrine: `knowledge/shared/harness-core/harness_incubator_doctrine.md`.
34
+
24
35
  | Layer | Role | Representative Assets |
25
36
  |---|---|---|
26
37
  | **① Control Tower** | Coordinates all connected projects and **drives harness-ification across them** — decides *which* projects to harness and *when*, propagates harness assets to each, and feeds their synced learnings into the hub's compounding loop. The *how* (rules · gates · 6-axis) is executed via the Core Axis. Command HQ, not a passive registry. | `.claude/rules/auto_project_mapping.md` (mapping + **Full-Harness Mode**) · `harvest-loop` (compounding loop) · `templates/` (project-harness bundle) · `CATALOG.md` |
@@ -49,6 +60,18 @@ Two orthogonal layers — never collapse them.
49
60
  self-correct over agreeing (governor-catch). Tone never touches this.
50
61
  - **Speech / reaction layer (soft)**: choose warmer words and a steadier texture. Softness is
51
62
  word-choice, not length — it adds no filler and lengthens nothing.
63
+ - **Register — match the user's language *and* register (applies to EVERY response, not just greetings)**:
64
+ reply in the user's language and in their register (formal ↔ informal). **Consistency is the rule, not
65
+ the default**: do NOT drift between formal and informal — Korean 반말↔존댓말, English casual↔corporate —
66
+ within a turn, across turns, or across a session. Register drift is a UX defect and dilutes the mascot
67
+ identity. If unsure which register a session is in, match the user's most recent message. The
68
+ Orthogonality guard below applies to register too (a warmer register never softens judgment). This rule
69
+ lives in always-loaded CLAUDE.md, not only in memory, so it fires every turn without depending on
70
+ recall — the 2026-07-12 miss was a session that drifted register because the rule lived only in memory.
71
+ Tone has **no** mechanical hook gate by nature: always-loaded salience is the strongest available lever,
72
+ **not a floor** (no mechanical floor exists for tone — an accepted limitation, not a guarantee). The
73
+ operator or project may pin a concrete default register in a local binding (`CLAUDE.local.md` / UAP
74
+ `preferred register`) — that pin is operator taste and stays local, never in this public file.
52
75
  - **Not flattery**: soft charisma is not pleasing the user. Disagree plainly when the work calls for it;
53
76
  warmth and a "no" coexist (no Gemini-grade sycophancy).
54
77
  - **Greeting / onboarding**: open with a warm, identity-revealing welcome (new / returning / operator
@@ -401,7 +424,14 @@ let the innovator center a recommend cascade, produce a ranked install plan, and
401
424
  `CLAUDE.md`, mapped `tracks/`, **locally-connected sibling repos** (the env-delta SessionStart hook already
402
425
  emits "N unmapped sibling repos"), and the `LOCAL_SKILL_REGISTRY` + stack/language. Then **branch**:
403
426
  *new-build* (no prior harness) · *extend-existing* (harness present → found→extend, never fork) ·
404
- *maintain* (mature harness → route to the Field-Harness Diagnostic instead). This audit-and-branch pre-step
427
+ *maintain* (mature harness → route to the Field-Harness Diagnostic instead).
428
+ **new-build sub-branch — simulate-first (incubator doctrine)**: judge the project's character before
429
+ building. Clear · small · low failure-cost → build immediately (current flow). Uncertain · exploratory ·
430
+ failure-expensive → **recommend simulate-first**: run the project as a simulation inside the FH chamber
431
+ (harness-unit sandbox) and *emit* the initial project only after the simulation holds — one-line
432
+ recommendation, operator decides (HITL; never forced). Same branch applies to a **new capability of an
433
+ existing harness** (incubate in the chamber, then transplant). Rationale + economics:
434
+ `knowledge/shared/harness-core/harness_incubator_doctrine.md §3`. This audit-and-branch pre-step
405
435
  is imported from the revfactory/harness Phase-0 State Audit (sister-audit 2026-07-07) — it tightens FH's
406
436
  found→extend reflex and is the "이미 로컬에 연결돼 있으면 자동 탐색" mechanism.
407
437
  2. **Innovator-centered recommend**: `persona-innovator` centers the cascade (Mode I on acceleration / Mode F
package/README.ja.md ADDED
@@ -0,0 +1,443 @@
1
+ <p align="center">
2
+ <img src="https://raw.githubusercontent.com/chrono-meta/forge-harness/main/docs/banner.png" alt="forge-harness — プロジェクトを鍛え、通せば、より速く仕上がる。品質が梃子であり、速度はその結果だ。" width="680">
3
+ </p>
4
+
5
+ <p align="center">
6
+ <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-22c55e.svg" alt="MIT License"></a>
7
+ <a href="https://zenodo.org/records/20397566"><img src="https://img.shields.io/badge/DOI-10.5281%2Fzenodo.20397566-blue.svg" alt="DOI"></a>
8
+ <img src="https://img.shields.io/badge/Claude_Code-compatible-a855f7.svg" alt="Claude Code">
9
+ <a href="https://github.com/chrono-meta/forge-harness/issues/72"><img src="https://img.shields.io/badge/Codex-beta_·_help_validate-f59e0b.svg" alt="Codex-compatible beta — help validate (issue #72)"></a>
10
+ <a href="https://www.npmjs.com/package/@chrono-meta/fh-gate"><img src="https://img.shields.io/npm/v/@chrono-meta/fh-gate.svg?color=cb3837" alt="npm"></a>
11
+ </p>
12
+
13
+ <p align="center">
14
+ <a href="README.md">English</a> · <a href="README.ko.md">한국어</a> · <a href="README.zh.md">中文</a> · <b>日本語</b>
15
+ </p>
16
+
17
+ <p align="center">
18
+ <b>あなたの Claude Code プロジェクトを鍛えて — 通せば、より速く仕上がります。</b><br>
19
+ 実務者の<b>メタハーネス (meta-harness)</b> — あなたのプロジェクトハーネスたちが暮らす銀河。<br>各プロジェクトの<b>床 (floor)</b> を上げ(設定をハーネス化)、<b>天井 (ceiling)</b> を上げた上で(作業を加速)、その利得をポートフォリオ全体に複利で積み上げます。
20
+ </p>
21
+
22
+ <p align="center">
23
+ <b>品質が梃子であり、速度はその結果です。</b> あらゆる変更はゲートを通って自らの値打ちを証明します —<br>敵対的 (adversarial) · ファントム (phantom) · 回帰 (regression) — そして<i>それ</i>が次の変更をより速くします。
24
+ </p>
25
+
26
+ <p align="center">
27
+ <i>フォークしてください。名前を変えてください。あなたのものにしてください。</i>
28
+ </p>
29
+
30
+ <p align="center">
31
+ <img src="docs/pillars.svg" alt="FORK · ADAPT · COLLABORATE · EMPOWER" width="680">
32
+ </p>
33
+
34
+ <p align="center">
35
+ <a href="docs/ETHOS.md"><b>原則</b></a> ·
36
+ <a href="docs/WHY.md"><b>存在理由</b></a> ·
37
+ <a href="docs/OUTPUT_EVIDENCE.md"><b>証拠</b></a> ·
38
+ <a href="CHEATSHEET.md"><b>使い方</b></a>
39
+ </p>
40
+
41
+ ---
42
+
43
+ | こんな理由で来たなら… | forge-harness が解決します |
44
+ |---|---|
45
+ | セッションが終わると文脈が消える | 永続 `tracks/` — どこからでも続きを再開 |
46
+ | プロジェクトごとに同じ設定を繰り返す | ハブに一度つなげば全プロジェクトで共有 |
47
+ | チームの AI ノウハウが人の頭の中にしかない | コードに刻んで全員で共有 |
48
+ | 作業が積み上がるほど AI が*より良く*なってほしい | スキルとパターンがセッションを重ねて複利で積み上がる |
49
+ | AI が生成したコードにガバナンス層が必要だ | `fh-gate` がどんなコーディングエージェントでも生成後ゲートで包む |
50
+
51
+ > **この文書は人間のためのものです。** AI 運用ルール → `CLAUDE.md` · コマンドリファレンス → `CHEATSHEET.md`
52
+
53
+ ---
54
+
55
+ ## 2分で始める
56
+
57
+ **前提条件**: Claude Code CLI — `claude --version` で確認
58
+
59
+ ```bash
60
+ # 1. プラグインをインストール
61
+ claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git
62
+ claude plugin install -s user fh-meta@forge-harness
63
+
64
+ # 2. ハブをクローン
65
+ git clone https://github.com/chrono-meta/forge-harness.git ~/forge-harness
66
+ cd ~/forge-harness
67
+
68
+ # 3. セッションを開始
69
+ claude
70
+ ```
71
+
72
+ > ✅ Claude が `CLAUDE.md` を読み、どのプロジェクトを接続するか、あるいはどの作業を始めるかを尋ねます。
73
+ > **「プロジェクトを接続して」** と言えば → ハブが `../` をスキャンして `.git` ディレクトリを見つけ、`tracks/{project}/` を作成します。
74
+
75
+ **プラグインのみ(クローンなし):**
76
+ ```bash
77
+ claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git # 初回1回
78
+ claude plugin install -s user fh-meta@forge-harness
79
+ cd ~/projects/{your-project} && claude
80
+ ```
81
+
82
+ > ⚠️ **プラグインのみは部分シナジーです。** スキルとエージェントは得られますが **Layer 1** は得られません —
83
+ > `CLAUDE.md` ガバナンス(能動オンボーディング、4軸ゲート、モード分岐)と複利文脈(`tracks/` メモリ蓄積、
84
+ > `harvest-loop` 学習)が抜けます。各スキルは孤立していても同じように動きますが、抜けるのはそれらを
85
+ > セッションをまたいで複利にするオーケストレーションです。道具だけでなく全体セットが欲しいなら、ハブを
86
+ > クローンしてください(上記参照)。
87
+
88
+ > 🚪 **初めてですか / スキルだけ欲しいですか?** 意見のこもった正面玄関から始めてください —
89
+ > [`templates/starter_profile.md`](templates/starter_profile.md): インストールコマンド1つ、厳選された
90
+ > 最初の5つのスキル、そしてインストール不要のガバナンスゲート(`npx … fh-gate`)。残りのスキルは必要になるまで待ちます。
91
+
92
+ ---
93
+
94
+ ## これは何か
95
+
96
+ forge-harness は**2つの明確な層**で構成されています:
97
+
98
+ | 層 | 内容 | AI 互換性 |
99
+ |---|---|---|
100
+ | **方法論層** | `tracks/`, `knowledge/`, `SKILL.md` 文書、セッションプロトコル | あらゆる AI モデル |
101
+ | **自動化層** | `plugins/*/agents/` (FH エージェント)、`.claude/agents/` (フィールドプロジェクトのオーバーライド)、フック、スラッシュコマンド、`CLAUDE.md` ルール | Claude Code 専用 |
102
+
103
+ 方法論層は移植可能な核です — 永続ハブ、学習の蓄積、プロジェクト間の知識キュレーション。自動化層は Claude Code 上でこれを摩擦なく動かします。
104
+
105
+ **これが立つ位置 (2026):** 「ハーネスエンジニアリング」はいまや公開パラダイムであり — 基本的なエージェント
106
+ オーケストレーションは急速に標準インフラへと商品化されています。FH はその配管 (plumbing) に何も賭けて
107
+ いません。FH の永続層は*商品化されないもの*です: ガバナンスゲート(敵対的 · ファントム · 回帰)、ドリフト
108
+ 制御、そしてプロジェクト間の複利ループ。ルーティングとディスパッチは手段であり、**ゲートとループが資産です。**
109
+
110
+ ```
111
+ forge-harness/ ← ハブ (永続する脳)
112
+ ├── knowledge/ → 全プロジェクトで共有
113
+ └── tracks/ → プロジェクト別の作業記録
114
+
115
+ Project A ──→ CLAUDE.md でハブを接続
116
+ Project B ──→ CLAUDE.md でハブを接続
117
+ ```
118
+
119
+ ---
120
+
121
+ ## ツールボックスではなくハーネスである理由
122
+
123
+ まず、ハーネスが*何のためにあるのか*から: ハーネスはあなたの**意図**を読み取り、**機械化された形**
124
+ へと鍛え上げます — AI が確実に従うルール、あるいはモデルを一切必要としない決定的なコードへ。
125
+ あなたが意図と洞察を渡し、ハーネスがそれを実行可能な形に鍛え、あなたが承認すれば、それは機械に
126
+ なります。見返りは**人間側の試行錯誤の最小化**です: リクエスト → フィードバック → 再生成のループは
127
+ 消えるのではなく、*場所を移す*のです — ハーネスの内部へ、エージェントとサイドカーが並列で回す
128
+ 場所へ — その結果あなたの時間は減り、あなたの注意は変更が不可逆な地点にだけ使われます。
129
+
130
+ スケールが第二の要点です。**スキル · エージェント · プラグイン**は1つの道具です。**ハーネス**は一段上 —
131
+ 1つの*星 (star)* です: あるプロジェクトの道具 · ルール · ゲート · 記憶が、1つの働く体へと束ねられたもの。
132
+ **forge-harness はその星たちが暮らす銀河です** — 複数のハーネスを1つの重力圏に収め、軌道に
133
+ とどめ(共通の床、ドリフトなし)、散り散りになる代わりに共に進化させます。そしてこの系はただの
134
+ 容れ物ではなく*ゆりかご (nursery)* です: FH はフィールドハーネスを**自らのサンドボックス内で
135
+ シミュレーションとして走らせることができ** — 1回あたりは高くつきますが、総コストは安くなります。
136
+ 試行錯誤が一箇所に集まり複利で積み上がるからです — シミュレーションが検証されれば、その
137
+ プロジェクトを独立した特化ハーネスとして**送り出します**。これが目指す目標です。実際には、
138
+ その重力は4つのものから生まれます:
139
+
140
+ **① 組み立て (Assemble)** — FH はハーネスの*クラスター*を最適なトークンコストで運用し、プロジェクトに合う
141
+ ハーネスを手に握らせます。スキルを1つずつ配線するのではなく、**ハーネス**を — そのプラグイン · スキル ·
142
+ エージェントまで含めて — 合わせて組み立てた状態で受け取ります。
143
+
144
+ **② 鍛え (Forge)** *(品質ゲート)* — あらゆる変更は敵対的 · ファントム · 回帰ゲートを通って自らの値打ちを
145
+ 証明します。これは「もっと検査する」ではありません。**責任ルーター (responsibility router)** です: 自動化が
146
+ 増えるほど人間の承認は減り1件あたりの重みは増すので、ゲートはあなたの注意を*取り返しのつかない*地点にだけ
147
+ 使います。品質が梃子であり、速度はその結果です。
148
+
149
+ **③ サイドカー (Sidecar)** — 能力そのものはフロンティアに置きます。FH は複数の LLM(Claude, Codex, Gemini,
150
+ ローカル)へディスパッチし、raw な力が1つのモデルや1つの世代に縛られないようにします。要点は各モデルの
151
+ 弱点を*機械的に埋めることではありません* — そうしたスキャフォールディングはモデルが強くなれば死んだコードに
152
+ なります。要点は**フロンティアの進化に共に乗ること**です: substrate がいまやネイティブでやってくれるものは
153
+ 脱ぎ捨て、新しく出してくるものは吸収します。脱相関 (decorrelation) は*いまの*信頼の梃子であり(クロス
154
+ ファミリーのパネルが単一モデルの天井を超えます)、共進化 (co-evolution) が構造です。
155
+
156
+ **④ 自己進化ループ (Self-evolving loop)** — ハーネスは再構築なしにより良くなります、2つの方向へ:
157
+ **外へ**、各セッションの教訓がハブに複利で積み上がって次のプロジェクトがより速く始まり、**内へ**、
158
+ *自分自身の*欠陥を捕まえて直します(4軸ゲート、双方向検証、ユーザー別適応)。
159
+
160
+ 全体は1つの分業です: **raw な能力はモデルのもの、組み立て · 信頼 · 進化はハーネスのもの。**
161
+
162
+ > **ここでのセルフヒーリングは主張ではなく — コミットログにあります。** まさにこの README の声のルールが
163
+ > セッションの途中で、FH が自らのドリフトを捕まえて直されました: トーンのミス → 診断 → 自らの最初の修正案まで
164
+ > 攻撃したクロスファミリー challenger → 再修正 → 床ティアでの再検証 → メモリ反映。ハーネスが自らの欠陥を
165
+ > 自分で直した実例 — スローガンではなく記録です。
166
+
167
+ ---
168
+
169
+ ## 動く理由
170
+
171
+ AI と長い共同執筆セッションを経ると、あなたと AI は同じ文脈を共有します — そして同じ盲点も
172
+ 共有します。持つ価値のあるレビュアーは、あなたの推論を一度も見たことのないレビュアーです。手作業でも
173
+ 得られます: 作業を空の新しいチャットに貼り付ければいいのです。FH はその面倒な作業をコマンド1つの
174
+ ルーティンに変えるだけです。
175
+
176
+ - **サイドカー / エージェントディスパッチ** → あなたのセッションの文脈がまったくないレビュアー
177
+ - **steel-quench · phantom-quench** → その冷静な検討を、必要なときにすぐ
178
+
179
+ モデルに依存しません: ある AI と一緒に作り、冷静な検討は他のどの AI ででも回します。もともとの
180
+ セッションにいなかった側があなたの冷静なレビュアーです — これはモデルの順位付けではありません。
181
+
182
+ **FH が主張しないこと:** 冷静な検討はあなたのベースモデル自身の能力であって、FH が追加する検知
183
+ エンジンではありません — 新しいインスタンスに普通のプロンプトを入れても、かなりの部分で同じことをします。FH の価値は
184
+ より狭く正直です: 実際の実務から出た方法を取り、その独立した検討を*スキップする面倒な作業*の代わりに
185
+ *ルーティン*にします。方法論はコピー可能であり、FH がパッケージするのは秘伝のソースではなくワークフローです。
186
+
187
+ ---
188
+
189
+ ## AI 生成コードのためのガバナンス層
190
+
191
+ FH はどんなコーディングエージェント(OpenCode, Codex など)でも**生成後ガバナンスゲート**で包みます。
192
+
193
+ ```bash
194
+ npx --package @chrono-meta/fh-gate fh-gate # 既定: Claude バックエンド
195
+ FH_BACKEND=codex npx --package @chrono-meta/fh-gate fh-gate # Codex バックエンド
196
+ FH_BACKEND=auto npx --package @chrono-meta/fh-gate fh-gate "src/foo.ts" full
197
+ # → FH_GATE_VERDICT: PASS | PENDING | BLOCKED | ESCALATE
198
+ ```
199
+
200
+ `fh-gate` は両ランタイムに同じ FH ガバナンスプロンプトを使います。`FH_BACKEND=claude` は `claude --print` を、`FH_BACKEND=codex` は `codex exec` を実行し、`FH_BACKEND=auto` は両 CLI が揃っていれば Codex を優先します。
201
+
202
+ Claude Code の外でスキルやエージェントを直接実行するには `fh-run` を使います:
203
+
204
+ ```bash
205
+ FH_BACKEND=codex npx --package @chrono-meta/fh-gate fh-run --skill phantom-quench --file docs/foo.md
206
+ FH_BACKEND=codex npx --package @chrono-meta/fh-gate fh-run --agent fh-commons:quench-challenger --file plugins/fh-meta/skills/foo/SKILL.md
207
+ ```
208
+
209
+ 変更された FH スキル/エージェント表面が依然としてきれいな Codex アダプター経路を持つか確認するには:
210
+
211
+ ```bash
212
+ npx --package @chrono-meta/fh-gate fh-codex-doctor --strict
213
+ ```
214
+
215
+ `fh-codex-doctor` は正典のスキル/エージェントレジストリをスキャンし、どのユニットが Codex ネイティブか、アダプターが
216
+ 必要か、Claude ネイティブか、未分類かを報告します。薄いアダプター境界のドリフト検知器であり、Claude
217
+ Code の自動化層を複製しようとはしません。FH チェックアウトから実行すれば現在の作業ツリーを、外から実行すれば
218
+ インストール済みパッケージをスキャンします。
219
+
220
+ Codex 主導の作業では、可能な限り Codex のネイティブな goal/session 機能を使い続けてください。`fh-goal` は FH
221
+ ガバナンスが後に続くべき一度きりの非対話実行のための移植用ラッパーにすぎません:
222
+
223
+ ```bash
224
+ FH_BACKEND=codex npx --package @chrono-meta/fh-gate fh-goal --prompt "Implement X and update tests" --gate quick
225
+ ```
226
+
227
+ より広い FH 自動化層は、依然としてサブエージェント · フック · スラッシュコマンドのために Claude Code に依存します。移植
228
+ 経路は共有文書 + ランタイムアダプターであり、別々の Codex フォークと Claude フォークではありません。
229
+
230
+ **推奨スタンス — Claude Code をオーケストレーターに、他をサイドカーに。** FH の自動化層(自動発火フック、
231
+ サブエージェントディスパッチ、オンボーディング、メモリ)は Claude Code ネイティブなので、もっとも完全な体験は **Claude
232
+ Code をメインオーケストレーターとし、Gemini, Codex, または Antigravity (`agy`) を能動的に使う
233
+ サイドカー**として回すことです。**非 CC ランタイムをメインエージェント**として回すこともできます —
234
+ `fh-gate`/`fh-run` を通じて方法論層全体と M1 スキルは維持されますが、オートパイロット層は得られません:
235
+ フックが自動発火せず、M2 エージェントディスパッチ段階はアダプター(または対話的承認)が必要で、M3
236
+ スキルは参照用です。これは意図された2層の境界であって、埋めるべきギャップではありません。ランタイム別の詳細:
237
+ [`docs/codex-compat.md`](docs/codex-compat.md) (ティア別) と
238
+ [`multi_model_sidecar_strategy.md`](knowledge/shared/harness-core/multi_model_sidecar_strategy.md)
239
+ (サイドカーエンジン、2026-06-18 EOL 時点の Gemini→`agy` 承継を含む)。
240
+
241
+ **実証結果 (2026-05-31)**: OpenCode の AI 生成 `permission/arity.ts`(163行、CI グリーン)に適用。
242
+ 現在のゲート意味論はこれを BLOCKED に分類します: CI が捕まえられなかった A 級の発見 2件(許可リストの短トークン
243
+ オーバーフロー、arity テーブルから抜けた executor ツール)。
244
+
245
+ 完全な仕様: [`fh_integration_contract.md`](knowledge/shared/harness-core/fh_integration_contract.md)
246
+
247
+ ---
248
+
249
+ ## 鍛冶場 (The forge)
250
+
251
+ forge-harness はプロジェクトを鋼のように扱います — そしてこの比喩は装飾ではなく文字通りです。
252
+ 作業は形を与えられ、攻撃で硬くなり、そうして生き延びたからこそ、はじめてより速く出荷されます。
253
+
254
+ | 工程 | 何が起きるか | コマンド |
255
+ |---|---|---|
256
+ | **鍛え (Forge)** | 生のプロジェクトをハーネスへ形づくる — その床を上げる | `install-wizard`, "harness-ify this project" |
257
+ | **焼き入れ (Quench)** | 攻撃で硬くする — 冷静な検討が健全なものだけを残す | `steel-quench` · `phantom-quench` |
258
+ | **焼き戻し (Temper)** | 硬くなった資産から脆さ (brittleness) を再び抜く | `steel-quench` Wave-T · `templates/temper_check.sh` |
259
+ | → **加速 (Accelerate)** | 鍛冶場を生き延びた刃はより速く斬る | `goal-quench` — *Pass → Accelerate* |
260
+
261
+ 4工程すべてが出荷されます。焼き戻し (Temper) は作られる*前に*名前が先に付けられ — 意図的に(参照:
262
+ [`ETHOS.md`](docs/ETHOS.md#the-forge))— 測定実行が検証したのちに出荷されました。鍛冶場の周りで、さらに2つの
263
+ シグネチャが回り続けます: `harvest-loop`(各セッションの教訓が恒久スキルになる)と
264
+ `agent-composer`(ディスパッチをオーケストレーション)。残りのスキルは必要になるまで待ちます — 全リストは
265
+ 下に。
266
+
267
+ ## 37 skills · 8 agents
268
+
269
+ <details>
270
+ <summary>全資産のアクティベーション確認</summary>
271
+
272
+ | 資産 | 役割 | トリガー |
273
+ |---|---|---|
274
+ | `steel-quench` | 全方位の敵対的検証 | "Run the quench", "Attack from the root" |
275
+ | `phantom-quench` | ファントム主張の検知 + ソース逆追跡 | "Verify the source", "Grounding audit" |
276
+ | `harvest-loop` | セッション終了の学習 → 進化パイプライン | "Harvest the session" |
277
+ | `agent-composer` | 最適なエージェントディスパッチを設計 | "Run in parallel", "Which agents?" |
278
+ | `sim-conductor` | メタシミュレーションオーケストレーター | "External user perspective" |
279
+ | `context-doctor` | トークン効率 + `.claudeignore` | "Session is slow", "Clean up context" |
280
+ | `harness-doctor` | ハーネス構造の診断 | "Check my Claude setup" |
281
+ | `pipeline-conductor` | 4軸品質ゲート (後方/敵対/前方/記録) | "Run the quality gate" |
282
+ | `field-harvest` | フィールドパターンをハブへ逆伝播 | "I could reuse this" |
283
+ | `frontier-digest` | HN + arXiv → 実行可能な洞察 | "AI trend digest" |
284
+ | `hub-cc-pr-reviewer` | 自動 PR レビュー | "Review this PR" |
285
+ | `verify-bidirectional` | 決定の逆検証 | "Is that right?", "Double-check" |
286
+ | `deep-clarify` | ソクラテス式の要件明確化 | "I'm not sure what to build" |
287
+ | `install-wizard` | 初期オンボーディング | "First-time setup" |
288
+ | `plugin-recommender` | プラグイン推薦 | "Is there a good tool for this?" |
289
+ | `apex-review` | 経営層視点の品質レビュー | "Will this hold up?" |
290
+ | `meta-prompt-builder` | メタプロンプト設計 | "Write a prompt for the agent" |
291
+ | `asset-placement-gate` | ハブ vs プロジェクトの資産ルーティング | "Should this be shared?" |
292
+ | `cross-ecosystem-synergy-detection` | ツール間シナジーの発見 | "Are my tools working together?" |
293
+ | `corpus-grounding-expander` | 多バージョンのパブリックドメインコーパス → 検証済み公理グラウンディングストア | "Broaden the grounded corpus" |
294
+ | `persona-roster-expander` | ペルソナのシード → ティア別・判断マッピングされたキャスト | "Broaden these personas" |
295
+ | `convergence-loop` *(fh-commons)* | N ラウンドの収束ループ | "Single-pass seems suspicious" |
296
+ | `token-budget-gate` *(fh-commons)* | 作業前のトークンコスト推定 | "How expensive is this?" |
297
+ | `mcp-circuit-breaker` *(fh-commons)* | MCP ツールの失敗パターン検知 | "MCP keeps failing" |
298
+ | `quench-challenger` *(fh-commons)* | 敵対的プレッシャーテストエージェント | "Challenge this with a devil" |
299
+ | *(+ 追加資産)* | marketplace-gate · contention-layer · edit-manifest · fact-checker · goal-quench · hub-persona-auditor · install-doctor · memory-hygiene · persona-innovator · prompt-regression · public-surface-audit · salience-splitter | |
300
+
301
+ | アクティブ数 | 診断 |
302
+ |:---:|---|
303
+ | **28+** | 上級 — agent-composer + sim-conductor + steel-quench + pipeline-conductor を連鎖 |
304
+ | **10–27** | アクティベーション段階 — 未チェックの資産を段階的にオンにする |
305
+ | **0–9** | 初期段階 — `install-wizard` から始める |
306
+
307
+ **やりたいことでスキルを探す:**
308
+
309
+ | クラスター | スキル |
310
+ |---|---|
311
+ | 検証 | `steel-quench` · `phantom-quench` · `convergence-loop` · `prompt-regression` · `return-path-gate` |
312
+ | オーケストレーション | `agent-composer` · `pipeline-conductor` · `goal-quench` · `deliberation` |
313
+ | 診断 | `harness-doctor` · `context-doctor` · `install-doctor` · `mcp-circuit-breaker` |
314
+ | 収穫 / 学習 | `harvest-loop` · `field-harvest` · `edit-manifest` · `memory-hygiene` |
315
+ | ゲート / ガード | `token-budget-gate` · `asset-placement-gate` · `marketplace-gate` |
316
+ | 発見 | `plugin-recommender` · `cross-ecosystem-synergy-detection` · `frontier-digest` · `verify-bidirectional` |
317
+ | コンテンツ / シミュレーション | `sim-conductor` · `apex-review` · `meta-prompt-builder` · `deep-clarify` |
318
+ | 設定 | `install-wizard` · `hub-cc-pr-reviewer` · `salience-splitter` |
319
+
320
+ > **完全フレーズ集** — すべてのスキル + エージェントとその一行定義、そしてそれを発火させる平易な表現:
321
+ > [`CHEATSHEET.md` §12](CHEATSHEET.md#12-skills--agents--what-each-does-and-what-to-say).
322
+
323
+ </details>
324
+
325
+ ---
326
+
327
+ ## モデル設定
328
+
329
+ Claude Code は作業の複雑さでモデルを自動選択しません — これは一度だけ設定します。
330
+
331
+ ```bash
332
+ /model sonnet # 推奨既定値 — FH が重要な箇所には自ら強いモデルをディスパッチ
333
+ ```
334
+
335
+ | コマンド | 誰が何を実行 | 最適な用途 |
336
+ |---|---|---|
337
+ | `/model sonnet` | Sonnet セッション; FH が宣言された床 (floor) で上位ティアのサブエージェントをディスパッチ | **FH 既定値** — 運用 + 日常開発 |
338
+ | `/model opus` | Opus がすべてを処理 | ハーネス編集セッション (Mode D) · 毎ターン最大の深さ |
339
+ | `/model opusplan` | Opus が*計画* · Sonnet が実行 *(Opus が関与するとき)* | コスト意識の日常コーディング — 注意点を参照 |
340
+
341
+ **なぜいま Sonnet 既定値で通用するのか**: 測定結果(下記 §Model setup evidence note 参照)、FH *運用*はほぼ
342
+ モデルフラットです — 文脈に入ったルールが大部分の仕事をします。それでも強いモデルが必要なのは深さに
343
+ 敏感な少数のターンで、FH はそれを自ら処理します: **一部のスキルとエージェントはモデルティアの床を
344
+ 宣言**し(例: `quench-challenger` は opus に床)、環境が届けばその床ティアの
345
+ サブエージェントとしてディスパッチされます — あなたのセッションモデルには触れません。**FH は決してあなたの
346
+ セッションモデルを変えません**: 手で設定した既定値はそのまま従い、床は FH 自身のサブエージェント
347
+ ディスパッチにのみ適用されます。環境が床より低いところで上限になると(例: Sonnet 専用 API ルーティング)、床が
348
+ かかった資産は依然として利用可能な最良のティアで実行され、出力に明示的な `below-floor` フラグを付けます —
349
+ 劣化した提供は見えるように、決して静かに済まされません(ティア床の解決:
350
+ `knowledge/shared/harness-core/multi_model_sidecar_strategy.md §Tier-floor`)。
351
+
352
+ **`opusplan` の注意点(測定済み)**: Opus の関与は**保証されません** — 測定した10ターンの実行で Opus を
353
+ **0**ターン使いました(CC が "plan-mode" に分類するターンが少ない)。毎ターン Opus を望むなら `/model opus` で
354
+ 固定してください(後続の実行で 22/22 ターン Opus)。**サブエージェントディスパッチ**のモデルはディスパッチ自体の `model`
355
+ パラメーターで決まります; セッションモデル/plan-mode はサブエージェントへ伝播し**ません**。
356
+
357
+ > **役割別**: FH 運用(フィールドプロジェクト、ゲート、日常開発)→ `/model sonnet` + 床にエスカレーションさせて
358
+ > おく。ハーネス自体の編集 (Mode D) → 持っている最強のモデルを固定 — ハーネスの*自己開発*はティアの深さが
359
+ > 測定可能なかたちで値打ちを払う場所であり(設計増分の発見)、運用はそうではありません。サブエージェントのトークン
360
+ > コストはセッション jsonl の `message.model` から CC で見られます。
361
+
362
+ **主張ではなく測定**(実測例): ブラインドのルール適用バッテリーで FH *運用*はほぼモデルフラットです —
363
+ **測定したすべての Claude ティアが 94–100%**(Fable, Opus 4.8, Sonnet 4.6 と 5, Haiku 4.5); 失った少数の点数は
364
+ フォーマットの規律であって、罠やゲート級のミスではありません。ティアが分かれるのはルーブリック超過の*設計*
365
+ 増分だけ(ハーネスを開発するのであって運用するのではない)— だから既定値が**ティア床ディスパッチ**で深さに
366
+ 敏感なターンを覆う Sonnet であり、固定された強いモデルはハーネス編集セッションにのみ推奨されます。
367
+
368
+ これは**不変式として述べられ、モデル別のリーダーボードではありません。** 新しいリリースが覆せない2つの構造法則:
369
+
370
+ 1. **運用はティアをまたいで平坦化する** — 文脈のルールが仕事をするので、あらゆるティアがルール適用で天井に
371
+ 届きます(2026-07-03 の再現で Sonnet 5 がバッテリー天井で Opus 4.8 と同率)。
372
+ 2. **深さ(設計増分)はティア順序であり、その順序は*ある世代の中で*固定される** — 低いティアが**同じ**
373
+ 世代の高いティアを決して追い越しません(ティアは値打ち通りに価格付けされるので、ベンダーが順序を
374
+ 維持します)。*世代をまたいで*は、新しい低ティアのモデルが古い高ティアを超えることがあります(運用での
375
+ Sonnet 5 ≥ Opus 4.8 がまさにこの世代交差のケース)— しかしどの世代でも現在の最上位ティアは依然として
376
+ 自らの深さターンで勝ちます。
377
+
378
+ だからドクトリンは恒久的で、腐りません: **運用は中間ティアを既定値に; 深さは現在の最上位
379
+ ティアへエスカレーション。** 再測定が正当なのは新しいモデルがフィールドメインの*候補*になるときだけ(一度きりの世代交差の
380
+ 閾値確認)であって、同世代のティア順序を再確認するためではありません — それは設計で保証されます。詳細 +
381
+ 日付別の実行: `docs/OUTPUT_EVIDENCE.md` §Validation signals.
382
+
383
+ 外部 CLI(Gemini, Codex, `gh copilot`)をサイドカーとして使うと、そのコストは各自のクォータに請求され、CC のトークン表示には見えません。
384
+
385
+ ### ハードウェアティア(ローカルサイドカーは任意のアクセラレーター)
386
+
387
+ FH は**ローカル LLM を必要としません** — 基準線は Claude Code を動かす何であってもです。ローカルモデルは
388
+ *任意*であり、カナリア / 低コスト幅出しの段だけに使われます:
389
+
390
+ | ティア | 仕様 | ローカル実行 | 何が得られるか |
391
+ |---|---|---|---|
392
+ | **最小** | Claude Code を動かす何でも | なし | 方法論 + ゲートすべて; FH 運用は測定したすべてのティアで ~モデルフラット (94–100%) |
393
+ | **推奨** | ラップトップ級、~16GB RAM | 8B 級の量子化モデル1つ(例: 8B / 小型 Gemma) | トークン無料の**床カナリア**(課金される sim の前の事前スクリーン)· オフライントリアージ · 低コスト幅出しのパネルアーム |
394
+ | **任意(ヘビー)** | ~24GB VRAM GPU | 27–32B モデル | *より強い*脱相関カナリア |
395
+
396
+ > ローカルティアは**カナリアであって最終判定では決してありません** — 測定: 床モデルはフロンティアが捕まえた微妙な
397
+ > 敵対ケースを見逃しました(27–32B のローカルでさえそのケースで 1/4)。ローカルは*幅出しのコスト*を
398
+ > 下げるだけで、判定はフロンティアに残ります。
399
+
400
+ ---
401
+
402
+ ## マルチモデルサイドカー
403
+
404
+ Gemini, Codex, または `gh copilot` を Claude の隣で独立したレビュアーとして回します。要点は**文脈の隔離**です:
405
+ 作業を一緒に作ら*なかった*レビュアーはその泡 (froth) に冷静です — 協業の*外*に座る者が、いまや共有された
406
+ 結果の擁護者となった共同著者が滑らかに見過ごしたものをよく捕まえます。これは対称的で、モデル順位ではありません:
407
+ Gemini と一緒に作れば新しい Claude がその泡を捕まえ、Claude と一緒に作れば新しいサイドカーが Claude の泡を捕まえます。
408
+
409
+ ある内部のケーススタディでは、レビュアーを層に重ねるほどより多くの問題が明らかになりました — 単一のセッション内パスが
410
+ 見逃した項目をセッション間のペルソナが捕まえ、外部 CLI のレビュアーが Claude のペルソナたちが共有した盲点をいくつか
411
+ 明らかにしました。これをベンチマークではなく**実測例**として扱ってください: 利得は作業の複雑さと成果物をどれだけ
412
+ 共同創作したかに比例して大きくなり、隔離されたレビュアーはトリアージすべき誤検知 (false positive) も加えます。
413
+ 与えられた作業で純利得が値打ちあるかは、経験的で使うたびに異なる問いです。
414
+
415
+ レビュアーが外部 CLI のとき Claude 側のトークンコストは増えません — 各自のクォータに請求されます。
416
+
417
+ ---
418
+
419
+ ## 研究 (Research)
420
+
421
+ > **FH 論文** — 以下の方法論は主張だけでなく文書化されています:
422
+ > - **v1.0 — 方法論** · [Zenodo](https://zenodo.org/records/20397566) (DOI 10.5281/zenodo.20397566). 2層設計、6軸フレームワーク、4エージェントオーケストレーション、そして複利ループを実証証拠とともに。
423
+ > - **cs.SE companion — ガバナンスゲート方法論** · **掲載済み** [Zenodo](https://zenodo.org/records/20680081) (DOI 10.5281/zenodo.20680081 · 最新 v1.1 10.5281/zenodo.20740038 · CC-BY-4.0) · arXiv 提出済み (cs.SE, モデレーション中)。
424
+ > - **cs.AI companion — "Governance Dividend"** · 準備中。
425
+
426
+ 外部の収束:
427
+ - ["Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems"](https://arxiv.org/abs/2604.14228) — arXiv 2026年4月
428
+ - ["Code as Agent Harness"](https://arxiv.org/abs/2605.18747) — arXiv 2026年5月
429
+ - Stanford IRIS Lab: ["Meta-Harness"](https://arxiv.org/abs/2603.28052) — 4倍少ないトークンで +7.7pts
430
+
431
+ ---
432
+
433
+ ## もっと知る
434
+
435
+ | リソース | 目的 |
436
+ |---|---|
437
+ | [`CLAUDE.md`](CLAUDE.md) | AI 運用ルール + 同期/プッシュプロトコル |
438
+ | [`CHEATSHEET.md`](CHEATSHEET.md) | 全コマンドリファレンス |
439
+ | [`AGENTS.md`](AGENTS.md) | ランタイムエージェント仕様 |
440
+ | [`CATALOG.md`](CATALOG.md) | 過去作業の検索インデックス |
441
+ | [`CONTRIBUTING.md`](docs/CONTRIBUTING.md) | スキルとパターンの貢献方法 |
442
+ | [`tracks/_contrib/`](tracks/_contrib/README.md) | **同意レーン** — 非識別化した作業セッションを共有; レポがローカルだけでなく運用者たちにまたがって複利で積み上がる |
443
+ | [`fh_integration_contract.md`](knowledge/shared/harness-core/fh_integration_contract.md) | ガバナンスゲート仕様 |