@chrono-meta/fh-gate 1.4.57 → 1.4.59

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -11,13 +11,13 @@
11
11
  "plugins": [
12
12
  {
13
13
  "name": "fh-meta",
14
- "version": "1.4.57",
14
+ "version": "1.4.59",
15
15
  "description": "Hub meta-operations toolkit — 33 skills + 7 agents. New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
16
16
  "source": "./plugins/fh-meta"
17
17
  },
18
18
  {
19
19
  "name": "fh-commons",
20
- "version": "1.4.57",
20
+ "version": "1.4.59",
21
21
  "description": "Project-agnostic utility skills — 4 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
22
22
  "source": "./plugins/fh-commons"
23
23
  }
package/CLAUDE.md CHANGED
@@ -82,6 +82,25 @@ Two orthogonal layers — never collapse them.
82
82
  judgment rigor — if a response reads as more agreeable, less verified, or hedged, the tone layer has
83
83
  leaked into the judgment layer and the response is wrong, not warm.
84
84
 
85
+ ## Envelope-Boundary Discipline — the reinvention-reflex counterweight
86
+
87
+ When the operator's input introduces something that does **not** fit an existing asset or category — a
88
+ novel insight, a specific case that resists the known boxes, a live path with no slot — the default
89
+ reflex is to **normalize** it ("we have that / that's like X") and pull it back inside the envelope.
90
+ That pull mis-scores the new as familiar and can extinguish what would become net-new. The entire asset
91
+ base leans toward normalization (no-reinvention gate · `asset-placement-gate` · "build only what adds
92
+ governance"), so this is its deliberate **counterweight** — and the reflex **strengthens with maturity**
93
+ (more boxes to pattern-match against), so the counterweight must be explicit, never assumed.
94
+
95
+ **Discipline**: at the boundary, do **not** normalize. Hold the unfamiliar unfamiliar; test what it
96
+ actually *is* — net-new? tool-shaped (→ possible EMIT) or judgment-shaped (→ doctrine)? — **before**
97
+ mapping it to a known asset. This is the meta-harness's growth point: it evolves by *not-collapsing the
98
+ unfamiliar*, not by adding machinery. The reflex fires **before** memory recall, so this lives
99
+ always-loaded, not only in memory. (Measured 2026-07-14, one session, 3×: two identities each collapsed
100
+ onto their single hardest sub-mechanism, and a failure from a **non-harness** run mapped onto a harness
101
+ metric — each read a live-but-incomplete thing as zero, each caught by the operator, not self-caught.
102
+ Detail: `[[feedback_reinvention_reflex_normalization_counterweight]]`.)
103
+
85
104
  ## New Project Onboarding
86
105
 
87
106
  > Detailed procedure: `knowledge/shared/rules/auto_project_mapping.md` (5-step mapping + §6 Full-Harness Mode)
@@ -204,7 +223,7 @@ No user request is needed — this is a mandatory autonomous step, not a proposa
204
223
  **Commit gate**: `git commit` on FH asset changes is hard-blocked by `templates/.git-hooks/pre-commit` until all required axes PASS. Hook installation (one-time): `git config core.hooksPath templates/.git-hooks && chmod +x templates/.git-hooks/pre-commit templates/.git-hooks/pre-push` (the same `core.hooksPath` also activates the **pre-push** Destructive-Op gate — see that section below).
205
224
 
206
225
  ```
207
- FH asset modified → Axis 1 (regression_guard.sh --pr {BRANCH})
226
+ FH asset modified → Axis 1 (templates/regression_guard.sh --pr {BRANCH})
208
227
  → Axis 2 (/steel-quench) → Axis 3 (/phantom-quench)
209
228
  → marker: tracks/_meta/.axes_23_passed_{branch}_{date}.marker
210
229
  (required fields: axis2-engine / axis2-model / floor-status / axis2-evidence;
@@ -281,7 +300,7 @@ no weak-local-judge regression of the judge-robustness principle (mechanical anc
281
300
 
282
301
  | Axis | Skill | What it catches |
283
302
  |---|---|---|
284
- | Backward | `regression_guard.sh` | Critical section loss, broken refs, syntax errors, line reduction |
303
+ | Backward | `templates/regression_guard.sh` | Critical section loss, broken refs, syntax errors, line reduction |
285
304
  | Adversarial | `steel-quench` | Trigger phrase collisions, design attack surface, over-engineered steps |
286
305
  | Forward | `phantom-quench` | Phantom references, paths that don't exist, stale external links |
287
306
  | Record | `edit-manifest` RECORD | Logs predicted impact — closes the predict-verify loop for future harvest-loop |
@@ -427,10 +446,18 @@ let the innovator center a recommend cascade, produce a ranked install plan, and
427
446
  *maintain* (mature harness → route to the Field-Harness Diagnostic instead).
428
447
  **new-build sub-branch — simulate-first (incubator doctrine)**: judge the project's character before
429
448
  building. Clear · small · low failure-cost → build immediately (current flow). Uncertain · exploratory ·
430
- failure-expensive → **recommend simulate-first**: run the project as a simulation inside the FH chamber
431
- (harness-unit sandbox) and *emit* the initial project only after the simulation holds — one-line
432
- recommendation, operator decides (HITL; never forced). Same branch applies to a **new capability of an
433
- existing harness** (incubate in the chamber, then transplant). Rationale + economics:
449
+ failure-expensive → **flag simulate-first as an option**: doctrine says such a project *should* be
450
+ chamber-simulated before emit. The chamber **run orchestration is now wired** (`scripts/chamber_run.sh`
451
+ an intent-driven, resumable 7-step runner: budget-entry cap, ≥3-blind-persona gate, Emission Gate,
452
+ G4 ledger auto-append; run #3 exercised it 2026-07-14). But a **live one-command autonomous
453
+ simulate→EMIT of a field harness is NOT yet a capability**: step-4 persona dispatch is human/Claude-driven
454
+ (bash cannot spawn the isolated Agents — the honest muscle boundary), the EMIT terminus is HITL, and
455
+ **EMIT has never fired — the ledger's real runs are 2/2 KILL** (the chamber to date *screens*, it has not
456
+ *birthed*). So today this branch = a one-line HITL recommendation to run the chamber (`chamber_run.sh`),
457
+ then fall back to Full-Harness Mode §6 for the actual onboarding; the runner gates and records a
458
+ human-driven run — it must **not** be presented as a push-button autonomous emit. The same branch applies
459
+ to a **new capability of an existing harness** — the incubate-in-chamber-then-transplant flow is likewise
460
+ run-orchestrated but not autonomously emitting today. Rationale + economics:
434
461
  `knowledge/shared/harness-core/harness_incubator_doctrine.md §3`. This audit-and-branch pre-step
435
462
  is imported from the revfactory/harness Phase-0 State Audit (sister-audit 2026-07-07) — it tightens FH's
436
463
  found→extend reflex and is the "이미 로컬에 연결돼 있으면 자동 탐색" mechanism.
@@ -451,6 +478,10 @@ let the innovator center a recommend cascade, produce a ranked install plan, and
451
478
  **Guards**: (a) **non-overwriting is inviolable** — the one thing both revfactory surfaces get wrong; FH
452
479
  proposes merge, never clobbers; (b) **no-reinvention** — Tier 0/1 first, scaffold only what adds governance;
453
480
  (c) **company residency** — discovery of a company sibling repo surfaces it, does not auto-map/leak it;
481
+ promoted to a machine field (`residency` on the skill registry, `fh_detail_protocols.md §1-c`) so any
482
+ derived recommendation naming a `company`/`operator-private` entry lands only in gitignored `tracks/_meta/`
483
+ or the private companion store, never a tracked public file (chamber run #7, 2026-07-14 — the guard was
484
+ prose-only and the field didn't exist);
454
485
  (d) **autonomy floor** — the discover/rank judgment is trusted at opus-tier+; below-floor, present the raw
455
486
  recommend and ask; (e) **once per door-entry**, not a per-turn nag. This is the door ③ (accelerate) engine
456
487
  and the new-project/harness-write path made autonomous — the operator asks once and the harness discovers,
@@ -744,6 +775,8 @@ Based on LOCAL_SKILL_REGISTRY (Step 1-c), **propose and connect skills from othe
744
775
 
745
776
  **Method**: create `tracks/_meta/fh_signal_{YYYY-MM-DD}_{source}.md` (1 file/session, append if same date+source). Structural candidates only — exclude typos and in-session-resolved issues.
746
777
 
778
+ **Chamber-candidate hook (feeds the discovery pipeline)**: when a signal is an *incubatable capability or project* (uncertain / exploratory / failure-expensive / high-reinvention-risk — a chamber-run candidate, not just a fix), add a `CHAMBER-CANDIDATE: <one-line description>` line to the signal file. `scripts/chamber_candidate_collect.sh` greps that convention across the 6 sources (harness-doctor · harvest-loop · fh-signal · field-harvest · frontier-digest · uap), dedups/ranks, screens for reinvention, and skips anything the G4 ledger already KILLed. Adoption is incremental — the queue is honestly sparse until sources emit the marker; the collector measures the real volume.
779
+
747
780
  ## Execution Tier Settings
748
781
 
749
782
  > **Full tier table + config**: `knowledge/shared/harness-core/fh_detail_protocols.md` — read when selecting a non-default tier.
@@ -795,10 +828,17 @@ Closing phrase detected ("wrap up", "done", "good work", "end session", etc.)
795
828
  → ③ Sync local/gitignored session state to your durable companion store, if you keep one
796
829
  → ④ Memory hygiene — update stale entries + record new session findings
797
830
  → ④-b npm freshness — if any npm-shipped asset changed (`package.json` `files[]`: skills · agents ·
798
- README · AGENTS.md · CLAUDE.md · CHEATSHEET), **propose republish**: version bump **in lockstep**
831
+ knowledge/ · docs/ · README · AGENTS.md · CLAUDE.md · CHEATSHEET · CATALOG.md): **first a Codex-entrypoint
832
+ drift check** — the script (`session_close_check.sh`) auto-*fires a candidate reminder* by cheap grep
833
+ (file co-occurrence, not topical parity), then **you judge**: does the changed CLAUDE.md/knowledge topic
834
+ actually mirror a section in `AGENTS.md` / `docs/codex-compat.md` (both files[]-shipped Codex entry
835
+ points)? sync it, else record `drift:none`. The grep flags; it does not determine — the parity call is
836
+ judged. Version lockstep invalidates the plugin.json *cache* but is **orthogonal** to
837
+ entry-point *content* — a version-only bump can ship a stale Codex entry point (gate-locality,
838
+ Codex side). Then **propose republish**: version bump **in lockstep**
799
839
  across `package.json` + every `.claude-plugin/plugin.json` + `marketplace.json` (single-source =
800
840
  `package.json`) → Pre-Publish gate → `npm publish` → `git tag vX.Y.Z` at publish. **Propose, don't
801
- auto-publish.** (Why lockstep — Codex caches on plugin.json version — + tag-drift caveat → §detail below.)
841
+ auto-publish.** (Why lockstep — Codex caches on plugin.json version — + drift-check + tag-drift caveat → §detail below.)
802
842
  → ④-c Handoff lifecycle (cross-machine continuity) — when a durable **result artifact lands** this
803
843
  session (mechanical hint: a new `*result*`/`*signal*`/`*_run_*` file in your companion store or
804
844
  `tracks/`), do two things: **(a) ④-c stamps** any `"run this/start here"` run-handoff whose
package/README.md CHANGED
@@ -224,6 +224,19 @@ The broader FH automation layer still depends on Claude Code for sub-agents, hoo
224
224
 
225
225
  **Empirical result (2026-05-31)**: Applied to OpenCode's AI-generated `permission/arity.ts` (163 lines, CI green). Current gate semantics classify this as BLOCKED: 2 A-grade findings CI didn't catch (short-token overflow in allowlist, executor tools absent from arity table).
226
226
 
227
+ **Does the method actually add anything? A measured check (2026-07-14).** We held the model fixed at a
228
+ mid-tier floor and varied only the review *method*, on unseen gate snippets with planted *default-toward-PASS*
229
+ (fail-open) holes. On eight subtle holes — authored by two other models so the test set wasn't tuned to our
230
+ method — a plain review caught 5/8 (and two of those "catches" were the wrong bug, i.e. false confidence);
231
+ the same model with FH's degrade-direction lens caught 6/8 with zero false alarms. The honest part: **both
232
+ single-model lanes missed the same two holes** (a falsy error-sentinel, and a separator-negation parse). A
233
+ different model family, same lens, caught both — so the FH *stack* (lens + cross-family + a mechanical
234
+ pre-screen) reaches 8/8. The takeaway isn't a headline score; it's that the value comes from the
235
+ **decorrelated stack**, because even a well-prompted single model has a correlated blind spot that only a
236
+ different family closes. The two missed classes are now caught mechanically (a lint pre-screen), one layer
237
+ earlier. Small sample (single draw); reps and harder holes are the stated next step. Method + full result:
238
+ [`ship_readiness_gate.md`](knowledge/shared/harness-core/ship_readiness_gate.md).
239
+
227
240
  Full spec: [`fh_integration_contract.md`](knowledge/shared/harness-core/fh_integration_contract.md)
228
241
 
229
242
  ---
@@ -159,6 +159,26 @@ package stale. **Tag drift caveat**: when a bump rides inside a functional commi
159
159
  commit), tag *that* commit — otherwise the version ships to npm untagged (e.g. 1.4.4/1.4.5 shipped
160
160
  untagged, backfilled 2026-06-08).
161
161
 
162
+ **Entry-point content drift (④-b drift-check — orthogonal to the version cache above).** The plugin.json
163
+ version keys the *cache path*, so bumping it forces Codex to refetch — that closes the **stale-cache**
164
+ axis. It does **not** close a second, orthogonal axis: `AGENTS.md` / `docs/codex-compat.md` are the
165
+ Codex-user entry points, and their *prose* must mirror whatever CLAUDE.md/knowledge change triggered the
166
+ republish. A version-only bump invalidates the cache yet still serves an AGENTS.md that never absorbed the
167
+ change — **version fresh, entry point stale** (the Codex-side face of `[[feedback_gate_locality_principle]]`:
168
+ a gate/pointer is only as fresh as the surface the actor actually reads). So ④-b greps whether the changed
169
+ topic touches a mirrored AGENTS.md/codex-compat section → sync it, else record `drift:none`. Mechanical
170
+ grep, ~0 cost. **Mechanically *emitted*, judged-*determined* (honest scope)**:
171
+ `scripts/session_close_check.sh` ④-b-drift auto-fires a drift-*candidate* warning when a shipped
172
+ CLAUDE.md/knowledge path changed but the Codex entry points (`AGENTS.md`/`docs/codex-compat`) did not.
173
+ What is mechanized is the *reminder* — it no longer depends on the runner remembering to look (that half
174
+ of the old "prose-only" gap is closed). What is **not** mechanized is the *parity determination*: the
175
+ script tests file **co-occurrence**, not topical parity, so it (a) can false-positive when the changed
176
+ path doesn't actually mirror an entry-point section, and (b) can false-negative if AGENTS.md was touched
177
+ for an unrelated reason in the same tag range. So the runner still judges each candidate (sync it, else
178
+ record `drift:none`) — the script flags, it does not *catch*. Origin: 2026-07-13 the close chain
179
+ lockstep-bumped v1.4.56/57 but only an operator question ("코덱스 호환성도 자동?") confirmed AGENTS.md was
180
+ clean — the chain never auto-checked it (`fh_signal_2026-07-13_self-dev` S3).
181
+
162
182
  ---
163
183
 
164
184
  ## §Session-Close-Handoff-Lifecycle
@@ -58,10 +58,26 @@ Then **fail-closed** (irreversible-ish: a silent empty overwrite blinds the bus)
58
58
  **and** the existing registry has >0 entries, do **not** overwrite — flag `⚠️ scan returned 0 (root=$ROOT);
59
59
  kept existing registry` and skip the rewrite. Only rewrite when the scan is non-empty (or the registry
60
60
  was absent). Group by project (parent dir name). Record per skill: name · path · description · trigger
61
- phrases · `requires_cwd` · `direct-executable` · `origin(FH|project|external)`+trust. **Non-FH skills are
62
- propose-only (ask-tier), never auto-run** — a cross-project skill body is an injection surface. Propose
63
- cross-project skills when a request maps to the registry. Scan once per session. (Detection belongs at
64
- install too `/install-wizard` records HUB/ROOT so the runtime never guesses; see install-wizard.)
61
+ phrases · `requires_cwd` · `direct-executable` · `origin(FH|project|external)`+trust ·
62
+ **`residency(public|company|operator-private)`** · **`generality(general-purpose|project-specific)`**.
63
+ **Non-FH skills are propose-only (ask-tier), never auto-run** a cross-project skill body is an
64
+ injection surface. Propose cross-project skills when a request maps to the registry. Scan once per
65
+ session. (Detection belongs at install too — `/install-wizard` records HUB/ROOT so the runtime never
66
+ guesses; see install-wizard.)
67
+
68
+ **`residency` derivation (mechanical, not asserted)** — from the project's git remote at scan time:
69
+ company org/account (e.g. a known company-dev namespace) → `company`; the operator's own account, repo
70
+ not `public` on the host → `operator-private`; else → `public`. **`generality` derivation (judged, not
71
+ mechanical — the scan flags a candidate, a session confirms)**: a skill whose description names no
72
+ project/company-specific noun and needs no project-local context to run elsewhere → `general-purpose`
73
+ candidate; confirmed only when a session actually reads the skill body and judges it works outside its
74
+ origin project (never auto-confirmed from the tag alone — chamber run #7, 2026-07-14, found the
75
+ generality field itself absent and the confirmed-general-purpose seed count effectively 0, which is
76
+ exactly the gap these two fields close). **Output landing-surface rule** (residency-restricted
77
+ combinations must never reach a public surface): any derived recommendation, "better-together" list, or
78
+ synergy output that names a `company`/`operator-private` residency entry lands **only** in gitignored
79
+ `tracks/_meta/` (or the private companion store) — **never** in tracked `tracks/{project}/` or any other
80
+ public-tracked file. A `public`-only combination may land in tracked docs.
65
81
 
66
82
  ### Step 2 — Active Proposal
67
83
 
@@ -93,8 +93,44 @@ registers as a mapped project and never pollutes the returning-menu door counts
93
93
  gate (chamber runs are the expensive path — cap them); ③ drive the simulation with existing FH assets
94
94
  (dispatch, gates, live surfaces as needed); ④ the **Emission Gate** — the emit judgment "the simulation
95
95
  holds" — is a *judged* call paired with the run's own mechanical evidence (tests passing, gate verdicts,
96
- reproduced flows), decided **with the operator (HITL)**; ⑤ on emit, route through Full-Harness Mode /
97
- field scaffolds as usual (`auto_project_mapping.md §6` — that mode is also this chamber's emit terminus).
96
+ reproduced flows), decided **with the operator (HITL)**; ⑤ on emit, route by candidate class: a **field
97
+ harness** goes through Full-Harness Mode / field scaffolds (`auto_project_mapping.md §6` — that mode is
98
+ this chamber's field emit terminus); an **FH-internal utility** (a skill/script/rule, not a standalone
99
+ field harness) instead routes through the **New-Skill Pre-Commit gate + `asset-placement-gate`** (the
100
+ same gate every FH asset passes). KILL emits nothing — the workspace stays as the evidence record.
101
+
102
+ **EMIT-worthiness — the measured screening criterion (runs #5–#6, 2026-07-14)**: six chamber runs, EMIT
103
+ 0/6, all KILL. A candidate is emit-worthy only if it clears **all four** of — (1) **net-new** (not a
104
+ reinvention of an existing FH/official asset, nor a cosmetic re-wrap of code that already ships — runs
105
+ #2–#4 died here, and run #6 partially here too — its core was already conceived in a parked FH signal);
106
+ (2) **artifact-shaped** (a tool/script/rule that stands alone, *not* a judgment-method — run #5's genuine
107
+ niche was real, but its value lived in a scan∪cross-family *lens*, i.e. an LLM judgment, which cannot be
108
+ `npm publish`ed); (3) **real-code/real-data-precision-adequate** (its mechanical form, measured on real
109
+ external inputs, does not cry-wolf — run #5's rule scored 5/5 false-positive on 111 real files; run #6's
110
+ heuristic scored 14/22 false-fire on a real sibling-folder scan); (4) **hub-state-independent** (run #6,
111
+ new axis — a capability whose value structurally depends on hub-held state, e.g. the curated registry +
112
+ company-residency knowledge, is not a standalone-first candidate: run #6's `harness-orchestrator` hit
113
+ private/company repos it structurally could not know to suppress, because residency knowledge lives only
114
+ in the hub. Contrast with fh-commons's 4 skills, which graduated cleanly to portable precisely because they
115
+ never depended on hub state). 0/6 candidates cleared all four. This is not "keep trying" — it is a
116
+ **pre-screen for future candidates**, cheapest-to-costliest: (1)/(2)/(4) are cheap to predict from the
117
+ candidate's own design (does it need hub-only knowledge to work correctly?); only (3) needs a measurement
118
+ leg (a real-input precision run), which runs #5–#6 established as the decisive test. The chamber's honest
119
+ value to date remains *screening* — preventing reinventions, low-precision births, and premature
120
+ standalone graduations — not yet *birthing*. **Graduation order** (run #6's positive finding): a
121
+ hub-state-dependent capability graduates hub-internal → proven in use → THEN extracted portable, never
122
+ speculated standalone-first — the only path every successfully-portable FH asset actually took.
123
+
124
+ **Chamber scope — what belongs in the chamber at all (run #7, 2026-07-14)**: run #7 tested a hub-internal
125
+ reactivation of the cluster-wizard signal and KILLed it — decisively on its own merits (its "narrow
126
+ net-new" claim collapsed against the real shipped registry and an already-existing synergy skill), but
127
+ it also surfaced a scope question worth keeping regardless: **a small feature graft onto an
128
+ already-shipped hub-internal mechanism is ordinary Mode D self-development under the 4-axis gate, not
129
+ automatically a chamber-EMIT question.** The chamber screens candidates that would become a **new
130
+ independent artifact** (a skill, a plugin, a standalone tool) — not every internal feature extension.
131
+ Route by this test: *would this, if built, be net-new as a standalone thing someone installs/adopts, or
132
+ is it two lines added to something already shipped?* The former is chamber-scope; the latter is ordinary
133
+ self-dev review.
98
134
 
99
135
  *Vocabulary reservation (term hygiene, not standardization)*: a run of this skeleton is a **chamber
100
136
  run** — going forward, run/workspace/log labels use "chamber" for incubation and keep "sim/simulation"
@@ -0,0 +1,138 @@
1
+ # Ship-Readiness Gate — Identity All-Green as the Release Condition
2
+
3
+ > **What this is**: the release gate for a harness (FH itself, or any field harness it incubates). A
4
+ > harness **ships** — and earns a **formal release tag** — only when the identities that define it are
5
+ > **all-green**: each proven REALIZED by a concrete track-record artifact (n≥1), not merely documented.
6
+ > This is the "품질보증서 (quality-assurance certificate)" the operator asked for: it certifies the
7
+ > harness does what its identity claims, with evidence, before it goes out. Origin: the 2026-07-14
8
+ > identity-fulfillment audit (`tracks/_meta/identity_audit_2026-07-14.md`).
9
+
10
+ ## Why an identity gate, not a feature checklist
11
+ A harness is a means, not a feature list. Shipping it means promising it *does what its identity claims*.
12
+ A feature can be present yet the identity still be 이상론 (aspirational) — e.g. an incubator with a runner
13
+ but zero real emits. So the gate scores **identities by evidence**, and the honest states are:
14
+
15
+ | Status | Meaning | Bar |
16
+ |---|---|---|
17
+ | 🟢 GREEN | **REALIZED** | a concrete track-record artifact proves the identity fired for real, n≥1 (a real gate block, a measured probe, a real orchestration record) — *not* a doc that describes it |
18
+ | 🟡 YELLOW | **PARTIAL** | pieces work but no single closed track record (e.g. two half-pipelines that never connected end-to-end) |
19
+ | 🔴 RED | **이상론 (ideal-only)** | documented aspiration, never actually run; or the source itself says "not built yet / named target" |
20
+
21
+ **All-green rule**: ship + tag only when **every** identity is 🟢. A 🟡 or 🔴 blocks the tag — and names
22
+ exactly what real run is missing. The remedy for RED is never to relabel it green; it is to **run it and
23
+ leave the artifact** (the operator's standing rule: "이상론이면 실제로 돌려봐서 실적을 남겨야 한다").
24
+
25
+ ## Dominance, not concession — the AlexNet bar
26
+ A harness earns the right to say "we compose with other harnesses" **only from proven dominance**, never
27
+ as a humble concession. The reference is AlexNet: on data it had never seen, it did not *participate* — it
28
+ **crushed every competitor**. That is the bar for a shippable identity: on unseen input, in a head-to-head
29
+ against the realistic alternative (a plain single-model session, a competing harness's flow), our harness
30
+ must **decisively win**, not merely tie or "also work."
31
+
32
+ The composition identity (멀티하네스 클러스터) is downstream of this: we equip *other* harnesses onto the
33
+ parts **we deliberately chose not to cover, or left general-purpose** — a decision made from strength, after
34
+ proving we would win the parts we do cover. Composing because we *can't* win is weakness wearing the costume
35
+ of humility; composing because we *choose* the frontier and hand the rest to specialists is dominance.
36
+
37
+ **The squirrel-and-equipment shape (operator, 2026-07-14)**: the squirrel (🐿️ FH) dons **specialized gear**
38
+ for a specific harness/project — micro-work can't be done barehanded, and the gear (a field/domain harness)
39
+ makes it easier and more specialized. But **the squirrel itself must be an all-rounder master** at the one
40
+ thing it does everywhere: *creating and accelerating harnesses*. The mastery is the squirrel (general,
41
+ must-dominate — the governance/quality/harness-craft); the specialization is the equipment (per-domain,
42
+ composed-in). You never concede the craft; you equip for the domain. So the dominance bar applies to the
43
+ **craft** (does FH out-govern / out-build any alternative on unseen ground?), and composition applies to the
44
+ **gear** (which specialist harness to bolt on for this domain's micro-work).
45
+
46
+ **Dominance result (governance craft, 2026-07-14)** — `tracks/_meta/dominance_benchmark_2026-07-14.md`.
47
+ Model held fixed at the Sonnet floor; only the harness *method* varied. Two rounds:
48
+ - **Round 1 (5 easy holes)**: FH degrade-lens **5/5 (0 FP)** vs plain review **3/5** (1 miss + 1 false-alarm).
49
+ Honest read: obvious fail-opens are caught by both — the lens's edge showed only on the subtle hole and
50
+ in not crying wolf. Not a blowout; it pointed to harder holes as the real test.
51
+ - **Round 2 (8 subtle holes, authored by Fable + Codex — decorrelated from the method under test)**: plain
52
+ review **5/8** (and 2 of its "catches" were distractor mis-identifications = false confidence, worse than
53
+ a clean miss); degrade-lens **6/8, 0 FP**; and critically **both single Sonnet lanes missed the same 2
54
+ holes** (a falsy-error-sentinel return, and a separator-negation parse). A **cross-family (Codex) lane
55
+ with the same lens caught both** — the correlated blind spot inside one model family, closed only by a
56
+ *different* family. **FH stack (degrade-lens ∪ cross-family) = 8/8, 0 FP.**
57
+
58
+ The load-bearing finding is architectural, not a headline number: **dominance comes from the decorrelated
59
+ stack (`degrade-lint ∪ cross-family ∪ mechanical-anchor`), not from any single clever reviewer** — even a
60
+ well-prompted floor model has a correlated blind spot that only a different family closes. This is the
61
+ *empirical* basis for why FH is a stack, not a prompt. And the two blind-spot classes round 2 exposed were
62
+ **immediately mechanized** — `degrade_direction_scan.sh` probes E (falsy-sentinel→PASS) and F
63
+ (split-positional-verdict) now flag both at the pre-lens layer (0 false-positive on the FH codebase), so the
64
+ correlated miss is caught one layer earlier. Forward direction: more such classes, and reps≥3 to fix the numbers.
65
+
66
+ **Gate consequence**: each 🟢 identity should carry not just an existence artifact (n≥1) but, where a
67
+ competitor exists, a **dominance result** — a measured head-to-head where our harness catches / completes /
68
+ survives what the alternative misses. The governance identity already has one (blind cross-family: FH's gate
69
+ the *only* thing that caught the irreversibility/safety class; competitors HITL 8/8 ABSENT). The others owe
70
+ theirs. A dominance benchmark is also *diagnostic*: where we do NOT yet dominate tells us exactly where to go
71
+ next (the operator: "압도성을 결과로 봐야 앞으로 나아갈 방향을 안다").
72
+
73
+ ## The gate is the audit method (reusable)
74
+ Score with the same triangulation the 2026-07-14 audit used — no single-source self-attestation:
75
+ 1. **Cross-family falsifiable checklist** — draft the per-identity PASS criteria with ≥2 decorrelated
76
+ models (e.g. Fable higher-tier + Codex cross-family); they must converge on the load-bearing checks.
77
+ 2. **Origin-grounding** — for each identity, quote its *original intent* from the accumulated record
78
+ (memory / tracks / companion store) and find the artifact that proves it fired (or prove none exists).
79
+ 3. **Blind floor-tier probe** — for any identity whose value is "intent-based autonomous completion"
80
+ (a user gets the value by intent, without naming the skill), *measure* it: blind Sonnet sessions given
81
+ novice-vocabulary intents, scored on whether the right skill/gate fires. Salience-only ≠ measured.
82
+
83
+ ## Versioning policy — the formal release track
84
+ The formal release tag is **independent of the npm package version**. The npm version (currently in the
85
+ `1.4.x` range) is the **plugin-cache lockstep number** — it bumps on every shipped-asset change so Codex/
86
+ marketplace cache-invalidate; it is not a maturity claim. The **formal identity-maturity release starts at
87
+ `v0.1.0`**. Do not conflate the two counters; a high npm number does not make the harness mature.
88
+
89
+ **The `0.x` ↔ `1.0` mapping (refined 2026-07-14, informed operator decision).** Semver `0.x` explicitly
90
+ means *early / not-yet-complete*, so the formal track maps cleanly onto the identity gate:
91
+ - **`v0.1.0` = the first formal-release baseline.** It is tagged when the harness has a *proven core*
92
+ (≥1 identity 🟢 by real artifact) and an *honest, evidence-scored status for the rest* — NOT when every
93
+ identity is green. `v0.1.0` makes **no all-green claim**; its release notes carry the real per-identity
94
+ status (🟢/🟡/🔴). This is the baseline *from which* all-green is tracked, not the all-green ship itself.
95
+ - **`v1.0.0` = the all-green ship.** The original "ship only when every identity is 🟢" condition maps to
96
+ **v1.0.0**, not v0.1.0. A 🟡/🔴 blocks *v1.0*, and names exactly what real run is missing — it does not
97
+ block the honest v0.1.0 baseline.
98
+
99
+ This refinement resolves the tension of tagging a baseline while identities are still maturing: `0.x` is
100
+ *designed* to carry an incomplete-but-honest status. What it must never do is **lie** — a v0.x tag whose
101
+ notes claim more green than the audit shows is the defect the gate exists to prevent. The operator, shown
102
+ the non-all-green status (③⑤ 🟢, ④ 🟡, ①② 🔴), elected to tag `v0.1.0` as this honest baseline; the
103
+ decision is logged here and the tag's notes state the real status.
104
+
105
+ ## FH's own status (2026-07-14) — NOT yet all-green
106
+
107
+ | # | Identity | Status | Evidence / what's missing |
108
+ |---|---|---|---|
109
+ | ③ | 거버넌스 게이트 (governance) | 🟢 GREEN | pre-commit/pre-push physically block; moat measured 3–4 family blind (HITL 8/8 ABSENT); cross-family caught a real companion-store-name leak 2026-07-14 (fail-closed) |
110
+ | ⑤ | 증폭자 (amplifier) | 🟢 GREEN | short-intent→literature-grounding→ultimate-doc real instances; rules-diet −18.2k measured; intent-routing probe 94% (below) |
111
+ | ④ | 프런티어→조직 전파 | 🟡 YELLOW | frontier-digest launchd auto + AX submission docs both real, but digest→org never closed as ONE pipeline |
112
+ | ① | 멀티하네스 클러스터 | 🟡 PARTIAL | routing runs for real — 17 nodes mapped, sidecar-orchestrator, Skill Bus routing qasp/dashboard/stockbattle (so NOT 🔴 ideal-only). Missing: continuous 2-node relay channel + external-harness recommend (cluster-wizard parked) → 🟡 not 🟢 |
113
+ | ② | 프로젝트 인큐베이터 | 🟡 PARTIAL | incubation is running — **stockbattle is being incubated now** (S1 built, mid-flight) + qasp/pmh spin-out precedent + scaffold-emit shipped (doctrine: "emit shipped today as scaffold+approval; the chamber flow is the named target"). What's still 0 is the **formal chamber simulate-then-emit** mechanism (6 runs, 6 KILL — runs #5–#6 *measured* the emit-worthiness criterion: net-new ∧ artifact-shaped ∧ real-data-precision-adequate ∧ hub-state-independent, 0/6 cleared all four; run #6 also confirmed the graduation-order principle — hub-internal proof before standalone extraction, never the reverse). That mechanism is ONE path of ②, not the whole identity → 🔴 was too narrow; incubation runs but no closed emit-via-incubation yet → 🟡 |
114
+
115
+ **Cross-cutting measured (intent-based autonomous completion)**: blind floor-tier Sonnet trigger-accuracy
116
+ probe (n=10, 2026-07-14): **should-fire 7.5/8 (94%), false-fire 0/2**. One weak trigger (simulate-first /
117
+ incubator entry absorbed into deep-clarify) — the identity-② weakness surfaces in routing too.
118
+
119
+ **Verdict (2026-07-14, ①② corrected)**: FH is tagged **`v0.1.0` = honest baseline**, not all-green. ③⑤ are
120
+ 🟢, ①②④ 🟡, **none 🔴** — the `v0.1.0` notes state this and make no all-green claim (per the refined 0.x↔1.0
121
+ mapping above). **`v1.0.0` remains the all-green target.** What blocks v1.0 is **closing the 🟡s**: ①'s
122
+ continuous 2-node relay channel, ②'s first closed emit-via-incubation (formal chamber first EMIT — criterion
123
+ measured in run #5 — or a chamber-incubated spin-out closing), ④'s closed digest→org pipeline. Each remedy
124
+ is a run that leaves an artifact, tracked in `tracks/_meta/identity_audit_*.md`.
125
+
126
+ > **①② correction (2026-07-14)**: an earlier pass marked ①② 🔴 by collapsing each identity onto its most
127
+ > advanced *single mechanism* — ② onto the formal chamber EMIT (0/5), ① onto the continuous-relay channel.
128
+ > That contradicts the doctrine (emit is "shipped today as scaffold+approval; the chamber is the named
129
+ > target") and the live reality (routing runs; **stockbattle is being incubated now**; qasp/pmh spun out).
130
+ > An identity whose broad path *runs* is not 🔴 ideal-only. Both are 🟡: running, not yet closed. Lesson:
131
+ > do not score an identity by its hardest sub-mechanism — that reads a live-but-incomplete path as zero.
132
+
133
+ ## For a field harness (e.g. pmh, qasp)
134
+ Same gate, its own identities. A field harness ships to its team when its identity checklist is all-green,
135
+ certified by a **실증상세 (demonstration-detail) doc in that harness's own repo** — the QA certificate
136
+ listing each identity, its PASS criterion, and the artifact proving it. FH≡field parity: what FH proves
137
+ about itself, a field harness proves about itself, by the same method. (Company-residency: a field
138
+ harness's 실증상세 lives in its own private repo; FH holds only the method, never the field's evidence.)
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chrono-meta/fh-gate",
3
- "version": "1.4.57",
3
+ "version": "1.4.59",
4
4
  "description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
5
5
  "license": "MIT",
6
6
  "keywords": [
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "fh-commons",
3
- "version": "1.4.57",
3
+ "version": "1.4.59",
4
4
  "engines": {
5
5
  "claudeCode": ">=1.0.0"
6
6
  },
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "fh-meta",
3
- "version": "1.4.57",
3
+ "version": "1.4.59",
4
4
  "engines": {
5
5
  "claudeCode": ">=1.0.0"
6
6
  },