@chrono-meta/fh-gate 2.6.0 → 2.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/rules/fh_4axis_gate.md +26 -3
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +24 -3
- package/CLAUDE.md +39 -12
- package/README.ja.md +32 -8
- package/README.ko.md +32 -8
- package/README.md +112 -15
- package/README.zh.md +28 -8
- package/knowledge/shared/harness-core/field_verdict_crossfamily_gate.md +19 -5
- package/knowledge/shared/harness-core/harness_incubator_doctrine.md +12 -2
- package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +32 -0
- package/knowledge/shared/harness-core/ship_readiness_gate.md +77 -8
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +90 -0
- package/knowledge/shared/rules/knowledge_layer_seam.md +1 -1
- package/package.json +13 -2
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +111 -0
- package/plugins/fh-meta/agents/persona-innovator.md +170 -0
- package/plugins/fh-meta/skills/public-surface-audit/SKILL_detail.md +16 -1
- package/plugins/fh-meta/skills/steel-quench/SKILL.md +97 -0
- package/scripts/adapters/peer_resolve.sh +58 -6
- package/scripts/cluster_capability_scan.sh +42 -13
- package/scripts/digest_landing_check.sh +142 -2
- package/scripts/fh_hub_identity.sh +83 -0
- package/scripts/fh_session_load.sh +53 -5
- package/scripts/fh_track_resolve.sh +114 -0
- package/scripts/field_canon_preload.sh +50 -5
- package/scripts/package_coverage_check.sh +8 -0
- package/scripts/prior_art_prompt.sh +168 -0
- package/scripts/psa_scan_lib.sh +201 -15
- package/scripts/residency_admission_check.sh +204 -0
- package/scripts/selfcheck.sh +88 -0
- package/scripts/test_adapter_lanes.sh +67 -2
- package/scripts/test_heavy_classifier_lanes.sh +144 -0
- package/scripts/test_marker_defense_lanes.sh +152 -0
- package/scripts/test_marker_soul_check_lanes.sh +211 -0
- package/scripts/test_prior_art_prompt_lanes.sh +128 -0
- package/scripts/test_psa_singlefile_lanes.sh +351 -1
- package/scripts/test_residency_admission_lanes.sh +60 -0
- package/scripts/test_track_resolve_lanes.sh +158 -0
- package/templates/.git-hooks/pre-commit +400 -4
- package/templates/.git-hooks/pre-push +17 -2
- package/templates/settings.PriorArt.snippet.json +15 -0
|
@@ -190,9 +190,32 @@ controls: n/a — no measurement in this delta (<reason>)
|
|
|
190
190
|
```
|
|
191
191
|
|
|
192
192
|
**`standpoint:`** — closed enum, canonical spec in
|
|
193
|
-
`knowledge/shared/harness-core/field_verdict_crossfamily_gate.md §7`. **Still validated
|
|
194
|
-
— zero hook lines, no fixture suite
|
|
195
|
-
|
|
193
|
+
`knowledge/shared/harness-core/field_verdict_crossfamily_gate.md §7`. 🟥 **CORRECTED 2026-08-20, then CORRECTED AGAIN the same hour.** This used to read "Still validated
|
|
194
|
+
by nothing — zero hook lines, no fixture suite", which is FALSE: `validate_standpoint_leg()` lives in
|
|
195
|
+
`templates/.git-hooks/pre-commit` (grep the function name — **line numbers are deliberately not cited
|
|
196
|
+
here; the first version of this fix cited `:798`/`:1575` and a commit landed the same hour that moved
|
|
197
|
+
them to `:878`/`:1665`**), and `scripts/test_marker_standpoint_lanes.sh` is wired through
|
|
198
|
+
`scripts/selfcheck.sh`.
|
|
199
|
+
🟥 **But the first correction over-shot, and cross-family caught that too.** It said "the value enum
|
|
200
|
+
IS closed and the grounds ARE required". Only the first half holds. Measured by varying ONE variable
|
|
201
|
+
at a time — the original known-pair varied two and mis-attributed the result:
|
|
202
|
+
|
|
203
|
+
```
|
|
204
|
+
standpoint: tier2(qasp) — ran <cmd>, saw <out> rc=0 ✅
|
|
205
|
+
standpoint: tier2(qasp) (no grounds) rc=0 ⚠️ warns, does NOT block
|
|
206
|
+
standpoint: tier2 (no parens) rc=1 ❌ not an enum member
|
|
207
|
+
standpoint: banana(qasp) rc=1 ❌ not an enum member
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
So: **the enum is closed and enforced; the execution grounds for `tier2`+ are ADVISORY.** A marker can
|
|
211
|
+
still record `tier2` without naming a command and pass with a warning — that is the real remaining
|
|
212
|
+
gap, and it is narrower than "validated by nothing" and wider than "grounds are required". Neither
|
|
213
|
+
earlier sentence was accurate, and the accurate one required varying one variable at a time.
|
|
214
|
+
🟥 **This same false claim stood in FIVE places**, not three — `CLAUDE.md` (twice), here, `AGENTS.md`,
|
|
215
|
+
and `field_verdict_crossfamily_gate.md`. Fixing one and stopping is the half-fix propagation failure;
|
|
216
|
+
fixing three and stopping was the same failure one round later. The question is never "is this
|
|
217
|
+
sentence wrong" but **"where else does this sentence live"** — and the answer came from a
|
|
218
|
+
different-family reviewer, not from me.
|
|
196
219
|
|
|
197
220
|
**`thirdparty:` — ⓓ3자 대면의 자기 필드 (2026-08-17 신설).** ⓑ가 `standpoint:` 를 갖는 것과
|
|
198
221
|
같은 형태다: `axes-run` 에는 포인터(`ⓓ=→thirdparty`)만 두고 값은 이 필드가 나른다.
|
|
@@ -11,13 +11,13 @@
|
|
|
11
11
|
"plugins": [
|
|
12
12
|
{
|
|
13
13
|
"name": "fh-meta",
|
|
14
|
-
"version": "2.
|
|
14
|
+
"version": "2.7.0",
|
|
15
15
|
"description": "New in 2.2.0: BREAKING (gate): chamber step 6 now reads ACTUAL.md, not BUDGET.md — an in-flight chamber run whose actual cost sits in BUDGET.md blocks until the ACTUAL: line moves to tracks/_chamber/<slug>/ACTUAL.md (the runner prints the path). Why: BUDGET.md's pre-verdict hash IS the ordering witness, and step 6 hard-blocked until that same file changed, so every run that reached COMPLETE necessarily mutated a witnessed artifact and verify returned TAMPERED — the chamber's promotion condition was unsatisfiable by construction, not by strictness. Two roles (immutable witness / post-verdict calibration sink) had collided in one file; each was correct alone, so neither side's code showed the conflict. Also: ko-tech-writer Step 2/4-b scans are now calibration-backed (known-pair fixtures + reproducible command, shipped) — discrimination is proven, 'zero residue' is explicitly NOT; chamber lane suite 12 -> 33 including the runner x witness seam no test covered; chamber_run.sh now teaches the two-commit discipline (gate hashes and verdict hash must land in separate commits/PRs — it previously advised the opposite). New in 2.1.0: BREAKING (gate): `crossfamily: declined` in an Axes 2-3 marker now requires grounds naming a record path that RESOLVES on disk — bare `declined`, and `declined` justified by author judgment, are blocked at commit. Remedy: cite where the operator decision lives (e.g. `.. — operator declined sidecars, per knowledge/shared/rules/operational_adaptation.md`), or use `DEGRADED_PANEL_UNUSED` if a panel was reachable and you chose not to recruit it — which is what author judgment actually is. `declined` was the only enum value with no grounds requirement; a cross-family review then broke the first (vocabulary-grep) fix three ways — self-validating on the value's own token, vacuous keyword passes, and over-blocking real declinations in natural prose — so the check asserts a resolvable record instead of words. Also: standpoint axis gains `tier1b` (a STATIC read of a target repo, executed nothing) plus a decide-in-order procedure, after blind floor-tier sims graded pure cold-reads as `tier2` three rounds running; steel-quench Wave 1's sixth angle (gate-locality) gains the output-template row it never had, so a mandatory angle stops being structurally unreportable; verify-bidirectional gains category 5 (prescriptive doctrine statement); Sister Asset Protocol gains an active-adoption trigger; new resident doctrine — Mechanization Boundary, Local Execution First, Skeleton-not-Muscle, Expedition track, and this package's versioning policy. Hub meta-operations toolkit — 35 skills + 7 agents. New in 2.0.1: harness-doctor cadence hook, portability lint wired into pre-commit, branch_claim.sh claim-count-vs-tree-count warning, louder confidentiality-scan fail-open notice, fh-gate.sh missing-package.json survival, identity ① reclassified 🟢 (cross-harness adapters + relay argument channel). New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
16
16
|
"source": "./plugins/fh-meta"
|
|
17
17
|
},
|
|
18
18
|
{
|
|
19
19
|
"name": "fh-commons",
|
|
20
|
-
"version": "2.
|
|
20
|
+
"version": "2.7.0",
|
|
21
21
|
"description": "Project-agnostic utility skills — 5 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate · ko-tech-writer) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
|
|
22
22
|
"source": "./plugins/fh-commons"
|
|
23
23
|
}
|
package/AGENTS.md
CHANGED
|
@@ -127,9 +127,14 @@ Because non-Claude runtimes do not auto-load Claude path rules, apply these rule
|
|
|
127
127
|
2 of the 4 circled-key markers on disk are dated 2026-08-10 and carry the OLD meanings. Aligning
|
|
128
128
|
the notation still helps going forward; it does not work backwards.
|
|
129
129
|
`standpoint:` remains the canonical field for ⓑ (the `axes-run` entry
|
|
130
|
-
is only a pointer to it, and a pointer at an empty field is blocked)
|
|
131
|
-
|
|
132
|
-
|
|
130
|
+
is only a pointer to it, and a pointer at an empty field is blocked). 🟥 **CORRECTED 2026-08-20.** This used to say
|
|
131
|
+
"its value enum is still validated by nothing" — FALSE. `validate_standpoint_leg()` in the
|
|
132
|
+
pre-commit hook enforces a closed enum. Measured by varying one variable at a time:
|
|
133
|
+
`standpoint: banana(qasp)` → **blocked (enum)** · `standpoint: tier2` without parens → **blocked
|
|
134
|
+
(enum)** · `standpoint: tier2(qasp)` with no execution grounds → **passes with a warning** ·
|
|
135
|
+
with grounds → **passes**. So the enum is enforced and the `tier2`+ execution grounds are
|
|
136
|
+
**advisory** — a first version of this correction said grounds were required, which over-shot.
|
|
137
|
+
The hook also cannot check whether `tier2` is **true**. Those two, together, are the gap. Format spec: `.claude/rules/fh_4axis_gate.md §Marker axis fields`.
|
|
133
138
|
(Two drift corrections landed here on 2026-08-17: first this sentence said "four" while its own
|
|
134
139
|
next clause described the +1 — caught by the session-close ④-b CC↔Codex parity check — and then
|
|
135
140
|
the machine layer moved to six the same day.)
|
|
@@ -144,6 +149,22 @@ Because non-Claude runtimes do not auto-load Claude path rules, apply these rule
|
|
|
144
149
|
`❌`. Execution is the half with no substitute; see `field_verdict_crossfamily_gate.md §7`. Read it before recording; §1-c
|
|
145
150
|
holds the sample limits — read it before citing. This record is self-attested and has no hook
|
|
146
151
|
behind it; it is closed by a different-family reader, not by writing it more carefully.
|
|
152
|
+
🟥 **`axis2-defense:` — a marker field added 2026-08-20 that a Codex-side author WILL hit.**
|
|
153
|
+
It is required when the marker records `floor-status: sonnet-floor` or `below-floor`, and it carries
|
|
154
|
+
three sub-answers (on one line, or across continuation lines — the hook reads both) about **your own findings and numbers**:
|
|
155
|
+
`axis2-defense: reproducibility=<exact command or file:line another session runs> fairness=<reps,
|
|
156
|
+
inputs and environment named for BOTH arms> estimation-layer=<per number: measured | estimate |
|
|
157
|
+
quotation; for a measurement, what showed the instrument works on this target>`.
|
|
158
|
+
The hook (`validate_defense_leg`) checks **presence · completeness · non-vacuity** — `ok`/`yes`/
|
|
159
|
+
`n/a` is rejected as a filled form rather than an answer, and all three sub-answers must exist.
|
|
160
|
+
It cannot check whether the answers are true. Why it exists: measured n=6 in the origin field, a
|
|
161
|
+
floor-tier pass runs the attack angles without defect but does **not** spontaneously ask these
|
|
162
|
+
three; that is a checklist gap, not a capability gap, and a harness whose behaviour depends on
|
|
163
|
+
which model drives it is defective by the Sonnet-Floor doctrine.
|
|
164
|
+
Fixtures: `scripts/test_marker_defense_lanes.sh`. Spec: `plugins/fh-meta/skills/steel-quench/SKILL.md`
|
|
165
|
+
§Wave 1-D. **Recorded here because a rule living in only one entry point is invisible to the other
|
|
166
|
+
runtime** — that is the gate-locality principle, and this line is it being applied rather than cited.
|
|
167
|
+
|
|
147
168
|
8. **Branch-surface claims:** GitHub branch protection is two independent layers — legacy
|
|
148
169
|
protection and rulesets coexist, and the strictest wins. Read both
|
|
149
170
|
`/repos/{owner}/{repo}/branches/{branch}/protection` and
|
package/CLAUDE.md
CHANGED
|
@@ -376,8 +376,18 @@ claim 이 `main` 인 동안 실제 HEAD 는 peer 브랜치였다 — 두 세션
|
|
|
376
376
|
**남은 잔여 둘, 이름으로 남긴다**: ⓐ 게이트는 **커밋 시점**이라 `switch -c` 사고 자체는 못 막는다
|
|
377
377
|
(그래서 위 두 줄이 여전히 사람 몫이다 — [[feedback_gate_binding_point_not_check_point]])
|
|
378
378
|
ⓑ 반대 방향 — 내가 브랜치를 쥔 동안 남이 되돌리고 커밋하면 **내 staged 파일이 index 에 살아
|
|
379
|
-
있다.**
|
|
380
|
-
|
|
379
|
+
있다.** 🟥 **2026-08-21 실제로 났고, 섞은 세션은 파일 지정 `add` 를 썼다**(초판은 여기 「파일
|
|
380
|
+
지정 add 를 썼기 때문에 안 섞였다」고 적었는데 거짓이다). 파일 지정 add 는 «내가 무엇을
|
|
381
|
+
**추가**하나»만 통제하고 **index 에 이미 있는 것은 못 막는다.** ⇒ `git add -A` 금지는 필요조건
|
|
382
|
+
이지 충분조건이 아니다. **커밋을 한 호출로 묶고 그 안에서 둘을 확인한다** — 확인은 점이고 위험은
|
|
383
|
+
구간이라, 사이에 턴이 끼면 창이 다시 생긴다(같은 날 두 세션이 시점을 각각 `switch` 직전/직후로
|
|
384
|
+
달리 골랐는데 **둘 다 뚫렸다**):
|
|
385
|
+
```
|
|
386
|
+
B=$(git branch --show-current); [ "$B" = "<내 브랜치>" ] || exit 1
|
|
387
|
+
git diff --cached --name-only # 내 것만 있나 — 남의 것은 restore --staged
|
|
388
|
+
git commit …
|
|
389
|
+
```
|
|
390
|
+
([[feedback_shared_checkout_ops_touch_others_work]] · [[feedback_gate_binding_point_not_check_point]])
|
|
381
391
|
|
|
382
392
|
**Integration branch is PR-only** (operator decision 2026-07-20). Never `git push origin main`
|
|
383
393
|
directly. Normal path: `git switch -c <branch>` → push the branch → `gh pr create` → after review
|
|
@@ -625,17 +635,24 @@ English with FH's own persona/viewpoint sense of "standpoint" (`fh-meta:beginner
|
|
|
625
635
|
`expert`) — a different axis (which persona reviews, not whose repo is ground truth); kept as-is,
|
|
626
636
|
not renamed, but do not conflate the two. 🟥 **CORRECTED 2026-08-20 — this paragraph used to say
|
|
627
637
|
`standpoint:` was "Prose-only today — no pre-commit hook or fixture suite validates this field yet".
|
|
628
|
-
That is FALSE and was false in this same file**: `validate_standpoint_leg()`
|
|
629
|
-
`templates/.git-hooks/pre-commit
|
|
630
|
-
`scripts/test_marker_standpoint_lanes.sh` is wired through `scripts/selfcheck.sh
|
|
638
|
+
That is FALSE and was false in this same file**: `validate_standpoint_leg()` lives in
|
|
639
|
+
`templates/.git-hooks/pre-commit` and is called from the marker block, and its fixture suite
|
|
640
|
+
`scripts/test_marker_standpoint_lanes.sh` is wired through `scripts/selfcheck.sh`. (🟥 **Grep the
|
|
641
|
+
names, do not trust line numbers** — the first version of this correction cited `:798`/`:1575` and a
|
|
642
|
+
commit landed the same hour that moved them to `:878`/`:1665`. A hardcoded anchor in prose is a
|
|
643
|
+
phantom waiting for the next edit.) §자기 대조
|
|
631
644
|
above already said so (PR #429), so **one file carried both claims at once** and a reader landed on
|
|
632
645
|
whichever they reached first. Found by the residency-ledger pass, not by a lane — no check compares
|
|
633
646
|
a rule's self-description against the machinery it describes, which is why a stale "we have not
|
|
634
647
|
built this yet" is the quietest form of drift: it reads as honest modesty and it suppresses use of a
|
|
635
|
-
control that already exists. **What is validated is
|
|
636
|
-
|
|
637
|
-
|
|
638
|
-
|
|
648
|
+
control that already exists. **What is validated is the ENUM** — measured by varying ONE variable at a
|
|
649
|
+
time, because the first version of this correction varied two and mis-attributed the result:
|
|
650
|
+
`banana(qasp)` → blocked (enum) · `tier2` without parens → blocked (enum) · `tier2(qasp)` with **no**
|
|
651
|
+
execution grounds → **passes with a warning** · with grounds → passes. 🟥 So the first fix's claim
|
|
652
|
+
that "grounds are non-empty" are checked **over-shot, and a different-family reviewer caught it**:
|
|
653
|
+
the `tier2`+ execution grounds are **advisory**. Two residuals remain and both are real — grounds are
|
|
654
|
+
not enforced, and whether `tier2` is *true* is still self-attested. What was wrong was only the claim
|
|
655
|
+
that nothing validated the field at all. Three artifacts, one carrying two
|
|
639
656
|
independent trials (forge-harness PR #368, a sibling field harness's PR #8 reps=3 and its
|
|
640
657
|
known-answer trial, qasp-dev PR #161 as adjacent corroboration) crossed this repo's own evidence
|
|
641
658
|
bar the same day this was formalized — including one caught by this session's own qasp PR #161
|
|
@@ -824,7 +841,8 @@ deleted branch was carrying).
|
|
|
824
841
|
**When this gate fires** — *before* any of: branch deletion (local or remote), history rewrite /
|
|
825
842
|
force-push, scrub of tracked history, bulk deletion of session records / tracks content.
|
|
826
843
|
|
|
827
|
-
1. **Enumerate (measured)
|
|
844
|
+
1. **Enumerate (measured)** — 🟥 **this step is run BY A HUMAN; no hook executes it** (measured
|
|
845
|
+
2026-08-20, residency-ledger pass): `bash templates/predelete_check.sh <repo> [base]` — per branch: commits
|
|
828
846
|
off base + unique paths. Verdicts: SAFE (fully merged) · CHECK (0 unique paths but commits off
|
|
829
847
|
base — shared files may hold *newer* content, e.g. an unmerged session card) · REVIEW (unique
|
|
830
848
|
paths — recovery mandatory).
|
|
@@ -845,6 +863,15 @@ prose gate is now stopped); it does **not** close the injected/adversarial one
|
|
|
845
863
|
or errors, this irreversible surface **fails closed** — the pre-push hook blocks (enumerate by hand or
|
|
846
864
|
take the explicit `DESTRUCTIVE_OP_OK=1` override); a tooling-down enumerate step never silently degrades
|
|
847
865
|
into "just delete it."
|
|
866
|
+
🟥 **Read that precisely — the floor is the HOOK, not this script.** `templates/.git-hooks/pre-push`
|
|
867
|
+
implements the per-ref verdict **inline**; it does not call `predelete_check.sh`. Measured 2026-08-20
|
|
868
|
+
(control: the same scan finds `session_close_check` wired in that hook): every in-repo reference to
|
|
869
|
+
`predelete_check.sh` is a **mention, not an execution** — `pre-push:408` lists the path inside a *grep
|
|
870
|
+
pattern*, `destructive_pre_gate.sh:191` *prints the command* as advisory text, and `selfcheck.sh:192`
|
|
871
|
+
runs `bash -n` on it. So the script is the **operator's enumerate tool**, and step 1 above is a human
|
|
872
|
+
step that the machinery reminds you of rather than performs. Stating it the other way round is the
|
|
873
|
+
"prose-invoked floor" the 4-axis marker spec calls M-tier when a rule claims a floor its script has
|
|
874
|
+
no caller for — this section does not make that claim, and this note keeps it from drifting into one.
|
|
848
875
|
|
|
849
876
|
> **Detail**: See `knowledge/shared/harness-core/claude_md_gate_details.md §Destructive-Op-Hook-Coverage`
|
|
850
877
|
> — the per-ref verdict mechanics, what the hook does/does not close (honest scope + adversarial residual),
|
|
@@ -1005,7 +1032,7 @@ Self-healing is not only FH-self-dev (Mode D 4-axis) and reactive (`verify-bidir
|
|
|
1005
1032
|
|
|
1006
1033
|
## Agent Dispatch Operation (FH cwd-Based)
|
|
1007
1034
|
|
|
1008
|
-
> **Runtime authority (canonical):** one explicit governor per context + capability-routed sidecars; sidecar findings are evidence candidates, not terminal verdicts, until source-closed by the governor *via a mechanical anchor* — never governor agreement alone. CC=action/governor · Codex=repo-grounded audit sidecar · Gemini/agy=breadth/multimodal sidecar · other runtimes=portable `AGENTS.md` entrypoint only. Full doctrine + Maintenance-Cost Rule: `knowledge/shared/harness-core/multi_model_sidecar_strategy.md §Runtime Authority`.
|
|
1035
|
+
> **Runtime authority (canonical):** one explicit governor per context + capability-routed sidecars; sidecar findings are evidence candidates, not terminal verdicts, until source-closed by the governor *via a mechanical anchor* — never governor agreement alone. 🟥 **A sidecar audits; it does not WRITE to the target tree** — findings and at-most a proposed patch as text, applied by the governor (measured 2026-08-21: an auditor sidecar edited the tree and its fix introduced a self-referential fail-open that 41 lanes passed). CC=action/governor · Codex=repo-grounded audit sidecar · Gemini/agy=breadth/multimodal sidecar · other runtimes=portable `AGENTS.md` entrypoint only. Full doctrine + Maintenance-Cost Rule: `knowledge/shared/harness-core/multi_model_sidecar_strategy.md §Runtime Authority`.
|
|
1009
1036
|
|
|
1010
1037
|
**Isolated delegation is a component of the identity, not an optional extra** (operator decision,
|
|
1011
1038
|
2026-08-08). FH/PMH are defined as governor + orchestrator; a harness that cannot dispatch is a
|
|
@@ -1317,7 +1344,7 @@ Closing phrase detected ("wrap up", "done", "good work", "end session", etc.)
|
|
|
1317
1344
|
|
|
1318
1345
|
| Bump | Reserved for (operator's own wording, 2026-08-17) |
|
|
1319
1346
|
|---|---|
|
|
1320
|
-
| **major** `+1.0.0` | **any one of three**: ⓐ **완전히 새로 지음** — rebuilt from scratch, not extended · ⓑ **정체성이 확립됨** — an identity of the five
|
|
1347
|
+
| **major** `+1.0.0` | **any one of three**: ⓐ **완전히 새로 지음** — rebuilt from scratch, not extended · ⓑ **정체성이 확립됨** — 🟥 **다섯이 «전부» 🟢** 인 순간이지 하나가 🟢 로 올라선 순간이 아니다(운영자 결정 2026-08-21). 초판은 *"an identity of the five … actually standing 🟢"* 였고 **「하나만 초록이어도 major」로 읽혔다** — 실제로 그날 ②가 🟢 로 판정되면서 3.0.0 후보로 올라왔고, 그 애매함이 그때 닫혔다. 🟥 그리고 **정체성 등급은 npm 이 나르는 신호가 아니다** — 그건 `identity-v*` 계보의 사건이고, npm 이 또 나르면 같은 날 고친 「두 계보 한 이름」 결함을 번호에서 재생산한다. ⇒ major-ⓑ 는 **`identity-v1.0.0` 과 같은 사건**을 가리킨다 · ⓒ **기능이 혁신적으로 변경되거나 늘어남** — a capability *class* appears or is replaced, not a capability instance. 🟥 **Never** for tightening a gate that already existed |
|
|
1321
1348
|
| **minor** `+0.1.0` | 미들급 — new assets, new gate lanes, doctrine that changes behavior; **including changes that break a consumer's gate acceptance**, which then carry a mandatory `BREAKING (gate):` line |
|
|
1322
1349
|
| **patch** `+0.0.1` | 트리비아급 — fixes, wiring, docs that change no behavior |
|
|
1323
1350
|
|
package/README.ja.md
CHANGED
|
@@ -18,24 +18,48 @@
|
|
|
18
18
|
</p>
|
|
19
19
|
|
|
20
20
|
<p align="center">
|
|
21
|
-
<
|
|
21
|
+
<b>やりたいことを頼んでください。同じ依頼が繰り返されたら、それ自体を作りましょうと先に提案します。</b>
|
|
22
22
|
</p>
|
|
23
23
|
|
|
24
24
|
<p align="center">
|
|
25
|
-
|
|
26
|
-
|
|
25
|
+
すでに Claude Code へ同じことを繰り返し伝えているはずです。走らせる検査、守るべきルール、
|
|
26
|
+
変更が満たすべき形。<br>
|
|
27
|
+
<b>forge-harness はそれを再利用できるものに変えます。</b> リポジトリの中に住み、自分から発火する
|
|
28
|
+
スキル・ゲート・エージェントとして。<br>
|
|
29
|
+
スキルはあえて汎用のままにし、使ううちにあなたの事例へ合わせて鍛えます。
|
|
30
|
+
<b>同じ形が何度も戻ってきたら、そこで出荷を提案します</b> — 独立したスキルとして、または独立したハーネスとして。
|
|
27
31
|
</p>
|
|
28
32
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## 2分で試せます — この文書を読み切る必要はありません
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git
|
|
39
|
+
claude plugin install -s user fh-meta@forge-harness
|
|
40
|
+
git clone https://github.com/chrono-meta/forge-harness.git ~/projects/forge-harness
|
|
41
|
+
cd ~/projects/forge-harness && claude
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**そして `hi` と入力してください。** 番号付きのメニューが出て、そこからはツールが案内します。
|
|
45
|
+
入口を選んでいくつか答えれば、インストールウィザードまで代わりに実行します。この線から下は
|
|
46
|
+
必要になったときに見る参考資料であり、始める前の宿題ではありません。
|
|
47
|
+
|
|
48
|
+
**増幅するもの**: 試行の回数です。試行錯誤があなたから離れ、並列で回ります。
|
|
49
|
+
**増幅しないもの**: モデルの天井です。ハーネスはモデルを自分の天井まで引き上げるだけです。
|
|
50
|
+
**確かめ方**: 自分の評価を公開します。五つのアイデンティティを正直に、
|
|
51
|
+
[リリース](https://github.com/chrono-meta/forge-harness/releases)ごとに記します。緑でない欄は、
|
|
52
|
+
まだ足りていない実際の実行が何かを名前で述べます。
|
|
53
|
+
|
|
54
|
+
---
|
|
32
55
|
|
|
33
56
|
<p align="center">
|
|
34
|
-
<
|
|
57
|
+
<img src="docs/pillars.svg" alt="FORK - ADAPT - COLLABORATE - EMPOWER" width="680">
|
|
35
58
|
</p>
|
|
36
59
|
|
|
37
60
|
<p align="center">
|
|
38
|
-
<
|
|
61
|
+
<b>品質が梃子であり、速度はその結果です。</b> <i>フォークしてください。名前を変えてください。あなたのものにしてください。</i><br>
|
|
62
|
+
<sub>役に立ったら ⭐ が他の人の発見につながります。</sub>
|
|
39
63
|
</p>
|
|
40
64
|
|
|
41
65
|
<p align="center">
|
package/README.ko.md
CHANGED
|
@@ -18,24 +18,48 @@
|
|
|
18
18
|
</p>
|
|
19
19
|
|
|
20
20
|
<p align="center">
|
|
21
|
-
<
|
|
21
|
+
<b>필요한 걸 시키세요. 같은 요청이 반복되면, 그걸 만들어 드리겠다고 먼저 제안합니다.</b>
|
|
22
22
|
</p>
|
|
23
23
|
|
|
24
24
|
<p align="center">
|
|
25
|
-
|
|
26
|
-
|
|
25
|
+
이미 Claude Code 에 같은 말을 반복하고 계실 겁니다. 돌려야 할 검사, 지켜야 할 규칙,
|
|
26
|
+
변경이 갖춰야 할 모양.<br>
|
|
27
|
+
<b>forge-harness 는 그걸 재사용 가능한 것으로 바꿉니다.</b> 저장소 안에 살면서 스스로 발화하는
|
|
28
|
+
스킬과 게이트와 에이전트로요.<br>
|
|
29
|
+
스킬은 일부러 범용으로 두고 쓰시는 동안 사례에 맞춰 벼립니다.
|
|
30
|
+
<b>같은 모양이 자꾸 돌아오면 그때 출하를 제안합니다</b> — 별도 스킬로, 또는 별도 하네스로.
|
|
27
31
|
</p>
|
|
28
32
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## 2분이면 됩니다 — 이 문서를 다 읽지 않으셔도 됩니다
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git
|
|
39
|
+
claude plugin install -s user fh-meta@forge-harness
|
|
40
|
+
git clone https://github.com/chrono-meta/forge-harness.git ~/projects/forge-harness
|
|
41
|
+
cd ~/projects/forge-harness && claude
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**그리고 `hi` 라고 입력하세요.** 번호가 붙은 메뉴가 뜨고 거기서부터는 도구가 안내합니다.
|
|
45
|
+
문을 고르고 몇 가지에 답하면 설치 마법사까지 대신 돌려 줍니다. 이 아래는 필요할 때 찾아보는
|
|
46
|
+
참고 자료이고, 시작하기 전에 읽어야 하는 숙제가 아닙니다.
|
|
47
|
+
|
|
48
|
+
**증폭하는 것**: 시도 횟수입니다. 시행착오가 당신에게서 떨어져 나가 병렬로 돕니다.
|
|
49
|
+
**증폭하지 않는 것**: 모델의 천장입니다. 하네스는 모델을 자기 천장까지 끌어올릴 뿐입니다.
|
|
50
|
+
**확인하는 법**: 자기 등급을 공개합니다. 다섯 정체성을 정직하게,
|
|
51
|
+
[릴리스](https://github.com/chrono-meta/forge-harness/releases)마다 적습니다. 초록이 아닌 칸은
|
|
52
|
+
아직 없는 실제 실행이 무엇인지 이름으로 말합니다.
|
|
53
|
+
|
|
54
|
+
---
|
|
32
55
|
|
|
33
56
|
<p align="center">
|
|
34
|
-
<
|
|
57
|
+
<img src="docs/pillars.svg" alt="FORK - ADAPT - COLLABORATE - EMPOWER" width="680">
|
|
35
58
|
</p>
|
|
36
59
|
|
|
37
60
|
<p align="center">
|
|
38
|
-
<
|
|
61
|
+
<b>품질이 지렛대이고 속도는 그 결과입니다.</b> <i>포크하세요. 이름을 바꾸세요. 당신의 것으로 만드세요.</i><br>
|
|
62
|
+
<sub>도움이 됐다면 ⭐ 하나가 다른 분이 찾는 데 도움이 됩니다.</sub>
|
|
39
63
|
</p>
|
|
40
64
|
|
|
41
65
|
<p align="center">
|
package/README.md
CHANGED
|
@@ -18,24 +18,48 @@
|
|
|
18
18
|
</p>
|
|
19
19
|
|
|
20
20
|
<p align="center">
|
|
21
|
-
<
|
|
21
|
+
<b>Ask it for things. When the asking repeats, it offers to build you the thing.</b>
|
|
22
22
|
</p>
|
|
23
23
|
|
|
24
24
|
<p align="center">
|
|
25
|
-
|
|
26
|
-
|
|
25
|
+
You already tell Claude Code the same things over and over — the checks to run, the rules to hold,
|
|
26
|
+
the shape a change has to have.<br>
|
|
27
|
+
<b>forge-harness turns that into something reusable</b>: skills, gates and agents that live in your
|
|
28
|
+
repo and fire on their own.<br>
|
|
29
|
+
Its skills stay general on purpose and get shaped to your case as you go.
|
|
30
|
+
<b>When one shape keeps coming back, it offers to ship it</b> as its own skill, or its own harness.
|
|
27
31
|
</p>
|
|
28
32
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## Try it in two minutes — you do not have to read this document
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git
|
|
39
|
+
claude plugin install -s user fh-meta@forge-harness
|
|
40
|
+
git clone https://github.com/chrono-meta/forge-harness.git ~/projects/forge-harness
|
|
41
|
+
cd ~/projects/forge-harness && claude
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**Then type `hi`.** A numbered menu appears and takes it from there — pick a door, answer a couple of
|
|
45
|
+
questions, and it runs the install wizard for you. Everything below this line is reference for when you
|
|
46
|
+
want it, not homework before you start.
|
|
47
|
+
|
|
48
|
+
**What it amplifies** — the number of attempts; trial and error moves off you and runs in parallel.
|
|
49
|
+
**What it does not** — the model's ceiling. A harness lifts a model to its own ceiling, not past it.
|
|
50
|
+
**How you can check** — it grades itself in public, five identities, in every
|
|
51
|
+
[release](https://github.com/chrono-meta/forge-harness/releases); the ones that are not green name
|
|
52
|
+
the real run still missing.
|
|
53
|
+
|
|
54
|
+
---
|
|
32
55
|
|
|
33
56
|
<p align="center">
|
|
34
|
-
<
|
|
57
|
+
<img src="docs/pillars.svg" alt="FORK - ADAPT - COLLABORATE - EMPOWER" width="680">
|
|
35
58
|
</p>
|
|
36
59
|
|
|
37
60
|
<p align="center">
|
|
38
|
-
<
|
|
61
|
+
<b>Quality is the lever; speed is the result.</b> <i>Fork it. Rename it. Make it yours.</i><br>
|
|
62
|
+
<sub>If this is useful, a star helps others find it.</sub>
|
|
39
63
|
</p>
|
|
40
64
|
|
|
41
65
|
<p align="center">
|
|
@@ -59,7 +83,7 @@
|
|
|
59
83
|
|
|
60
84
|
---
|
|
61
85
|
|
|
62
|
-
##
|
|
86
|
+
## Requirements
|
|
63
87
|
|
|
64
88
|
**Prerequisite**: Claude Code CLI — verify with `claude --version`
|
|
65
89
|
|
|
@@ -76,11 +100,8 @@ the single place a new machine can learn it. That is an improvement over nowhere
|
|
|
76
100
|
python3 -m pip install --user pyyaml # verify: python3 -c 'import yaml; print(yaml.__version__)'
|
|
77
101
|
```
|
|
78
102
|
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
machine's own `python3` did not. The gate was never bypassed — it passed, and the pass simply was not
|
|
82
|
-
portable. Every verdict from that gate now prints the interpreter and PyYAML version it used, so a
|
|
83
|
-
green states what produced it instead of leaving the reader to assume.
|
|
103
|
+
Every verdict from that gate prints the interpreter and PyYAML version it used, so a green
|
|
104
|
+
states what produced it.
|
|
84
105
|
|
|
85
106
|
</details>
|
|
86
107
|
|
|
@@ -141,6 +162,36 @@ cd ~/projects/{your-project} && claude
|
|
|
141
162
|
|
|
142
163
|
---
|
|
143
164
|
|
|
165
|
+
## Two version numbers, and they measure different things
|
|
166
|
+
|
|
167
|
+
This repo publishes **two counters**, deliberately. Conflating them is the single most common way to
|
|
168
|
+
misread the project's status, so they are named here rather than only in the canon.
|
|
169
|
+
|
|
170
|
+
| Counter | Where you see it | What it means |
|
|
171
|
+
|---|---|---|
|
|
172
|
+
| **Package version** (currently **2.7.0**) | npm, the plugin manifests, `git tag v2.x` | *what you install.* Ordinary release numbering: fixes → patch, new assets and gate lanes → minor, a capability **class** appearing or the thing being rebuilt → major |
|
|
173
|
+
| **Identity-maturity release** (currently **identity-v0.4.0**) | the GitHub **Releases** page | *how far along the harness is.* `0.x` carries an incomplete-but-honest status **by design**; **the all-green ship is reserved for `identity-v1.0.0`** — every one of the five identities at 🟢, none 🔵/🟡/🔴 |
|
|
174
|
+
|
|
175
|
+
🟥 **A high package number does not mean maturity.** `2.7.0` is not "ahead of" `identity-v0.4.0`; they are not on
|
|
176
|
+
the same scale. The maturity track is deliberately allowed to sit at `0.x` while the package ships and
|
|
177
|
+
improves, because the thing `0.x` refuses to do is **lie** — it says out loud that not every identity has
|
|
178
|
+
cleared its bar yet, and each release names exactly which real run is still missing.
|
|
179
|
+
|
|
180
|
+
⚠️ **Fixed, and the wart is left on the record**: the two counters used to share one `vX.Y.Z` git-tag
|
|
181
|
+
namespace, and only the maturity track had GitHub *Release* objects — so the Releases page showed
|
|
182
|
+
`v0.3.0` as "Latest" while the shipped package was `2.6.0`. Two layers under one name is a defect this
|
|
183
|
+
project keeps finding in its own gates; here it was in its own version numbers. The maturity track now
|
|
184
|
+
carries its own `identity-v*` prefix (first such release: `identity-v0.4.0`, 2026-08-21). 🟥 **Not** by also publishing the package
|
|
185
|
+
track here — that was tried on 2026-08-21 and reverted the same hour: GitHub gives exactly **one**
|
|
186
|
+
"Latest" badge, so two tracks on one page compete for it, and whichever holds it defines what the repo
|
|
187
|
+
says it is. Putting the package number there pushed the maturity claim — the honest core — below it.
|
|
188
|
+
**The Releases page carries the maturity track; what the package shipped is carried by
|
|
189
|
+
[CHANGELOG](plugins/fh-meta/CHANGELOG.md) and the registry.** Existing tags are left alone — renaming them is an irreversible operation on a public surface, and the
|
|
190
|
+
[Destructive-Op gate](knowledge/shared/harness-core/claude_md_gate_details.md) applies to us too.
|
|
191
|
+
|
|
192
|
+
Full rules for what each grade requires:
|
|
193
|
+
[`ship_readiness_gate.md`](knowledge/shared/harness-core/ship_readiness_gate.md).
|
|
194
|
+
|
|
144
195
|
## What it is
|
|
145
196
|
|
|
146
197
|
forge-harness is structured as **two distinct layers**:
|
|
@@ -321,11 +372,26 @@ That is why the column that matters most below is *what it gets*:
|
|
|
321
372
|
|---|---|---|---|
|
|
322
373
|
| **ⓐ Different family** | the diff + the author's framing | the **implementation** is wrong | a reviewer from another model family (`auto-decorrelation`) |
|
|
323
374
|
| **ⓑ Standpoint** | the diff + **the target harness's own canon** | **whether the rule you cited actually says that** | run the diff from that harness's own repo and rules ([`§7`](knowledge/shared/harness-core/field_verdict_crossfamily_gate.md)) |
|
|
324
|
-
| **ⓒ Isolated grounding** | the sentences the author wrote + the tree as it stands now | the **claim** is wrong | someone who did not write it re-measures what it says |
|
|
375
|
+
| **ⓒ Isolated grounding** | the sentences the author wrote — their claims *and* what they **declared before starting** — + the tree as it stands now | the **claim** is wrong · the delta does not match what was declared | someone who did not write it re-measures what it says; for the pre-declaration, a gate that reads the stated success definition back against the delta |
|
|
325
376
|
| **ⓓ Third-party encounter** | the problem + **someone else's codebase** | **is this already solved** · where your change touches someone else's repo | look at the same problem in an unrelated third repo |
|
|
326
377
|
| **ⓔ First real use** | one real target | the **way you are measuring** is wrong — the instrument's instrument | run it once against one real target and check the result by hand |
|
|
327
378
|
| **ⓕ Revert and observe** | the tree with the wiring deleted | the **anchor** is wrong — the check is decorative | delete the thing it guards and confirm *that specific* check goes red |
|
|
328
379
|
|
|
380
|
+
> **ⓒ widened on 2026-08-21, and how it widened is the more useful part.** Every commit marker in
|
|
381
|
+
> this repo has been required since 2026-08-09 to carry the author's own pre-declaration — *what
|
|
382
|
+
> counts as success* and *what I will not do* — written before designing. Measured with a control
|
|
383
|
+
> that day: **nothing read it.** Zero lines of consuming code anywhere, while the sibling fields
|
|
384
|
+
> were checked in 21 places; the gate spec did not even name it. On the real corpus, **37 of 98
|
|
385
|
+
> markers carried no such line at all** — including a panel-reviewed one with 28 lanes and every
|
|
386
|
+
> other field filled. The axes all looked *outward* (the diff, the target repo, prior art, the
|
|
387
|
+
> artifact); none looked at the record's own mandatory field. A slot with no consumer always
|
|
388
|
+
> reports "done", because presence is doing the judging.
|
|
389
|
+
>
|
|
390
|
+
> The fix was not a seventh axis. ⓒ already receives *the sentences the author wrote plus the tree
|
|
391
|
+
> as it stands* — which is, word for word, what a pre-declaration check receives. Tense (declared
|
|
392
|
+
> beforehand vs claimed afterwards) is a **posture**, like adversariality, not an axis. Minting a
|
|
393
|
+
> new one would have repeated the exact error the blind reclassification above found.
|
|
394
|
+
|
|
329
395
|
**You do not run all six every time, and that is the design** — do not multiply them, **choose**:
|
|
330
396
|
|
|
331
397
|
```
|
|
@@ -354,6 +420,37 @@ with tools**, the boundary blurs — an outside judgment held that "the store is
|
|
|
354
420
|
Conversely, "a rule another repo retired long ago" **cannot be fetched by any tool** — there is no reason
|
|
355
421
|
to have access to that project's review history in the first place. That is where ⓓ remains.
|
|
356
422
|
|
|
423
|
+
**Where a rule lives — and why the always-loaded layer does not have to grow forever.**
|
|
424
|
+
|
|
425
|
+
A harness learns by writing rules down. The obvious place is the always-loaded file every session
|
|
426
|
+
reads, and that file only ever gets longer. Left there, the reasoning ends in a corner: *a harness
|
|
427
|
+
that keeps learning keeps getting more expensive to start.*
|
|
428
|
+
|
|
429
|
+
It does not, because a rule has **three possible seats**, and the right one is decided by **when the
|
|
430
|
+
rule has to fire**:
|
|
431
|
+
|
|
432
|
+
| Seat | Fires | Costs | Fits |
|
|
433
|
+
|---|---|---|---|
|
|
434
|
+
| **Always-loaded** | before you act | every session, every turn | rules whose trigger is an *intention* — tone, "don't normalize the unfamiliar", "prove the instrument works here". Nothing can hook an intention, so salience is the only layer |
|
|
435
|
+
| **The gate's own error message** | at the moment you act | **nothing** | rules whose trigger is an *action*. The message that blocks you also teaches the form: `Write, before the design: success = «…». never = «…».` |
|
|
436
|
+
| **The hook** | after you act | nothing | properties of a record — present · typed · attributable · non-vacuous |
|
|
437
|
+
|
|
438
|
+
The middle seat is the one that usually goes unused, and it is free. It is
|
|
439
|
+
[gate-locality](knowledge/shared/harness-core/gate_locality_principle.md) applied to salience: the actor reads it exactly where the
|
|
440
|
+
action happens, so it does not have to be carried all session to be there when needed.
|
|
441
|
+
|
|
442
|
+
🟥 **It is a third layer, not a replacement — and the honest limit is that it only fires on failure.**
|
|
443
|
+
Someone who gets it right never sees it. So mechanizing a rule does **not** shrink the resident layer:
|
|
444
|
+
measured on the very change described above, the machine grew by 480 lines and the always-loaded prose
|
|
445
|
+
by **zero**, and that is correct. The prose has to reach the author *before* they design; the hook
|
|
446
|
+
catches its absence *after*. A backstop cannot substitute for salience that must fire earlier.
|
|
447
|
+
|
|
448
|
+
⚠️ And the threshold that would tell you the resident layer is "too big" is, in this repo, **not
|
|
449
|
+
grounded** — the numbers in our own doctor skill were introduced without a single line justifying the
|
|
450
|
+
cutpoints, and one of them was set to a value the target already exceeded on the day it landed. We are
|
|
451
|
+
re-deriving them rather than trimming toward a number nobody can defend. Cutting resident text toward
|
|
452
|
+
an unjustified target buys fail-open with the savings.
|
|
453
|
+
|
|
357
454
|
> 🟥 **Limits to read before citing this**: the six-axis table is **n=1** (one artifact · one session ·
|
|
358
455
|
> one author). Whether the axes' non-overlap is structural or an accident of that day is **unmeasured**.
|
|
359
456
|
> And when the author's self-scoring was stripped out — 16 findings handed, **with their provenance
|
package/README.zh.md
CHANGED
|
@@ -18,24 +18,44 @@
|
|
|
18
18
|
</p>
|
|
19
19
|
|
|
20
20
|
<p align="center">
|
|
21
|
-
<
|
|
21
|
+
<b>有需要就交给它。同一个请求反复出现时,它会主动提议把那件事本身做出来。</b>
|
|
22
22
|
</p>
|
|
23
23
|
|
|
24
24
|
<p align="center">
|
|
25
|
-
|
|
26
|
-
|
|
25
|
+
你大概已经在对 Claude Code 反复说同样的话:要跑的检查、要守的规则、一次变更该有的样子。<br>
|
|
26
|
+
<b>forge-harness 把这些变成可复用的东西</b>:住在你仓库里、会自己触发的技能、闸门和 agent。<br>
|
|
27
|
+
技能刻意保持通用,在使用过程中按你的场景当场锻造。
|
|
28
|
+
<b>当同一种形状反复回来,它就提议出货</b> —— 作为独立技能,或独立框架。
|
|
27
29
|
</p>
|
|
28
30
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
## 两分钟就能试 —— 你不必读完这份文档
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git
|
|
37
|
+
claude plugin install -s user fh-meta@forge-harness
|
|
38
|
+
git clone https://github.com/chrono-meta/forge-harness.git ~/projects/forge-harness
|
|
39
|
+
cd ~/projects/forge-harness && claude
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
**然后输入 `hi`。** 会出现一个带编号的菜单,之后由工具引导你:选一个入口,回答几个问题,
|
|
43
|
+
它会替你运行安装向导。这条线以下是需要时再查的参考资料,而不是开始前的作业。
|
|
44
|
+
|
|
45
|
+
**它放大什么**:尝试的次数。试错从你身上移开,并行运行。
|
|
46
|
+
**它不放大什么**:模型的天花板。框架只把模型抬到它自己的天花板,不会更高。
|
|
47
|
+
**如何验证**:它公开自己的评级。五个身份,如实记录在每一次
|
|
48
|
+
[发布](https://github.com/chrono-meta/forge-harness/releases)中。未变绿的那些,会指名说出还缺哪一次真实运行。
|
|
49
|
+
|
|
50
|
+
---
|
|
32
51
|
|
|
33
52
|
<p align="center">
|
|
34
|
-
<
|
|
53
|
+
<img src="docs/pillars.svg" alt="FORK - ADAPT - COLLABORATE - EMPOWER" width="680">
|
|
35
54
|
</p>
|
|
36
55
|
|
|
37
56
|
<p align="center">
|
|
38
|
-
<
|
|
57
|
+
<b>质量是杠杆,速度是结果。</b> <i>Fork 它。改名。让它成为你的。</i><br>
|
|
58
|
+
<sub>如果这对你有用,⭐ 一下能帮助更多人发现它。</sub>
|
|
39
59
|
</p>
|
|
40
60
|
|
|
41
61
|
<p align="center">
|
|
@@ -469,11 +469,25 @@ own. `tier2b` is the honest reachable rung for that pairing; do not inflate a `t
|
|
|
469
469
|
`tier3`, and do not undersell it to `tier2` either — it is a distinct, real, if operator-correlated,
|
|
470
470
|
data point.
|
|
471
471
|
|
|
472
|
-
**Mechanization status —
|
|
473
|
-
|
|
474
|
-
|
|
475
|
-
|
|
476
|
-
|
|
472
|
+
**Mechanization status — 🟥 CORRECTED 2026-08-20.** This paragraph used to say `standpoint:` was
|
|
473
|
+
"prose-only today" with "**no value-enum validation and no fixture suite**". **Both halves are false**
|
|
474
|
+
and had been for a while: `validate_standpoint_leg()` lives in `templates/.git-hooks/pre-commit` and
|
|
475
|
+
`scripts/test_marker_standpoint_lanes.sh` is wired through `scripts/selfcheck.sh`. Measured by varying
|
|
476
|
+
one variable at a time — `banana(qasp)` → **blocked (enum)** · `tier2` without parens → **blocked
|
|
477
|
+
(enum)** · `tier2(qasp)` with no execution grounds → **passes with a warning** · with grounds →
|
|
478
|
+
**passes**.
|
|
479
|
+
|
|
480
|
+
**What is actually true, stated at the right width**: the enum IS closed and enforced; the `tier2`+
|
|
481
|
+
execution grounds are **advisory** (a thin `tier2` records and warns, it does not block); and nothing
|
|
482
|
+
checks whether the recorded value is *true*. The old sentence collapsed all three into "no validation",
|
|
483
|
+
which suppresses use of a control that exists — the quietest kind of drift, because it reads as
|
|
484
|
+
honest modesty.
|
|
485
|
+
|
|
486
|
+
🟥 **This same false claim stood in FIVE places** (`CLAUDE.md` ×2, `.claude/rules/fh_4axis_gate.md`,
|
|
487
|
+
`AGENTS.md`, here). The first repair fixed one; the second fixed three and still missed this file and
|
|
488
|
+
over-stated the fix. Both misses were surfaced by a **different-family reviewer**, not by a lane and
|
|
489
|
+
not by the author — there is still no check that compares a rule's self-description against the
|
|
490
|
+
machinery it describes.
|
|
477
491
|
|
|
478
492
|
🟥 **Two sentences that stood here were STALE and are corrected (2026-08-17, re-measured — a
|
|
479
493
|
cross-family reviewer flagged the second, the first fell out of checking it).** They read
|