@chrono-meta/fh-gate 2.6.0 → 2.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/.claude/rules/fh_4axis_gate.md +26 -3
  2. package/.claude-plugin/marketplace.json +2 -2
  3. package/AGENTS.md +28 -3
  4. package/CLAUDE.md +146 -168
  5. package/README.ja.md +144 -24
  6. package/README.ko.md +135 -22
  7. package/README.md +204 -27
  8. package/README.zh.md +126 -21
  9. package/docs/ETHOS.md +10 -3
  10. package/knowledge/shared/dialogue/ai_dialogue_playbook.md +131 -0
  11. package/knowledge/shared/harness-core/claude_md_gate_details.md +90 -0
  12. package/knowledge/shared/harness-core/dispatch_conditional_prohibition.md +75 -0
  13. package/knowledge/shared/harness-core/fh_three_layer_canon.md +77 -3
  14. package/knowledge/shared/harness-core/field_verdict_crossfamily_gate.md +92 -5
  15. package/knowledge/shared/harness-core/harness_incubator_doctrine.md +12 -2
  16. package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +32 -0
  17. package/knowledge/shared/harness-core/ship_readiness_gate.md +77 -8
  18. package/knowledge/shared/learnings/subagent_invocations_log.yaml +196 -0
  19. package/knowledge/shared/rules/knowledge_layer_seam.md +1 -1
  20. package/knowledge/shared/rules/multi_session_close_protocol.md +7 -3
  21. package/package.json +15 -2
  22. package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
  23. package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
  24. package/plugins/fh-meta/CHANGELOG.md +157 -0
  25. package/plugins/fh-meta/agents/persona-innovator.md +170 -0
  26. package/plugins/fh-meta/skills/cross-ecosystem-synergy-detection/SKILL.md +37 -0
  27. package/plugins/fh-meta/skills/public-surface-audit/SKILL_detail.md +16 -1
  28. package/plugins/fh-meta/skills/steel-quench/SKILL.md +97 -0
  29. package/scripts/adapters/peer_resolve.sh +58 -6
  30. package/scripts/branch_claim.sh +3 -1
  31. package/scripts/cluster_capability_scan.sh +42 -13
  32. package/scripts/digest_landing_check.sh +142 -2
  33. package/scripts/fh_hub_identity.sh +83 -0
  34. package/scripts/fh_session_load.sh +53 -5
  35. package/scripts/fh_track_resolve.sh +114 -0
  36. package/scripts/field_canon_preload.sh +50 -5
  37. package/scripts/listing_watch.sh +187 -0
  38. package/scripts/package_coverage_check.sh +8 -0
  39. package/scripts/prior_art_prompt.sh +168 -0
  40. package/scripts/psa_scan_lib.sh +201 -15
  41. package/scripts/residency_admission_check.sh +204 -0
  42. package/scripts/selfcheck.sh +103 -0
  43. package/scripts/test_adapter_lanes.sh +67 -2
  44. package/scripts/test_heavy_classifier_lanes.sh +144 -0
  45. package/scripts/test_listing_watch_lanes.sh +162 -0
  46. package/scripts/test_marker_defense_lanes.sh +152 -0
  47. package/scripts/test_marker_soul_check_lanes.sh +211 -0
  48. package/scripts/test_prior_art_prompt_lanes.sh +128 -0
  49. package/scripts/test_psa_singlefile_lanes.sh +351 -1
  50. package/scripts/test_residency_admission_lanes.sh +60 -0
  51. package/scripts/test_track_resolve_lanes.sh +158 -0
  52. package/scripts/test_wizard_snippet_merge_lanes.sh +40 -0
  53. package/templates/.git-hooks/pre-commit +400 -4
  54. package/templates/.git-hooks/pre-push +17 -2
  55. package/templates/settings.PriorArt.snippet.json +44 -0
package/README.md CHANGED
@@ -18,24 +18,69 @@
18
18
  </p>
19
19
 
20
20
  <p align="center">
21
- <sub>If this is useful, a helps others find it.</sub>
21
+ <b>A meta-harness with the quality gates built in.</b>
22
22
  </p>
23
23
 
24
24
  <p align="center">
25
- <b>Forge your Claude Code projects pass them through, they come out faster.</b><br>
26
- A practitioner's <b>meta-harness</b> — the galaxy your project harnesses live in.<br>It raises each project's <b>floor</b> (harness-ify the setup) and <b>ceiling</b> (accelerate the work), then compounds the gains across your whole portfolio.
25
+ Projects, skills, harnesses building them, checking them, speeding them up: you ask for it here.<br>
26
+ It does not simply hand the result back. It puts the work past several checks that fail in
27
+ <i>different</i> ways first.<br>
28
+ <b>And when the same request keeps coming back, it builds you the harness that does it for you.</b>
27
29
  </p>
28
30
 
29
31
  <p align="center">
30
- <b>Quality is the lever; speed is the result.</b> Every change earns its way through the gates —<br>adversarial · phantom · regression — and <i>that</i> is what makes the next change faster.
32
+ You already tell Claude Code the same things over and over the checks to run, the rules to hold,
33
+ the shape a change has to have. That is the part that becomes reusable, and it keeps its general
34
+ form on purpose so it can be shaped to your case as you go.<br>
35
+ <sub>What grows is the number of attempts: the trial and error comes off you and runs in parallel.</sub>
31
36
  </p>
32
37
 
38
+ ---
39
+
40
+ ## Try it in two minutes — you do not have to read this document
41
+
42
+ ```bash
43
+ claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git
44
+ claude plugin install -s user fh-meta@forge-harness
45
+ git clone https://github.com/chrono-meta/forge-harness.git ~/projects/forge-harness
46
+ cd ~/projects/forge-harness && claude
47
+ ```
48
+
49
+ **Then type `hi`** — or a greeting in whatever language you actually think in: `안녕`, `こんにちは`,
50
+ `你好`, `hola`, `bonjour`. **Any of them opens the menu**, and it will try to answer in the language you
51
+ used. A numbered menu appears and takes it from there — pick a door, answer a couple of questions, and
52
+ it runs the install wizard for you.
53
+
54
+ <sub>🟥 <b>Honest about that last part</b>: matching your language is a prose rule with nothing
55
+ mechanical behind it, so it does not always hold. Measured blind at the floor tier on 2026-08-21, on a
56
+ <b>clean clone like the one you just made</b>: Chinese and Korean greetings both came back fully
57
+ translated, door labels included. On the maintainer's own machine — which pins a default language —
58
+ Chinese landed only 1 time in 5, which is why this note exists at all. What is still shaky is whether
59
+ the menu fires: one greeting variant produced no menu. If it answers in the wrong language, or skips
60
+ the menu, just say so — it will switch. Written up in <code>CLAUDE.md</code> §Voice/Tone rather than
61
+ smoothed over.</sub>
62
+
63
+ Everything below this line is reference for when you want it, not homework before you start.
64
+
65
+ - **What it amplifies** — the number of attempts. Trial and error moves off you and runs in parallel.
66
+
67
+ - **What it does not** — the model's ceiling. A harness lifts a model to its own ceiling, not past it.
68
+
69
+ - **How you can check** — it grades itself in public. It names **five things it claims to be**
70
+ (harness cluster · project incubator · governance gate · frontier → org propagation · amplifier) and grades each
71
+ one honestly in every [release](https://github.com/chrono-meta/forge-harness/releases). Any square
72
+ that is not green names the real run that is still missing.
73
+ <sub>Each of the five is spelled out further down, under «The five identities».</sub>
74
+
75
+ ---
76
+
33
77
  <p align="center">
34
- <i>Fork it. Rename it. Make it yours.</i>
78
+ <img src="docs/pillars.svg" alt="HARNESS - FORGE - ACCELERATE - COMPOUND" width="680">
35
79
  </p>
36
80
 
37
81
  <p align="center">
38
- <img src="docs/pillars.svg" alt="FORK · ADAPT · COLLABORATE · EMPOWER" width="680">
82
+ <b>Quality is the lever; speed is the result.</b><br>
83
+ <sub>If this is useful, a star helps others find it.</sub>
39
84
  </p>
40
85
 
41
86
  <p align="center">
@@ -59,7 +104,7 @@
59
104
 
60
105
  ---
61
106
 
62
- ## Get started in 2 minutes
107
+ ## Requirements
63
108
 
64
109
  **Prerequisite**: Claude Code CLI — verify with `claude --version`
65
110
 
@@ -76,11 +121,8 @@ the single place a new machine can learn it. That is an improvement over nowhere
76
121
  python3 -m pip install --user pyyaml # verify: python3 -c 'import yaml; print(yaml.__version__)'
77
122
  ```
78
123
 
79
- Why this is called out rather than left implicit: a release once shipped green from a session whose
80
- `python3` happened to resolve to an **unrelated project's virtualenv** that had PyYAML, while the
81
- machine's own `python3` did not. The gate was never bypassed — it passed, and the pass simply was not
82
- portable. Every verdict from that gate now prints the interpreter and PyYAML version it used, so a
83
- green states what produced it instead of leaving the reader to assume.
124
+ Every verdict from that gate prints the interpreter and PyYAML version it used, so a green
125
+ states what produced it.
84
126
 
85
127
  </details>
86
128
 
@@ -97,7 +139,9 @@ cd ~/projects/forge-harness
97
139
  claude
98
140
  ```
99
141
 
100
- > ✅ Then **type a greeting ("hi")** — the 🐿️ door menu appears on a typed greeting, not on launch alone.
142
+ > ✅ Then **type a greeting** — `hi`, `안녕`, `こんにちは`, `你好`, whatever you actually think in; it
143
+ > tries to reply in that language (see the note above — it is right more often than not, not always).
144
+ > The 🐿️ door menu appears on a *typed* greeting, not on launch alone.
101
145
  > Say **"Connect a project"** → hub scans `../`, finds `.git` directories, creates `tracks/{project}/`.
102
146
  > For full initial setup (hooks · gates · baseline — each item individually approved, declining is
103
147
  > respected and recorded), ask for **`/install-wizard`**.
@@ -106,8 +150,8 @@ claude
106
150
 
107
151
  **Your first 15 minutes** — what success looks like, and what to do with it:
108
152
 
109
- 1. You'll know setup worked when a greeting ("hi") shows the 🐿️ door menu, and "Connect a project"
110
- creates `tracks/{your-project}/`.
153
+ 1. You'll know setup worked when a greeting in any language — shows the 🐿️ door menu, and
154
+ "Connect a project" creates `tracks/{your-project}/`.
111
155
  2. Then grab an immediate win in the same session: say **"accelerate this project"** (ranked plan of
112
156
  skills/plugins worth wiring, install-gated) or **"run /context-doctor"** (token-waste scan).
113
157
  3. One honest note: FH's core payoff is **compounding** — session records, harvested learnings,
@@ -141,6 +185,36 @@ cd ~/projects/{your-project} && claude
141
185
 
142
186
  ---
143
187
 
188
+ ## Two version numbers, and they measure different things
189
+
190
+ This repo publishes **two counters**, deliberately. Conflating them is the single most common way to
191
+ misread the project's status, so they are named here rather than only in the canon.
192
+
193
+ | Counter | Where you see it | What it means |
194
+ |---|---|---|
195
+ | **Package version** (currently **2.7.0**) | npm, the plugin manifests, `git tag v2.x` | *what you install.* Ordinary release numbering: fixes → patch, new assets and gate lanes → minor, a capability **class** appearing or the thing being rebuilt → major |
196
+ | **Identity-maturity release** (currently **identity-v0.4.0**) | the GitHub **Releases** page | *how far along the harness is.* `0.x` carries an incomplete-but-honest status **by design**; **the all-green ship is reserved for `identity-v1.0.0`** — every one of the five identities at 🟢, none 🔵/🟡/🔴 |
197
+
198
+ 🟥 **A high package number does not mean maturity.** `2.7.0` is not "ahead of" `identity-v0.4.0`; they are not on
199
+ the same scale. The maturity track is deliberately allowed to sit at `0.x` while the package ships and
200
+ improves, because the thing `0.x` refuses to do is **lie** — it says out loud that not every identity has
201
+ cleared its bar yet, and each release names exactly which real run is still missing.
202
+
203
+ ⚠️ **Fixed, and the wart is left on the record**: the two counters used to share one `vX.Y.Z` git-tag
204
+ namespace, and only the maturity track had GitHub *Release* objects — so the Releases page showed
205
+ `v0.3.0` as "Latest" while the shipped package was `2.6.0`. Two layers under one name is a defect this
206
+ project keeps finding in its own gates; here it was in its own version numbers. The maturity track now
207
+ carries its own `identity-v*` prefix (first such release: `identity-v0.4.0`, 2026-08-21). 🟥 **Not** by also publishing the package
208
+ track here — that was tried on 2026-08-21 and reverted the same hour: GitHub gives exactly **one**
209
+ "Latest" badge, so two tracks on one page compete for it, and whichever holds it defines what the repo
210
+ says it is. Putting the package number there pushed the maturity claim — the honest core — below it.
211
+ **The Releases page carries the maturity track; what the package shipped is carried by
212
+ [CHANGELOG](plugins/fh-meta/CHANGELOG.md) and the registry.** Existing tags are left alone — renaming them is an irreversible operation on a public surface, and the
213
+ [Destructive-Op gate](knowledge/shared/harness-core/claude_md_gate_details.md) applies to us too.
214
+
215
+ Full rules for what each grade requires:
216
+ [`ship_readiness_gate.md`](knowledge/shared/harness-core/ship_readiness_gate.md).
217
+
144
218
  ## What it is
145
219
 
146
220
  forge-harness is structured as **two distinct layers**:
@@ -296,11 +370,25 @@ harness is actually used.
296
370
 
297
371
  **The three-stage process** — this is an *order of investment*, not a menu:
298
372
 
373
+ ```mermaid
374
+ flowchart LR
375
+ S["① Circuit<br/>before design"]
376
+ P["② Parallel decorrelation<br/>in the middle"]
377
+ B["③ Burn it down<br/>on six axes"]
378
+ A(["⟹ It accelerates"])
379
+ S --> P --> B --> A
380
+ style A fill:#0f766e,stroke:#0f766e,color:#fff
381
+ ```
382
+
383
+ **Speed is the arrow at the end, not a fourth box.** And ② is **two dials, turned separately** —
384
+ *decorrelation* (reacts to blind-spot risk: a different model family ⓐ, a different standpoint ⓑ)
385
+ and *parallelism* (reacts to surface size: can one context hold it?). Do not multiply them, choose.
386
+
299
387
  ```
300
388
  ① Circuit before design the judgment circuit goes in FIRST — success · leaning · out-of-scope ·
301
389
  never-do — not written up afterwards as a record of what you did
302
390
 
303
- Decorrelate in the split the work into checks that fail DIFFERENTLY and run them at once.
391
+ Parallel decorrelation split the work into checks that fail DIFFERENTLY and run them at once.
304
392
  middle, to accelerate Choose which differences matter — a second reviewer of the same kind is
305
393
  not decorrelation, it is the same blind spot twice. Parallelism has no
306
394
  direction of its own; the judgment circuit from ① is what picks.
@@ -311,6 +399,12 @@ harness is actually used.
311
399
  riding it does not make that axis see what it cannot see
312
400
  ```
313
401
 
402
+ > 📖 **From here down is for whoever wants to go further — it is not needed to start using this.**
403
+ > If you came to install it and get going, the two-minute section at the top is the whole job; you can
404
+ > stop here and come back when a check surprises you. What follows is the reasoning behind the gates,
405
+ > written for someone already running the hub. Unfamiliar words →
406
+ > [`GLOSSARY.md`](knowledge/shared/GLOSSARY.md).
407
+
314
408
  **The six verification axes** — where "we reviewed it" usually turns out to mean only the first of them.
315
409
 
316
410
  🟥 **Axes are not divided by *how adversarial* they are. They are divided by *what they were given*.**
@@ -321,11 +415,42 @@ That is why the column that matters most below is *what it gets*:
321
415
  |---|---|---|---|
322
416
  | **ⓐ Different family** | the diff + the author's framing | the **implementation** is wrong | a reviewer from another model family (`auto-decorrelation`) |
323
417
  | **ⓑ Standpoint** | the diff + **the target harness's own canon** | **whether the rule you cited actually says that** | run the diff from that harness's own repo and rules ([`§7`](knowledge/shared/harness-core/field_verdict_crossfamily_gate.md)) |
324
- | **ⓒ Isolated grounding** | the sentences the author wrote + the tree as it stands now | the **claim** is wrong | someone who did not write it re-measures what it says |
418
+ | **ⓒ Isolated grounding** | the sentences the author wrote — their claims *and* what they **declared before starting** — + the tree as it stands now | the **claim** is wrong · the delta does not match what was declared | someone who did not write it re-measures what it says; for the pre-declaration, a gate that reads the stated success definition back against the delta |
325
419
  | **ⓓ Third-party encounter** | the problem + **someone else's codebase** | **is this already solved** · where your change touches someone else's repo | look at the same problem in an unrelated third repo |
326
420
  | **ⓔ First real use** | one real target | the **way you are measuring** is wrong — the instrument's instrument | run it once against one real target and check the result by hand |
327
421
  | **ⓕ Revert and observe** | the tree with the wiring deleted | the **anchor** is wrong — the check is decorative | delete the thing it guards and confirm *that specific* check goes red |
328
422
 
423
+ > **ⓒ widened on 2026-08-21, and how it widened is the more useful part.** Every commit marker in
424
+ > this repo has been required since 2026-08-09 to carry the author's own pre-declaration — *what
425
+ > counts as success* and *what I will not do* — written before designing. Measured with a control
426
+ > that day: **nothing read it.** Zero lines of consuming code anywhere, while the sibling fields
427
+ > were checked in 21 places; the gate spec did not even name it. On the real corpus, **37 of 98
428
+ > markers carried no such line at all** — including a panel-reviewed one with 28 lanes and every
429
+ > other field filled. The axes all looked *outward* (the diff, the target repo, prior art, the
430
+ > artifact); none looked at the record's own mandatory field. A slot with no consumer always
431
+ > reports "done", because presence is doing the judging.
432
+ >
433
+ > The fix was not a seventh axis. ⓒ already receives *the sentences the author wrote plus the tree
434
+ > as it stands* — which is, word for word, what a pre-declaration check receives. Tense (declared
435
+ > beforehand vs claimed afterwards) is a **posture**, like adversariality, not an axis. Minting a
436
+ > new one would have repeated the exact error the blind reclassification above found.
437
+
438
+ > **A peer session counts on ⓑ and ⓓ — decided 2026-08-21, and *which* peer you ask is the whole
439
+ > trick.** A parallel session of this same harness, running hot on a different branch of the work, is
440
+ > not a copy of you. At the point it got hot it has genuinely grown a second face: a real standpoint
441
+ > (ⓑ) and a real someone-else's-codebase (ⓓ). So peer findings are recorded on those two axes — no
442
+ > seventh axis was minted, and the marker's `axes-run` alphabet did not change.
443
+ >
444
+ > This is why the ⓓ row above says *someone else's codebase* rather than *another company's repo*:
445
+ > the boundary that matters is **whose working context produced the judgment**, not whose GitHub org
446
+ > owns the files. A peer hot on a different branch is across that boundary; a subagent you spawned
447
+ > from this context is not, however different the repo it reads.
448
+ >
449
+ > 🟥 The corollary is the part that bites. **Ask the peer about the axis it actually got hot on.**
450
+ > Anywhere else it wears your face and decorrelates nothing — same input, same blind spot. And a
451
+ > subagent cannot stand in for it, nor can re-reading your own work: the second face comes from that
452
+ > session having *done* different work, and a prompt cannot manufacture it.
453
+
329
454
  **You do not run all six every time, and that is the design** — do not multiply them, **choose**:
330
455
 
331
456
  ```
@@ -348,12 +473,44 @@ itself.
348
473
  **Why this is not superseded by base-model advances** — an axis is defined by its **input**, not by the
349
474
  *reviewer's ability*. A stronger model still **cannot see information it was not given.** Scaffolding
350
475
  sheds as models improve, but **input-boundary decorrelation does not**, and a single author cannot, by
351
- definition, step outside their own input. 🟥 Honest edge: if the agent **fetches more input by itself
352
- with tools**, the boundary blurs — an outside judgment held that "the store is never used in full" and
476
+ definition, step outside their own input.
477
+
478
+ 🟥 **Honest edge**: if the agent **fetches more input by itself with tools**, the boundary blurs — an outside judgment held that "the store is never used in full" and
353
479
  "the swallowed exception" are catchable by ⓐ and ⓒ as well, since those reviewers grep for themselves.
354
480
  Conversely, "a rule another repo retired long ago" **cannot be fetched by any tool** — there is no reason
355
481
  to have access to that project's review history in the first place. That is where ⓓ remains.
356
482
 
483
+ **Where a rule lives — and why the always-loaded layer does not have to grow forever.**
484
+
485
+ A harness learns by writing rules down. The obvious place is the always-loaded file every session
486
+ reads, and that file only ever gets longer. Left there, the reasoning ends in a corner: *a harness
487
+ that keeps learning keeps getting more expensive to start.*
488
+
489
+ It does not, because a rule has **three possible seats**, and the right one is decided by **when the
490
+ rule has to fire**:
491
+
492
+ | Seat | Fires | Costs | Fits |
493
+ |---|---|---|---|
494
+ | **Always-loaded** | before you act | every session, every turn | rules whose trigger is an *intention* — tone, "don't normalize the unfamiliar", "prove the instrument works here". Nothing can hook an intention, so salience is the only layer |
495
+ | **The gate's own error message** | at the moment you act | **nothing** | rules whose trigger is an *action*. The message that blocks you also teaches the form: `Write, before the design: success = «…». never = «…».` |
496
+ | **The hook** | after you act | nothing | properties of a record — present · typed · attributable · non-vacuous |
497
+
498
+ The middle seat is the one that usually goes unused, and it is free. It is
499
+ [gate-locality](knowledge/shared/harness-core/gate_locality_principle.md) applied to salience: the actor reads it exactly where the
500
+ action happens, so it does not have to be carried all session to be there when needed.
501
+
502
+ 🟥 **It is a third layer, not a replacement — and the honest limit is that it only fires on failure.**
503
+ Someone who gets it right never sees it. So mechanizing a rule does **not** shrink the resident layer:
504
+ measured on the very change described above, the machine grew by 480 lines and the always-loaded prose
505
+ by **zero**, and that is correct. The prose has to reach the author *before* they design; the hook
506
+ catches its absence *after*. A backstop cannot substitute for salience that must fire earlier.
507
+
508
+ ⚠️ And the threshold that would tell you the resident layer is "too big" is, in this repo, **not
509
+ grounded** — the numbers in our own doctor skill were introduced without a single line justifying the
510
+ cutpoints, and one of them was set to a value the target already exceeded on the day it landed. We are
511
+ re-deriving them rather than trimming toward a number nobody can defend. Cutting resident text toward
512
+ an unjustified target buys fail-open with the savings.
513
+
357
514
  > 🟥 **Limits to read before citing this**: the six-axis table is **n=1** (one artifact · one session ·
358
515
  > one author). Whether the axes' non-overlap is structural or an accident of that day is **unmeasured**.
359
516
  > And when the author's self-scoring was stripped out — 16 findings handed, **with their provenance
@@ -457,22 +614,42 @@ Full spec: [`fh_integration_contract.md`](knowledge/shared/harness-core/fh_integ
457
614
 
458
615
  ---
459
616
 
460
- ## The forge
617
+ ## The forge — where the name comes from
461
618
 
462
- forge-harness treats a project like steel and the metaphor is literal, not decoration. Work is shaped,
619
+ forge-harness treats a project like steel, and the metaphor is literal, not decoration. Work is shaped,
463
620
  hardened by attack, and only then does it ship faster, for having survived.
464
621
 
465
- | Movement | What happens | The commands |
622
+ > 🟥 **This is the name's origin, not a procedure — do not count it.** The smith's words below are a
623
+ > vocabulary, not a stage list. They are a **different layer** from the three-stage process, the four
624
+ > engines, the four-axis commit gate and the six verification axes above; nothing here lines up with a
625
+ > number in any of those, and reading it as a fifth numbered set is the one mistake this section can
626
+ > cause.
627
+
628
+ **Three at the anvil:**
629
+
630
+ | The smith's word | What it means here | The commands |
466
631
  |---|---|---|
467
632
  | **Forge** | shape the raw project into a harness — raise its floor | `install-wizard`, "harness-ify this project" |
468
633
  | **Quench** | harden it by attack — the cold pass leaves standing only what is sound | `steel-quench` · `phantom-quench` |
469
634
  | **Temper** | take the brittleness back out of the hardened asset | `steel-quench` Wave-T · `templates/temper_check.sh` |
470
- | → **Accelerate** | a blade that survived the forge cuts faster | `goal-quench` — *Pass → Accelerate* |
471
635
 
472
- All four movements ship. Temper was named before it was built — deliberately (see
473
- [`ETHOS.md`](docs/ETHOS.md#the-forge)) — and shipped once measurement runs validated it. Around the forge,
474
- two more signatures keep it running: `harvest-loop` (each session's lessons become permanent skills) and
475
- `agent-composer` (orchestrate the dispatch). The other skills wait until you need them — full list below.
636
+ ```mermaid
637
+ flowchart LR
638
+ F["Forge"] --> Q["Quench"] --> T["Temper"] --> A(["⟹ Accelerate"])
639
+ style A fill:#0f766e,stroke:#0f766e,color:#fff
640
+ ```
641
+
642
+ **⟹ And then it accelerates.** A blade that survived the forge cuts faster — `goal-quench`,
643
+ *Pass → Accelerate*. Speed is what the three above **produce**; it is not a fourth thing you do.
644
+ That is the same sentence as the tagline at the top of this page, said in the smith's words:
645
+ quality is the lever, speed is the result.
646
+
647
+ Every command named above ships today. Temper was named before it was built — deliberately (see
648
+ [`ETHOS.md`](docs/ETHOS.md#the-forge)) — and shipped once measurement runs validated it.
649
+
650
+ Around the forge, two more signatures keep it running: `harvest-loop` (each session's lessons become
651
+ permanent skills) and `agent-composer` (orchestrate the dispatch). The other skills wait until you need
652
+ them — full list below.
476
653
 
477
654
  ## 40 skills · 8 agents
478
655
 
package/README.zh.md CHANGED
@@ -18,24 +18,65 @@
18
18
  </p>
19
19
 
20
20
  <p align="center">
21
- <sub>如果这对你有用,⭐ 一下能帮助更多人发现它。</sub>
21
+ <b>这是一个把质量门禁内置进去的元框架(meta-harness)。</b>
22
22
  </p>
23
23
 
24
24
  <p align="center">
25
- <b>锻造你的 Claude Code 项目 —— 让它通过,它会更快出炉。</b><br>
26
- 一个实践者的 <b>元框架 (meta-harness)</b> —— 你的项目框架们所栖居的星系。<br>它抬高每个项目的 <b>下限 (floor)</b>(把设置框架化)和 <b>上限 (ceiling)</b>(加速工作),再把这些收益在你的整个项目组合中复利累积。
25
+ 项目、技能、框架 —— 造出来、验过、再加速:这些都在这里交代。<br>
26
+ 它不会就这么把结果递还给你,而是先让这份工作穿过好几道会以 <i>不同方式</i> 失败的检查。<br>
27
+ <b>而当同一个请求反复回来,它就替你造出那个专门干这件事的框架。</b>
27
28
  </p>
28
29
 
29
30
  <p align="center">
30
- <b>品质是杠杆,速度是结果。</b> 每一次变更都要挣得通过门禁的资格 ——<br>对抗 (adversarial) · 幽灵 (phantom) · 回归 (regression) —— 而 <i>正是这一点</i>让下一次变更更快。
31
+ 你大概已经在对 Claude Code 反复说同样的话:要跑的检查、要守的规则、一次变更该有的样子。
32
+ 变得可复用的正是这一部分,而它刻意保持通用的形态,好在使用过程中按你的场景当场锻造。<br>
33
+ <sub>长起来的是尝试的次数:试错从你身上剥离,并行地跑。</sub>
31
34
  </p>
32
35
 
36
+ ---
37
+
38
+ ## 两分钟就能试 —— 你不必读完这份文档
39
+
40
+ ```bash
41
+ claude plugin marketplace add https://github.com/chrono-meta/forge-harness.git
42
+ claude plugin install -s user fh-meta@forge-harness
43
+ git clone https://github.com/chrono-meta/forge-harness.git ~/projects/forge-harness
44
+ cd ~/projects/forge-harness && claude
45
+ ```
46
+
47
+ **然后输入 `你好`** —— 或者用你真正在用的那门语言打招呼:`hi`、`안녕`、`こんにちは`、`hola`、
48
+ `bonjour`。**不管你用哪种语言问候,菜单都会打开**,并且它会尝试用那种语言回复。会出现一个带编号
49
+ 的菜单,之后由工具引导你:选一个入口,回答几个问题,它会替你运行安装向导。
50
+
51
+ <sub>🟥 <b>关于最后这一点,说实话</b>:语言对齐是一条 <b>背后没有任何机械兜底</b> 的散文规则,所以
52
+ 它并不总是成立。2026-08-21 在 floor 档做的盲测,跑在 <b>一个跟你刚建的那份一样的干净克隆</b> 上:
53
+ 中文和韩文的问候都完整地用同一种语言回来了,连门的标签都翻了。而在维护者自己那台
54
+ <b>把默认语言钉死了的机器</b> 上,中文 <b>5 次里只中了 1 次</b> —— 这条注脚之所以存在,就是因为它。
55
+ 现在仍然不稳的是 <b>菜单会不会弹出来</b>:有一种问候写法就没能唤出菜单。如果它回错了语言、或者
56
+ 没给你菜单,直接说一声,它就会切过来。这条残留被如实写在 <code>CLAUDE.md</code> §Voice/Tone 里,
57
+ 而不是被抹平。</sub>
58
+
59
+ 这条线以下是需要时再查的参考资料,而不是开始前的作业。
60
+
61
+ - **它放大什么** —— 尝试的次数。试错从你身上移开,并行运行。
62
+
63
+ - **它不放大什么** —— 模型的天花板。框架只把模型抬到它自己的天花板,不会再往上推。
64
+
65
+ - **如何验证** —— 它公开自己的评级。它把 **自己声称是什么** 名列为五项(框架集群 · 项目孵化器 ·
66
+ 治理门禁 · 前沿 → 组织传导 · 放大器),并在每一次
67
+ [发布](https://github.com/chrono-meta/forge-harness/releases)里逐项如实打分。未变绿的那些,
68
+ 会指名说出还缺哪一次真实运行。<br>
69
+ <sub>这五项各自是什么,在下面的《五重身份》一节里展开。</sub>
70
+
71
+ ---
72
+
33
73
  <p align="center">
34
- <i>Fork 它。改名。让它成为你的。</i>
74
+ <img src="docs/pillars.svg" alt="HARNESS - FORGE - ACCELERATE - COMPOUND" width="680">
35
75
  </p>
36
76
 
37
77
  <p align="center">
38
- <img src="docs/pillars.svg" alt="FORK · ADAPT · COLLABORATE · EMPOWER" width="680">
78
+ <b>质量是杠杆,速度是结果。</b><br>
79
+ <sub>如果这对你有用,⭐ 一下能帮助更多人发现它。</sub>
39
80
  </p>
40
81
 
41
82
  <p align="center">
@@ -95,7 +136,9 @@ cd ~/projects/forge-harness
95
136
  claude
96
137
  ```
97
138
 
98
- > ✅ 然后 **打一句招呼("hi")** —— 🐿️ 门菜单是在你打出招呼时出现的,光是启动不会出现。
139
+ > ✅ 然后 **打一句招呼** —— `你好`、`hi`、`안녕`、`こんにちは`,用你习惯的语言就行,它会尝试用
140
+ > 那种语言回复(见上面的注记 —— 多数时候对,但不是每次)。
141
+ > 🐿️ 门菜单是在你 *打出* 招呼时出现的,光是启动不会出现。
99
142
  > 说 **"连接一个项目"** → 中枢扫描 `../`,找到 `.git` 目录,创建 `tracks/{project}/`。
100
143
  > 想做完整的初始设置(hooks · 门禁 · 基线 —— 每一项单独批准,拒绝会被尊重并记录),
101
144
  > 请要 **`/install-wizard`**。
@@ -104,8 +147,8 @@ claude
104
147
 
105
148
  **你的头 15 分钟** —— 成功长什么样,以及拿它做什么:
106
149
 
107
- 1. 当一句招呼("hi")能让 🐿️ 门菜单出现、而"连接一个项目"能建出 `tracks/{your-project}/` 时,
108
- 你就知道设置成功了。
150
+ 1. 当一句招呼(任何语言都可以)能让 🐿️ 门菜单出现、而"连接一个项目"能建出
151
+ `tracks/{your-project}/` 时,你就知道设置成功了。
109
152
  2. 然后在同一个会话里拿下一个即时收益:说 **"加速这个项目"**(一份值得接线的技能/插件排序方案,
110
153
  安装要过门禁),或者 **"跑一下 /context-doctor"**(token 浪费扫描)。
111
154
  3. 一条诚实说明:FH 的核心回报是 **复利累积** —— 会话记录、收割来的学习、跨会话记忆。它从
@@ -276,11 +319,24 @@ Project B ──→ 在 CLAUDE.md 中连接中枢
276
319
 
277
320
  **三段工序** —— 这是一个 *投入的顺序*,不是一份菜单:
278
321
 
322
+ ```mermaid
323
+ flowchart LR
324
+ S["① 立坐标系<br/>设计之前"]
325
+ P["② 并行去相关<br/>中段"]
326
+ B["③ 烧一遍<br/>六条轴上"]
327
+ A(["⟹ 于是它快起来"])
328
+ S --> P --> B --> A
329
+ style A fill:#0f766e,stroke:#0f766e,color:#fff
330
+ ```
331
+
332
+ **速度是末尾那支箭,不是第四个方框。** 而 ② 是 **两个分开拧的旋钮** —— *去相关*(应对盲点风险:
333
+ 换一个模型家族 ⓐ、换一个立场 ⓑ)与 *并行*(应对表面大小:一个上下文装得下吗)。别做乘法,要挑。
334
+
279
335
  ```
280
336
  ① 设计之前先立坐标系 判断坐标系「最先」进场 —— 成功 · 偏向 · 不在范围 · 绝不做 ——
281
337
  而不是事后补写成一份"我做了什么"的记录
282
338
 
283
- 中段做去相关,用来加速 把工作拆成会以「不同方式」失败的检查,然后一次性跑掉。要挑「哪些
339
+ 并行去相关 把工作拆成会以「不同方式」失败的检查,然后一次性跑掉。要挑「哪些
284
340
  差异算数」—— 再来一位同一种类的审阅者不是去相关,那是把同一个盲点
285
341
  看两遍。并行本身没有方向,挑方向的是 ① 里那套判断坐标系。
286
342
  这是一种「工作方式」,不是 ③ 里那道收尾检查。
@@ -290,6 +346,11 @@ Project B ──→ 在 CLAUDE.md 中连接中枢
290
346
  看不见的东西照样看不见
291
347
  ```
292
348
 
349
+ > 📖 **从这里往下,是给想再深入一点的人看的 —— 开始用它并不需要这些。**
350
+ > 如果你是来把它装上、直接开跑的,最上面那个「两分钟」小节就是全部;你可以在这里停住,等哪天
351
+ > 某道检查让你觉得意外了再回来。下面讲的是这些门禁为什么长成这样,是写给 **已经在跑这套中枢的人**
352
+ > 看的。碰到不认识的词 → [`GLOSSARY.md`](knowledge/shared/GLOSSARY.md)。
353
+
293
354
  **六条验证轴** —— 所谓"我们审过了",往往到头来只做了其中第一条。
294
355
 
295
356
  🟥 **轴不是按「有多对抗」来分的,是按「它拿到了什么」来分的。** 拿到的东西一样,你堆多少位审阅者,
@@ -297,13 +358,39 @@ Project B ──→ 在 CLAUDE.md 中连接中枢
297
358
 
298
359
  | 轴 | **它拿到什么** | 它抓到什么 | 典型手段 |
299
360
  |---|---|---|---|
300
- | **ⓐ 不同家族** | diff + 作者的框架叙述 | **实现** 错了 | 换一个模型家族的审阅者(`auto-decorrelation`) |
301
- | **ⓑ 立场 (standpoint)** | diff + **目标框架自己的正典** | **你引用的那条规约是不是真这么说** | 在那个框架自己的仓库与规则里去跑这份 diff ([`§7`](knowledge/shared/harness-core/field_verdict_crossfamily_gate.md)) |
302
- | **ⓒ 隔离接地** | 作者写下的那些句子 + 此刻的这棵树 | **主张** 错了 | 找一个没写过它的人,把它说的重新测一遍 |
361
+ | **ⓐ 不同家族** | 变更 + 作者的框架叙述 | **实现** 错了 | 换一个模型家族的审阅者(`auto-decorrelation`) |
362
+ | **ⓑ 立场 (standpoint)** | 变更 + **目标框架自己的正典** | **你引用的那条规约是不是真这么说** | 在那个框架自己的仓库与规则里去跑这份变更 ([`§7`](knowledge/shared/harness-core/field_verdict_crossfamily_gate.md)) |
363
+ | **ⓒ 隔离接地** | 作者写下的那些句子 —— 他的主张,*以及* 他在 **动手之前声明过的东西** —— + 此刻的这棵树 | **主张** 错了 · 增量与当初声明的对不上 | 找一个没写过它的人,把它说的重新测一遍。至于事前声明,那是一道把写下的成功定义拿回来对照增量的门禁 |
303
364
  | **ⓓ 第三方对面** | 问题 + **别人的代码库** | **这是不是早就被解过了** · 你的变更会碰到别人仓库的哪里 | 在一个无关的第三方仓库里看同一个问题 |
304
365
  | **ⓔ 首次真实使用** | 一个实打实的目标 | **你测量的方式** 错了 —— 量具的量具 | 拿一个真实目标真跑一次,然后动手核对那份结果 |
305
366
  | **ⓕ 撤回并观察** | 把接线删掉之后的那棵树 | **锚** 错了 —— 那道检查是装饰 | 把它所守护的东西删掉,确认 *正是那一条* 检查变红 |
306
367
 
368
+ > **ⓒ 在 2026-08-21 被拓宽了,而「怎么拓宽的」才是更有用的那一半。** 本仓库的提交标记自
369
+ > 2026-08-09 起就要求写上作者自己的 **事前声明** —— *什么算成功* 与 *什么绝不做* —— 而且必须写在
370
+ > 设计 *之前*。那天带着对照组一测:**没有任何代码在读它。** 消费它的代码整整零行,而兄弟字段在
371
+ > 21 处被检查;门禁规格甚至没有点过它的名。在真实语料里,**98 份标记中有 37 份根本没有那一行**
372
+ > —— 其中还包括一份跑过 28 条 lane、别的字段全部填满、并通过了评审面板的标记。这些轴全都朝
373
+ > *外* 看 —— 变更、目标仓库、先例、产物。**没有一条轴回头看那份记录自己的必填字段。**
374
+ > 一个没有消费方的槽位永远报告「完成」,**因为是「存在」本身在替你做判定。**
375
+ >
376
+ > 修法不是加第七条轴。ⓒ 本来就拿到 *作者写下的那些句子 + 此刻的这棵树*,而这与一次事前声明检查
377
+ > 所拿到的东西 **逐字相同**。时态(先声明的,还是事后主张的)和对抗性一样,是一种 **姿态**,
378
+ > 不是一条轴。真去新铸一条,就等于把上面那次盲判重分类抓到的错误再犯一遍。
379
+
380
+ > **同侪会话记在 ⓑ 和 ⓓ 上 —— 2026-08-21 决定,而「问哪一个同侪」才是全部诀窍。** 同一个框架分叉
381
+ > 出去、正在另一条支线上跑得火热的并行会话,不是你的副本。在它跑热的那一点上,它确实长出了
382
+ > **第二张脸**:一个货真价实的立场(ⓑ),一个货真价实的「别人的代码库」(ⓓ)。所以同侪的判定记在
383
+ > 这两条轴上 —— 没有新铸第七条轴,标记里的 `axes-run` 字母表也没有变。
384
+ >
385
+ > 上面那张表的 ⓓ 行之所以写的是「**别人的代码库**」而不是「别家公司的仓库」,原因就在这里:真正
386
+ > 起作用的边界是 **那份判定是在谁的工作脉络里产生的**,而不是这些文件挂在谁的 GitHub 组织名下。
387
+ > 一个在另一条支线上跑热的同侪,在这条边界 **之外**;而你从当前这个脉络里派发出去的子 agent,
388
+ > 不管它去读的是多么不相干的仓库,都仍然在这条边界 **之内**。
389
+ >
390
+ > 🟥 咬人的是那条推论。**要问同侪的,是它真正跑热的那条轴。** 出了这个范围,它戴的就是你的脸,
391
+ > 问了也去不了相关 —— 同样的输入,同样的盲点。而且子 agent 顶替不了它,重读一遍自己的文字也
392
+ > 顶替不了:那第二张脸来自那个会话真的 *做了* 不一样的事,靠提示词是造不出来的。
393
+
307
394
  **你不必每次都把六条跑满,这正是设计** —— 别做乘法,要 **挑**:
308
395
 
309
396
  ```
@@ -323,7 +410,9 @@ Project B ──→ 在 CLAUDE.md 中连接中枢
323
410
 
324
411
  **为什么这一套不会被基础模型的进步取代** —— 一条轴是由 **输入** 定义的,不是由 *审阅者的能力*
325
412
  定义的。模型再强,**「没拿到的信息」它依然看不见。** 脚手架会随着模型变好而脱落,但 **输入边界上的
326
- 去相关不会脱落**,而单个作者按定义就走不出自己的输入。🟥 诚实的边缘地带:如果 agent **自己用工具
413
+ 去相关不会脱落**,而单个作者按定义就走不出自己的输入。
414
+
415
+ 🟥 **诚实的边缘地带**:如果 agent **自己用工具
327
416
  去取更多输入**,这条边界就会变模糊 —— 外部判定确实认为「整个存储从未被完整使用」和「异常被吞掉」
328
417
  这两条 ⓐ · ⓒ 也能抓到(因为它们会自己 grep)。反过来,「别人的仓库当年废弃掉的某条规则」
329
418
  **用工具也取不到** —— 你压根没有理由去访问那个项目的评审历史。ⓓ 留下来的位置就在那里。
@@ -441,21 +530,37 @@ hooks 不会自动触发,M2 的 agent 派发步骤需要适配器(或交互
441
530
 
442
531
  ---
443
532
 
444
- ## 大锻炉 (The forge)
533
+ ## 大锻炉 (The forge) —— 名字的由来
445
534
 
446
- forge-harness 把项目当作钢来对待 —— 而这个隐喻是字面的,不是装饰。工作被塑形、以攻击淬硬,
535
+ forge-harness 把项目当作钢来对待,而这个隐喻是字面的,不是装饰。工作被塑形、以攻击淬硬,
447
536
  唯有如此才更快出炉,因为它挺过了考验。
448
537
 
449
- | 工序 | 发生了什么 | 命令 |
538
+ > 🟥 **这是名字的由来,不是流程 —— 不要数它。** 下面那些铁匠的用词是一套 **词汇**,不是阶段清单。
539
+ > 它与上文的 **三段工序 · 四大引擎 · 四轴门禁 · 六轴验证是不同的层**,不对应其中任何一个数字;
540
+ > 把它读成「第五个带编号的集合」,是本节唯一可能引起的误会,所以在这里先掐断。
541
+
542
+ **铁砧上的三道:**
543
+
544
+ | 铁匠的用词 | 在这里是什么意思 | 命令 |
450
545
  |---|---|---|
451
546
  | **锻造 (Forge)** | 把生坯项目塑形为框架 —— 抬高其下限 | `install-wizard`、"把这个项目框架化" |
452
547
  | **淬火 (Quench)** | 以攻击将其淬硬 —— 冷审阅只让健全的东西留存 | `steel-quench` · `phantom-quench` |
453
548
  | **回火 (Temper)** | 把淬硬资产里的脆性 (brittleness) 再退掉 | `steel-quench` Wave-T · `templates/temper_check.sh` |
454
- | → **加速 (Accelerate)** | 一把挺过锻炉的刀刃切得更快 | `goal-quench` —— *Pass → Accelerate* |
455
549
 
456
- 四道工序全部出货。回火 (Temper) 在被造出 *之前* 就先起了名字 —— 刻意为之(见
457
- [`ETHOS.md`](docs/ETHOS.md#the-forge))—— 并在测量运行验证之后出货。围绕这座锻炉,还有两个
458
- 签名部件让它持续运转:`harvest-loop`(每次会话的教训成为永久技能)与
550
+ ```mermaid
551
+ flowchart LR
552
+ F["锻造 Forge"] --> Q["淬火 Quench"] --> T["回火 Temper"] --> A(["⟹ 加速 Accelerate"])
553
+ style A fill:#0f766e,stroke:#0f766e,color:#fff
554
+ ```
555
+
556
+ **⟹ 然后,它自己就快起来了。** 一把挺过锻炉的刀刃切得更快 —— `goal-quench`,*Pass → Accelerate*。
557
+ 速度是上面这三道 **锻出来的东西**,不是你另外再做的第四件事。这和本页最上面那句标语说的是同一件事,
558
+ 只不过换成了铁匠的用词:质量是杠杆,速度是结果。
559
+
560
+ 上面点到名的命令,如今全部已经出货。回火 (Temper) 在被造出 *之前* 就先起了名字 —— 刻意为之
561
+ (见 [`ETHOS.md`](docs/ETHOS.md#the-forge))—— 并在测量运行验证之后出货。
562
+
563
+ 围绕这座锻炉,还有两个签名部件让它持续运转:`harvest-loop`(每次会话的教训成为永久技能)与
459
564
  `agent-composer`(编排派发)。其余技能等你需要时再出场 —— 完整清单见下。
460
565
 
461
566
  ## 40 skills · 8 agents
package/docs/ETHOS.md CHANGED
@@ -18,14 +18,21 @@ Everything below is **copyable**. None of it is a secret. The principles are the
18
18
  FH treats a project like steel — heat, shape, and shock, in named movements. The metaphor is literal,
19
19
  not decoration:
20
20
 
21
- | Movement | What it does | Today |
21
+ **Three at the anvil:**
22
+
23
+ | The smith's word | What it does | Today |
22
24
  |---|---|---|
23
25
  | **Forge** | shape the raw project into a harness — raise its floor | `install-wizard`, harness-ify |
24
26
  | **Quench** | harden it by attack — the cold pass leaves standing only what is sound | `steel-quench`, `phantom-quench` |
25
27
  | **Temper** | take the brittleness back out of the hardened asset | `steel-quench` **Wave-T** + `templates/temper_check.sh` |
26
- | → **Accelerate** | a blade that survived the forge cuts faster — pass, then run | `goal-quench` · *Pass → Accelerate* |
27
28
 
28
- All four movements now ship. **Temper** spent its first months named-but-unbuiltdeliberately, per
29
+ **⟹ And then it accelerates.** A blade that survived the forge cuts faster `goal-quench`,
30
+ *Pass → Accelerate*. Three at the anvil, and speed is what those three **produce**; it is not a
31
+ fourth thing you do at the anvil. Quality is the lever, speed is the result.
32
+
33
+ Every command named above ships today. (These are the smith's words — a **vocabulary**, not a stage
34
+ list, and a different layer from the three-stage process, the four engines, the four-axis commit gate
35
+ and the six verification axes. Do not count them against those.) **Temper** spent its first months named-but-unbuilt — deliberately, per
29
36
  principle 5 — and shipped only after measurement runs on independent quench convergences validated that
30
37
  the check flags over-hardening without punishing simplification. Quenched steel is hard but brittle;
31
38
  no smith ships it un-tempered, and now neither does FH: after convergence, Wave-T measures the complexity