@chrono-meta/fh-gate 1.4.89 → 1.4.91
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/judgment_circuits.txt +14 -0
- package/.claude/rules/fh_4axis_gate.md +7 -0
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +25 -0
- package/CLAUDE.md +202 -12
- package/knowledge/shared/harness-core/dispatch_conditional_prohibition.md +105 -0
- package/knowledge/shared/harness-core/fh_three_layer_canon.md +307 -0
- package/knowledge/shared/harness-core/harness_incubator_doctrine.md +100 -0
- package/knowledge/shared/harness-core/onboarding_acceleration_autopilot.md +3 -1
- package/knowledge/shared/harness-core/ship_readiness_gate.md +181 -13
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +640 -0
- package/knowledge/shared/rules/multi_session_close_protocol.md +118 -0
- package/package.json +23 -2
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +56 -8
- package/plugins/fh-meta/skills/install-wizard/SKILL.md +33 -0
- package/plugins/fh-meta/skills/install-wizard/SKILL_detail.md +104 -15
- package/scripts/branch_claim.sh +266 -0
- package/scripts/chamber_run.sh +64 -2
- package/scripts/chamber_witness.sh +439 -0
- package/scripts/compaction_probe.sh +456 -0
- package/scripts/digest_landing_check.sh +385 -0
- package/scripts/directional_diff_gate.sh +459 -0
- package/scripts/fh_env_delta_scan.sh +15 -0
- package/scripts/fh_session_load.sh +34 -1
- package/scripts/field_canon_preload.sh +129 -0
- package/scripts/hook_source_lib.sh +41 -0
- package/scripts/judgment_circuit_lint.sh +239 -0
- package/scripts/novelty_claim_check.sh +193 -0
- package/scripts/relay_channel.sh +645 -0
- package/scripts/reviewer_capability_corpus.tsv +124 -0
- package/scripts/selfcheck.sh +106 -0
- package/scripts/session_close_check.sh +134 -1
- package/scripts/test_branch_claim_lanes.sh +231 -0
- package/scripts/test_dispatch_log_lanes.sh +35 -1
- package/scripts/test_field_canon_lanes.sh +142 -0
- package/scripts/test_hook_source_gate_lanes.sh +81 -0
- package/scripts/test_marker_crossfamily_lanes.sh +132 -0
- package/scripts/test_marker_floor_lanes.sh +9 -8
- package/scripts/test_relay_channel_lanes.sh +583 -0
- package/scripts/test_reviewer_capability_conformance.sh +173 -0
- package/scripts/test_wizard_snippet_merge_lanes.sh +104 -11
- package/scripts/utterance_landing_check.sh +209 -0
- package/templates/.git-hooks/pre-commit +297 -13
- package/templates/settings.Compaction.snippet.json +56 -0
- package/templates/settings.FieldCanon.snippet.json +51 -0
|
@@ -0,0 +1,307 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: fh-three-layer-canon
|
|
3
|
+
description: FH 를 설명하는 3층 정본 — 3단 공정(엔진을 벼리는 순서) · 4대 엔진(그 능력) · 5대 정체성(사람이 실제로 쓰는 기능). 층별 명명 규칙과 4축 검증의 정의를 포함한다.
|
|
4
|
+
date: 2026-08-09
|
|
5
|
+
tags: [canon, three-layer, engines, identity, verification-axes]
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# FH 3층 정본 — 공정 · 엔진 · 정체성
|
|
9
|
+
|
|
10
|
+
> **왜 이 문서가 필요했나**: 세 층이 **따로 살고 있었다.** 5대 정체성만
|
|
11
|
+
> `ship_readiness_gate.md` 에 있었고, **3단 공정과 4대 엔진은 공개 지식층에 0건** —
|
|
12
|
+
> gitignored `tracks/` 와 개인 메모리에만 있었다(2026-08-09 실측). 즉 FH 를 설명하는
|
|
13
|
+
> 뼈대의 3분의 2를 **다른 세션·다른 런타임이 읽을 수 없었다.** 이건 내용 부족이 아니라
|
|
14
|
+
> `[[gate_locality_principle]]` 문제다 — 필요한 곳에 없으면 없는 것이다.
|
|
15
|
+
> 셋을 한 곳에서 다루는 문서는 그때까지 **존재하지 않았다**(no-reinvention 스캔 확인).
|
|
16
|
+
|
|
17
|
+
## 0. 세 층은 무엇이 다른가 — **명사가 층을 실어야 한다**
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
5대 정체성 사람이 FH 를 돌릴 때 **쓸 수 있는 기능** (표면 — 무엇을 얻나)
|
|
21
|
+
↑ 받친다
|
|
22
|
+
4대 엔진 그 기능을 가능하게 하는 **능력** (능력 — 무엇을 할 수 있나)
|
|
23
|
+
↑ 낳는다
|
|
24
|
+
3단 공정 그 엔진을 **벼리는 순서** (공정 — 어떻게 만들어지나)
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
**명명 규칙(2026-08-09 운영자 확정)**: `공정 / 엔진 / 정체성`.
|
|
28
|
+
이전 표기 `3단 개발법 · 4개 엔진 · 5대 요소` 의 문제는 단어 품질이 아니라 **층이 안 들린다**는
|
|
29
|
+
것이었다 — 「엔진」과 「요소」가 사실상 동의어라 듣는 쪽이 어느 층인지 구분하지 못한다.
|
|
30
|
+
바뀐 것은 두 단어뿐이다: `개발법 → 공정`(개발만이 아니라 검증까지 덮는다) ·
|
|
31
|
+
`요소 → 정체성`(부품이 아니라 존재다). **「엔진」은 유지** — 기존 PR·카드·메모리 인용이 그대로 산다.
|
|
32
|
+
|
|
33
|
+
### ⚠️ 순수 스택이 아니다 — 그리고 그게 결함이 아닌 이유
|
|
34
|
+
|
|
35
|
+
3단의 **1단계(영혼)** 와 **3단계(4축 태우기)** 는 엔진 ①②와 **같은 재료**다.
|
|
36
|
+
아래층이 위층을 쓰므로 엄밀한 계층이 아니다. 이 모순은 **주어가 다르다**는 것으로 풀린다:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
4대 엔진 = FH 가 **사용자 일**에 쓰는 능력
|
|
40
|
+
3단 공정 = FH 가 **자기 엔진을 벼릴 때** 쓰는 순서
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
밖에서 빌려온 방법론이면 엔진과 무관했을 것이다. **재료가 겹친다는 것은 도그푸딩의 지문**이고,
|
|
44
|
+
오히려 3단이 실재한다는 증거다. "깔려 있다" 를 *"엔진과 무관한 아래층"* 으로 읽으면 과장이다.
|
|
45
|
+
|
|
46
|
+
---
|
|
47
|
+
|
|
48
|
+
## §1 — 3단 공정 (엔진을 벼리는 순서)
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
① 초기 영혼 심지(판단 회로)를 먼저 심는다. 무엇이 성공 · 어디로 기움 ·
|
|
52
|
+
범위 밖 · 절대 안 함
|
|
53
|
+
② 중간 탈상관 가속화 축을 골라 동시에 친다. **곱하지 말고 골라라** — 병렬화 자체엔
|
|
54
|
+
방향이 없다(축은 판단 회로가 고른다)
|
|
55
|
+
③ 마무리 4축 태우기 §1-a 의 네 축으로 태운다. **적대검증 하나가 아니다**
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
**이건 도구 목록이 아니라 투입 순서다.** 축 없이 병렬부터 던지면 노이즈가 늘고,
|
|
59
|
+
태우기부터 하면 태울 게 없다. 근본 명제(운영자, 2026-08-09):
|
|
60
|
+
|
|
61
|
+
> **"하네스는 방법론이고 지능은 이미 가지고 있다. 방법론만 바꾸면 승리한다."**
|
|
62
|
+
> 병력 = 지능(이미 있다, 변수가 아니다) · 병법 = 하네스(배치와 순서)
|
|
63
|
+
|
|
64
|
+
실측: 모델 고정 · 방법만 혼자→cross-family → **자력 적발 0 → 11건**(2026-08-09).
|
|
65
|
+
선례: `dominance_benchmark_2026-07-14` 가 *"Sonnet floor 고정, method 만 변화"* 로
|
|
66
|
+
5/8 → 6/8 → **8/8** 계단을 냈다.
|
|
67
|
+
|
|
68
|
+
### §1-b — 스코프 (2026-08-09 운영자 정식화 + 같은 날 3스코프 실측)
|
|
69
|
+
|
|
70
|
+
> **운영자**: *"3단은 스코프가 자유자재다. 작은 스코프는 세션, 큰 스코프는 필드 하네스 개발,
|
|
71
|
+
> 그 위가 메타하네스 자기 자신."*
|
|
72
|
+
|
|
73
|
+
**대체로 맞다. 다만 실측하면 세 가지가 더 붙는다.**
|
|
74
|
+
|
|
75
|
+
#### ⓐ 자유자재가 아니라 — **하한이 있다**
|
|
76
|
+
|
|
77
|
+
판별 기준은 **크기가 아니라 «성공 정의가 설계를 바꾸는가»** 다.
|
|
78
|
+
|
|
79
|
+
```
|
|
80
|
+
gitignore 한 줄 수리 ① 영혼 불요 답이 하나다. 심으면 오버헤드
|
|
81
|
+
branch-claim 게이트 ① 영혼 필수 «무엇을 성공으로 볼 것인가»가 설계를 결정한다
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
**실측(2026-08-09)**: 후자에서 ①을 건너뛴 대가가 나왔다 — *"성공 = 오늘 그 사고가
|
|
85
|
+
재현될 때 차단된다"* 를 검증 가능한 문장으로 안 썼고, 그래서 **트리 단위 설계가
|
|
86
|
+
자기 성공 기준을 즉시 실패한다는 것**(A가 switch 후 claim 하면 B가 통과)을
|
|
87
|
+
**적대검증 4라운드 뒤에** 알았다. 심지 한 줄이 짓기 전에 잡았을 것이다.
|
|
88
|
+
|
|
89
|
+
#### ⓑ 스코프는 나란하지 않다 — **중첩되고, 서로를 먹인다**
|
|
90
|
+
|
|
91
|
+
```
|
|
92
|
+
안쪽 ③ ◀──빌린다── 바깥 기계 게이트를 검증한 4축 훅은 **세션/저장소 스코프의 것**이다.
|
|
93
|
+
게이트가 자기 힘으로 자기를 검증한 게 아니다 [강한 증거]
|
|
94
|
+
바깥 ② ──먹인다──▶ 안쪽 ② 두 축이 서로를 검증했다(각자 상대 주장을 반증) [약한 증거]
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
**강한 쪽부터**: 안쪽 스코프는 **자기보다 바깥의 기계에 의존해서** ③을 완수한다.
|
|
98
|
+
오늘 branch-claim 게이트를 통과시킨 4축 훅·마커·앵커는 전부 저장소 스코프 자산이다.
|
|
99
|
+
즉 «메타하네스 자기 스코프» 는 독립 실행 단위가 아니다 — **바깥이 없으면 ③이 성립하지 않는다.**
|
|
100
|
+
|
|
101
|
+
⚠️ **약한 쪽은 약하다고 적는다.** 초판은 *"세션 축 분리로 **예산**이 생겨 3렌즈를 띄웠다"*
|
|
102
|
+
라고 썼는데 **과장이다** — 두 세션은 예산이 독립이고, 축 분리가 준 것은 예산이 아니라
|
|
103
|
+
**컨텍스트 집중**이다. 게다가 그 분리는 *계획된 탈상관*이 아니라 **원래 다른 두 작업**이었고,
|
|
104
|
+
상호 검증은 계획이 아니라 **창발**이었다. 「바깥 ②가 안쪽 ②를 먹인다」는 **가설이지 실측이 아니다.**
|
|
105
|
+
|
|
106
|
+
즉 «작은/큰»의 병렬 배열은 아니되, **확인된 의존은 아래→위 한 방향뿐**이다.
|
|
107
|
+
|
|
108
|
+
#### ⓒ ②의 효과는 둘이고, **둘째가 더 크다**
|
|
109
|
+
|
|
110
|
+
> **운영자**: *"병렬 탈상관 가속화를 하여 개발을 빠르게 해낸다(**뒷단의 부담을 덜어낼 수도 있다**)."*
|
|
111
|
+
|
|
112
|
+
실측: 3렌즈가 잡은 **M급 9건 + 콜드리드 HARD 5건**이 없었다면 그 14건은 **직렬이고
|
|
113
|
+
시간이 정해진 리허설**로 몰렸을 것이다. 거기서 나오면 못 고친다.
|
|
114
|
+
`fh_signal_2026-08-07_shift-left-governance` §④ 의 *"뒷단 부하는 앞단 부재의 함수다"* 와
|
|
115
|
+
같은 명제이고, 이 괄호가 그것을 ②의 정의 안으로 넣는다.
|
|
116
|
+
|
|
117
|
+
⚠️ **가속화 ≠ shift-left.** 앞당기는 것과 탈상관하는 것은 다른 축이다. 오늘은 둘이 같이
|
|
118
|
+
갔을 뿐이고, 한 문장으로 묶으면 *"병렬로 돌렸으니 앞당긴 것"* 이라는 오독이 나온다.
|
|
119
|
+
|
|
120
|
+
#### 🟥 ⓓ 실측 — **자기 스코프가 가장 나빴다**
|
|
121
|
+
|
|
122
|
+
같은 날 한 세션에서 세 스코프가 다 돌았고, 품질이 이렇게 갈렸다:
|
|
123
|
+
|
|
124
|
+
| 스코프 | ① 영혼 | ② 탈상관 | ③ 4축 |
|
|
125
|
+
|---|---|---|---|
|
|
126
|
+
| **A 세션** (두 축 병렬) | 🟡 프레임은 있고 «절대 안 함» 없음 | 🟡 **창발이지 설계가 아니다** — 원래 다른 두 작업이었고 상호검증이 결과로 생겼다 | ✅ peer 상호검증 · 마감 스윕이 미착지 결정 1건 적발 |
|
|
127
|
+
| **B 산출물** (발표 원고) | ✅✅ 인용금지 원장 = «절대 안 함» | ✅✅ 3렌즈를 **실패 모드로** 골랐다 | ✅ ⓐ 다른계열 · ⓑ 첫실사용 · ⓒ 격리감사 4건 반증 |
|
|
128
|
+
| **C 메타하네스 자기** (게이트·배선) | 🟥 **안 심었다** | 🟡 사후 검증으로 씀(«중간»이 아니다) | 🟡 ⓐⓑⓓ 함 · **ⓒ 안 함** |
|
|
129
|
+
|
|
130
|
+
**FH 가 남의 것엔 3단을 잘 쓰고 자기 것엔 못 썼다.** §0 이 *"재료가 겹친다는 것은
|
|
131
|
+
도그푸딩의 지문"* 이라고 적었는데, 실측은 한 겹 더 말한다 — **겹치면 건너뛰기도 쉽다.**
|
|
132
|
+
자기 코드는 *"내가 아니까"* 라는 감각이 ①을 먼저 지운다.
|
|
133
|
+
|
|
134
|
+
> **운영자(2026-08-09)**: *"아직 RC 이긴 한데, RC 인 만큼 스스로가 잘 굴려야 초록으로도
|
|
135
|
+
> 갈 수 있겠지."* → **C 스코프의 품질이 승급 경로다.** 남에게 잘 쓰는 것으로는 안 오른다.
|
|
136
|
+
> ⚠️ 등급 어휘(🔵/🟢)는 **§3 정체성 층에 붙지 이 공정 층에 붙지 않는다**(정본 =
|
|
137
|
+
> `ship_readiness_gate.md`). 위 문장은 «공정을 자기에게 적용한 품질이 정체성 승급을
|
|
138
|
+
> 좌우한다» 로 읽어라 — 공정 자체에 등급을 매기는 것이 아니다.
|
|
139
|
+
|
|
140
|
+
**근거 등급 (정직하게)**: 위 표는 **한 세션 n=1** 이다. 하한 판별 기준(ⓐ)의 근거는
|
|
141
|
+
**사례 2개**(gitignore vs 게이트)뿐이다. ⓓ 의 «자기 스코프가 가장 나쁘다» 가 구조적인지
|
|
142
|
+
그날 우연인지는 **미측정** — 다음 세션들에서 같은 표를 채워 봐야 안다.
|
|
143
|
+
|
|
144
|
+
### §1-a — 마무리 «4축» (2026-08-09 실측으로 정의)
|
|
145
|
+
|
|
146
|
+
3단계를 오래 «적대검증» 이라고 불렀으나, 그건 **넷 중 하나**다. 각 축은 **다른 것이 틀린
|
|
147
|
+
경우**를 잡고 **서로를 대체하지 못한다**:
|
|
148
|
+
|
|
149
|
+
| 축 | 무엇이 틀린 경우 | 대표 계기 |
|
|
150
|
+
|---|---|---|
|
|
151
|
+
| **ⓐ 다른 계열** | **구현**이 틀렸다 | cross-family 적대검증(`auto-decorrelation`) |
|
|
152
|
+
| **ⓑ 첫 실사용** | **재는 방식**이 틀렸다 | 실물 대상에 한 번 돌려본다 |
|
|
153
|
+
| **ⓒ 기록 그라운딩** | **주장**이 틀렸다 | 격리 감사가 문서의 수치·인용·기전을 재측정 |
|
|
154
|
+
| **ⓓ 되돌림 실측** | **앵커**가 틀렸다 | 배선을 지우고 대응 레인이 빨개지는지 본다 |
|
|
155
|
+
|
|
156
|
+
**2026-08-09 실측(한 산출물, 같은 날)**: ⓐ 18건 · ⓑ 2건 · ⓒ 11건 · ⓓ 15건.
|
|
157
|
+
그리고 **ⓐ 4라운드가 ⓒ 를 0건 잡았고, ⓒ 가 ⓓ 를 0건 잡았다.**
|
|
158
|
+
ⓑ 는 *"계기의 계기가 틀린"* 경우라 ⓐ·ⓓ 가 **구조적으로** 못 본다.
|
|
159
|
+
|
|
160
|
+
⚠️ **한 문장으로 묶지 마라 — 처방이 갈린다**: ⓒ=기록 재감사 · ⓓ=앵커 되돌림 ·
|
|
161
|
+
계기 유효성=known-pair 보정. 두 병렬 세션이 이 결론에 독립 수렴했다.
|
|
162
|
+
비용 경계: 넷을 매번 다 돌리는 축이 아니다. **곱하지 말고 실패 모드에 맞춰 골라라**
|
|
163
|
+
(`[[feedback_decorrelation_not_fanout_cost_boundary]]`).
|
|
164
|
+
|
|
165
|
+
|
|
166
|
+
### §1-c — 자기 대조를 **상시 의무**로 만든 근거 (그리고 그 표본의 한계)
|
|
167
|
+
|
|
168
|
+
`CLAUDE.md` 가 «3층 자기 대조는 상시 의무 · 트리거는 발화가 아니라 자산 접촉» 을 상주로 싣는다.
|
|
169
|
+
그 규칙은 여기서 근거를 가져간다. **인용 전에 이 절의 «표본 한계» 를 같이 읽어라** — 규칙 본문은
|
|
170
|
+
짧고, 짧은 만큼 근거보다 세게 들린다.
|
|
171
|
+
|
|
172
|
+
**왜 트리거를 옮겼나**: §1-b ⓓ 의 A/B/C 표가 «자기 스코프가 가장 나빴다» 를 보였고, 그 미스는
|
|
173
|
+
**규칙을 몰라서 난 것이 아니다** — 같은 세션이 아침에 규칙을 읽고 저녁에 어겼다. 살리언스는 이미
|
|
174
|
+
최대였다. 그러면 처방이 「더 세게 읽어라」가 될 수 없고, 남는 지렛대는 **트리거를 «누가 물었나»
|
|
175
|
+
에서 «지금 자산을 건드리고 있다» 로 옮기는 것**이다. 후자는 4축 게이트가 이미 쓰는 트리거라
|
|
176
|
+
새 기계층이 필요 없다.
|
|
177
|
+
|
|
178
|
+
**축보다 컨트롤이 먼저다 (실측 n=4, 2026-08-09 병렬 세션)**:
|
|
179
|
+
| | |
|
|
180
|
+
|---|---|
|
|
181
|
+
| 측정 시도 | 4 |
|
|
182
|
+
| **컨트롤 없이 죽은 측정** | **2** |
|
|
183
|
+
| 최악 사례 | 컨트롤이 없었으면 *"운영자 발화 **30건** 전건 미착지"* 로 보고할 뻔했다 — **실제는 전건 착지** |
|
|
184
|
+
|
|
185
|
+
> 🟥 **이 표의 「30」은 정정본이다 — 초판은 `39` 였고 그것도 계기 결함이었다** (2026-08-09,
|
|
186
|
+
> 원 측정 세션이 격리 재측정으로 자진 철회). `type=user` 442건을 `origin.kind` 로 완전 분할하니
|
|
187
|
+
> **human 30 · peer 메시지 6 · task-notification 3** 이었다 — 파서가 `type=user` 를 사람 발화로
|
|
188
|
+
>읽었는데 **다른 AI 세션이 보낸 메시지와 시스템 알림이 같은 타입·같은 문자열 content** 라 섞여
|
|
189
|
+
> 들어왔다. **랜덤 오차가 아니라 범주 오류**이고, 나쁜 쪽은 커버리지가 30% 부풀려지는 것보다
|
|
190
|
+
> **AI 발화와 사람 의도가 한 통에 섞이는 것**이다.
|
|
191
|
+
> **방향은 살아남았다** — 30건도 전건 착지다. 죽은 것은 **분모**다.
|
|
192
|
+
> ★ **이 정정 자체가 이 절의 논거다.** 「컨트롤이 잡았다」던 사례의 분모를 잡은 것은 컨트롤이
|
|
193
|
+
> 아니라 **격리 재측정**(ⓒ 기록 그라운딩)이었다 — 축이 서로를 대체하지 못한다는 §1-a 의 주장이
|
|
194
|
+
> 자기 근거 위에서 한 번 더 성립했다.
|
|
195
|
+
> ⚠️ **전파 경계**: 이 값은 `CLAUDE.md` 상주 블록에는 **실려 있지 않다** — cross-family 지적으로
|
|
196
|
+
> 상주를 31줄→14줄로 줄일 때 사례가 빠졌기 때문이다. 축약이 오염 범위를 좁혔다(의도한 효과는
|
|
197
|
+
> 아니었다). 원장·PR 본문 등 다른 표면에 `39` 가 남았는지는 **표면별로 확인해야 한다**
|
|
198
|
+
> (`[[feedback_half_fix_propagation_boundary]]`).
|
|
199
|
+
> 🟡 **`n=4` 는 그 시점 값이다.** 원 측정 세션이 이후 더 늘었다고 보고했으나 **이 문서는 그
|
|
200
|
+
> 수치를 검증하지 않았고, 그래서 싣지 않는다** — 검증 전 인용이 방금 정정한 그 경로다
|
|
201
|
+
> (Instrument-Calibration §publish-order: *"publish = 어떤 형태로든 처음 말한 순간"*).
|
|
202
|
+
|
|
203
|
+
죽은 것은 대상이 아니라 **세는 줄**이었다(`[[feedback_absence_measurement_needs_control]]`).
|
|
204
|
+
그래서 규칙은 「4축을 돌려라」가 아니라 **「축을 돌렸다의 최소 증거 = 컨트롤이 살아 있는 실행
|
|
205
|
+
출력」** 을 요구한다. 축을 늘리는 것보다 각 축이 살아 있는지 보는 것이 먼저다.
|
|
206
|
+
|
|
207
|
+
**🟥 표본 한계 — 규칙이 이 표본보다 넓게 적용된다**
|
|
208
|
+
|
|
209
|
+
| 주장 | 표본 | 적용 범위 |
|
|
210
|
+
|---|---|---|
|
|
211
|
+
| 자기 스코프가 가장 나쁘다 | **세션 1개** (§1-b ⓓ) | 모든 FH/PMH 자산 변경 |
|
|
212
|
+
| 컨트롤 없는 측정이 절반 죽는다 | **n=4**, 한 세션 | 모든 축, 모든 install |
|
|
213
|
+
| 자평은 자력 적발률 0 | **분모 없는 관측** | 규칙 전체의 최약점 판정 |
|
|
214
|
+
|
|
215
|
+
이 비대칭은 cross-family 리뷰(2026-08-09, codex)가 지목한 것이고 **닫히지 않았다.** 규칙을 유지하는
|
|
216
|
+
근거는 표본 크기가 아니라 **비용 비대칭**이다 — 대조를 한 번 더 도는 비용은 작고, 안 돌아서
|
|
217
|
+
설계가 4라운드 뒤에 반증되는 비용은 크다. 표본이 커지면 이 절을 갱신하라.
|
|
218
|
+
|
|
219
|
+
**🟥 안 닫힌 것 셋** — 규칙 본문에도 그대로 적혀 있다.
|
|
220
|
+
- **자평이다.** 지금 기계 앵커가 붙는 축은 **ⓓ 뿐**이다(되돌려서 적색인지는 실행 결과다).
|
|
221
|
+
⚠️ 다만 이것은 **현 상태이지 원리적 한계가 아니다** — ⓐ(구현)는 테스트로, ⓒ(주장)는 출처
|
|
222
|
+
검사로 기계화할 수 있다(codex 지목). 「ⓓ만 가능」으로 읽지 마라.
|
|
223
|
+
- **게임 가능하다.** 축을 임의로 고르고 「안 고른 이유」만 적으면 형식은 충족된다. 최소 증거
|
|
224
|
+
요구가 그 폭을 좁히지만 없애지 못한다.
|
|
225
|
+
- **훅이 없다.** 트리거가 *의도*라 커밋 훅이 못 잡는다. 재발 실측이 기계화 임계 미만이라
|
|
226
|
+
**일부러 안 지었다**(`[[feedback_mechanize_at_repetition_prose_before]]`).
|
|
227
|
+
|
|
228
|
+
닫는 방향은 **cross-family 가 그 마커를 읽는 것**이다. 세션이 자기 채점을 더 성실히 하는 것은
|
|
229
|
+
같은 축이라 구조적으로 못 잡는다(`[[feedback_citing_a_rule_is_not_obeying_it]]`).
|
|
230
|
+
|
|
231
|
+
---
|
|
232
|
+
|
|
233
|
+
## §2 — 4대 엔진 (그 능력)
|
|
234
|
+
|
|
235
|
+
게이트 표의 `engine` 열이 **이미 이 넷을 기계적으로 들고 있었다** — 이 문서는 그것에
|
|
236
|
+
이름을 붙인 것이지 새 이론이 아니다.
|
|
237
|
+
|
|
238
|
+
| 엔진 | 게이트 표기 | 받치는 정체성 |
|
|
239
|
+
|---|---|---|
|
|
240
|
+
| **영혼(심지)** | `judgment-circuit` | ⑤ 증폭자 · ② 인큐베이터 |
|
|
241
|
+
| **품질 게이트** | `ship-gate` | ③ 거버넌스 게이트 |
|
|
242
|
+
| **맥락 유지** | `context-continuity` | ① 클러스터 · ② 인큐베이터 |
|
|
243
|
+
| **질문하기** | `external-grounding` | ④ 프런티어→조직 전파 |
|
|
244
|
+
|
|
245
|
+
- **영혼 = 정체성 선언이 아니라 판단의 좌표계**다(무엇이 성공 · 어디로 기움 · 범위 밖 ·
|
|
246
|
+
절대 안 함). 페르소나 선언은 105런 실측에서 **약모델 순손실**이므로 넣지 않는다.
|
|
247
|
+
형식 검사기: `scripts/judgment_circuit_lint.sh`.
|
|
248
|
+
- **영혼은 한 번에 안 만들어진다** — FH 는 **씨앗 초안**만 주고, 실제로 채워지는 것은
|
|
249
|
+
현장 하네스가 사용자에 의해 무수히 사용되면서다. 등급 대응:
|
|
250
|
+
**씨앗 출하 = 🔵 RC · 사용 축적 = 🟢 REALIZED**.
|
|
251
|
+
- **질문하기 ↔ external-grounding 은 같은 것의 양면**이다: 밖에 묻는다 = 밖에 근거를 둔다.
|
|
252
|
+
⚠️ 이 엔진의 **수용 절반이 비어 있다** — 묻는 경로는 있고 *들어온 통찰을 자산화하는*
|
|
253
|
+
경로가 약하다(`[[feedback_operator_insight_must_become_asset]]`).
|
|
254
|
+
|
|
255
|
+
---
|
|
256
|
+
|
|
257
|
+
## §3 — 5대 정체성 (사람이 실제로 쓰는 기능)
|
|
258
|
+
|
|
259
|
+
**등급 표의 정본은 `ship_readiness_gate.md` 다.** 여기 중복해서 적지 않는다 —
|
|
260
|
+
두 곳에 같은 등급을 두면 반드시 한쪽이 stale 이 된다(2026-08-09 에 그 결함을 실제로
|
|
261
|
+
세 번 겪었다: 표만 고치고 요약 안 고침 → 요약만 고치고 등급 칸 안 고침).
|
|
262
|
+
|
|
263
|
+
| # | 정체성 | 사람이 얻는 것 |
|
|
264
|
+
|---|---|---|
|
|
265
|
+
| ① | 멀티하네스 클러스터 | 한 태스크를 여러 하네스에 태우고 그 사이에서 거버넌스가 계산된다 |
|
|
266
|
+
| ② | 프로젝트 인큐베이터 | 새 하네스가 **태어난 자리에서 걷는** 상태로 나온다 |
|
|
267
|
+
| ③ | 거버넌스 게이트 | 못 나갈 것이 기계적으로 막힌다 |
|
|
268
|
+
| ④ | 프런티어→조직 전파 | 바깥에서 온 것이 조직 안까지 착지한다 |
|
|
269
|
+
| ⑤ | 증폭자 | 짧은 의도가 최종 산출물까지 벼려진다 |
|
|
270
|
+
|
|
271
|
+
**등급 어휘**: `🔴 이상론` → `🟡 부분` → `🔵 RC`(실험실에서 섰다) → `🟢 REALIZED`(밖에서 걸었다).
|
|
272
|
+
**RC 는 구현 성공이면 된다** — 문턱을 높이지 마라. 🟢 는 실사용 축적을 요구한다.
|
|
273
|
+
|
|
274
|
+
### §3-a — 정체성은 **별도 층이 아니라 스킬·에이전트에 퍼져 있다** (운영자, 2026-08-09)
|
|
275
|
+
|
|
276
|
+
> *"5개 요소는 자연스럽게 스킬과 에이전트를 포함한 하네스 자체 기능으로 퍼져 있는 것"*
|
|
277
|
+
|
|
278
|
+
이건 정체성이 스킬 위에 **덧씌운 분류**가 아니라, **스킬이 뭉치는 모양의 이름**이라는 뜻이다.
|
|
279
|
+
§2 에서 게이트의 `engine` 열이 이미 4엔진을 들고 있었던 것과 같은 형태 — 있는 것의 이름이다.
|
|
280
|
+
|
|
281
|
+
**그리고 이 관점은 공짜로 검출기 하나를 준다**: 어느 정체성에도 안 붙는 스킬 =
|
|
282
|
+
**고아 유닛**이고, `CLAUDE.md §Identity` 가 메타하네스의 적신호로 명시한 것이 정확히 그것이다
|
|
283
|
+
(*"red flags are orphaned, redundant, and decorative units"*). 정체성 매핑은 분류가 아니라
|
|
284
|
+
**자산 감사 계기**다.
|
|
285
|
+
|
|
286
|
+
🟡 **전수 대응은 미측정이다 — 눈으로 배정하지 마라.** 실측 규모는 **스킬 38 · 에이전트 8**
|
|
287
|
+
(2026-08-09). 명확한 것도 많지만 **실제로 저항하는 항목이 있다**: `video-ingest`(⑤ 증폭자인가
|
|
288
|
+
④ external-grounding 인가) · `mcp-circuit-breaker` · `token-budget-gate` — 인프라성이라
|
|
289
|
+
다섯 중 어디에도 자연스럽게 안 붙는다. 그 저항이 **고아 신호일 수도, 분류 축이 모자란
|
|
290
|
+
신호일 수도** 있고, 둘은 처방이 정반대다. 손으로 배정해 표를 채우면 그 구분이 사라진다.
|
|
291
|
+
|
|
292
|
+
**다음 세션 후보(미착수)**: 스킬별 정체성 선언을 frontmatter 로 받고 미선언·미분류를
|
|
293
|
+
세는 검사기. 선언은 저자가 하고 계기는 **세기만** 한다 — 계기가 분류를 대신 하면
|
|
294
|
+
`[[feedback_not_found_is_not_zero_family]]` 로 되돌아간다.
|
|
295
|
+
|
|
296
|
+
⚠️ **⑤ 의 🟢 는 범위 한정 가능성이 열려 있다**(2026-08-09 관찰, 미반영): 그 근거가 전부
|
|
297
|
+
*"사용자 산출물"* 대상에서 측정됐고, `⑤ × 하네스 자신` · `⑤ × 낳는 하네스` 는 미측정이다.
|
|
298
|
+
등급을 내리자는 게 아니라 **무엇에 대해 초록인지 명시하자는 것**이며, 등급 관련이라
|
|
299
|
+
**운영자 확인 후** 반영한다.
|
|
300
|
+
|
|
301
|
+
---
|
|
302
|
+
|
|
303
|
+
## 4. 이 문서가 하지 않는 것
|
|
304
|
+
|
|
305
|
+
- **등급 판정을 하지 않는다** — `ship_readiness_gate.md` 가 정본이다.
|
|
306
|
+
- **4축을 매번 다 돌리라고 하지 않는다** — 축은 실패 모드에 맞춰 고른다.
|
|
307
|
+
- **3층이 깔끔한 계층이라고 주장하지 않는다** — §0 의 자기참조 절을 반드시 같이 읽어라.
|
|
@@ -157,6 +157,106 @@ authored-case baseline (single-draw per case; reps waived per measurement-integr
|
|
|
157
157
|
first draw matched expected — see the 2026-07-13 subagent-invocations log entry), not a calibrated
|
|
158
158
|
accuracy estimate.
|
|
159
159
|
|
|
160
|
+
### 3-a. What is born, and what it must be able to do on day one (operator-forged, 2026-08-09)
|
|
161
|
+
|
|
162
|
+
**We are not raising a person. We are shipping a harness that does one thing well.** The founding
|
|
163
|
+
image is a calf or a foal: it is born in a laboratory sense — brand new, thin, nowhere near an adult
|
|
164
|
+
— but it **stands and walks in the place where it was born.** That is the incubator's bar, and it is
|
|
165
|
+
much lower and much clearer than "finished."
|
|
166
|
+
|
|
167
|
+
```
|
|
168
|
+
depth / density altricial — like an infant. Filled in only by real use. Takes a long time.
|
|
169
|
+
basic locomotion precocial — like a calf. Works from the moment it is set down.
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
The two axes are independent, and confusing them is what made this look far away. Aiming at an adult
|
|
173
|
+
(a complete judgment circuit at birth) is **not merely slow — it is unreachable**, because density is
|
|
174
|
+
supplied by usage that has not happened yet. Aiming at a calf is reachable today.
|
|
175
|
+
|
|
176
|
+
**Operational form**: *born walking* = on the first run, with the user adding nothing, the thing
|
|
177
|
+
produces something useful. This maps onto the existing rungs without inventing a scale —
|
|
178
|
+
`🔵 RC` = it stood up in the lab; `🟢 REALIZED` = it walked outside.
|
|
179
|
+
|
|
180
|
+
**The opposite of this doctrine has a name we already use: `built-but-not-wired`.** A harness that
|
|
181
|
+
was born but does not walk is one whose parts exist and whose call sites do not — measured instances
|
|
182
|
+
exist (a field harness with a judgment-circuit file and **zero callers**; a sibling meta-harness with
|
|
183
|
+
none at all). So *"born walking"* is not a metaphor about vitality; it is the engineering claim that
|
|
184
|
+
**wiring is part of the birth**, not a follow-up task. Being born and running are different events,
|
|
185
|
+
and the incubator is answerable for the second.
|
|
186
|
+
|
|
187
|
+
#### What the seed contains — a coordinate system, not a declaration
|
|
188
|
+
|
|
189
|
+
The seed is **not** an identity sentence. `"You are a world-class QA expert"` is an artifact of the
|
|
190
|
+
prompt-engineering era and is actively harmful here: the 105-run measurement scored a bare identity
|
|
191
|
+
declaration as a **net loss on the weak tier** (removing it recovered +0.67), while a judgment
|
|
192
|
+
circuit gained on the frontier tier. Told only *what it is*, a newborn harness still does not know
|
|
193
|
+
what to do, and the gaps show up as arbitrary decisions.
|
|
194
|
+
|
|
195
|
+
What a newborn actually needs is closer to *how to see, how to walk, how to speak*:
|
|
196
|
+
|
|
197
|
+
| Layer | What it fixes | Note |
|
|
198
|
+
|---|---|---|
|
|
199
|
+
| **Seeing** | what counts as a signal at all | inputs — without this the circuit has nothing to run on |
|
|
200
|
+
| **Judging** | success · which way to lean under uncertainty · out of scope · never | = the `judgment-circuit` definition |
|
|
201
|
+
| **Speaking** | how it reports, what shape its output takes | outputs |
|
|
202
|
+
| **Walking** | how it actually executes | wiring, call sites |
|
|
203
|
+
|
|
204
|
+
Shipping the middle layer alone is the common failure: the circuit is present and has no input or
|
|
205
|
+
output attached, which is exactly the zero-caller symptom above. **Form is machine-checkable today**
|
|
206
|
+
(`scripts/judgment_circuit_lint.sh` — branch rules, self-sealing, conflict resolution, lean
|
|
207
|
+
direction, mandated shape; FH's own `CLAUDE.md` measures `CIRCUIT 4/5`). Density is not, and should
|
|
208
|
+
not be given a scale yet — see 3-a-2.
|
|
209
|
+
|
|
210
|
+
#### 3-a-1. Field ⊥ meta — and meta is out of this incubator's scope
|
|
211
|
+
|
|
212
|
+
The two kinds have **opposite profiles**, which is why one method cannot birth both:
|
|
213
|
+
|
|
214
|
+
```
|
|
215
|
+
field harness hard to birth (design · seed · wiring) │ walks on day one precocial
|
|
216
|
+
meta harness easy to birth (a declaration starts one) │ needs endless tending altricial
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
A meta-harness cannot clear a bar that reads *"walks on day one"* — not because it is worse, but
|
|
220
|
+
because unbounded growth is its point. FH itself is the standing evidence: it is tended continuously,
|
|
221
|
+
by design. **Therefore a meta-harness candidate is not a chamber candidate.**
|
|
222
|
+
|
|
223
|
+
⚠️ **Retrospective observation, not a finding.** Re-reading the run ledger along this axis: of the
|
|
224
|
+
KILLed candidates, those aimed at chamber-internal metering, hub-internal orchestration, cluster
|
|
225
|
+
wizardry and org relay are all **meta**-shaped, while the single EMIT (`forge-wiki`) is **field**-shaped
|
|
226
|
+
— a tool that does one thing. If that holds, several KILLs were not the chamber being strict but
|
|
227
|
+
**the wrong kind of candidate entering it**. The classification was made *after the fact* by the same
|
|
228
|
+
session that proposed the axis, over n=9; it is a hypothesis to pre-register and predict against, not
|
|
229
|
+
a result. The way to test it is to fix the classification first and call the next run before its
|
|
230
|
+
verdict — the ordering witness now makes that checkable.
|
|
231
|
+
|
|
232
|
+
**Proposed consequence (not yet applied)**: give the chamber's entry reason a field/meta axis and
|
|
233
|
+
retire meta candidates as **`NOT-APPLICABLE`** rather than `KILL`. Today both land in the same bucket,
|
|
234
|
+
so when the ledger says *"the chamber screens"* it is summing two different events.
|
|
235
|
+
|
|
236
|
+
#### 3-a-2. Density is measured by comparison, never by an absolute scale
|
|
237
|
+
|
|
238
|
+
Density — how filled-in a circuit is — has no honest unit. Counting clauses, counting cases, or
|
|
239
|
+
counting how often the circuit answers all measure different things, and picking one invites the
|
|
240
|
+
failure where a metric scores presence instead of the relation it was meant to capture. The way
|
|
241
|
+
around it is to **not define the unit**: clone versions (seed only / seed + some usage / seed + more)
|
|
242
|
+
and run them against one task set **in parallel**, then read the *shape of the curve* rather than any
|
|
243
|
+
version's score. **Where the curve flattens is the interesting point** — that plateau is the practical
|
|
244
|
+
floor for "enough of a soul."
|
|
245
|
+
|
|
246
|
+
Two conditions carry over from the decorrelation work: the clones must be **independent** (run
|
|
247
|
+
sequentially in one context and the earlier one bleeds into the later), and the **scorer must be a
|
|
248
|
+
different party than the forger**.
|
|
249
|
+
|
|
250
|
+
🟥 **Named limit — synthesized history yields synthesized density.** Density is defined as accruing
|
|
251
|
+
from *real* use; injecting simulated usage into a clone measures something else, and it can fill in a
|
|
252
|
+
different direction than real use would. So the question this experiment can answer is narrowed on
|
|
253
|
+
purpose: **not** "does the soul grow?" but **"if it grows, where does it plateau?"** That is enough
|
|
254
|
+
for a minimum-condition verdict and does not overclaim.
|
|
255
|
+
|
|
256
|
+
Internal version comparison gives a **growth curve**; comparison against an outside harness gives an
|
|
257
|
+
**absolute position**. Both are relative, but their reference points differ — and they can share one
|
|
258
|
+
task set, which lets a standing external dominance pre-registration ride along instead of waiting.
|
|
259
|
+
|
|
160
260
|
### 3-b. The nursery also verifies what it births
|
|
161
261
|
|
|
162
262
|
The incubator's arc does not end at emission: FH **reviews, accelerates, and verifies** the harnesses
|
|
@@ -30,7 +30,9 @@ install plan, and gate every install.**
|
|
|
30
30
|
But a **live one-command autonomous simulate→EMIT of a field harness is NOT yet a capability**:
|
|
31
31
|
step-4 persona dispatch is human/Claude-driven (bash cannot spawn the isolated Agents — the
|
|
32
32
|
honest muscle boundary), the EMIT terminus is HITL, and **EMIT has never fired — the ledger's
|
|
33
|
-
real runs are honest KILLs** (the chamber
|
|
33
|
+
real runs are honest KILLs** (the chamber overwhelmingly *screens* — 9 full runs, 8 KILL, 1 EMIT, hand-counted 2026-08-08; it has
|
|
34
|
+
birthed once, but that run left no intent/budget/persona artifacts so the formal flow is not what
|
|
35
|
+
produced it, and the first end-to-end formal run KILLed). So today this
|
|
34
36
|
branch = a one-line HITL recommendation to run the chamber (`chamber_run.sh`), then fall back to
|
|
35
37
|
Full-Harness Mode §6 (`auto_project_mapping.md`) for the actual onboarding; the runner gates and
|
|
36
38
|
records a human-driven run — it must **not** be presented as a push-button autonomous emit. The
|