@chrono-meta/fh-gate 2.0.1 → 2.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/rules/fh_4axis_gate.md +77 -1
- package/.claude-plugin/marketplace.json +3 -3
- package/AGENTS.md +26 -3
- package/CLAUDE.md +240 -4
- package/knowledge/shared/harness-core/capability_composition_contract.md +68 -0
- package/knowledge/shared/harness-core/fh_three_layer_canon.md +38 -12
- package/knowledge/shared/harness-core/field_verdict_crossfamily_gate.md +115 -3
- package/knowledge/shared/harness-core/harness_incubator_doctrine.md +82 -0
- package/knowledge/shared/harness-core/harness_terminal_correlation_and_recommendations.md +85 -10
- package/knowledge/shared/harness-core/ship_readiness_gate.md +241 -1
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +102 -0
- package/knowledge/shared/rules/sister_asset_protocol.md +12 -0
- package/package.json +7 -2
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-commons/agents/quench-challenger.md +6 -1
- package/plugins/fh-commons/skills/ko-tech-writer/fixtures/known_negative.md +19 -0
- package/plugins/fh-commons/skills/ko-tech-writer/fixtures/known_positive.md +20 -0
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +88 -0
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +14 -2
- package/plugins/fh-meta/skills/steel-quench/SKILL_detail.md +13 -0
- package/plugins/fh-meta/skills/verify-bidirectional/SKILL.md +34 -1
- package/scripts/capability_registry_check.sh +75 -12
- package/scripts/chamber_run.sh +39 -5
- package/scripts/chamber_witness.sh +25 -0
- package/scripts/fh-gate.sh +16 -1
- package/scripts/ko_tech_writer_calibrate.py +127 -0
- package/scripts/package_coverage_check.sh +10 -0
- package/scripts/prepublish_scope_note.sh +139 -0
- package/scripts/publish_freshness_check.sh +21 -2
- package/scripts/selfcheck.sh +32 -6
- package/scripts/test_fh_gate_regressions.sh +21 -9
- package/scripts/test_ko_tech_writer_lanes.sh +61 -0
- package/scripts/test_marker_axes_run_lanes.sh +87 -3
- package/scripts/test_marker_crossfamily_lanes.sh +31 -1
- package/templates/.git-hooks/pre-commit +225 -10
- package/templates/subagent-tally-hook.json +2 -2
|
@@ -71,8 +71,11 @@ tags: [canon, three-layer, engines, identity, verification-axes]
|
|
|
71
71
|
> | ⓔ | — | 첫 실사용 |
|
|
72
72
|
> | ⓕ | — | 되돌림 실측 |
|
|
73
73
|
>
|
|
74
|
-
>
|
|
75
|
-
>
|
|
74
|
+
> ✅ **2026-08-17 부로 마커도 이 배열이다** — 날짜 ≥ 2026-08-17 인 마커는 `axes-run` 이
|
|
75
|
+
> **기호 여섯 글자(ⓐ~ⓕ)**를 요구하고, 그 이전 마커는 옛 ASCII 넷을 그대로 쓴다(소급 강제 없음).
|
|
76
|
+
> **표기법이 곧 배열 선언**이라 감사자가 grep 한 번으로 판별한다. 남은 어긋남은 **하나뿐** —
|
|
77
|
+
> `standpoint:` 값의 enum 을 **검증하는 코드가 아직 0줄**이다. 아래 「옛 배열」 서술은 유예일
|
|
78
|
+
> 이전 마커를 읽을 때만 유효한 이력이다. 옛 배열로 쓴 줄을
|
|
76
79
|
> 기계가 **조용히 다른 축으로 읽으며 오류를 내지 않는다.** 확장은 기존 마커 전부를 무효화하므로
|
|
77
80
|
> **운영자 결정 사항으로 열려 있다** — 그때까지 마커를 쓸 때는 **옛 배열**을 쓰고, 산문으로
|
|
78
81
|
> 논할 때는 §1-a-2 의 6축을 쓴다. 이 어긋남을 **알고** 쓰는 것이 현 규율이다.
|
|
@@ -197,8 +200,9 @@ measurement_needs_control]])에 있다. 표본이 지지하는 건 "바깥 하
|
|
|
197
200
|
> → 현 **ⓔ**, 옛 ⓓ(되돌림) → 현 **ⓕ**. 표는 §1 요약줄 아래에 있다. 이 절의 문자를 그대로
|
|
198
201
|
> 6축 문맥에 옮기면 **조용히 다른 축을 지목하게 된다.**
|
|
199
202
|
>
|
|
200
|
-
> ⚠️
|
|
201
|
-
>
|
|
203
|
+
> ⚠️ **이 절의 배열은 2026-08-17 이전 날짜의 마커에만 유효하다** — 그 날짜부터 마커도 §1-a-2 의
|
|
204
|
+
> 6축(기호 키)을 요구한다. 옛 마커를 **읽을** 때는 여기가 맞고, 새로 **쓸** 때는 §1-a-2 가 맞다.
|
|
205
|
+
> 산문/기계 어긋남은 해소됐고, 남은 잔여는 `standpoint:` 값 검증이 없다는 것 하나다.
|
|
202
206
|
|
|
203
207
|
3단계를 오래 «적대검증» 이라고 불렀으나, 그건 **넷 중 하나**다. 각 축은 **다른 것이 틀린
|
|
204
208
|
경우**를 잡고 **서로를 대체하지 못한다**:
|
|
@@ -255,16 +259,38 @@ measurement_needs_control]])에 있다. 표본이 지지하는 건 "바깥 하
|
|
|
255
259
|
다른 축으로 넘어갔다). 반대로 「남의 레포가 과거에 폐기한 규칙」은 **도구로도 못 가져온다** —
|
|
256
260
|
그 프로젝트의 리뷰 이력에 접근할 이유가 없기 때문이다. 거기가 ⓓ가 남는 자리다.
|
|
257
261
|
|
|
258
|
-
|
|
262
|
+
✅ **기계층이 6축으로 확장됐다 (2026-08-17, 운영자 승인).** 아래는 그 상태와 **남은 잔여 하나**다.
|
|
259
263
|
```
|
|
260
|
-
axes-run:
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
+
axes-run: ⓐ~ⓕ 마커 날짜 ≥ 2026-08-17 이면 pre-commit 이 **기호 여섯 글자를 요구**한다
|
|
265
|
+
axes-run: a/b/c/d 그 이전 날짜의 마커는 옛 네 글자 그대로 (소급 강제 없음)
|
|
266
|
+
standpoint: ⓑ 는 **자기 필드**가 정본이고, `axes-run` 은 `ⓑ=→standpoint` 포인터만 든다
|
|
267
|
+
(이중 기록 회피). 포인터가 있는데 그 줄이 비면 **죽은 포인터로 차단**
|
|
268
|
+
⚠️ 🟥 **값 자체의 enum 검증은 여전히 0줄** — 이게 남은 잔여다
|
|
269
|
+
ⓓ 3자 대면 ✅ 기록할 자리가 생겼다 (`ⓓ=`)
|
|
264
270
|
```
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
271
|
+
🟥 **판별자는 마커 파일명의 날짜다**(`< 2026-08-17` = 옛 4축). 표기를 가른 이유는 미관이 아니라
|
|
272
|
+
실측이다: 두 배열은 **같은 글자가 다른 축을 가리킨다**(옛 `b`=첫실사용 → 지금 **ⓔ**, 옛 `d`=되돌림
|
|
273
|
+
→ 지금 **ⓕ**). 같은 글자를 쓰면서 의미만 바꾸면 **조용히 다른 축으로 읽히고 오류가 안 난다** —
|
|
274
|
+
`not found ≠ 0` 의 형제다.
|
|
275
|
+
|
|
276
|
+
⚠️ **다만 「표기법이 배열을 선언한다」는 명제는 반증됐다** — 이 절의 초판이 그렇게 적었고, 같은 날
|
|
277
|
+
ⓓ3자대면 축이 코퍼스를 손으로 세어 반례를 냈다: `axes-run` 보유 **53건 중 기호 키 4 · 혼용 1**,
|
|
278
|
+
그리고 **기호 4건 중 2건이 2026-08-10 자이면서 옛 4축 의미**(`ⓑ 첫실사용` · `ⓓ 되돌림`)다. 훅은
|
|
279
|
+
그 셋을 안 읽으니 커밋에는 영향이 없고, 틀린 답을 받는 쪽은 **감사자**다. 표기 정렬은 «앞으로 쓸
|
|
280
|
+
때 옛 줄을 복붙하지 않게» 하는 값이지 **소급 판별자가 아니다**.
|
|
281
|
+
|
|
282
|
+
⚠️ 그리고 **날짜 비교 자체가 프로덕션에서 도달 불가 분기**다 — 훅 호출부가 마커 경로를 `${TODAY}`
|
|
283
|
+
로 구성하므로 `mdate` 는 항상 오늘이다. **기존 마커를 지킨 것은 그 상수가 아니라 경로 구성**이고,
|
|
284
|
+
8일 전 `crossfamily:` 확장이 컷오프 없이 성립한 이유도 같다. 숨기지 않고 적는다.
|
|
285
|
+
|
|
286
|
+
**승인의 전제였던 「기존 마커 전부 무효화」는 반증됐다(2026-08-17 실측).** 훅은
|
|
287
|
+
`.axes_23_passed_{브랜치}_{오늘}.marker` **한 개만** 검증하고 과거 마커 재검증 경로가 없다:
|
|
288
|
+
디스크 마커 **190개** 중 커밋이 막히는 건수는 **0**, 실제 비용은 유예일 이후 **51건이 「어느
|
|
289
|
+
배열인지 읽을 수 없다」**는 것 — 파손이 아니라 **미측정**이다. 그래서 처방이 「고친다」가 아니라
|
|
290
|
+
**「어느 배열인지 말하게 만든다」**(표기 정렬)가 됐다. 이식성은 bash 3.2 + BSD grep, 4개 locale
|
|
291
|
+
× 6키 전부 HIT · known-negative 0 으로 선행 실측했다(Linux/GNU grep 은 CI 가 첫 측정).
|
|
292
|
+
형식 정본 = `.claude/rules/fh_4axis_gate.md §Marker axis fields` · 픽스처 25레인 =
|
|
293
|
+
`scripts/test_marker_axes_run_lanes.sh`.
|
|
268
294
|
|
|
269
295
|
> **한계 — 인용 전에 읽어라**: **n=1**(한 산출물 · 한 세션 · 한 저자). 축의 «비중복»이 구조적인지
|
|
270
296
|
> 그날 우연인지는 **미측정**이다. 축은 **사전 등록되지 않았다**(ⓓ는 세션 도중에 생겼고 그 뒤에
|
|
@@ -221,13 +221,71 @@ marker, never replacing it. Closed enum, same discipline as `crossfamily:`'s thr
|
|
|
221
221
|
not / did not / did not look* split (a free-prose field would let an unrun check read as a clean
|
|
222
222
|
pass):
|
|
223
223
|
|
|
224
|
+
**🟥 DECIDE IN THIS ORDER — first match wins. Do not pick by matching a description.**
|
|
225
|
+
|
|
226
|
+
```
|
|
227
|
+
Q1. Did anything EXECUTE in the target's repo — a command, a script, a suite?
|
|
228
|
+
NO, I only read files → tier1b(<harness>) STOP.
|
|
229
|
+
NO, I did not touch its repo → tier1 STOP.
|
|
230
|
+
YES → continue to Q2
|
|
231
|
+
⚠️ tier2 AND tier2b BOTH REQUIRE EXECUTION. This question is first precisely so the
|
|
232
|
+
next one cannot be used to reason backwards into "tier2 must be the non-executing rung."
|
|
233
|
+
Q2. Was the target's own LOCAL / gitignored wiring visible (settings, consent bindings,
|
|
234
|
+
node-local state) — i.e. its real installed runtime, not a bare clone?
|
|
235
|
+
NO (bare clone, tracked content only) → tier2(<harness>)
|
|
236
|
+
YES (the target's real runtime) → tier2b(<harness>)
|
|
237
|
+
Q3. Was it run by a DIFFERENT operator of the target harness, not you?
|
|
238
|
+
YES → tier3(<harness>) (supersedes Q2)
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
**Why the procedure exists rather than more definition.** `tier2` vs `tier2b` is **wiring
|
|
242
|
+
visibility**, NOT execution-vs-reading — both execute. But `tier2b`'s gloss names "the target's
|
|
243
|
+
real runtime," which reads as *"tier2b is the execution rung"*, and a reader then infers that
|
|
244
|
+
`tier2` must therefore be the non-executing one.
|
|
245
|
+
|
|
246
|
+
🟥 **RETRACTED (2026-08-17) — the sim evidence formerly cited here is withdrawn.** This paragraph
|
|
247
|
+
read: *"after all three were corrected to say EXECUTED CODE, two independent blind Sonnet reps STILL
|
|
248
|
+
graded a pure cold-read `tier2`"*, and quoted one rep's reasoning verbatim. Those runs had
|
|
249
|
+
**`tool_uses: 0`** — the agents never opened a file, so the quoted reasoning is a cold guess about
|
|
250
|
+
text it did not read, and the grades measure nothing
|
|
251
|
+
(`tracks/_meta/fh_completed_2026-08-16.md:690`). The live re-run **inverted** the result at
|
|
252
|
+
**reps=1**, below this repo's `reps>=3` bar. **Neither direction is established**; do not restore
|
|
253
|
+
the numbers and do not cite the inversion either.
|
|
254
|
+
|
|
255
|
+
What remains, and it is enough to justify ordering the questions: the enum's *wording* really does
|
|
256
|
+
place the execution claim on `tier2b`'s line, so a reader can reach "then `tier2` is the
|
|
257
|
+
non-executing one" **by the text alone** — that is a property of the text, checkable by reading it,
|
|
258
|
+
and it needs no sim. Ordering the questions removes the inference instead of arguing with it.
|
|
259
|
+
|
|
224
260
|
```
|
|
225
261
|
tier1 content-only review — no standpoint decorrelation (the default
|
|
226
262
|
unless upgraded; NOT itself a failure, most changes have no target
|
|
227
263
|
standpoint to borrow)
|
|
228
|
-
|
|
229
|
-
(
|
|
230
|
-
|
|
264
|
+
tier1b(<target-harness>) STATIC standpoint read — the reviewer read the TARGET's own files
|
|
265
|
+
(cold, from the target's repo, without the author's framing) and
|
|
266
|
+
adjudicated the change against them, but executed NOTHING. Added
|
|
267
|
+
2026-08-16 because its absence was actively harmful: a run of exactly
|
|
268
|
+
this shape was recorded as `tier2`, since tier1 undersold it and no
|
|
269
|
+
nearer value existed. A missing rung does not stay empty — it gets
|
|
270
|
+
filled by the next one up. Real but weak: see «execution is the
|
|
271
|
+
load-bearing half» below before crediting it.
|
|
272
|
+
tier2(<target-harness>) peer-simulated — the reviewer EXECUTED CODE in the TARGET's own repo
|
|
273
|
+
(a local clone, real content) and observed the result.
|
|
274
|
+
🟥 DISCRIMINATOR — «reading the target's real files is NOT this rung».
|
|
275
|
+
🟥 RETRACTED (2026-08-17): the «two blind Sonnet sims both graded a
|
|
276
|
+
cold-read tier2 … 0/2» measurement that stood here is WITHDRAWN —
|
|
277
|
+
tool_uses: 0, the agents never opened a file, so the grades and the
|
|
278
|
+
quoted reasoning measure nothing (fh_completed_2026-08-16.md:690).
|
|
279
|
+
The live re-run INVERTED it at reps=1, below this repo's reps>=3 bar.
|
|
280
|
+
Neither direction is established; restore no number here.
|
|
281
|
+
The AMBIGUITY it was cited for is still checkable WITHOUT a sim: the
|
|
282
|
+
earlier wording said «instantiated/ran the target's own repo», and
|
|
283
|
+
«ran» admits «operated within / engaged with», which a read satisfies.
|
|
284
|
+
That is a property of the text — read the two lines and see it.
|
|
285
|
+
THE TEST, and it is mechanical: **name the command you executed and
|
|
286
|
+
the output you observed.** Cannot name one → `tier1b`, always. An
|
|
287
|
+
agent that read files, however cold and however many, executed
|
|
288
|
+
nothing. Closes shared-body-path defects; a BARE clone cannot see
|
|
231
289
|
the target's gitignored local wiring (settings, consent bindings,
|
|
232
290
|
node-local state) — that gap is inherent to a clone, not a defect in
|
|
233
291
|
a given run. Named exception, not a loophole: if the reviewer's own
|
|
@@ -259,6 +317,60 @@ DEGRADED_NOT_RUN did not — target was accessible, standpoint revi
|
|
|
259
317
|
UNKNOWN did not look — applicability itself was never assessed
|
|
260
318
|
```
|
|
261
319
|
|
|
320
|
+
**🟥 Execution is the load-bearing half — a static standpoint read is largely subsumed by the other
|
|
321
|
+
axes (operator decision, 2026-08-16).** The reason to pay for a standpoint at all is not that
|
|
322
|
+
someone re-read the diff from a different chair; it is that **the target harness was actually made
|
|
323
|
+
to run.** Operator's framing, verbatim: *"그 하네스의 입장에서 정적리뷰하는 것만으로도 뭔가 잡을
|
|
324
|
+
수야 있겠지만 그건 다른 검증축으로도 커버가 아마 가능하지 않을까. 진짜로 중요한 건, 그 하네스
|
|
325
|
+
입장에서 돌려봐서 구동이 되는지를 로컬에서 완벽하게 확인하는 것."* A static read competes with
|
|
326
|
+
cross-family review (§3) and isolated grounding for the same defect classes and mostly loses —
|
|
327
|
+
those axes are cheaper and already routine. Execution has no substitute, because the class it
|
|
328
|
+
catches is *unreachable by reading*: the target's own environment differs (absent files, different
|
|
329
|
+
resolution order, a lane that has never once run there).
|
|
330
|
+
|
|
331
|
+
**Measured the same day this was written, on one delta (pmh-dev PR #72)**:
|
|
332
|
+
|
|
333
|
+
| Arm | Found |
|
|
334
|
+
|---|---|
|
|
335
|
+
| `tier1b` static standpoint read (isolated agent, target's own files, cold) | **1** — a 2-tier-vs-3-tier root-path resolution mismatch |
|
|
336
|
+
| Running the target's own `selfcheck.sh` to completion, locally | **2 more**, both invisible to any read: a test fixture keying on a file that does not exist in that repo, and an npm-shipping check whose premise is inapplicable there. One of them printed **neither `FAIL` nor `❌`** anywhere in its output (it used its own vocabulary, `INSTRUMENT ERROR`) and needed a `bash -x` trace to locate — a defect that is *structurally* undiscoverable by reading, since the reader must already know which string to look for. |
|
|
337
|
+
|
|
338
|
+
n=1 delta, same operator, same session — reported as a directional observation, not a rate. It is
|
|
339
|
+
recorded because it is the first time the two arms were run *separately on the same change* and
|
|
340
|
+
their yields could be attributed. ⚠️ **Do not read the table as "static review is worthless"** — it
|
|
341
|
+
found a real defect that shipped a fix. Read it as: *static is the half that has substitutes;
|
|
342
|
+
execution is the half that does not.*
|
|
343
|
+
|
|
344
|
+
**🟥 What dynamic standpoint review DOES and DOES NOT subsume (operator + governor, agreed
|
|
345
|
+
2026-08-16 after both arms were measured).** Operator's proposal: *"FH가 기여하는 다른 레포들도 다
|
|
346
|
+
마찬가지일 테니, 그쪽 입장에서의 동적 리뷰는 소넷을 상주시켜 돌려보는 거지. 이건 굳이 짓지 않아도
|
|
347
|
+
동적 입장리뷰만 잘 시킨다면 알아서 커버될 거라고 생각해."* Agreed, with one boundary — and the
|
|
348
|
+
agreement is evidenced, not deferential:
|
|
349
|
+
|
|
350
|
+
- **SUBSUMES: building per-target instruments.** Do not write a new scanner for each contributed
|
|
351
|
+
repo, language, or defect class. Execution is **stack-agnostic**; an instrument is not. Measured
|
|
352
|
+
the same day: `degrade_direction_scan.sh` is bound to shell/python *and* to verdict vocabulary —
|
|
353
|
+
a faithful Python port of a real upstream defect scores CLEAN, because the fail-open value was
|
|
354
|
+
ordinary *data*, not a verdict token. An instrument carries a scope boundary into every repo it
|
|
355
|
+
visits; running the target's own suite does not.
|
|
356
|
+
- **DOES NOT SUBSUME: adversarial reading.** **Execution is a detector, not a generator.** It
|
|
357
|
+
answers *does this break · does it fire · does the environment differ*; it cannot answer *is
|
|
358
|
+
there a defect class nobody has instrumented yet*. The upstream `clawd-on-desk` PR #888 defect
|
|
359
|
+
was **not** found by running that repo's tests — it was found by reading and conceiving the
|
|
360
|
+
unreadable-file case, after which a test was written. And where the target ships no runnable
|
|
361
|
+
suite, execution has nothing to run at all.
|
|
362
|
+
|
|
363
|
+
**The measured split, same delta, same day**: static read → **1** finding · execution → **2** more.
|
|
364
|
+
Neither arm was zero, which is the whole result. So both run: reading generates the hypothesis,
|
|
365
|
+
execution confirms or refutes it, **and the hypothesis that survives becomes a test left behind in
|
|
366
|
+
the target** — which is exactly what our own upstream contribution did (the survivor-lane pattern,
|
|
367
|
+
`tracks/_meta/fh_signal_2026-08-16_clawd-survivor-lane-air.md`).
|
|
368
|
+
|
|
369
|
+
**Consequence for the marker**: a `tier2`/`tier2b`/`tier3` claim asserts that something was RUN. If
|
|
370
|
+
the review only read, the honest value is `tier1b` — and since `tier1b` is explicitly the weak rung,
|
|
371
|
+
recording it truthfully is what surfaces that the execution arm is still owed. (This rule exists
|
|
372
|
+
because it was broken on the day it was written: see the `tier1b` entry above.)
|
|
373
|
+
|
|
262
374
|
**Residual this enum split names rather than hides**: for a harness pair with one shared human
|
|
263
375
|
operator (this repo and a sibling field harness the same operator also runs), `tier3` is either
|
|
264
376
|
unreachable or collapses into "the same author ran it in the other repo" — which is exactly the
|
|
@@ -277,6 +277,88 @@ recommended path is *simulate inside the chamber first, then emit the initial mo
|
|
|
277
277
|
build-immediately. (Wired as a recommendation branch in `CLAUDE.md §Onboarding / Acceleration
|
|
278
278
|
Autopilot`; build-immediately remains correct for clear, small, low-failure-cost projects.)
|
|
279
279
|
|
|
280
|
+
### 3-c. The EMIT bar — «walks on its own» (operator-forged, 2026-08-16)
|
|
281
|
+
|
|
282
|
+
> **Relationship to §3's EMIT-worthiness criterion — two different questions, do not merge them.**
|
|
283
|
+
> That one screens candidates *going in*: is this worth emitting at all (net-new · artifact-shaped ·
|
|
284
|
+
> precision-adequate). This one judges a candidate *coming out*: is it ready to stand. A candidate can
|
|
285
|
+
> clear the first and fail this one (worth building, not yet alive), or clear this one and fail the
|
|
286
|
+
> first (alive, but a reinvention). Both are required; neither substitutes.
|
|
287
|
+
|
|
288
|
+
Until now the chamber's EMIT threshold was **implicit**, and that is why its record (9 runs,
|
|
289
|
+
8 KILL, 1 EMIT) could not be read: with no stated bar, "screened well" and "over-screened" are
|
|
290
|
+
the same picture. An instrument whose output is almost always one value is indistinguishable
|
|
291
|
+
from an instrument stuck on that value — the dead-control signature this repo keeps re-finding.
|
|
292
|
+
The bar is therefore stated, and stated **low on purpose**:
|
|
293
|
+
|
|
294
|
+
> **A candidate EMITs when it can stand on its own — not when it is good.**
|
|
295
|
+
>
|
|
296
|
+
> ① It **runs standalone**, with no defect severe enough to prevent execution.
|
|
297
|
+
> ② Every necessary function **fires at least weakly**. Firing, not performing.
|
|
298
|
+
> ③ It is in a state where it can **fill itself in through back-and-forth with a human**.
|
|
299
|
+
> ④ **Weakness is not a KILL.** Only inability-to-run is.
|
|
300
|
+
|
|
301
|
+
Newborn-foal semantics: it staggers, and it is standing. Shortcomings are expected and are the
|
|
302
|
+
*point* — an emitted harness is raised, not delivered finished.
|
|
303
|
+
|
|
304
|
+
**The mechanical discriminator, and it is already in hand.** ② is not a judgment call, because
|
|
305
|
+
"fired weakly" and "did not fire" separate cleanly at the exit boundary:
|
|
306
|
+
|
|
307
|
+
```
|
|
308
|
+
dead — dies before doing anything e.g. a gate that aborts on its own version line (rc=1)
|
|
309
|
+
alive — ran and returned a TYPED verdict e.g. the same gate returning HARNESS_ERROR (rc=10)
|
|
310
|
+
```
|
|
311
|
+
|
|
312
|
+
`HARNESS_ERROR` is a *bad* result and it satisfies ②: the unit ran and said why it could not
|
|
313
|
+
finish. `rc=1`-before-anything is not a weak firing, it is absence. This is the same
|
|
314
|
+
`not found ≠ 0` line the rest of this repo runs on, applied to birth.
|
|
315
|
+
|
|
316
|
+
**Where the judgment must happen: outside the author's environment.** Measured 2026-08-16 — a
|
|
317
|
+
gate suite read 31/31 green in the harness that wrote it and was structurally dead in the sibling
|
|
318
|
+
that inherited it, because the author's repo happened to supply a file the code assumed. An EMIT
|
|
319
|
+
verdict rendered inside the incubator's own environment cannot see that class at all. So the EMIT
|
|
320
|
+
check is run the way a recipient would run it: a clean checkout, none of the author's local state,
|
|
321
|
+
each declared function exercised once.
|
|
322
|
+
|
|
323
|
+
**Calibration owed, not claimed.** This bar makes the chamber testable for the first time; it does
|
|
324
|
+
not retroactively validate the existing ledger. The 8 KILLs were decided against an unstated bar,
|
|
325
|
+
so whether each died on "cannot run" (correct) or on "weak" (over-screen) is **unknown from the
|
|
326
|
+
record**. Known-pair to run before the ledger is cited as evidence of anything: a known-positive
|
|
327
|
+
(a field harness the operator judges already stands on its own) must EMIT; a ledger KILL whose
|
|
328
|
+
recorded reason is explicitly *cannot run* must KILL. If the known-positive is KILLed, the
|
|
329
|
+
instrument over-screens and every prior KILL is weakened by that much.
|
|
330
|
+
|
|
331
|
+
### 3-d. What separates incubating from boosting — the AUTHORITY, not the effort (operator, 2026-08-16)
|
|
332
|
+
|
|
333
|
+
Operator, verbatim: *"부스팅보다 인큐베이터의 장점은 그 레포 전체(또는 출하하기 전 상태의 플젝)를
|
|
334
|
+
FH가 감싸서 **모든 것을 만들어내고 재설계한다**에 있어. 그래서 모든 것을 빠삭히 알아야 하고,
|
|
335
|
+
그러면서도 **모든 것을 조작할 수 있는 권한**이 있는 거야. 부스팅과는 다른 지점이 이것이기도 하다."*
|
|
336
|
+
|
|
337
|
+
The two identities are not «small help» vs «big help». They differ in **scope of authority over the
|
|
338
|
+
target**, and the two properties are a matched pair — neither is optional:
|
|
339
|
+
|
|
340
|
+
| | Ⓑ Project Booster | ② Project Incubator |
|
|
341
|
+
|---|---|---|
|
|
342
|
+
| Touches | the point that was asked about | **the whole repo — including redesigning what already works** |
|
|
343
|
+
| Must know | enough for that point | **the entire target, thoroughly** |
|
|
344
|
+
| Reads the codebase | on demand, to answer | **as a precondition, before proposing anything** |
|
|
345
|
+
|
|
346
|
+
**Why the pair is load-bearing**: total authority without total knowledge is vandalism, and total
|
|
347
|
+
knowledge without authority is a booster that reads a lot. An incubation session that has not read
|
|
348
|
+
the target's design canon, gate code, and existing verification surfaces has not *earned* the second
|
|
349
|
+
column — and the correct move there is to read first, not to scope down to a booster-shaped change
|
|
350
|
+
and call it incubation.
|
|
351
|
+
|
|
352
|
+
**Where this bit, same day**: the-bible's ⓐ run scoped itself to «invariant-preserving» — correct as
|
|
353
|
+
a *choice the operator made for that request*, and it was then mistaken for the identity's own reach.
|
|
354
|
+
It is not. ⓐ/ⓑ is the operator setting the blast radius for one job; the incubator's standing
|
|
355
|
+
authority is the full repo either way. A session that reads a per-request scope as the identity's
|
|
356
|
+
ceiling will never propose the redesign that was the point of incubating.
|
|
357
|
+
|
|
358
|
+
**Relationship to 3-a's «day one» bar**: 3-a says what the *born thing* must do. This says what the
|
|
359
|
+
*nursery* is allowed to do to it while it is still inside. The first is an exit condition; the second
|
|
360
|
+
is a working posture.
|
|
361
|
+
|
|
280
362
|
## 4. Compose ∪ disrupt — two operating modes over other harnesses
|
|
281
363
|
|
|
282
364
|
| Mode | What | FH mechanism |
|
|
@@ -12,7 +12,7 @@ tags: [harness, terminal, cmux, orca, architecture, ergonomics, governor-pattern
|
|
|
12
12
|
본 분석 보고서는 **AI 에이전트 하네스(Agent Harness)의 발달 수준과 터미널 환경(Multiplexer / Sandbox) 간의 상관관계 및 직교적 계층 구조**를 규명하고, 신뢰성 기반의 최적 운용 모델을 제안합니다.
|
|
13
13
|
|
|
14
14
|
* **핵심 명제**: **"신뢰된 로컬 개발 환경에서는, 터미널 인프라에 의존하지 않는 얇은 하네스(Thin Harness)로 충분하다."** (전칭 «가장 뛰어난» 을 뺐다 — 비교 모집단을 잰 적이 없다)
|
|
15
|
-
* **결론**: 에이전트 자체의 오케스트레이션 및 자가 검증(**4축 게이트** — Axis 1 `regression_guard.sh` · Axis 2 `steel-quench` · Axis 3 `phantom-quench` · Axis 4 `edit-manifest`, FH 자산 변경 시
|
|
15
|
+
* **결론**: 에이전트 자체의 오케스트레이션 및 자가 검증(**4축 게이트** — Axis 1 `regression_guard.sh` · Axis 2 `steel-quench` · Axis 3 `phantom-quench` · Axis 4 `edit-manifest`, FH 자산 변경 시 커밋 경계에서 훅으로 강제 — ⚠️ **단 그 훅은 기본 활성이 아니다**: `core.hooksPath` 를 배선해야 발동하며, 갓 클론한 레포에서는 `templates/` 안의 템플릿일 뿐이다)이 성숙한 환경에서는 터미널 멀티플렉서(`cmux`)를 **가벼운 UI 껍데기**로 활용하고 **거버너(Governor) 에이전트 위임 모델**을 적용하는 것이 인지 부하를 최소화합니다. 단, 비신뢰 코드 실행 및 파괴적 부작용이 수반되는 과제는 **위험도 기반 샌드박싱(Risk-Driven Sandboxing)**에 따라 **실제 커널 경계를 가진 샌드박스**(Firecracker · Kata · gVisor · Docker Sandboxes 계열)를 하부에 계층화하여 방어합니다. 🟥 초판은 이 자리에 `Orca` 를 적었고 그것은 틀렸다(§4 최종추천 2 · §명명된 잔여 1).
|
|
16
16
|
|
|
17
17
|
---
|
|
18
18
|
|
|
@@ -20,7 +20,7 @@ tags: [harness, terminal, cmux, orca, architecture, ergonomics, governor-pattern
|
|
|
20
20
|
|
|
21
21
|
과거 AI 에이전트 운용 초기에는 에이전트의 오탐, 환각, 호스트 파일시스템 오염을 막기 위해 **외부 인프라(터미널 멀티플렉서, 무거운 Docker/VM 샌드박스)**로 에이전트를 감싸고 통제했습니다.
|
|
22
22
|
|
|
23
|
-
그러나 메타 하네스(`forge-harness`) 체계가 도입되면서, 에이전트 스스로 하위 작업을 생성·위임하고, 검증(4축 게이트)과 마감 피어 동기화(`knowledge/shared/rules/multi_session_close_protocol.md` — **마감 순서와 peer 델타 append 규율**이지, 메모리 색인 락이 아니다)를 소프트웨어적 통제 레이어에서 수행할 수 있게 되었습니다.
|
|
23
|
+
그러나 메타 하네스(`forge-harness`) 체계가 도입되면서, 에이전트 스스로 하위 작업을 생성·위임하고, 검증(4축 게이트)과 마감 피어 동기화(`knowledge/shared/rules/multi_session_close_protocol.md` **(FH 전용 자산 — 형제 하네스에는 없다. 동기화 시 dangling 참조가 된다)** — **마감 순서와 peer 델타 append 규율**이지, 메모리 색인 락이 아니다)를 소프트웨어적 통제 레이어에서 수행할 수 있게 되었습니다.
|
|
24
24
|
|
|
25
25
|
이에 따라 **개발자 UX를 위한 UI 레이어(`cmux`)**, **거버넌스·오케스트레이션 레이어(`FH`)**, 그리고 **OS/커널 보안 레이어**의 역할 분담 및 선택 기준을 규정할 필요성이 제기되었습니다. 🟥 **초판은 세 번째 자리에 `Orca` 를 놓았으나 Orca 는 보안 레이어가 아니다** — Stably AI 의 멀티에이전트 오케스트레이터 데스크톱 앱이고 격리 기전이 **git worktree** 라, `cmux`·`FH` 와 **같은 층의 경쟁재**다.
|
|
26
26
|
|
|
@@ -44,7 +44,7 @@ tags: [harness, terminal, cmux, orca, architecture, ergonomics, governor-pattern
|
|
|
44
44
|
| 비교 축 | Execution Sandbox (Firecracker · gVisor · Docker Sandboxes) | UI/오케스트레이터 (`cmux` · `Orca`) | Protocol-Native Harness (FH 거버너 위임) |
|
|
45
45
|
|---|---|---|---|
|
|
46
46
|
| **계층 역할** | OS/커널 레벨의 물리적 보안 및 자원 격리 | 시각적 탭/창 관리를 위한 UI 껍데기 | 지능적 오케스트레이션 및 소스 앵커링 |
|
|
47
|
-
| **격리 범위** | 파일시스템, 네트워크 포트, 커널, CPU/RAM 쿼터 | 🟥 **«없음»
|
|
47
|
+
| **격리 범위** | 파일시스템, 네트워크 포트, 커널, CPU/RAM 쿼터 | 🟥 **«없음» 은 틀렸다. 다만 «worktree 격리 제공»도 과하다** — 벤더 페이지(<https://cmux.com/>, 2026-08-16 열람) 가 확인해 주는 것은 **SSH 원격 워크스페이스 · 탭별 브랜치/작업디렉터리 · 세션 복원**이고, 오히려 *"strict worktree isolation 을 강제하기보다 동시 작업을 정리하는 데 중점"* 이라고 읽힌다. «local/worktree/SSH 격리»는 **2차 출처발**이며 벤더 페이지에서 확인되지 않았다. ⚠️ 어느 쪽이든 **격리 강도는 미실측** | 서브에이전트 컨텍스트 격리 (별도 컨텍스트 윈도우·요약 반환). **worktree 는 FH 의 기본 격리 수단이 아니다** |
|
|
48
48
|
| **주요 장점** | 비신뢰 코드 실행 시 호스트 시스템 보호(완전 격리는 아니다 — MicroVM 탈출 사례가 실존한다) | 낮은 인지 오버헤드, 빠른 로컬 파일 접근 | 편향 격리(저자 추론을 못 본 채 평가), 컨텍스트 보존, 검증 게이트 |
|
|
49
49
|
| **주요 단점** | 기동 오버헤드, 피어 감지 및 볼륨 바인딩 마찰 | 자원 경합 및 직접적인 호스트 부작용 무방비 | 에이전트의 거버넌스 능력 필요. 🟥 **worktree 를 쓰면 게이트가 깨진다** (아래 §3.2 주의) |
|
|
50
50
|
| **선택 기준** | **비신뢰 코드, 패키지 설치, 파괴적 셸 실행** | **신뢰된 로컬 환경 (1~3개 상위 세션)** | **모든 에이전트 오케스트레이션 및 검증** |
|
|
@@ -70,8 +70,15 @@ tags: [harness, terminal, cmux, orca, architecture, ergonomics, governor-pattern
|
|
|
70
70
|
> 🟥 **worktree 로 위임하지 마라 — 이건 FH 가 실측으로 반대하는 경로다.** `CLAUDE.md §Agent Dispatch
|
|
71
71
|
> Operation` 이 정본이다: ⓐ 문서가 설치를 지시하는 **상대경로 `core.hooksPath`** 형태에서는
|
|
72
72
|
> worktree 안의 훅 사본을 고치면 그 worktree 의 게이트가 무력화된다(실측 `rc=0` — 마커 없는 FH 자산
|
|
73
|
-
> 커밋이 통과). ⓑ 그와 무관하게
|
|
74
|
-
>
|
|
73
|
+
> 커밋이 통과). ⓑ 그와 무관하게 **게이트 증거가 gitignored 라 worktree 로 안 따라간다** — 다만
|
|
74
|
+
> **어디까지 사라지는지는 레포마다 다르다.** 최소 공통은 **Axis 2–3 마커**이고, Axis 4 매니페스트는
|
|
75
|
+
> 레포-국소다. 실측(2026-08-16): FH 는 `.gitignore` 가 `tracks/**` 전면이라 **매니페스트도 사라진다**
|
|
76
|
+
> (`check-ignore` 히트, tracked 0). 형제 하네스 PMH 는 `tracks/*` + `!tracks/_meta/` 라
|
|
77
|
+
> **매니페스트가 추적되고**, 마커만 사라진다. 결론(«worktree 에서 FH 자산을 커밋하지 마라»)은 양쪽
|
|
78
|
+
> 모두에서 성립하지만, **그 이유를 «둘 다 부재»로 적으면 PMH 독자는 자기 레포에서 반증하고 마커
|
|
79
|
+
> 경고까지 같이 버린다.** 각자 확인할 것: `git check-ignore -v tracks/_meta/edit_manifest.yaml`.
|
|
80
|
+
> 🟥 **이 문단은 PMH 입장리뷰가 정정했다** — FH 안에서 읽으면 보편 명제로 보이고, 입장을 바꿔야
|
|
81
|
+
> 국소성이 드러난다. 계열 다양성(codex)은 같은 레포 안에서 읽으므로 이 클래스를 못 잡았다.
|
|
75
82
|
> ⇒ **FH 자산 커밋은 표준 세션에서 한다.** worktree 는 «자율 하위 격리»의 수단이 아니라 게이트 무결성
|
|
76
83
|
> 위험이며, 이 문서의 초판은 그것을 강점으로 서술했다(임포트 심사에서 정정).
|
|
77
84
|
|
|
@@ -96,9 +103,23 @@ graph TD
|
|
|
96
103
|
|
|
97
104
|
## 4. 위험도 기반 샌드박싱 의사결정 트리 (Risk-Driven Decision Tree)
|
|
98
105
|
|
|
106
|
+
🟥 **초판 트리에는 뿌리 분기가 «비신뢰 코드인가» 하나뿐이었고, 그건 개인 머신 입장에서만 충분하다.**
|
|
107
|
+
residency 가 걸린 환경(조직 내부·규제·고객 데이터)에서는 **평범한 작업이 전부 그 분기에서 NO 로 떨어져**
|
|
108
|
+
«얇은 하네스 + UI 층» 으로 직행한다 — 정작 그 환경의 구속조건은 코드 신뢰도가 아니라 **데이터 유출**
|
|
109
|
+
인데도. 형제 하네스 입장리뷰가 지목했고, 뿌리 분기를 하나 앞에 세운다.
|
|
110
|
+
|
|
99
111
|
```
|
|
100
112
|
[ 새로운 작업 오더 ]
|
|
101
113
|
│
|
|
114
|
+
Is the data residency-bound? ← 신설 (입장리뷰)
|
|
115
|
+
(조직 내부·고객·규제 데이터가 관여하는가)
|
|
116
|
+
│
|
|
117
|
+
┌────────────┴────────────┐
|
|
118
|
+
YES NO
|
|
119
|
+
│ │
|
|
120
|
+
[ 환경 승인 도구만 · 도구의 telemetry· │
|
|
121
|
+
상태파일 기록 표면을 먼저 실측 ] │
|
|
122
|
+
│
|
|
102
123
|
Is Execution Dangerous / Untrusted?
|
|
103
124
|
(비신뢰 코드, 외부 패키지, 파괴적 셸)
|
|
104
125
|
│
|
|
@@ -120,7 +141,9 @@ graph TD
|
|
|
120
141
|
### 최종 요약 추천
|
|
121
142
|
1. **신뢰된 로컬 환경 (Trusted Local Setup)**: **`Meta-Harness (forge-harness) + cmux` (가벼운 UI 껍데기 1~3개 탭)**
|
|
122
143
|
* 메타 하네스가 자가 검증(4축 게이트)과 거버너 오케스트레이션, 서브에이전트 **컨텍스트 격리**를 수행하므로, 인프라 오버헤드가 없는 `cmux` 를 **잠정 권고**합니다(이 운영자 구성에서의 판단, n=1 · cmux 미실측). 여기서 "Thin Harness"는 인프라 샌드박스 대비 소프트웨어 프로토콜 중심이라는 최소 기준을 의미하며, `forge-harness`는 이 계층의 최상위 메타 하네스로 작동합니다.
|
|
123
|
-
* ⚠️ **범위
|
|
144
|
+
* ⚠️ **범위 한정 ① OS**: 벤더 페이지 기준(2026-08-16) **배포되는 빌드는 Mac 이고 그 외 플랫폼은 waitlist** 다 — «macOS 전용 설계»가 아니라 «현재 Mac 만 출시»가 정확한 서술이다(초판·1차 정정 모두 이걸 «전용»으로 단정했다). 리눅스/윈도우 노드의 UI 층 선택은 다루지 않는다(미조사).
|
|
145
|
+
* ⚠️ **범위 한정 ② 런타임**: Claude Code 세션을 전제한다. 다른 런타임(OpenCode 등)에서의 적합성은 **판단 불가 — 조사한 적 없다.**
|
|
146
|
+
* 🟥 **범위 한정 ③ residency — 이 문서의 자기모순이었다.** 초판은 §5 에서 **이전 도구가 프롬프트 원문·툴콜을 로컬 상태파일에 기록하더라는 것을 residency 표면으로 지목해 놓고**, 바로 그 다음 절에서 새 UI 층을 **동일 점검 없이** 권고했다. UI 층은 **그 세션의 모든 프롬프트를 보는 구성요소**다. ⇒ **residency 가 걸린 노드에서는 도입 전에 telemetry·상태파일 기록 표면을 실측하고 승인받는다**(§5 와 같은 점검). «macOS 전용»은 맞는 단서지만 **틀린 축의 단서**였다 — 입장리뷰가 지목.
|
|
124
147
|
2. **비신뢰 및 파괴적 과제 (Untrusted Execution)**: **실제 커널/하이퍼바이저 샌드박스** — Firecracker · Kata Containers · gVisor · Docker Sandboxes 계열
|
|
125
148
|
* 파괴적 셸 명령, 비신뢰 외부 코드 실행 시에는 커널 경계를 가진 샌드박스를 하부에 배치해 호스트 노출을 크게 줄입니다(«완벽»이 아니다 — 커널/하이퍼바이저 탈출은 실존하는 위협 클래스다).
|
|
126
149
|
* 🟥 **초판은 여기에 `Orca` 를 «물리 샌드박스»로 적었고 그건 틀렸다. 이 정정이 이 문서에서 가장
|
|
@@ -188,6 +211,18 @@ graph TD
|
|
|
188
211
|
현재 방어선은 파일 소유 분리라는 *규율*뿐이다.
|
|
189
212
|
5. **부록의 심사 이력은 저자 런타임 자기신고**이며 재현 불가하다(아래 부록 자체 주석 참조).
|
|
190
213
|
6. **표본 n=1** — 단일 운영자·단일 머신 구성에서의 판단이다.
|
|
214
|
+
7. **형제 하네스가 아직 옛 입장을 들고 있다.** PMH 정본(`CLAUDE.md` 상주층)은 «에이전트뷰 기본
|
|
215
|
+
운용»을 선언하고 있어, 이 문서 §3.2 와 **정면으로 모순**된다. 이 문서를 동기화하면 그쪽에
|
|
216
|
+
상충하는 상주 지시가 둘 생기고 tiebreaker 가 없다. ⇒ **문서만 옮기지 말고 그쪽 정본의
|
|
217
|
+
carve-out 을 같은 변경에서 처리해야 한다.** 이 PR 범위 밖이며 미해결.
|
|
218
|
+
8. **런타임 범위 미조사** — cmux × 비-Claude-Code 런타임 조합은 판단하지 않았다.
|
|
219
|
+
9. **외부 1차 출처가 레포에 스냅샷되어 있지 않다.** Orca·cmux 인용은 열람일만 기록돼 있고
|
|
220
|
+
본문 스냅샷·커밋해시·아카이브 링크가 없다. 이 레포만 받은 사람은 **정정의 근거를 확인할 수
|
|
221
|
+
없다** — 초판이 틀렸다는 것도, 정정판이 맞다는 것도. (fresh-clone arm 지목)
|
|
222
|
+
10. **게이트 강제는 배선 조건부다.** `core.hooksPath` 미배선 클론에서는 훅이 `templates/` 안
|
|
223
|
+
템플릿일 뿐이고 아무것도 막지 않는다. 초판 요약은 이를 무조건적 기계 사실로 적었다.
|
|
224
|
+
11. **§5 는 이 레포 안에서 검증 불가다** — 4행 «2026-07-12 실측» 표와 인용된 운영자 결론의
|
|
225
|
+
출처가 전부 비공개 저장소에 있다. 라벨은 붙었으나 그건 갭을 드러낼 뿐 닫지 않는다.
|
|
191
226
|
|
|
192
227
|
---
|
|
193
228
|
|
|
@@ -197,18 +232,58 @@ graph TD
|
|
|
197
232
|
|---|---|---|
|
|
198
233
|
| **Step 0 레지스터** | 독자용 아키텍처 기술문서 (문어 존댓말/명료체) | mandatory-pass |
|
|
199
234
|
| **Step 1 캘리브레이션** | FH Knowledge Core 정본 표본(`knowledge/shared/harness-core/fh_ecosystem_positioning.md`) 서식 및 톤 대조 완료 | mandatory-pass |
|
|
200
|
-
| **Step 2 문체 규율** | 번역투·조각문 5종 스캔 «잔여 0건 (양성 컨트롤 동반)» 주장 |
|
|
235
|
+
| **Step 2 문체 규율** | 번역투·조각문 5종 스캔 «잔여 0건 (양성 컨트롤 동반)» 주장 | 🟡 **부분 해소 (2026-08-17)** — 계기의 **판별력**은 이제 재현 가능하다: `bash scripts/test_ko_tech_writer_lanes.sh`(9레인, 픽스처 `plugins/fh-commons/skills/ko-tech-writer/fixtures/known_{positive,negative}.md`). 5클래스 전부 «양성 ≥1 · 음성 0» 으로 갈린다. 🟥 **그러나 이 행의 원 주장(«이 문서에서 잔여 0건»)은 여전히 미검증**이다 — 판별력과 잔여는 다른 명제이고, 후자는 문서마다 따로 재야 한다. 아래 §Step2-calibration 참조 |
|
|
201
236
|
| **Step 3 정직 수위** | 저자 내부 집계 규율 서술 제거 및 독자 의사결정 기반 정보 보존 | judged |
|
|
202
|
-
| **Step 4 수치·주장 게이트** | 수치 전수 추출 + «전칭 단정 스캔 잔여 0건» 주장 | 🟥 `UNCALIBRATED — 자기반증`: 그 «0건» 시점에 본문에 전칭 단정이 **3건 살아 있었다**(«완벽 보호» ×2, «가장 뛰어난 하네스») — 임포트 심사가 손으로 잡아 정정했다. 계기는 초록인데 대상을 안 쟀다 |
|
|
237
|
+
| **Step 4 수치·주장 게이트** | 수치 전수 추출 + «전칭 단정 스캔 잔여 0건» 주장 | 🟡 **부분 해소 (2026-08-17)** — 전칭 단정 후보 검출 2패턴(어휘형·부정형)의 판별력이 같은 스위트로 재현된다(양성 2·2건 / 음성 0·0건). 🟥 **자기반증 사실 자체는 그대로 남는다** — 아래 원 판정 유지. 그리고 «수치 전수 추출» 쪽은 **미보정**이다(이 스위트가 안 다룬다). ↓ 원 판정: 🟥 `UNCALIBRATED — 자기반증`: 그 «0건» 시점에 본문에 전칭 단정이 **3건 살아 있었다**(«완벽 보호» ×2, «가장 뛰어난 하네스») — 임포트 심사가 손으로 잡아 정정했다. 계기는 초록인데 대상을 안 쟀다 |
|
|
203
238
|
| **Step 5 지각 QA** | Mermaid 다이어그램 노드 레이블 및 Decision Tree ASCII 렌더링 시각 확인 완료 | judged |
|
|
204
239
|
| **글로벌 인프라 심사** | 논리 격리 vs 보안 샌드박싱 분리, 자원 경합 및 Context Budgeting 피드백 반영 | 🟥 `LOCAL-ONLY ATTESTATION — UNVERIFIED`: 저자 런타임(Antigravity) 측 심사이고, 짝으로 적혀 있던 `research` 서브에이전트는 **이 레포에 존재하지 않는다**(등록 에이전트 8종 중 없음). FH 안에서 재현 불가 |
|
|
205
240
|
| **적대적 공격 심사** | 직교적 3계층 모델 재정립, $N_{human}$ vs $M_{subagent}$ 인지 분리, 샌드박싱 조건 개고 | 🟥 `LOCAL-ONLY ATTESTATION — UNVERIFIED`: `challenger` 는 실재하는 FH 에이전트지만(`plugins/fh-meta/agents/challenger.md`), 이 행이 가리키는 실행의 마커·로그가 없다. **이름의 실재는 실행의 증거가 아니다** |
|
|
206
241
|
|
|
242
|
+
<a name="step2-calibration"></a>
|
|
243
|
+
### §Step2-calibration — 2026-08-17, 무엇이 해소됐고 무엇이 안 됐나
|
|
244
|
+
|
|
245
|
+
**기원**: 챔버 런 #12(`prosody-lens`, KILL)가 이 두 행을 **배출 판정의 결정적 근거**로 인용했다 —
|
|
246
|
+
*"같은 계열 계기가 미보정인데 하나 더 짓는 것은 재발명이자 미보정 계기의 증식이다."*
|
|
247
|
+
운영자 결정(2026-08-17): **새 계기보다 이 부채가 먼저.**
|
|
248
|
+
|
|
249
|
+
```
|
|
250
|
+
재현 커맨드 bash scripts/test_ko_tech_writer_lanes.sh
|
|
251
|
+
픽스처 plugins/fh-commons/skills/ko-tech-writer/fixtures/known_positive.md
|
|
252
|
+
plugins/fh-commons/skills/ko-tech-writer/fixtures/known_negative.md
|
|
253
|
+
결과 9 레인 전건 통과 — 5클래스 + 전칭 2패턴 + META 컨트롤 2
|
|
254
|
+
엔진 ripgrep 15.1.0 (SKILL.md 가 rg 로 고정한 그 엔진)
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
**해소된 것**: 「컨트롤이 무엇이었는지·재현 커맨드·출력이 하나도 없다」 — 셋 다 생겼다.
|
|
258
|
+
각 클래스가 **양성 ≥1 · 음성 0** 으로 갈린다(존재 확인이 아니라 판별 확인).
|
|
259
|
+
|
|
260
|
+
🟥 **해소되지 않은 것 — 축소하지 않는다**
|
|
261
|
+
1. **원 주장(«이 문서에서 잔여 0건»)은 여전히 미검증.** 판별력과 잔여는 다른 명제다.
|
|
262
|
+
이 스위트를 근거로 «잔여 0» 을 주장하면 강등 사유가 그대로 재발한다 — 스위트 자신이
|
|
263
|
+
출력 말미에 그렇게 인쇄한다.
|
|
264
|
+
2. **Step 4 의 자기반증 사실은 그대로 남는다**(«0건» 시점에 전칭 단정 3건 생존).
|
|
265
|
+
3. **«수치 전수 추출»은 이 스위트가 안 다룬다** — Step 4 의 절반은 여전히 미보정.
|
|
266
|
+
4. 🟥 **부수 발견 — 정본의 분류가 과장이다.** SKILL.md 는 *"앞 다섯 줄은 기계 검출 가능"*
|
|
267
|
+
이라 적었는데, **실제로 grep 을 싣고 있는 것은 C1(줄표)·C5(소유 직역) 둘뿐**이고
|
|
268
|
+
C2·C3·C4 는 산문 힌트다. 특히 **C4 조각문은 패턴이 없다**(«서술어 없는 마침»).
|
|
269
|
+
첫 캘리브레이션이 그걸 드러냈다 — 내가 임의로 지은 C4 패턴이 **양성을 0건으로 놓쳤다.**
|
|
270
|
+
지금 실린 C4 는 **닫힌 명사 어휘 목록**이라 recall 이 낮고, 넓히려면 어휘 추가가 아니라
|
|
271
|
+
형태소 분석이 필요하다. **낮다는 사실을 스크립트 주석에 적었다.**
|
|
272
|
+
5. 이 스위트는 `rg` 부재 시 **SKIP 이 아니라 rc=10** 으로 끝난다(미측정을 통과로 렌더 금지).
|
|
273
|
+
⚠️ 실측 계기: 이 개발 머신의 `grep` 은 **ugrep 7.5.0**(GNU 아님)이라 한글 word 경계가
|
|
274
|
+
갈릴 수 있다 — SKILL.md 가 엔진을 `rg` 로 못 박은 이유가 여기서 실증된다.
|
|
275
|
+
|
|
207
276
|
> 🟥 **이 부록 전체의 지위**: 저자 런타임의 **자기신고**이며, 위 «잔여 0건»·«PASS» 는 아티팩트로
|
|
208
277
|
> 뒷받침되지 않는다. FH 자기 규율상 이것은 증거가 아니라 저자의 주장이다
|
|
209
278
|
> (`fh_4axis_gate.md §Reviewer-visible evidence` 의 degrade 라벨을 그대로 적용).
|
|
210
|
-
> **이 문서에 대한
|
|
211
|
-
>
|
|
279
|
+
> **이 문서에 대한 심사는 임포트 시점의 4축 게이트에서 이뤄졌다.** 🟥 **그런데 그 기록을 «커밋
|
|
280
|
+
> 이력과 `tracks/_meta/edit_manifest.yaml` 에 있다»고 가리키는 것은, 이 레포를 받는 사람에게는
|
|
281
|
+
> 죽은 포인터다** — `tracks/**` 가 gitignored 라 매니페스트도 마커도 **공개 레포에 애초에 실리지
|
|
282
|
+
> 않는다.** fresh-clone 입장 arm 이 지목했다: 부록의 자기신고를 «증거 아님»으로 강등해 놓고, 그
|
|
283
|
+
> 대체물로 **수령자에게 구조적으로 비어 있는 곳**을 가리켰다. 즉 **강등만 참이고 대체는 거짓**이었다.
|
|
284
|
+
> ⇒ 리뷰어가 실제로 닿을 수 있는 것은 **PR 본문의 증거 캡슐**(PR #402)과 이 파일의 커밋 메시지뿐이며,
|
|
285
|
+
> 마커·매니페스트는 저자 머신에만 존재한다. 이것은 `fh_4axis_gate.md §Reviewer-visible evidence`
|
|
286
|
+
> 가 이미 명명한 구조적 갭이고, 여기서 닫히지 않는다.
|
|
212
287
|
|
|
213
288
|
---
|
|
214
289
|
*Authored by an Antigravity (Gemini-family) runtime, 2026-08-15; imported into the FH Knowledge Core and reviewed under the 4-axis gate on 2026-08-16.*
|