@chrono-meta/fh-gate 2.5.0 → 2.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +11 -0
- package/CATALOG.md +41 -0
- package/CLAUDE.md +45 -10
- package/README.ja.md +4 -3
- package/README.ko.md +4 -3
- package/README.md +4 -3
- package/README.zh.md +4 -3
- package/docs/USER_GUIDE.md +118 -0
- package/docs/platform_sustainability.md +174 -0
- package/knowledge/shared/GLOSSARY.md +26 -1
- package/knowledge/shared/harness-core/fh_detail_protocols.md +44 -2
- package/knowledge/shared/harness-core/harness_incubator_doctrine.md +117 -0
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +50 -0
- package/package.json +5 -2
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +37 -0
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +44 -0
- package/plugins/fh-meta/skills/fh/SKILL.md +31 -0
- package/plugins/fh-meta/skills/goal-quench/SKILL_detail.md +1 -1
- package/plugins/fh-meta/skills/harness-doctor/SKILL.md +7 -1
- package/plugins/fh-meta/skills/harvest-loop/SKILL.md +32 -0
- package/scripts/adapters/fixtures/mate_agent_boundary_known_negative.md +31 -0
- package/scripts/adapters/fixtures/mate_agent_boundary_known_positive.md +56 -0
- package/scripts/cluster_capability_scan.sh +18 -5
- package/scripts/digest_landing_check.sh +39 -8
- package/scripts/package_coverage_check.sh +9 -0
- package/scripts/selfcheck.sh +29 -1
- package/scripts/test_satellite_publish_gate_lanes.sh +0 -339
|
@@ -90,7 +90,12 @@ Identity marker: every greeting response opens with **🐿️ then an identity-r
|
|
|
90
90
|
**Branch test (mechanical — local state only)**: returning = session files exist (any `tracks/**/session_*.md` or `tracks/_meta/*.md` beyond `.gitkeep`) **OR** mapped project tracks exist (`tracks/{name}/` dirs — **any underscore-prefixed dir doesn't count** (`tracks/_*`, general rule not a closed list: `_meta`/`_audit`/`_contrib`/`_chamber`…); covers mapped-but-not-yet-synced users). **Never infer the branch from git log or CATALOG residue** — a fresh clone carries full commit history but zero session files: it is a NEW install (origin: fresh-clone sonnet sim rendered the returning menu off commit messages, `fh_signal_2026-06-11` FP8).
|
|
91
91
|
|
|
92
92
|
**New user** (neither condition holds — fresh clone/install): 2-door starter, never the returning menu —
|
|
93
|
-
> 🐿️ **Welcome to FH.** *Looks like you're new here!
|
|
93
|
+
> 🐿️ **Welcome to FH.** *Looks like you're new here! What would you like to do?*
|
|
94
|
+
> - **① Create your first project** — guided
|
|
95
|
+
> - **② Map an existing project**
|
|
96
|
+
> - **📖 Read the guide / ask me anything**
|
|
97
|
+
>
|
|
98
|
+
> *…and I can run `/install-wizard` to finish initial setup.*
|
|
94
99
|
|
|
95
100
|
- **① Create your first project** → Step 3-0 (guided: name → `tracks/` → `.claudeignore` → cascade)
|
|
96
101
|
- **② Map an existing project** → `auto_project_mapping.md`; after a successful mapping, offer the §6 Full-Harness promotion prompt
|
|
@@ -100,8 +105,33 @@ Identity marker: every greeting response opens with **🐿️ then an identity-r
|
|
|
100
105
|
> 🐿️ **Welcome to FH.** *forge-harness is a tool hub for rapidly setting up Claude Code projects. It supports plugin recommendations, project setup, and harness diagnostics. What would you like to work on?*
|
|
101
106
|
|
|
102
107
|
**Returning user** (branch test above) — open with the fixed 4-door menu (the doors are stable; the contents are composed live). A summary copy lives in CLAUDE.md §Active Onboarding — keep branch tests and door labels in sync when editing:
|
|
103
|
-
> 🐿️ **Welcome back to FH.** *What would you like to start
|
|
108
|
+
> 🐿️ **Welcome back to FH.** *What would you like to start?*
|
|
109
|
+
> - **① Map a project**
|
|
110
|
+
> - **② Create a new project**
|
|
111
|
+
> - **③ Accelerate or diagnose a mapped project** (work · Full-Harness · skills/agents/plugins · 진단) — {field candidates}
|
|
112
|
+
> - **④ Cross-project synergy**
|
|
113
|
+
> - **📖 Guide / Q&A**
|
|
104
114
|
>
|
|
115
|
+
🟥 **One door per line — never join them with `·` into a single run-on line** (operator, 2026-08-20).
|
|
116
|
+
A `·`-joined menu wraps at an arbitrary terminal width, so the reader cannot see where one door ends
|
|
117
|
+
and the next begins. The vertical list satisfies **G-GREET-02** (🐿️ + welcome on the SAME line),
|
|
118
|
+
**G-GREET-03** (fixed 4-door set) and **G-GREET-05** (welcome literals) unchanged — those probes pin
|
|
119
|
+
the door *set*, the *literals*, and the *welcome line*, **not the menu's line count**. Layout is the
|
|
120
|
+
render layer; the probes are the verdict layer. The 🔧 developer door is likewise its own row at the
|
|
121
|
+
bottom of the list, never appended to another line.
|
|
122
|
+
|
|
123
|
+
**📖 door (unnumbered, always rendered)** — opens `docs/USER_GUIDE.md` and takes FH-usage questions.
|
|
124
|
+
🟥 **Do not renumber.** ①–④ are the fixed set; 🔧 was the only unnumbered exception and 📖 joins it at
|
|
125
|
+
that layer. A guide is *reference before work*, not *the start of work*, so it does not belong in the
|
|
126
|
+
numbered set. **G-GREET-03 (fixed 4-door) and G-GREET-05 (welcome literals) both stay satisfied** — no
|
|
127
|
+
number was added and no welcome phrase was touched.
|
|
128
|
+
🟥 **"Open" means path + a 3-line table of contents FIRST.** Never dump the file inline — burning tokens
|
|
129
|
+
every session is precisely what this door exists to avoid. Then, and only then, branch on `uname -s`
|
|
130
|
+
to *suggest* an opener (Darwin→`open` · Linux→`xdg-open` · MINGW/MSYS→`start`); if none exists, skip
|
|
131
|
+
silently — the path already landed, so nothing is lost. ⚠️ `open` is macOS-only and FH ships via npm,
|
|
132
|
+
so it must never be the default path. Operating detail (allowed corpus · say "not found" when absent)
|
|
133
|
+
lives in `/fh` Step 3.5.
|
|
134
|
+
|
|
105
135
|
> (When **FH-dev state exists** — the operator — the welcome line is **"The FH operator — good to see you."** in place of "Welcome back to FH.")
|
|
106
136
|
|
|
107
137
|
- **① Map a project** → routes to `auto_project_mapping.md`; after a successful mapping, offer the §6 Full-Harness promotion prompt
|
|
@@ -170,6 +200,18 @@ priority: high|medium|low
|
|
|
170
200
|
---
|
|
171
201
|
# FH Improvement Signal — {date} ({source})
|
|
172
202
|
|
|
203
|
+
## Session Retrospective ← 마감 회고로 생성된 신호만. 상한 8줄. 이벤트 신호는 이 절 없음
|
|
204
|
+
- 정정: {운영자가 나를 정정한 건수} — 각 한 줄, 무엇을 어떻게 틀렸나
|
|
205
|
+
- 자력 {N} / 타력 {M} — 타력은 **잡은 축 이름**으로 (레인 · 되돌림 · cross-family · 첫실사용 · ⓓ · 운영자 · CI)
|
|
206
|
+
- 안 돌린 축: {이름, 또는 「없음」}
|
|
207
|
+
- 반복: {이번이 N번째인 실수 — memory 키 또는 「신규」}
|
|
208
|
+
|
|
209
|
+
🟥 **등급을 적지 마라.** «세션이 잘 됐다/못 됐다» 는 자평이고 게임 가능하다. 적는 것은 **사건과
|
|
210
|
+
그것을 잡은 축의 이름**뿐이다. 판정을 안 적으면 자평할 대상이 없다. 「정정 건수」는 트랜스크립트
|
|
211
|
+
사실이지 판단이 아니라서 게임이 어렵고, 자력/타력 분리는 «내가 다 잡았다» 를 쓰기 불편하게 만든다.
|
|
212
|
+
⚠️ 셋 다 자평을 **어렵게** 할 뿐 **닫지 않는다** — 닫히는 것은 peer 나 cross-family 가 이 회고를
|
|
213
|
+
읽을 때이고 그건 이 형식의 범위 밖이다.
|
|
214
|
+
|
|
173
215
|
## Friction Point
|
|
174
216
|
-
|
|
175
217
|
|
|
@@ -530,6 +530,123 @@ ceiling will never propose the redesign that was the point of incubating.
|
|
|
530
530
|
*nursery* is allowed to do to it while it is still inside. The first is an exit condition; the second
|
|
531
531
|
is a working posture.
|
|
532
532
|
|
|
533
|
+
### 3-e. Two moves observed in one incubation session — «채운다» is not the whole job (operator-approved, 2026-08-19)
|
|
534
|
+
|
|
535
|
+
⚠️ **Scope, up front**: this section records a **failure taxonomy measured on one field harness in
|
|
536
|
+
one session**, and the move it names. It is **not** a claim about what incubation essentially *is*.
|
|
537
|
+
A cross-family review of the first draft killed that framing, and it was right to
|
|
538
|
+
(§3-e-PROVENANCE below).
|
|
539
|
+
|
|
540
|
+
Operator's framing going in: *"qasp가 아직 남은 점들이 있으나 fh나 pmh의 인큐베이팅 모드로
|
|
541
|
+
사용시(pmh가 실시간 검증 및 보강)에는 부족한 점을 메꾸어서 의도한 기능들이 … 발휘 가능한
|
|
542
|
+
상태들이라고 나는 보고있어. 인큐베이터 기능의 확장이지."*
|
|
543
|
+
|
|
544
|
+
That holds **for the class of gap it names**. The session then measured a class outside it:
|
|
545
|
+
|
|
546
|
+
| | |
|
|
547
|
+
|---|---|
|
|
548
|
+
| ① **채운다** | The target **declares** it cannot do something — `UNCALIBRATED`, an unwired slot, a domain lock. The incubator supplies the capability at more tokens, more time. |
|
|
549
|
+
| ② **선언되게 만든다** | The target produces **plausible output that is wrong**. ① has nothing to grab. The move here is to *manufacture the declaration* — or, more often than expected, **to move an existing declaration into the path someone reads.** |
|
|
550
|
+
|
|
551
|
+
#### 🟥 The five cases split three ways — and the first draft got this wrong
|
|
552
|
+
|
|
553
|
+
The draft asserted *"none of the five signals a shortfall upstream."* **That is false, and the
|
|
554
|
+
falsification came from this repo's own artifacts** (cross-family review flagged it; source check
|
|
555
|
+
confirmed and went further than the reviewer did):
|
|
556
|
+
|
|
557
|
+
```
|
|
558
|
+
진짜 무음 — 채널이 없었다 "면제"·"무료" 자기억제 2건
|
|
559
|
+
(scanner did not exist until the session built it)
|
|
560
|
+
신호는 있는데 안 읽혔다 라우팅 고아 — caller_zero_baseline.json 이 2건
|
|
561
|
+
ORPHAN 으로 기록 중이었다(tracked · CI 가 읽는다)
|
|
562
|
+
핸드오프 — 선행 문서가 "후속이 갱신본" 이라는
|
|
563
|
+
포인터를 남겼고, 그 후속이 침묵했다
|
|
564
|
+
채널을 만들었더니 선언됐다 p6 분류표 0/6 — `_UNMAPPED_SEEN` 관측 채널이 1건
|
|
565
|
+
하루 전 커밋으로 도입돼 리포트까지 배선됐고,
|
|
566
|
+
그 다음 날 수리됐다
|
|
567
|
+
```
|
|
568
|
+
|
|
569
|
+
**The third row is the load-bearing one, and the draft had it backwards** — it cited p6 as an example
|
|
570
|
+
of silence when p6 is the **existence proof of ② succeeding**: someone built the channel, the silent
|
|
571
|
+
thing declared itself, ① closed it the next day. One case, one day apart, in-repo.
|
|
572
|
+
|
|
573
|
+
⇒ So ② is not one move but two, and the second is more common than expected:
|
|
574
|
+
|
|
575
|
+
- **(a) build the channel where none exists** — the 2 self-suppression cases
|
|
576
|
+
- **(b) put the channel where the reader is** — the 2 unread-signal cases. This is
|
|
577
|
+
`gate_locality_principle.md` applied to *declarations* rather than to gates: a declaration nobody
|
|
578
|
+
reads is not a declaration. **Neither of those two was fixed by adding information; both were fixed
|
|
579
|
+
by moving where it surfaces.**
|
|
580
|
+
|
|
581
|
+
#### What was actually built, stated narrowly
|
|
582
|
+
|
|
583
|
+
Function ② was realized on that session as **authoring-time static instruments** —
|
|
584
|
+
`scripts/self_suppression_scan.py` (trigger vocabulary swallowed by negation vocabulary),
|
|
585
|
+
`scripts/domain_coupling_scan.py` (hardcoded domain vocabulary), plus anchors that go red when the
|
|
586
|
+
instrument itself dies.
|
|
587
|
+
|
|
588
|
+
🟥 **This is not a runtime declaration mechanism and must not be read as one.** Nothing there makes a
|
|
589
|
+
*running* audit announce «this output is confidently wrong». What was demonstrated: **an incubator
|
|
590
|
+
holding §3-d's whole-repo authority built instruments the target had not built for itself.** Whether
|
|
591
|
+
the target *could* have is not established — the observed reason was ordinary (nobody had looked),
|
|
592
|
+
not structural.
|
|
593
|
+
|
|
594
|
+
#### Ordering: conditional, not a law
|
|
595
|
+
|
|
596
|
+
The draft said *"② is prior to ①"*. Weakened deliberately: **where a defect is undeclared, ① has no
|
|
597
|
+
target, so ② has to come first for that defect.** That is a statement about one defect's handling
|
|
598
|
+
order, not a phase ordering for incubation, and one session cannot support the stronger reading.
|
|
599
|
+
|
|
600
|
+
Likewise the draft's *"the second is worse than the first"* (silent-wrong vs crash) is **withdrawn** —
|
|
601
|
+
severity depends on the operating context, and nothing here measured it.
|
|
602
|
+
|
|
603
|
+
#### 🟥 ② presupposes decorrelated input — this belongs in the definition, not the caveats
|
|
604
|
+
|
|
605
|
+
Self-detection on that session was **2 of 16**, and both of the two were *reading something already
|
|
606
|
+
written down* (a baseline file that said `ORPHAN`; a loop that finished suspiciously fast), not
|
|
607
|
+
inference. The rest came from cross-family review (two families, zero overlap), revert probes, the
|
|
608
|
+
full suite, a caller-zero ratchet, a peer session, and a company session.
|
|
609
|
+
|
|
610
|
+
⇒ **② as described is a property of an incubation regime that has decorrelated review attached, not
|
|
611
|
+
of a lone incubation session.** A single session should not be assumed able to run it.
|
|
612
|
+
|
|
613
|
+
🟥 **And «decorrelated» is decided by what you SEND, not by which family you send it to.** A panel
|
|
614
|
+
that all receives the same payload keeps the same blind spot no matter how many reviewers sign it.
|
|
615
|
+
The payload for ② is **the frozen diff *and* the change's own marker** — the diff carries the code,
|
|
616
|
+
the marker carries the *claims* (which axes ran, which controls were alive, what residual is
|
|
617
|
+
admitted). A reviewer who never sees the marker cannot catch «an axis asserted with no trace» or
|
|
618
|
+
«a residual implied but unlisted», because those defects are not in the code.
|
|
619
|
+
|
|
620
|
+
**Measured on this section itself, the same day**: the grounding arm was sent a **5-item fact list**
|
|
621
|
+
instead of repo access, and consequently reported every well-grounded statement outside that list as
|
|
622
|
+
invention — one usable finding out of six. The family was fine; **the payload was the defect.**
|
|
623
|
+
⇒ [[feedback_decorrelation_axis_is_what_you_send]]. The plumbing that supplies this is
|
|
624
|
+
`auto-decorrelation` Step 5 (payload = frozen diff + marker), scoped to **consistency**, not honesty
|
|
625
|
+
— asking whether the record matches the diff is a channel check; asking whether the record is *true*
|
|
626
|
+
is a verdict, and §Mechanization Boundary keeps verdicts out of code.
|
|
627
|
+
|
|
628
|
+
#### The failure mode ② carries
|
|
629
|
+
|
|
630
|
+
The instruments were wrong **seven times** in that session. Six produced **plausible wrong numbers**;
|
|
631
|
+
one died loudly (`No such file`, `rc=1`) and was caught instantly. Once, a control was the only thing
|
|
632
|
+
standing between «the corpus has none of this shape» and the truth, «the detection surface does not
|
|
633
|
+
see that shape».
|
|
634
|
+
|
|
635
|
+
⇒ **An instrument built to expose silent wrongness is itself a source of silent wrongness.** The
|
|
636
|
+
prescription measured there: **put the control in the same commit as the instrument.** Where that was
|
|
637
|
+
done the error surfaced immediately; where it was not, it surfaced afterwards, if at all.
|
|
638
|
+
|
|
639
|
+
#### §3-e-PROVENANCE — what this section's own review caught
|
|
640
|
+
|
|
641
|
+
- **cross-family (codex, gpt-5.5)**: killed the «essence of incubation» framing, the
|
|
642
|
+
«none of the five» claim, the «prior to» ordering, and the «worse than» severity rule. All four
|
|
643
|
+
were adopted. Self-detection on this section: **0**.
|
|
644
|
+
- **grounding pass (agy, gemini-3.1-pro)**: mostly measured **the author's own dispatch error** — it
|
|
645
|
+
was given a 5-item fact list rather than repo access, so every well-grounded statement outside that
|
|
646
|
+
list was reported as invention. One finding survived: *"all of them wrong"* overstated p6
|
|
647
|
+
(a finding whose true category happened to be the fallback would not be wrong), now softened.
|
|
648
|
+
🟥 Recorded because the lesson is about **the reviewer's input scope**, not about the reviewer.
|
|
649
|
+
|
|
533
650
|
## 4. Compose ∪ disrupt — two operating modes over other harnesses
|
|
534
651
|
|
|
535
652
|
| Mode | What | FH mechanism |
|
|
@@ -2302,3 +2302,53 @@
|
|
|
2302
2302
|
finding: "OpenWiki(2026-08-17 발표, MIT) — 우리가 같은 날 지은 위성과 **정면으로 겹치는 선행자산**. 자막 원본 영어 트랙으로 음성 100% 흡수, 🟥 슬라이드 0%(그래서 토큰 감소 수치는 **미확보** — 발표자 본인이 'I don't have it here'). 채택 1건: **PR 로 낸다, 직접 쓰지 않는다**(우리 «AI 는 PR 을 제안한다» 규칙을 위성에 적용하는 걸 남이 먼저 검증해 준 형태). 기각 3건을 사유와 함께: 델타게이트=정의역 불일치(저쪽 델타는 레포 안, 위성은 바깥을 긁는다) · log.md 사람용 분리=닫는 결함 미확정 · quickstart/OKF=이미 있음. 🟥 반증 1건: 저쪽이 «에이전트만 읽을 거라 생각했다»가 현장에서 즉시 깨졌다 — 위성 노드를 에이전트용으로만 최적화하지 말라는 실측 반례."
|
|
2303
2303
|
note: "🟥 이 arm 은 «막히면 안 지어낸다» 지시를 지켰고 인용마다 «영상 발화 ✅ / 슬라이드 추정 ❌» 를 갈랐다 — 그 덕에 eval 수치(20개 중 7~8 → 9~10)에 발표자 본인의 'about' 단서가 보존됐고, 벤치마크명은 자동캡션 전사가 불안정해 «DeepSWE 추정, 미확인» 으로 남았다. 인용 전 확인이 필요한 항목이 그대로 표시된 채 왔다는 것이 이 위임의 실제 값이다."
|
|
2304
2304
|
cost: 138,973 tokens (subagent_tokens · tool_uses 5 · 105s)
|
|
2305
|
+
|
|
2306
|
+
- date: 2026-08-19
|
|
2307
|
+
agent: general-purpose (isolated)
|
|
2308
|
+
model: opus (orchestrator) / 위임 기본 티어
|
|
2309
|
+
purpose: "v2.5.0 게시 직전 Pre-Publish 게이트 3항(코드 보안 패스) — 출하되는 diff 만 대상으로, 토큰 스캔이 아니라 «코드 거동»"
|
|
2310
|
+
prompt_summary: "출하 파일 목록을 명시해 범위를 못 박고(마크다운·JSON 제외) · 위협모델을 «소비자가 훅과 셸을 자기 레포에서 실행한다» 로 고정 · 확신도 8 미만 폐기 · 이미 방어가 있으면 결함 아님(주장 전에 그 줄 주변을 읽어라) · 없으면 «없다» 한 줄"
|
|
2311
|
+
outcome: accepted
|
|
2312
|
+
finding: "🟥 **실 결함 2건, 둘 다 재현했다** — 내가 이번 릴리스에서 **새로 출하하는** 레인 2개가 고정 `/tmp` 경로(`/tmp/.r4out` · `/tmp/.lw_*` · `/tmp/.lg_*`)라 공유 `/tmp` 에서 **심볼릭 링크 선점으로 임의 파일 truncate+overwrite**(CWE-377). 게다가 `selfcheck.sh` 에 배선돼 있어 소비자가 셀프체크만 돌려도 발동한다. 나머지 5개 파일은 «없다» 로 냈고, 그 판단마다 근거를 댔다(EVIDENCE_ROOT 폴백은 fail-closed 유지 · `DLC_*` 비인용은 사용자 소유 env · selfcheck 판정이 종료코드)."
|
|
2313
|
+
note: "🟥 이 축은 **레인·되돌림·타계열 diff 리뷰 셋이 다 못 잡은 것**을 잡았다. 그 셋은 «게이트가 옳게 판정하나» 를 봤고 이건 «스크립트가 자기 실행 중 무엇을 쓰나» 다 — 같은 diff 를 봐도 **묻는 것이 다르면 다른 사각을 본다**. 값은 «많이 잡는다» 가 아니라 «내가 방금 만든 것을 다른 질문으로 본다» 쪽에서 나왔다."
|
|
2314
|
+
cost: 153,401 tokens (subagent_tokens · tool_uses 12 · 205s)
|
|
2315
|
+
|
|
2316
|
+
- date: 2026-08-19
|
|
2317
|
+
agent: codex/gpt-5.5 + agy(gemini) 사이드카 (cross-family 패널)
|
|
2318
|
+
model: gpt-5.5 · gemini
|
|
2319
|
+
purpose: "R1~R4 델타의 cross-family 적대 리뷰 — 축을 갈라서 보냈다"
|
|
2320
|
+
prompt_summary: "🟥 **받는 것을 다르게** 했다: codex=diff 전문(fail-open·리댁션 누수·셸 함정·장식 앵커) · agy=**주장 목록**(코드 아님, C1~C8 을 참/거짓/미검증으로) · 둘 다 «미검증을 0 이나 통과로 접지 마라» · 공유 체크아웃이라 git stash/checkout 금지"
|
|
2321
|
+
outcome: accepted
|
|
2322
|
+
finding: "codex 4건 중 **3 채택**(파일명에 든 토큰이 로그에 남음 · 유도 컨트롤이 ERE 로 쓰여 레포명 메타문자면 오탐 — 실측은 더 나빴다, 글로빙까지 타서 컨트롤 4개로 확장 · pathspec 공백 분리) **1 실측 기각**(빈/손상 오버라이드는 기존 기계가 이미 «게이트 INACTIVE» 로 시끄럽게 막는다 — 레인 E3/E3b 로 고정). agy 는 경로 잔존을 독립 재현했고 **`git log --all` 로 «실제 유출 0»** 을 확인해 줬다."
|
|
2323
|
+
note: "🟥 **자력 적발 0 인 항목이 2건**(F2 파일명 토큰 · F3 정규식 컨트롤) — 둘 다 내가 «닫았다» 고 적은 **뒤에** 나왔다. ⚠️ 첫 agy 런은 타임아웃(부분 출력)이라 **좁혀서 재발주**했고 그 값을 썼다 — 죽은 팔의 침묵을 «없다» 로 읽지 않았다."
|
|
2324
|
+
cost: codex 88,598 tokens · agy 미보고(0 으로 렌더하지 말 것)
|
|
2325
|
+
|
|
2326
|
+
- date: 2026-08-19
|
|
2327
|
+
agent: general-purpose (isolated)
|
|
2328
|
+
model: opus (orchestrator) / 위임 기본 티어
|
|
2329
|
+
purpose: "6축 중 **ⓓ 3자대면**만 UNKNOWN 이라 실제로 돌렸다 — 두 절반(선행자산 · 남의 하네스 의자)"
|
|
2330
|
+
prompt_summary: "🟥 두 절반을 명시적으로 갈랐다. ① 「새롭다」는 주장을 **넷으로 쪼개** 각 조각의 선행자산을 찾고 겹침/비겹침을 갈라라 · «없다» 는 찾아보고 없을 때만, 못 찾아본 건 blocked · ② **gstack 의 자기 규율을 먼저 읽고**(FH 어휘로 정규화 금지) 그 운영자 입장에서 위성 도입을 판정 · 🟥 «FH 가 이미 아는 잔여를 다시 말하는 것은 값이 없다 — 그 레포 규율에서만 나오는 것을 대라»"
|
|
2331
|
+
outcome: accepted
|
|
2332
|
+
finding: "🟥 **FH 가 몰랐던 것 넷, 그중 둘은 설계 결함이다.** ⓐ **위성이 미등록 redaction sink** — gstack 은 「타 모델 dispatch」를 이미 sink 로 분류·스캔하는데 위성 경로는 그 스캐너를 한 줄도 안 탄다. 🟥 FH 의 publish_gate 는 **산출물**을 스캔하고 **입력**은 아무도 안 본다(방향이 반대다) ⓑ 무인 acceptEdits 의 폭발반경이 **레포 밖** — 그쪽 `.claude/skills/gstack` 심링크가 라이브 글로벌 설치본이라 동시 실행 중인 남의 CC 세션을 깬다 ⓒ **AI 에게 period 로 닫힌 파일**(ETHOS.md)이 있고 제안조차 위반 — 프로필 스키마에 「금지 파일」이 필수여야 한다 ⓓ 산출 자리를 발명할 필요 없음(`~/.gstack-dev/plans/` 가 이미 의미론이 같다) + 그 레포는 «git status 오염 = 사건». 절반 ① 은 `checked(겹침 있음)` — Renovate(구조) · Sourcegraph Agentic Batch Changes(«repository specific instructions», 2026-06 상용) · gh-aw(착지 형태) 가 각 조각을 선행한다. **net-new 가 아니라 조합**이고 그 사실이 「새롭다」의 강도를 깎는다."
|
|
2333
|
+
note: "🟥 **회수분이 따로 있다**: gstack 의 «짓기 전에 검색» 규율은 절반 ①(선행자산)을 절반 ②의 **입력**으로 요구한다 — FH 는 둘을 따로 돌렸는데 대상의 의자에 앉으면 한 문서다. ⓓ 의 두 질문이 «둘»이라는 것이 FH 쪽 구성이지 보편이 아니라는 뜻이고, 이건 6축 정의 자체에 닿는다. ⚠️ 절반 ① 은 스니펫 기반이고 1차 문서 전수 직독이 아니다 — Sourcegraph 의 그 문구가 「레포가 쓴 것」인지 「code graph 생성」인지 안 갈렸고, 그 한 줄이 갈리면 겹침이 «부분»에서 «전면»으로 바뀐다. **미확인으로 남겼다.** 신호 = tracks/_meta/fh_signal_2026-08-19_satellite-thirdparty-axis.md"
|
|
2334
|
+
cost: 157,378 tokens (subagent_tokens · tool_uses 20 · 313s)
|
|
2335
|
+
|
|
2336
|
+
- date: 2026-08-19
|
|
2337
|
+
agent: claude-code-guide · general-purpose ×2 · general-purpose(sonnet, blind sim) · codex sidecar
|
|
2338
|
+
model: opus (orchestrator) / sonnet (blind sim) / gpt-5.5 (sidecar)
|
|
2339
|
+
purpose: "위성 프로필 스키마 게이트 + 온보딩·마감회고 설계 + 게이트 적대검증 + 플로어 티어 발화 확인 — 한 세션 5건을 클래스로 묶어 1엔트리"
|
|
2340
|
+
prompt_summary: "① 헤드리스 권한 의미론을 **1차 출처 직독**으로(추측 금지, 확인 불가는 명시) ② 마감 회고 상설화 설계 — 기존 자산 커버리지부터 갈라라 ③ 온보딩 가이드/Q&A 설계 — «이미 되는 것»이면 그렇게 결론내라 ④ **시험 대상을 안 알려준** 블라인드 4상황 sim ⑤ cross-family: fail-open 구멍·되돌림 우회를 찾아라"
|
|
2341
|
+
outcome: accepted
|
|
2342
|
+
finding: "🟥 **cross-family 가 S급 1건** — `FD_PROFILE` 미설정으로 게이트를 통째로 끌 수 있었고 **내가 쓴 레인이 그 우회를 초록으로 고정**하고 있었다(자력 적발 0). A급 4건도 전부 실물: acceptEdits 가 `sed·mv` 를 자동승인해 `sed -i` 로 금지파일 수정 가능 · 절대경로가 `.//abs` 로 조립 · 레인이 JSON 구조를 안 보고 문자열만 grep · 「egress 방화벽」 주장 과장. ★ **블라인드 sonnet sim 4/4 발화**(문 렌더 · 인라인 덤프 거부 · Linux 분기 · 없는 스킬에 「0건 근거」)이면서 **진짜 결함 1건 적발** — CLAUDE.md 서문이 신규자를 CHEATSHEET 로 오라우팅. ★ 설계 2건은 둘 다 **「새 자산 신설 없음」으로 수렴**했고, 그중 하나는 요청받은 형식(install-wizard 3부 동의)이 **이 용도엔 오용**이라고 반증했다."
|
|
2343
|
+
note: "🟥 **블라인드 sim 은 계기 생존을 확인했다**(tool_uses=7) — 2026-08-16 의 `tool_uses: 0` 죽은 sim 과 구분된다. 그러나 **reps=1 로 이 레포 자체 바(reps≥3) 미달**이고, 「발화했다」는 관측 1회지 일반화가 아니다. ⚠️ 훅 tally 는 51 인데 이 엔트리 포함 4엔트리다 — tally 는 재시도·내부 스폰까지 세므로 **1:1 대응이 아니다**(개별 기록이 아니라 클래스 집계로 남긴다)."
|
|
2344
|
+
cost: "claude-code-guide 124,771 · 설계 138,542 + 159,983 · 블라인드 sim 140,094 · codex 64,443 tokens"
|
|
2345
|
+
|
|
2346
|
+
- date: 2026-08-20
|
|
2347
|
+
agent: general-purpose ×2 (harness-doctor · pmh-dev 답습) · codex sidecar · agy sidecar · headless sonnet ×6 (블라인드 sim 양팔)
|
|
2348
|
+
model: opus (orchestrator) / gpt-5.5 (codex) / gemini-3.1-pro (agy) / sonnet (sim)
|
|
2349
|
+
purpose: "2.6.0 재출하 준비 — 소비자 설치 selfcheck 복구 · 30일 구조 진단 · pmh-dev 선행분 조사 · 메뉴 세로화 검증. 한 세션 5클래스를 1엔트리로 묶는다"
|
|
2350
|
+
prompt_summary: "① harness-doctor 30일 캐던스 — known-pair 보정 후 숫자, 못 잰 칸은 UNMEASURED 로 ② pmh-dev 3분류(net-new / PMH 앞섬 / FH 앞섬), residency 가 organization-private 인 항목은 인용 금지 ③ codex: diff 를 받고 **수리를 반증**해라(fail-open 인가 · 결박이 남았나) ④ agy: 코드 아닌 **기록의 주장**을 검증해라(표본이 결론을 지탱하나 · 미측정을 0으로 렌더했나) ⑤ 시험 대상 안 알려준 블라인드 렌더 sim, ARM/CONTROL 각 3회"
|
|
2351
|
+
outcome: accepted
|
|
2352
|
+
finding: "🟥 **cross-family 지적 3건 전부 실재·전부 채택·자력 적발 0.** codex(diff 축)가 내 수리의 반쪽을 잡았다 — L13c 라벨을 유도로 바꿨는데 **생산자는 리터럴을 낸다**, 즉 수리가 divergent-normalizer 를 새로 만들고 있었다(되돌림 + known-pair 재현: 상수 20/0 · 유도 19/1). agy(주장 축)가 둘 — ③문 자기모순, 그리고 **detail 파일의 신규사용자 문이 안 따라온 반쪽-픽스**(내 확인 grep 이 따옴표 탓에 거짓 「없음」을 냈다). 🟥 **두 계열 지적이 하나도 안 겹쳤다(코드 1 / 주장 2)** — 축 분리의 값이 실측으로 나온 자리다. harness-doctor: FAIL(M-tier 1 = 상주 140,217자) · S 9 · R 4, «못 잰 것» 8항 명시. pmh-dev: 흡수 후보 3(Wave 1-D · HEAVY 분류기 레인 · Step 0.35), 그중 **A1 은 PMH 문서가 없는 기계를 있다고 적었다**(`axis2-defense` 훅 히트 0, 컨트롤 `crossfamily` 21) — 산문만 옮기면 팬텀 기계-주장을 들여온다. 블라인드 sim ARM 3/3 vs CONTROL 3/3 완전 분리."
|
|
2353
|
+
note: "🟥 **계기가 못 보는 자리를 탈상관이 봤다.** 소비자-완주 계기는 「돌아가나」를 재지 「고친 게 옳은가」를 못 잰다 — SKIP 으로 빠지는 팔은 그 계기가 **구조적으로 안 보는** 자리이고 codex 지적이 정확히 거기 있었다. ⚠️ codex 세션이 **최종 산문 판정 없이 끝났다**(4턴, 마지막이 도구 출력) — 지적은 중간 턴에서 건졌다. 「지적 없음」이 아니라 **부분 산출**로 계상한다. ⚠️ 훅 tally 는 67 인데 이 엔트리 포함 소수다 — tally 는 재시도·내부 스폰까지 세므로 1:1 대응이 아니다(클래스 집계)."
|
|
2354
|
+
cost: "harness-doctor 196,158 · pmh 답습 186,940 tokens (subagent_tokens) · codex/agy/sim 은 CLI 라 토큰 미노출 = UNMEASURED"
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@chrono-meta/fh-gate",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.6.0",
|
|
4
4
|
"description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"keywords": [
|
|
@@ -53,6 +53,7 @@
|
|
|
53
53
|
"CLAUDE.md",
|
|
54
54
|
".claude/registry/agent_cards.json",
|
|
55
55
|
".claude/registry/README.md",
|
|
56
|
+
"docs/USER_GUIDE.md",
|
|
56
57
|
"docs/CONTRIBUTING.md",
|
|
57
58
|
"bin/fh-codex-doctor.js",
|
|
58
59
|
"bin/fh-gate.js",
|
|
@@ -88,12 +89,15 @@
|
|
|
88
89
|
"scripts/adapters/qasp_web_rules.sh",
|
|
89
90
|
"scripts/adapters/fixtures/qasp_web_rules_known_positive.json",
|
|
90
91
|
"scripts/adapters/fixtures/qasp_web_rules_known_negative.json",
|
|
92
|
+
"scripts/adapters/fixtures/mate_agent_boundary_known_positive.md",
|
|
93
|
+
"scripts/adapters/fixtures/mate_agent_boundary_known_negative.md",
|
|
91
94
|
"scripts/test_adapter_lanes.sh",
|
|
92
95
|
"scripts/test_fh_gate_regressions.sh",
|
|
93
96
|
"templates/local_fh_context.md",
|
|
94
97
|
"docs/ETHOS.md",
|
|
95
98
|
"docs/WHY.md",
|
|
96
99
|
"docs/OUTPUT_EVIDENCE.md",
|
|
100
|
+
"docs/platform_sustainability.md",
|
|
97
101
|
"knowledge/shared/GLOSSARY.md",
|
|
98
102
|
"knowledge/shared/patterns",
|
|
99
103
|
"knowledge/shared/plugin-catalog",
|
|
@@ -129,7 +133,6 @@
|
|
|
129
133
|
"scripts/chamber_candidate_collect.sh",
|
|
130
134
|
"scripts/chamber_witness.sh",
|
|
131
135
|
"scripts/digest_landing_check.sh",
|
|
132
|
-
"scripts/test_satellite_publish_gate_lanes.sh",
|
|
133
136
|
"scripts/relay_channel.sh",
|
|
134
137
|
"scripts/test_relay_channel_lanes.sh",
|
|
135
138
|
"scripts/fh_session_load.sh",
|
|
@@ -10,6 +10,43 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/)
|
|
|
10
10
|
|
|
11
11
|
## Plugin Level
|
|
12
12
|
|
|
13
|
+
### [2.6.0] — 2026-08-20
|
|
14
|
+
|
|
15
|
+
**배포본이 소비자 설치에서 `SELFCHECK: FAIL` 이었다 — 그리고 원인 넷이 전부 «계기가 저자의 머신에 결박» 이었다.** 이 릴리스의 중심은 새 기능이 아니라 그 복구다. 발견 경로는 재출하 준비 중의 손 실행이다: `npm pack` → 추출 → **`node_modules/@chrono-meta/fh-gate` 실경로에서 완주**. 레지스트리에서 받은 **실물 2.5.1** 로도 재현했다(컨트롤).
|
|
16
|
+
|
|
17
|
+
- **위성 레인 둘이 주체 없이 출하됐다.** `test_satellite_publish_gate_lanes.sh`(2.5.1 에 이미 실림) · `test_satellite_profile_schema_lanes.sh`(이번 범위에 추가됨)의 주체는 `frontier_digest_daily.sh` 인데 그건 의도적으로 출하 대상이 아니다(소비자 계정으로 `claude` CLI 를 태운다). 실측: publish_gate **4 passed / 21 failed**, `rc=127`. 🟥 **통과한 쪽이 더 나빴다** — "dispatch 자체가 안 일어남 ✅" 은 러너가 **없어서** 통과한 거짓 초록이다. 선행 사례(`test_frontier_digest_retry.sh`, *"Anchor follows subject"*)대로 `ACCEPTED_ABSENT` 로 내리고, `selfcheck.sh` 는 **주체** 부재를 보고 `_absent_subject_verdict` 로 위임한다 → 이름 있는 SKIP
|
|
18
|
+
- **`digest_landing_check --self-test` 가 폴더 이름에 결박돼 있었다.** 컨트롤이 `basename "$FH"` 로 유도되는데 픽스처 카드가 리터럴 `forge-harness` 를 담고 있어, **디렉터리 이름이 `forge-harness` 일 때만** 컨트롤이 살았다. known-pair(같은 바이트, 이름만 교체): `forge-harness/` **15/15 PASS** · `some-consumer-app/` **5/15 FAIL**. 레인별로 컨트롤을 픽스처 토큰에 고정했다. 🟥 «10 을 기대하는» 레인들도 고정했다 — 고정 전에도 10 을 냈지만 **의도한 사유가 아니라 컨트롤 사망** 때문이었다. 🟥 N1·★N-ctl·★N-ctl-re 는 일부러 유도/사망을 주장하므로 **고정하지 않았다**
|
|
19
|
+
- **`cluster_capability_scan` L13c — 전제만 결박이었다.** `discover` 가 `tracks/`(gitignored, 배포물에 구조적 부재)를 전제하므로, 부재는 **FAIL 이 아니라 이름 있는 SKIP**(미측정 ≠ 0건)으로 낸다. 🟥 **초판은 여기서 하나를 더 «고쳤고», 그게 틀렸다** — 기대 문자열 `^forge-harness\(hub\)` 를 폴더명 결박으로 읽고 `basename` 유도로 바꿨는데, **생산자(`:128`)는 그 리터럴을 낸다.** 유도로 바꾸면 이름이 다른 트리에서 기대와 산출이 갈려 **거짓 FAIL** 이 된다 — 이 파일이 자기 주석에서 경고하는 divergent-normalizer 를 수리가 새로 만든 꼴이다. 되돌렸다. 그 라벨은 디렉터리 이름이 아니라 **허브의 상수 식별자**이고, 양쪽이 같은 상수를 쓰는 한 결박이 아니다. **cross-family(codex/gpt-5.5)가 잡았다 — 자력 적발 0.** known-pair 로 재현: 이름이 다른 트리 + `tracks/` 존재에서 상수판 **20 PASS / 0 FAIL** · 유도판 **19 PASS / 1 FAIL**. 내 소비자 테스트는 그 팔에 **구조적으로 못 닿았다**(거기선 L13c 가 SKIP 이라)
|
|
20
|
+
- **mate 어댑터의 known-pair 픽스처가 안 실렸다.** 게이트는 출하되는데 보정쌍이 소비자 머신에 없어 `HARNESS_ERROR(10)`. `mate_agent_boundary_known_{positive,negative}.md` 를 `files[]` 에 추가 — 이건 **출하하는 쪽**이 맞다(주체가 이미 출하되므로)
|
|
21
|
+
- 검증: 수리 후 소비자 설치 **`SELFCHECK: PASS` (rc=0)**. 🟥 중간에 내 계기가 한 번 틀렸다 — `grep '^FAIL'` 로 세어 «FAIL=0 인데 FAIL» 이라는 가짜 모순을 만들었다. 이 스위트들은 `❌` 로 찍는다
|
|
22
|
+
|
|
23
|
+
**온보딩 메뉴를 세로로 편다.** `·` 로 이어붙인 한 줄 메뉴는 터미널 폭에서 임의로 접혀 문 경계가 안 보인다(운영자 지적). `G-GREET-02`(🐿️+환영문 같은 줄)·`G-GREET-03`(고정 4문)·`G-GREET-05`(문구 리터럴) **셋 다 불변** — 그 프로브들이 박은 것은 문 집합·리터럴·환영문 줄이지 메뉴의 줄 수가 아니다. 플로어 티어 블라인드 sim(레포 밖 cwd·헤드리스·reps=3): **ARM 3/3 세로 · CONTROL 3/3 가로**, 같은 실행에서 G-GREET-02 도 3/3 유지.
|
|
24
|
+
|
|
25
|
+
**BREAKING 없음 — 그리고 그 판정을 적어둔다.** 카드는 «위성 런이 프로필 미선언이면 막힌다» 를 `BREAKING (gate):` 후보로 올려뒀는데, 그 게이트가 사는 `frontier_digest_daily.sh` 가 **출하 대상이 아니라** 소비자의 게이트 수용은 안 바뀐다. 오늘 수리는 FAIL→PASS 라 완화 방향이다.
|
|
26
|
+
|
|
27
|
+
🟥 **이 릴리스가 스스로 낸 교훈**: 결함 넷 중 **셋은 진단이 맞았고 하나는 틀렸는데, 틀린 하나를 자력으로는 못 잡았다.** 소비자-설치 완주라는 계기는 「돌아가나」를 재지 「고친 게 옳은가」를 못 잰다 — SKIP 으로 빠지는 팔은 그 계기가 구조적으로 안 보는 자리다. 잡은 것은 **다른 계열에 diff 를 보낸 것**이다.
|
|
28
|
+
|
|
29
|
+
**같은 릴리스에 함께 나가는 것 — `harness-doctor` 30일 캐던스 수리 (#468).** 진단이 낸 것 중 **지금 출하물에 살아 있던 것**만 골랐다.
|
|
30
|
+
- **팬텀 「500줄 / 16스킬」 임계 제거** (`docs/platform_sustainability.md`, 이번에 출하 대상에 편입). 실제는 1,414줄 / 40스킬이고, `harness-doctor` 는 meta 타깃에서 줄수 행을 «판정 아님»으로 **비활성화**하지 큰 숫자로 갈지 않는다. 🟥 덤이 본체보다 컸다 — 이 팬텀의 사후분석이 *"`500` 은 이 파일 어디에도 없다(grep 0 hits)"* → «런이 지어냈다» 로 결론냈는데 **그 grep 이 자기 파일만 봤다.** 교훈은 «환각했다»가 아니라 **«부재 검사를 틀린 코퍼스에 돌렸다»**
|
|
31
|
+
- **CATALOG 미등재 정본 7건 등재** (컨트롤 `harness_6axis_framework`=2, 대상 7건 전부 0). 그중 `fh_three_layer_canon.md` 는 CLAUDE.md 가 **필독**으로 지정한 문서다 — 색인이 못 찾는 필수 문서는 파일명을 이미 아는 세션이 아닌 한 부재와 구별되지 않는다
|
|
32
|
+
- **`[[wikilink]]` 규약 선언** (`knowledge/shared/GLOSSARY.md` 신설 절). 출하 `.md` 156개에 안 풀리는 타깃 **59 / 출현 98**, 선언이 아무 데도 없었다. 🟥 이건 **깨진 참조가 아니라 출처 표시**다 — 고치려 들지 마라. 메모리 스토어는 운영자별·세션 스코프라 vendoring 은 수리가 아니라 residency 위반이다
|
|
33
|
+
- 🟥 **`CLAUDE.md` 가 자기 기계를 거짓 서술하고 있었다.** *"standpoint 는 Prose-only today — 검증하는 pre-commit 훅도 픽스처도 없다"* → `validate_standpoint_leg()` 는 `pre-commit:798` 정의 · `:1575` 호출이고 `test_marker_standpoint_lanes.sh` 는 `selfcheck.sh:531` 배선이다. **같은 파일이 정반대도 적고 있었다.** 이 계열은 레인이 구조적으로 못 잡는다 — 규칙의 **자기서술**을 그 규칙이 서술하는 **기계**와 대조하는 검사가 없다. 그리고 stale 한 «아직 안 지었다» 는 가장 조용한 드리프트다: 정직한 겸손처럼 읽히면서 **이미 있는 컨트롤의 사용을 억제한다**
|
|
34
|
+
- **M-1(상주 140k)은 안 닫았다.** 상주 원장 32절 전수 결과 **M·S·R 전부 0** — 절 단위 레버가 없다는 것이 실측이다. 유일한 레버는 capability-level merge(후보 6묶음)이고, **그 병합이 실제로 문자를 줄이는지는 UNMEASURED**
|
|
35
|
+
|
|
36
|
+
**명시 잔여**: `package_coverage_check.sh` 는 「참조된 경로가 출하되나」를 보지 「**출하된 레인의 주체가 출하되나**」를 안 본다 — 네 결함 중 셋을 rc=0 으로 통과시켰다. 그 갭은 이번에 안 닫았다.
|
|
37
|
+
|
|
38
|
+
### [2.5.1] — 2026-08-19
|
|
39
|
+
|
|
40
|
+
**진입점 패리티 — 규칙이 출하돼야 발화한다.** 2.5.0 직후 `④-b` 드리프트 검사가 **AGENTS.md 에 공유 체크아웃 규율이 없다**를 냈다. CLAUDE.md 에는 있고 Codex 진입점에는 없는 상태였고, 그 상태에서는 비-Claude 런타임에 그 규칙이 **보이지 않는다**(gate-locality). 이 릴리스는 그 한 항목을 소비자에게 실제로 보내기 위한 것이다.
|
|
41
|
+
|
|
42
|
+
- **`AGENTS.md` 항목 9 — 공유 체크아웃**: 스위치 전 `git branch --show-current` · 자른 직후 `git log main..HEAD` · **claim 파일은 lock 이 아니라 스냅샷**(남이 브랜치를 옮겨도 안 바뀐다) · 공유 체크아웃에서 `git add -A`/`git stash` 는 트리 전체에 닿는다. 🟥 이 레포에서 두 세션이 오늘 실제로 그걸로 부딪혔다. 기계층(`pre-commit` 의 `branch_claim` 게이트)은 **커밋 시점**이라 그 앞의 `git switch` 사고는 못 막는다
|
|
43
|
+
- 검증: 플로어 티어 블라인드 sim(무엇을 시험하는지 안 알림, reps=3) **arm 3/3 vs control 0/3**. 🟥 control 한 런이 죽어 있어(stdin 경고만) 재실행해 유효 3수를 채웠다 — 죽은 팔을 0 으로 계상하면 n 을 부풀린다
|
|
44
|
+
- 🟥 **잔여: 비-Claude 런타임에서 실제로 안 돌렸다.** sim 은 Claude 로 돌렸다 — 「따라와지는가」는 쟀고 「Codex 가 실제로 그렇게 하는가」는 못 쟀다
|
|
45
|
+
|
|
46
|
+
**함께 나가는 것 (데이터)**: `subagent_invocations_log.yaml` 3건 — 게시 전 보안 패스 · cross-family 패널 · **ⓓ 3자대면**. 마지막 것이 위성 설계에서 **미등록 redaction sink**(입력 무검사)와 **레포 밖 폭발반경**을 찾았고, 처방 5건은 **전부 미착수**다(`tracks/` 신호에 기록, 출하 대상 아님).
|
|
47
|
+
|
|
48
|
+
**BREAKING 없음.** 문서·데이터만이고 게이트 수용은 안 바뀐다.
|
|
49
|
+
|
|
13
50
|
### [2.5.0] — 2026-08-18
|
|
14
51
|
|
|
15
52
|
**정체성 ④(프런티어→조직 전파)의 «조직 = 레포» 축.** 다이제스트 러너가 FH 한 곳만 대상으로 지어져 있어 ④ 가 «자기 소비» 에 머물렀다. 대상 축을 열고, 그 산출이 **복제가 아니라 응용**이 되도록 프로필을 싣고, 공개 표면과 착지를 각각 기계로 잰다.
|
|
@@ -202,6 +202,50 @@ re-checks** it against the artifact (phantom-quench back-trace) before accepting
|
|
|
202
202
|
source-grounded is **dropped, not judged**. No sidecar-only verdict (no weak-local-judge regression).
|
|
203
203
|
Local 4090 = **canary tier** (evidence-of, never terminal verdict).
|
|
204
204
|
|
|
205
|
+
🟥 **What you SEND decides which axis you get — put the MARKER in the payload, not just the diff**
|
|
206
|
+
(2026-08-19). Sidecars have been receiving the *diff* only. That buys a review of the **code** and
|
|
207
|
+
buys nothing on the **record**: the Axes 2+3 marker is self-attested, and this repo has already
|
|
208
|
+
written down that the closing move is *"cross-family reading that marker"* — then never wired the
|
|
209
|
+
marker into the thing that recruits cross-family. Measured twice: (a) a release marker's false
|
|
210
|
+
`not-applicable` passed the new typed lanes untouched, because **form was correct**; (b) this skill's
|
|
211
|
+
own dispatch on 2026-08-19 sent a diff and no marker, so the round could not have caught a false
|
|
212
|
+
axis claim even in principle.
|
|
213
|
+
|
|
214
|
+
**So the payload is two parts**: the frozen diff **and** the marker for this change — **as it stands
|
|
215
|
+
at dispatch time**: `axes-run:` · `controls:` · `standpoint:` · `thirdparty:` · `residual:`.
|
|
216
|
+
|
|
217
|
+
🟥 **`crossfamily:` is NOT in the payload — it is this round's OUTPUT.** Step 6 below says every rung
|
|
218
|
+
*emits* that value and "the rung is not done until its verdict is recorded", so requiring it in the
|
|
219
|
+
thing you send is circular: you would be shipping a field this dispatch has not produced yet. Send the
|
|
220
|
+
marker with that line **absent or `UNKNOWN`**, and fill it from the result. (Caught by a blind
|
|
221
|
+
floor-tier sim of this very paragraph, 2026-08-19 — the first draft listed `crossfamily:` among the
|
|
222
|
+
payload fields and no reader could have satisfied it.) Ask the sidecar one extra
|
|
223
|
+
question: *"does the record match the diff — is any axis claimed that the diff shows no trace of, and
|
|
224
|
+
is any residual missing that the diff implies?"*
|
|
225
|
+
|
|
226
|
+
⚠️ **Scope, deliberately narrow.** This asks whether the record is **consistent with the artifact**.
|
|
227
|
+
It does NOT ask the sidecar to score honesty — «is this marker truthful» is a *conclusion*, and
|
|
228
|
+
§Mechanization Boundary forbids freezing that into machinery. Consistency is checkable from two
|
|
229
|
+
documents; honesty is not.
|
|
230
|
+
⚠️ **You cannot observe that the sidecar READ it — say so rather than implying otherwise.** Putting the
|
|
231
|
+
marker in the payload and the sidecar actually using it are different events, and nothing here
|
|
232
|
+
distinguishes them: a silently-ignored marker looks exactly like a marker that was read and raised
|
|
233
|
+
nothing. The cheap partial anchor is to require the return to **quote one marker line it checked** —
|
|
234
|
+
a reply that quotes nothing did not demonstrably read it. That is evidence-of-reading, not proof, and
|
|
235
|
+
it does not close §Mechanization Boundary's named self-attestation residual. Three independent
|
|
236
|
+
readers flagged this same gap on the day the paragraph was written (a peer session, and two blind
|
|
237
|
+
floor-tier sims), which is why it is stated here instead of left to the reader to notice.
|
|
238
|
+
|
|
239
|
+
⚠️ **Prose, not a check** — measured recurrence is 2, below this repo's own N≥3 bar
|
|
240
|
+
(`[[feedback_mechanize_at_repetition_prose_before]]`). On the third occurrence, mechanize it here.
|
|
241
|
+
⚠️ **Residency still governs**: a marker can name company assets. Sanitize before any external-family
|
|
242
|
+
dispatch, exactly as with the diff — the marker is not exempt because it is metadata.
|
|
243
|
+
|
|
244
|
+
*Origin*: sister-asset read of `raphaelchristi/harness-evolver`'s `harness-critic` agent, whose whole
|
|
245
|
+
role is auditing the **evaluator** rather than the artifact. Its detection signatures (score jumps,
|
|
246
|
+
suspiciously fast convergence) do **not** port — FH markers carry no score — but the *target* does.
|
|
247
|
+
Full assessment incl. what was deliberately not imported: `tracks/_audit/proposal_2026-08-19_sister_harness-evolver.md`.
|
|
248
|
+
|
|
205
249
|
**Target freeze — a prior drop reason, checked before any of the above (2026-08-17).** Grounding a
|
|
206
250
|
finding against *today's* tree proves nothing if the sidecar reviewed *yesterday's*. Pin before
|
|
207
251
|
dispatch and verify on return:
|
|
@@ -46,6 +46,37 @@ the session card and CATALOG, exactly as the greeting path would.
|
|
|
46
46
|
End with "pick a door, say a phrase, or just state your task". Do not auto-run anything — this
|
|
47
47
|
command is a map, not a dispatcher.
|
|
48
48
|
|
|
49
|
+
|
|
50
|
+
## Step 3.5 — 가이드 · Q&A (📖 문을 골랐을 때만 진입)
|
|
51
|
+
|
|
52
|
+
🟥 **Q&A 는 net-new 기능이 아니다.** FH 에 대해 묻고 답하는 것은 이미 된다(CLAUDE.md 가 상주라
|
|
53
|
+
세션이 문·게이트·스킬을 안다). 이 절이 더하는 것은 **하나뿐**이다 — 「무엇을 근거로 답하나,
|
|
54
|
+
그리고 없으면 없다고 말한다」는 **계약**. 그래서 새 스킬을 만들지 않고 여기 붙인다.
|
|
55
|
+
|
|
56
|
+
**ⓐ 가이드** — `docs/USER_GUIDE.md` 의 **경로와 3줄 목차**를 출력한다. 승인하면 플랫폼 opener 를
|
|
57
|
+
제안한다(`uname -s`: Darwin→`open` · Linux→`xdg-open` · MINGW/MSYS→`start`). opener 가 없으면
|
|
58
|
+
조용히 건너뛴다 — 경로는 이미 나갔으므로 손실이 없다.
|
|
59
|
+
🟥 **전문 인라인 출력 금지.** 자동 실행도 안 한다.
|
|
60
|
+
|
|
61
|
+
**ⓑ Q&A** — 1문 1답. 근거는 아래 코퍼스 **안에서만** 찾는다:
|
|
62
|
+
|
|
63
|
+
```
|
|
64
|
+
1 docs/USER_GUIDE.md 사용법 · FAQ
|
|
65
|
+
2 CHEATSHEET.md 명령 · 트리거 문구
|
|
66
|
+
3 knowledge/shared/GLOSSARY.md 용어
|
|
67
|
+
4 README.md §Get started/§Learn more 설치 · 진입 경로
|
|
68
|
+
5 CATALOG.md 「예전에 뭐 했지」
|
|
69
|
+
6 설치된 SKILL.md frontmatter 「무슨 스킬 있어」
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
**degrade — 코퍼스에 없으면**
|
|
73
|
+
🟥 **지어내지 않는다.** 「못 찾음 — 코퍼스 N개를 봤고 여기엔 없다」 + **다음 한 걸음**(어느 파일을
|
|
74
|
+
열지 · `/install-doctor` 같은 실제 진단 경로)을 준다. 「아마 이럴 것이다」 형태의 답은 **금지**다.
|
|
75
|
+
이건 §Instrument Calibration 의 «미측정을 0으로 렌더하지 않는다» 와 같은 규율이다.
|
|
76
|
+
답마다 `file:line` 을 단다 — 근거 없는 문장은 팬텀이다.
|
|
77
|
+
|
|
78
|
+
⚠️ 코퍼스 밖 질문(도메인 작업 · 코드)은 이 절이 아니라 평소대로 처리한다. Q&A 는 **FH 사용법**용이다.
|
|
79
|
+
|
|
49
80
|
## Done When
|
|
50
81
|
|
|
51
82
|
| Condition | Check class |
|
|
@@ -214,7 +214,7 @@ fi
|
|
|
214
214
|
No Stop hook is required. After the Codex goal/session completes, resolve changed files with git and run:
|
|
215
215
|
|
|
216
216
|
```bash
|
|
217
|
-
FH_BACKEND=codex npx @chrono-meta/fh-gate "{changed-files}" quick codex-goal
|
|
217
|
+
FH_BACKEND=codex npx --package @chrono-meta/fh-gate fh-gate "{changed-files}" quick codex-goal
|
|
218
218
|
```
|
|
219
219
|
|
|
220
220
|
For high-stakes or external-facing work, use `full` instead of `quick`. Treat `BLOCKED` or `ESCALATE` the same as the Claude path: fix and re-run the gate, or surface the decision to the user.
|
|
@@ -149,7 +149,13 @@ running tool itself). Full procedure: `knowledge/shared/harness-core/measurement
|
|
|
149
149
|
`M-n · <verbatim row text or its threshold> · <measured value>`. If you cannot quote the row, **you do not
|
|
150
150
|
have a finding** — downgrade to an observation. Origin (2026-07-20, instrument defect n+4, the *fourth* in
|
|
151
151
|
a single run): a run fired `M-1 · CLAUDE.md 816 lines — exceeds the FH threshold of 500`. The string `500`
|
|
152
|
-
|
|
152
|
+
did not occur anywhere in this file at the time (a later grep finds 2 — both in this narrative, added by
|
|
153
|
+
this very post-mortem; the self-referential claim drifted and is corrected here rather than re-asserted).
|
|
154
|
+
🟥 **And that grep was scoped wrong**: it searched this file only. The 500 was NOT invented from nothing —
|
|
155
|
+
`docs/platform_sustainability.md` carried it as a live standard until 2026-08-20, contradicting this row.
|
|
156
|
+
The lesson is not "the run hallucinated" but **"the absence check was run against the wrong corpus"**
|
|
157
|
+
([[feedback_absence_measurement_needs_control]] — a no-hit grep is evidence only when its scope covers
|
|
158
|
+
where the thing would actually live). The meta-harness row it claimed to read says raw
|
|
153
159
|
line count is **"Not a verdict."** So the run invented a threshold *and* fired a row this skill explicitly
|
|
154
160
|
disables for meta-harnesses — the exact recurrence of the 2026-07-15 inversion documented below, which had
|
|
155
161
|
already been patched *in the skill*. The patch held; the **run** ignored it. A verdict grounded in a
|
|
@@ -54,6 +54,38 @@ Session end
|
|
|
54
54
|
│ edit-manifest VERIFY: check pending predictions in edit_manifest.yaml
|
|
55
55
|
│ memory-hygiene scan: staleness check on memory/*.md entries (skip if < 7 days)
|
|
56
56
|
│
|
|
57
|
+
[Step 0-d] Session Retrospective (close_retro 가 granted 일 때만)
|
|
58
|
+
🟥 새 스킬을 만들지 않는다 — 회고 산출은 그대로 Step 2(contention-layer) → Step 3 → 3.5(등급)
|
|
59
|
+
로 **이미 있는 파이프라인**을 탄다. 「개선포인트 정리」가 거기서 공짜로 붙는다.
|
|
60
|
+
|
|
61
|
+
진입 조건 (기계적, 이 순서로):
|
|
62
|
+
1. `tracks/_meta/user_adaptation_profile.md` frontmatter 의 `close_retro`
|
|
63
|
+
granted → 실행
|
|
64
|
+
declined → 건너뛴다. **다시 묻지 않는다**(operational_adaptation.md no-re-nag)
|
|
65
|
+
unset / UAP 부재 / ephemeral → **실행하지 않고, 지어내지도 않는다.**
|
|
66
|
+
운영자 맥락이면(= `CLAUDE.local.md` 존재, `psa_detect_operator_context()` 와 같은
|
|
67
|
+
판별자) **최초 마감 1회만** 제안하고 답을 UAP 에 기록한다. 그 이상 묻지 않는다.
|
|
68
|
+
🟥 `CLAUDE.local.md` 존재가 판별자인 이유는 그것이 **gitignored** 라 fresh clone·CI 체크아웃·
|
|
69
|
+
워크트리 어디에도 안 따라오기 때문이다. 기각된 후보: 「tracks/_meta 가 비어있지 않음」 —
|
|
70
|
+
`.gitkeep` 이 tracked 라 **fresh clone 이 만족시킨다**(psa_scan_lib.sh 주석의 실측).
|
|
71
|
+
🟥 **운영자 맥락이라고 자동 granted 가 아니다.** 파일이 존재해서 승인되는 게 아니라
|
|
72
|
+
**운영자 발화가 인용돼 기록되는 순간** 승인된다. 전자는 세션이 자기 권한을 넓히는 형태다.
|
|
73
|
+
|
|
74
|
+
산출: 오늘자 `tracks/_meta/fh_signal_{date}_{source}.md` 에 `retro: close` 를 달고
|
|
75
|
+
`## Session Retrospective` 4필드를 채운다(형식 정본 = fh_detail_protocols.md).
|
|
76
|
+
「없음」·「0」도 유효한 값이고, **비우는 것만 안 된다.**
|
|
77
|
+
|
|
78
|
+
Done When:
|
|
79
|
+
+ 오늘자 signal 파일이 존재하고 `retro: close` 를 달고 있다 — mandatory-pass
|
|
80
|
+
+ Session Retrospective 4필드가 전부 채워졌다 — mandatory-pass
|
|
81
|
+
+ 정정 건수는 **세었지 회상하지 않았다** — measured
|
|
82
|
+
+ FH Registration Candidate 가 비면 그 이유가 한 줄 적혀 있다 — mandatory-pass
|
|
83
|
+
|
|
84
|
+
⚠️ `session_close_check.sh` 에 기계 검사를 **붙이지 않는다.** 값싸게 붙일 수 있는 형태는
|
|
85
|
+
«오늘자 signal 이 있나» 인데 그건 *존재*만 재고 *회고했나*는 못 잰다 — CLAUDE.md 가 ④ 메모리
|
|
86
|
+
위생에 대해 같은 이유로 검사를 뺐다(«a check that can be satisfied without doing the work
|
|
87
|
+
is a decoration»). 같은 판단을 그대로 적용한다.
|
|
88
|
+
|
|
57
89
|
[Step 0] Regression Guard
|
|
58
90
|
│ Check: does anything from this session conflict with or regress a validated skill?
|
|
59
91
|
│ → Regression detected: flag, route to contention-layer
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# [CALIBRATION FIXTURE — known-negative] 정보경계 불변식이 **빠진** 합성 명세서
|
|
2
|
+
|
|
3
|
+
> 🟥 **이것은 실행되는 에이전트 정의가 아니다.** `scripts/blackbox_invariant_check.sh` 의
|
|
4
|
+
> 캘리브레이션 쌍 중 **막혀야 하는 쪽**(expect = FAIL, exit 1)이다. 짝은
|
|
5
|
+
> `agent_boundary_known_positive.md`.
|
|
6
|
+
>
|
|
7
|
+
> 이 파일이 필요한 이유: 양성만 세우면 「불변식이 살아있다」와 「검사기가 아무것도 안 본다」가
|
|
8
|
+
> 같은 초록으로 보인다. 음성 arm 은 계기에 **판별력이 있는지**를 재는 자리다.
|
|
9
|
+
>
|
|
10
|
+
> **의도적으로 위반한다** — 불변식 문장들을 전부 뺐다. 실 식별자는 담지 않는다(그건 T1 이
|
|
11
|
+
> 잡는 별개의 위반이고, 여기서 그것까지 트리거하면 어느 축이 막았는지 흐려진다).
|
|
12
|
+
|
|
13
|
+
## 입력 인터페이스
|
|
14
|
+
|
|
15
|
+
외부 findings 를 받는다. `finding_id` 는 불투명 식별자다(어휘 제약 없음).
|
|
16
|
+
`location` 은 특정 접두어가 보이면 폐기한다(denylist).
|
|
17
|
+
|
|
18
|
+
## 투영 규칙
|
|
19
|
+
|
|
20
|
+
주입 파이프가 민감 정보를 지운다. 받은 입력은 그대로 담아 하위로 전달한다.
|
|
21
|
+
금지 대상은 이 레포 밖의 파일이다.
|
|
22
|
+
|
|
23
|
+
## 안전 규칙
|
|
24
|
+
|
|
25
|
+
1. 판정 근거는 인용 가능한 증거로 남긴다.
|
|
26
|
+
2. 증거 없는 추정은 신뢰도를 올리지 않는다.
|
|
27
|
+
3. 실패를 성공으로 삼키지 않는다.
|
|
28
|
+
4. 재현 절차 없이 수정안을 제시하지 않는다.
|
|
29
|
+
5. 도메인 데이터는 로그에 남기지 않는다.
|
|
30
|
+
6. Read·Grep·Glob 으로 금지된 소스를 열지 않는다.
|
|
31
|
+
7. 신뢰도는 근거의 강도에 따라 정한다.
|