@chrono-meta/fh-gate 2.5.1 → 2.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/rules/fh_4axis_gate.md +26 -3
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +24 -3
- package/CATALOG.md +41 -0
- package/CLAUDE.md +77 -15
- package/README.ja.md +36 -11
- package/README.ko.md +36 -11
- package/README.md +116 -18
- package/README.zh.md +32 -11
- package/docs/USER_GUIDE.md +118 -0
- package/docs/platform_sustainability.md +174 -0
- package/knowledge/shared/GLOSSARY.md +26 -1
- package/knowledge/shared/harness-core/fh_detail_protocols.md +44 -2
- package/knowledge/shared/harness-core/field_verdict_crossfamily_gate.md +19 -5
- package/knowledge/shared/harness-core/harness_incubator_doctrine.md +129 -2
- package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +32 -0
- package/knowledge/shared/harness-core/ship_readiness_gate.md +77 -8
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +110 -0
- package/knowledge/shared/rules/knowledge_layer_seam.md +1 -1
- package/package.json +17 -3
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +136 -0
- package/plugins/fh-meta/agents/persona-innovator.md +170 -0
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +44 -0
- package/plugins/fh-meta/skills/fh/SKILL.md +31 -0
- package/plugins/fh-meta/skills/goal-quench/SKILL_detail.md +1 -1
- package/plugins/fh-meta/skills/harness-doctor/SKILL.md +7 -1
- package/plugins/fh-meta/skills/harvest-loop/SKILL.md +32 -0
- package/plugins/fh-meta/skills/public-surface-audit/SKILL_detail.md +16 -1
- package/plugins/fh-meta/skills/steel-quench/SKILL.md +97 -0
- package/scripts/adapters/fixtures/mate_agent_boundary_known_negative.md +31 -0
- package/scripts/adapters/fixtures/mate_agent_boundary_known_positive.md +56 -0
- package/scripts/adapters/peer_resolve.sh +58 -6
- package/scripts/cluster_capability_scan.sh +60 -18
- package/scripts/digest_landing_check.sh +181 -10
- package/scripts/fh_hub_identity.sh +83 -0
- package/scripts/fh_session_load.sh +53 -5
- package/scripts/fh_track_resolve.sh +114 -0
- package/scripts/field_canon_preload.sh +50 -5
- package/scripts/package_coverage_check.sh +17 -0
- package/scripts/prior_art_prompt.sh +168 -0
- package/scripts/psa_scan_lib.sh +201 -15
- package/scripts/residency_admission_check.sh +204 -0
- package/scripts/selfcheck.sh +117 -1
- package/scripts/test_adapter_lanes.sh +67 -2
- package/scripts/test_heavy_classifier_lanes.sh +144 -0
- package/scripts/test_marker_defense_lanes.sh +152 -0
- package/scripts/test_marker_soul_check_lanes.sh +211 -0
- package/scripts/test_prior_art_prompt_lanes.sh +128 -0
- package/scripts/test_psa_singlefile_lanes.sh +351 -1
- package/scripts/test_residency_admission_lanes.sh +60 -0
- package/scripts/test_track_resolve_lanes.sh +158 -0
- package/templates/.git-hooks/pre-commit +400 -4
- package/templates/.git-hooks/pre-push +17 -2
- package/templates/settings.PriorArt.snippet.json +15 -0
- package/scripts/test_satellite_publish_gate_lanes.sh +0 -339
|
@@ -2292,6 +2292,24 @@
|
|
|
2292
2292
|
finding: "6건 중 확정 2 · 기각 2 · 반쪽 1 · 자기모순 1. 🟥 **초판이 이 팔을 `partial` 로 채점했는데 운영자가 정정했다 — 발견기를 정밀도로 채점한 오류다.** 탈상관은 오라클이 아니라 **값싼 발견기**이고(다른 사고방식을 그대로 써서 다른 각도로 묻는 것이 가장 싸고 빠르다), **발견이 어렵고 채점이 싸다.** 그러면 오탐 4건은 실패가 아니라 정상 작동이고, 채점 기준은 «거버너가 혼자서는 못 낸 것을 표면화했는가» 다. 그 기준으로는 통과한다 — 아래 ⓐ가 정확히 그것이다. ✅ 확정: ⓐ 「15만 토큰 압축 임계」가 **근거 없는 하드코딩 예시값**(발표자가 슬라이드 코드의 >150000 을 읽은 것) — 🟥 거버너가 이미 자기 보고에 기준선처럼 적어 넣은 뒤였고 대조 없이는 카드까지 갔다 ⓑ cron 기반 문서 갱신이 커밋-시점 게이트 결속을 깬다. ❌ 기각: OpenWiki 자기모순 주장은 agents.md(포인터)와 log.md(변경로그)를 합친 것 · compact 가 frontmatter/[[link]] 그래프를 파손한다는 주장은 컨텍스트 압축이 디스크 파일을 안 고친다는 점에서 기전 오류(인접 위험은 실재)."
|
|
2293
2293
|
note: "채점 기준 정정의 함의: 이 원장에서 `partial`/`rejected` 는 **발견기의 정밀도**를 뜻하지 않는다 — 그건 정상 비용이다. 🟥 사이드카가 자기 자신과 모순했다 — §1 에서 「툴 4~5개는 화자 자체 모순」이라 적고 §3 에서 같은 숫자를 「공식 규격에서 기인, 근거 있음」으로 승격. 원문에 그 근거가 없고 Anthropic 실제 스펙 대조도 안 했다(미검증을 검증으로 렌더). 인용 6개는 기계 대조로 5/6 원문 확인 + 1건은 표기 변형(슬랙↔슬렉)뿐. ⇒ 값은 「많이 잡는다」가 아니라 **「거버너가 이미 쓴 것 중 틀린 걸 잡는다」** 쪽에서 나왔다."
|
|
2294
2294
|
cost: 미측정 (agy 토큰 미보고 — 0 으로 렌더하지 말 것)
|
|
2295
|
+
- date: 2026-08-18
|
|
2296
|
+
agent: general-purpose (입장/standpoint 리뷰어, qasp 유지보수자 입장)
|
|
2297
|
+
model: opus (orchestrator)
|
|
2298
|
+
purpose: "qasp-dev PR #183 — 저자(거버너)가 쓴 §B4-후속2 조사 문서를 «대상 하네스의 입장» 에서 검증. 계열 축이 아니라 **그라운드 트루스 출처를 바꾸는 축**"
|
|
2299
|
+
prompt_summary: "tier2 요건을 프롬프트에 못 박았다 — «읽기만 하면 tier1b». 실행 4건을 명시 지시: 토큰 158종 수확 재현 · payment/insurance 가 VALUE 인지 · p6 오버라이드 부재를 grep 컨트롤과 함께 재현 · 스위트 완주. 「발견 없으면 없다고 적어라, 다만 위 4개를 안 돌리고 없다고 적으면 그건 리뷰가 아니다」"
|
|
2300
|
+
outcome: accepted
|
|
2301
|
+
finding: "VERDICT=REQUEST-CHANGES. **S 2건이 PR 제목이 내건 두 주장을 정면으로 반증했고 자력 적발 0.** ⓐ 귀속 교란 — 「후보 1→3」을 계기 오염 탓으로 돌렸으나 재주행은 (a)토큰 158→45 수리와 (b)검출면 확대(분기조건→+딕트키)를 **동시에** 바꿨다. 2x2 로 갈라보니 1→3 은 전적으로 (b) 이고 (a) 의 효과는 **반대 방향**(오탐 7→1). ⓑ p6_menu_tree 를 「이식성 약결합」으로 적었으나 RiskCategory 6멤버 × 표 키 실행 대조 = **0/6 히트 = 현재 도메인 100% 라이브 결함**. A 4건(인용금지 라벨이 숫자에서 178~410줄 떨어져 구조적으로 못 만남 · 절 순서 역전 + 「B4 상태」 중복 모순 · ⓑ 재현 갭 126 미기재 · 면책이 미검토 2행을 안 덮음)도 전부 타당. 부수로 라이브 버그 2건 추가 발견(auditor_v2 의 없는 enum 멤버 참조 → AttributeError · html 렌디션이 정반대 상태 표기)."
|
|
2302
|
+
note: "🟥 **이 축의 값은 «실행» 에서 나왔다.** S·A 6건 중 정적 읽기로 나올 수 있는 것은 거의 없었다 — enum 대조·import 그래프·git log --follow·pytest 완주가 각각 근거였다. `field_verdict_crossfamily_gate.md §7` 의 «execution is the load-bearing half» 가 이 케이스로 지지된다. 반대로 **저자(나)는 같은 파일을 몇 시간 들여다보고도 6건 중 0건을 자력 적발했다** — 「읽는 자가 저자가 아니어야 닫힌다」의 재현."
|
|
2303
|
+
cost: 165,326 subagent tokens · 34 tool_uses · 807s
|
|
2304
|
+
- date: 2026-08-18
|
|
2305
|
+
agent: general-purpose (입장/standpoint 리뷰어, qasp 유지보수자 입장)
|
|
2306
|
+
model: opus (orchestrator)
|
|
2307
|
+
purpose: "qasp-dev PR #184 — **실제 코드 변경**의 입장 리뷰. #183(문서)에서 같은 축이 값을 냈으므로 코드에도 붙였다"
|
|
2308
|
+
prompt_summary: "실행 5건 명시 지시(스위트 완주+skip 목록 · fail-before 를 **코드 경로로도** 재현 · 「동작 무변화」 주장 검증 · 리포트 실물 확인 · caller-zero 기준선 미갱신 확인). 🟥 «저자는 주석 한 줄로만 증명했다 — 실제 import/호출로 심었을 때도 잡히나» 를 콕 집어 물었다. 작업은 **워크트리**에서, 본체 체크아웃은 다른 세션 것이라 switch 금지를 명시"
|
|
2309
|
+
outcome: accepted
|
|
2310
|
+
finding: "VERDICT=REQUEST-CHANGES. **S 2건이 또 헤드라인(「동작은 안 바꾼다」)을 반증했고 자력 적발 0.** ⓐ `_UNMAPPED_SEEN` 이 프로세스 전역이고 `run()` 이 초기화를 안 해서, p6.run 2회 실행 시 **강등 0건인 2회차 리포트가 1회차 강등을 자기 것으로 신고** — PR 이 스스로 명시한 컨트롤(「강등 0건이면 절이 안 나온다」)이 fresh process 에서만 참이었다. 이 PR 이 없애려는 병과 **같은 얼굴**. ⓑ `SERVICE_ORG_MAP` 별칭 전환이 문서화된 확장 경로(`docs/GHE_INTEGRATION.md`)를 **무음 no-op** 으로 바꿨다 — main 판/PR 판 대조 실행으로 `['new-org']` vs `[]`, 예외도 로그도 없음. A 5건: `profile.available` 분기 누락으로 absent/unparseable 을 한 값으로 접음(같은 PR 안에서 p6 와 ghe_mapper 두 소비처가 갈림) · 새 관측 채널에 읽는 코드·committed 테스트 0 · **레인 docstring 이 권한 우회법이 그대로 저차단 구멍**(`getattr(m,\"exit_code_\"+\"for_verdict\")` 로 실제 교차 호출이 통과) · 죽은 포인터 · 리포트 배너의 프로덕션 소비처 0. B 4건 중 하나는 **안내 문구가 사용자를 회귀로 인도**(「추가한다」인데 구현은 replace 라 한 키만 적으면 내장 12키 소멸)."
|
|
2311
|
+
note: "🟥 **탈상관 축이 서로 다른 것을 잡는다는 실증.** #183 과 #184 는 같은 형식·같은 축인데 잡은 결함의 종류가 겹치지 않았다(문서=귀속/라벨, 코드=상태누적/무음 no-op). 그리고 이 리뷰가 요구한 «코드 경로로도 심어봐라» 한 줄이 저차단 구멍을 열었다 — **저자의 fail-before 는 주석으로만 증명돼 있었다.** 짝: 수리 중 `ast.literal_eval` 이 문자열 덧셈을 지원 안 하는데 `except Exception: pass` 로 삼켜 우회가 통과했고, 같은 커밋에 넣은 판별력 테스트가 **자기 수리 안에서 재발한 무음 강등**을 잡았다."
|
|
2312
|
+
cost: 194,462 subagent tokens · 62 tool_uses · 976s
|
|
2295
2313
|
|
|
2296
2314
|
- date: 2026-08-18
|
|
2297
2315
|
agent: general-purpose (isolated)
|
|
@@ -2332,3 +2350,95 @@
|
|
|
2332
2350
|
finding: "🟥 **FH 가 몰랐던 것 넷, 그중 둘은 설계 결함이다.** ⓐ **위성이 미등록 redaction sink** — gstack 은 「타 모델 dispatch」를 이미 sink 로 분류·스캔하는데 위성 경로는 그 스캐너를 한 줄도 안 탄다. 🟥 FH 의 publish_gate 는 **산출물**을 스캔하고 **입력**은 아무도 안 본다(방향이 반대다) ⓑ 무인 acceptEdits 의 폭발반경이 **레포 밖** — 그쪽 `.claude/skills/gstack` 심링크가 라이브 글로벌 설치본이라 동시 실행 중인 남의 CC 세션을 깬다 ⓒ **AI 에게 period 로 닫힌 파일**(ETHOS.md)이 있고 제안조차 위반 — 프로필 스키마에 「금지 파일」이 필수여야 한다 ⓓ 산출 자리를 발명할 필요 없음(`~/.gstack-dev/plans/` 가 이미 의미론이 같다) + 그 레포는 «git status 오염 = 사건». 절반 ① 은 `checked(겹침 있음)` — Renovate(구조) · Sourcegraph Agentic Batch Changes(«repository specific instructions», 2026-06 상용) · gh-aw(착지 형태) 가 각 조각을 선행한다. **net-new 가 아니라 조합**이고 그 사실이 「새롭다」의 강도를 깎는다."
|
|
2333
2351
|
note: "🟥 **회수분이 따로 있다**: gstack 의 «짓기 전에 검색» 규율은 절반 ①(선행자산)을 절반 ②의 **입력**으로 요구한다 — FH 는 둘을 따로 돌렸는데 대상의 의자에 앉으면 한 문서다. ⓓ 의 두 질문이 «둘»이라는 것이 FH 쪽 구성이지 보편이 아니라는 뜻이고, 이건 6축 정의 자체에 닿는다. ⚠️ 절반 ① 은 스니펫 기반이고 1차 문서 전수 직독이 아니다 — Sourcegraph 의 그 문구가 「레포가 쓴 것」인지 「code graph 생성」인지 안 갈렸고, 그 한 줄이 갈리면 겹침이 «부분»에서 «전면»으로 바뀐다. **미확인으로 남겼다.** 신호 = tracks/_meta/fh_signal_2026-08-19_satellite-thirdparty-axis.md"
|
|
2334
2352
|
cost: 157,378 tokens (subagent_tokens · tool_uses 20 · 313s)
|
|
2353
|
+
|
|
2354
|
+
- date: 2026-08-19
|
|
2355
|
+
agent: claude-code-guide · general-purpose ×2 · general-purpose(sonnet, blind sim) · codex sidecar
|
|
2356
|
+
model: opus (orchestrator) / sonnet (blind sim) / gpt-5.5 (sidecar)
|
|
2357
|
+
purpose: "위성 프로필 스키마 게이트 + 온보딩·마감회고 설계 + 게이트 적대검증 + 플로어 티어 발화 확인 — 한 세션 5건을 클래스로 묶어 1엔트리"
|
|
2358
|
+
prompt_summary: "① 헤드리스 권한 의미론을 **1차 출처 직독**으로(추측 금지, 확인 불가는 명시) ② 마감 회고 상설화 설계 — 기존 자산 커버리지부터 갈라라 ③ 온보딩 가이드/Q&A 설계 — «이미 되는 것»이면 그렇게 결론내라 ④ **시험 대상을 안 알려준** 블라인드 4상황 sim ⑤ cross-family: fail-open 구멍·되돌림 우회를 찾아라"
|
|
2359
|
+
outcome: accepted
|
|
2360
|
+
finding: "🟥 **cross-family 가 S급 1건** — `FD_PROFILE` 미설정으로 게이트를 통째로 끌 수 있었고 **내가 쓴 레인이 그 우회를 초록으로 고정**하고 있었다(자력 적발 0). A급 4건도 전부 실물: acceptEdits 가 `sed·mv` 를 자동승인해 `sed -i` 로 금지파일 수정 가능 · 절대경로가 `.//abs` 로 조립 · 레인이 JSON 구조를 안 보고 문자열만 grep · 「egress 방화벽」 주장 과장. ★ **블라인드 sonnet sim 4/4 발화**(문 렌더 · 인라인 덤프 거부 · Linux 분기 · 없는 스킬에 「0건 근거」)이면서 **진짜 결함 1건 적발** — CLAUDE.md 서문이 신규자를 CHEATSHEET 로 오라우팅. ★ 설계 2건은 둘 다 **「새 자산 신설 없음」으로 수렴**했고, 그중 하나는 요청받은 형식(install-wizard 3부 동의)이 **이 용도엔 오용**이라고 반증했다."
|
|
2361
|
+
note: "🟥 **블라인드 sim 은 계기 생존을 확인했다**(tool_uses=7) — 2026-08-16 의 `tool_uses: 0` 죽은 sim 과 구분된다. 그러나 **reps=1 로 이 레포 자체 바(reps≥3) 미달**이고, 「발화했다」는 관측 1회지 일반화가 아니다. ⚠️ 훅 tally 는 51 인데 이 엔트리 포함 4엔트리다 — tally 는 재시도·내부 스폰까지 세므로 **1:1 대응이 아니다**(개별 기록이 아니라 클래스 집계로 남긴다)."
|
|
2362
|
+
cost: "claude-code-guide 124,771 · 설계 138,542 + 159,983 · 블라인드 sim 140,094 · codex 64,443 tokens"
|
|
2363
|
+
|
|
2364
|
+
- date: 2026-08-20
|
|
2365
|
+
agent: general-purpose ×2 (harness-doctor · pmh-dev 답습) · codex sidecar · agy sidecar · headless sonnet ×6 (블라인드 sim 양팔)
|
|
2366
|
+
model: opus (orchestrator) / gpt-5.5 (codex) / gemini-3.1-pro (agy) / sonnet (sim)
|
|
2367
|
+
purpose: "2.6.0 재출하 준비 — 소비자 설치 selfcheck 복구 · 30일 구조 진단 · pmh-dev 선행분 조사 · 메뉴 세로화 검증. 한 세션 5클래스를 1엔트리로 묶는다"
|
|
2368
|
+
prompt_summary: "① harness-doctor 30일 캐던스 — known-pair 보정 후 숫자, 못 잰 칸은 UNMEASURED 로 ② pmh-dev 3분류(net-new / PMH 앞섬 / FH 앞섬), residency 가 organization-private 인 항목은 인용 금지 ③ codex: diff 를 받고 **수리를 반증**해라(fail-open 인가 · 결박이 남았나) ④ agy: 코드 아닌 **기록의 주장**을 검증해라(표본이 결론을 지탱하나 · 미측정을 0으로 렌더했나) ⑤ 시험 대상 안 알려준 블라인드 렌더 sim, ARM/CONTROL 각 3회"
|
|
2369
|
+
outcome: accepted
|
|
2370
|
+
finding: "🟥 **cross-family 지적 3건 전부 실재·전부 채택·자력 적발 0.** codex(diff 축)가 내 수리의 반쪽을 잡았다 — L13c 라벨을 유도로 바꿨는데 **생산자는 리터럴을 낸다**, 즉 수리가 divergent-normalizer 를 새로 만들고 있었다(되돌림 + known-pair 재현: 상수 20/0 · 유도 19/1). agy(주장 축)가 둘 — ③문 자기모순, 그리고 **detail 파일의 신규사용자 문이 안 따라온 반쪽-픽스**(내 확인 grep 이 따옴표 탓에 거짓 「없음」을 냈다). 🟥 **두 계열 지적이 하나도 안 겹쳤다(코드 1 / 주장 2)** — 축 분리의 값이 실측으로 나온 자리다. harness-doctor: FAIL(M-tier 1 = 상주 140,217자) · S 9 · R 4, «못 잰 것» 8항 명시. pmh-dev: 흡수 후보 3(Wave 1-D · HEAVY 분류기 레인 · Step 0.35), 그중 **A1 은 PMH 문서가 없는 기계를 있다고 적었다**(`axis2-defense` 훅 히트 0, 컨트롤 `crossfamily` 21) — 산문만 옮기면 팬텀 기계-주장을 들여온다. 블라인드 sim ARM 3/3 vs CONTROL 3/3 완전 분리."
|
|
2371
|
+
note: "🟥 **계기가 못 보는 자리를 탈상관이 봤다.** 소비자-완주 계기는 「돌아가나」를 재지 「고친 게 옳은가」를 못 잰다 — SKIP 으로 빠지는 팔은 그 계기가 **구조적으로 안 보는** 자리이고 codex 지적이 정확히 거기 있었다. ⚠️ codex 세션이 **최종 산문 판정 없이 끝났다**(4턴, 마지막이 도구 출력) — 지적은 중간 턴에서 건졌다. 「지적 없음」이 아니라 **부분 산출**로 계상한다. ⚠️ 훅 tally 는 67 인데 이 엔트리 포함 소수다 — tally 는 재시도·내부 스폰까지 세므로 1:1 대응이 아니다(클래스 집계)."
|
|
2372
|
+
cost: "harness-doctor 196,158 · pmh 답습 186,940 tokens (subagent_tokens) · codex/agy/sim 은 CLI 라 토큰 미노출 = UNMEASURED"
|
|
2373
|
+
|
|
2374
|
+
- date: 2026-08-20
|
|
2375
|
+
agent: general-purpose ×2 (Pre-Publish 코드 보안 · CLAUDE.md 상주 원장) · codex sidecar ×3 (재출하 diff · Wave 1-D 훅 레그 · standpoint 정정문)
|
|
2376
|
+
model: opus (orchestrator) / gpt-5.5 (codex)
|
|
2377
|
+
purpose: "2.6.0 배포 후속 — 비가역 표면 보안 패스 · M-1 레버 탐색 · 신설 훅 레그와 문서 정정의 적대검증. 앞 엔트리(같은 날)와 분리한 이유는 **성격이 다르기 때문**이다: 그쪽은 조사, 이쪽은 **내 산출물에 대한 반증 발주**다"
|
|
2378
|
+
prompt_summary: "① 출하 실행코드가 소비자 머신에서 임의실행·삭제·송신·자격증명 접근을 하나(비가역 표면이니 확신 없으면 UNMEASURED) ② CLAUDE.md 절마다 트리거 클래스 × 백스톱 배선을 grep 으로 확인해 등급표를 채워라, 크기로 등급 매기지 마라 ③ 내 수리를 반증해라 — fail-open 인가, 과차단인가, 결박이 남았나 ④ 내 훅 레그를 반증해라 — 파싱 구멍, 호출되지 않는 경로, 레인이 주장하는 것을 실제로 재나 ⑤ 내 정정문을 반증해라 — 과교정인가, 놓친 전파 자리가 있나"
|
|
2379
|
+
outcome: accepted
|
|
2380
|
+
finding: "🟥 **cross-family 지적 8건, 전부 실재, 전부 채택, 자력 적발 0.** ③ 「L13c 라벨 유도는 반쪽 — 생산자가 리터럴을 낸다」 → 내 수리가 divergent-normalizer 를 새로 만들고 있었다(되돌림 known-pair 20/0 vs 19/1). ④ 3건 — 호출을 LOADBEARING 블록에 둬서 문서가 말한 «floor tiers only» 와 실제 조건이 달랐다 · **레인이 호출부를 안 쟀다**(함수를 추출해 직접 부르므로 훅의 호출을 지워도 초록인데 문서는 «pre-commit runs this» 라 주장 — **내가 sibling 을 비판한 그 결함을 한 층 위에서 재생산**) · presence 가드가 rc 기준 장식. ⑤ 4건 — **내 정정이 과교정**(`tier2(qasp)` 는 근거 없이도 rc=0 경고로 통과하는데 «근거 필수」라 적었다; 내 known-pair 가 **두 변수를 섞어** rc=1 을 틀린 원인에 귀속) · 전파 자리가 3곳이 아니라 **6곳** · AGENTS.md 의 «one line» 이 훅보다 좁다 · **내가 적은 줄번호가 내 직전 커밋 때문에 이미 밀려 있었다**. ① 보안 PASS(MED 2 명시) — `bypassPermissions` 경로가 실물 tarball 에서 도달 불가임을 **실행으로** 확인, 단 «의존성 부재라는 우연이지 명시 게이트가 아니다」. ② **M·S·R 전부 0** — 절 단위 레버 부재가 실측 결론이고 capability-level merge 만 남는다."
|
|
2381
|
+
note: "🟥 **오늘 값을 낸 축은 계기가 아니라 탈상관과 되돌림이다.** 소비자 완주(계기)는 «돌아가나」를 재지 «고친 게 옳은가」를 못 쟀다 — SKIP 으로 빠지는 팔이 그 계기의 구조적 사각이었고 codex 지적이 정확히 거기 있었다. 되돌림 프로브는 별도로 4건을 자력 적발했다(장식 가드 2 · 내 픽스처 순서 버그 1 · 앵커 격리 확인). ⚠️ codex 세션 하나가 **최종 산문 판정 없이** 끝나 중간 턴에서 지적을 건졌다 — 「지적 없음」이 아니라 **부분 산출**로 계상한다. ⚠️ 훅 tally 124 vs 엔트리 2 는 1:1 이 아니다(재시도·내부 스폰 포함) — 클래스 집계다."
|
|
2382
|
+
cost: "보안 183,520 · 상주원장 144,614 tokens (subagent_tokens) · codex ×3 는 CLI 라 미노출 = UNMEASURED"
|
|
2383
|
+
|
|
2384
|
+
- date: 2026-08-20
|
|
2385
|
+
agent: fh-meta:beginner ×2 · fh-meta:main-player ×2 · fh-meta:challenger ×2 · fh-meta:expert ×2 (챔버 런 #14 · #15, 각 런 4개)
|
|
2386
|
+
model: opus (전부)
|
|
2387
|
+
purpose: "인큐베이터 챔버 두 런의 step-4 블라인드 페르소나 + ⓓ3자대면 큐레이션. 런 #14 `multisurface-reading-harness` → EMIT(발표 준비 하네스), 런 #15 `interslide-dependency-graph` → EMIT ∧ **WITNESSED**"
|
|
2388
|
+
prompt_summary: "① 후보 정본을 냉독하고 어디서 이해가 깨지는지(경로·줄번호) ② 실사용 가치 + **명시 배정된 반대편 변호**(#14 는 「도구로 선다」 쪽, #15 는 「기계로 못 선다」 쪽 — 역할과 페르소나를 탈상관) ③ 적대검증: known-negative 가 만들어지나·미탐 방향·상위 하네스와의 중복 ④ 외부 선행 조사(URL 인용 필수, 확인 못 한 것은 미확인 표기). 🟥 **INTENT 는 주지 않았다** — #15 부터는 사전 봉인 예측을 워크스페이스 **밖**(비공개 컴패니언 스토어)에 두고 INTENT 엔 sha256 만"
|
|
2389
|
+
outcome: accepted
|
|
2390
|
+
finding: "🟥 **설계를 준 것이 검증보다 컸다 — 이게 오늘의 net-new 관측이다.** #15 에서 거버너의 계기가 known-positive 를 **미탐**하고 있었는데, 블라인드 냉독이 실물 8건을 손분류해 **판별자 자체를 줬다**(«트리거는 회고 부사가 아니라 «이 어구가 이 장에서 처음 정의되는가»»). 그 규칙으로 갈아끼우자 known-pair 통과. ★ **이미 배출한 하네스의 결함 2건을 페르소나가 잡았다**: L1 이 `>` 를 무조건 주석으로 떨궈 **원고 낭독을 통째로 안 보면서 초록**이었다(판정 0건→12건) · 은퇴 선언을 헤딩에서만 찾아 죽은 장 4개를 live 로 계상. **거버너 자력 적발 0**, 둘 다 실행으로 확정. ★ #14: 「인스턴스 1개」가 실물 대조로 **반증**(덱 밖 5개, 소급 귀속 아님 — 그 코드 저자가 렌즈 이전에 자기 언어로 같은 메커니즘을 적었다) · ⓓ3자대면 첫 실행이 §3 을 «거의 전량 선행»(RTE·BX/lens·XLIFF·Jupytext·pandoc·token-drift)으로 뒤집었고, 그 결과가 **출구가 아니라 재료**가 됐다(운영자 결정). ★ #15 expert: 정밀도 선행이 설계를 바꿨다 — bridging 자동해소 **F1 26~30**, discourse deixis **21.5** ⇒ 자동 «판정»으로 설계하면 미검출이 지배한다(`not_found_is_not_zero_family`). 추출=재현율 · 판정=사람으로 범위 재조정."
|
|
2391
|
+
note: "🟥 **#14 에서 challenger 가 INTENT.md 를 읽고 자진 신고했다** — 워크스페이스가 페르소나와 같은 레포에 있으니 «주지 마라» 라는 산문은 격리를 못 만든다. **격리는 산문이 아니라 경로다.** #15 에서 봉인을 워크스페이스 밖으로 빼자 아무도 못 읽었다(처방 닫힘). ⚠️ 그리고 그 신고가 없었으면 «독립 수렴 3건» 으로 오계상됐을 것이다 — 자진 신고에 의존하는 구조라 다음에도 잡힌다는 보장이 없다. ⚠️ K2(도구/판단)에서 두 페르소나가 **반대 결론**을 냈고, 그 갈림 자체를 «안 닫힌 축» 으로 기록했다(한쪽으로 접지 않았다)."
|
|
2392
|
+
cost: "런 #14: 508,284 · 런 #15: 620,426 tokens (subagent_tokens, 기계 출처) · 거버너 = **UNMEASURED**(세션이 자기 소비를 못 읽는다) · 🟥 **합계 안 적는다** — 미측정 칸을 0 으로 접는 것이다. #15 는 추정 450k 대비 +38%(CAP 700k 이내), #14 는 추정 단위가 달라 UNCALIBRATED 였고 그 교훈이 #15 의 단위 정정으로 갔다"
|
|
2393
|
+
|
|
2394
|
+
- date: 2026-08-20
|
|
2395
|
+
agent: codex/gpt-5.6-terra (cross-family sidecar, headless)
|
|
2396
|
+
caller: FH hub session (air node)
|
|
2397
|
+
purpose: adversarial refutation of the sync-to-be.sh destination-newer abort-message fix (load-bearing data-loss guard)
|
|
2398
|
+
reps: 1
|
|
2399
|
+
outcome: accepted
|
|
2400
|
+
evidence: "4 findings, all reproduced and all fixed. 1 HIGH (the fix planted a fresh dead pointer at the file-level guard site: sync_file's only call site is CLAUDE.local.md, which sync-from-be.sh refuses by name), 2 MED (unquoted printed command; overclaimed 'discriminator'), 1 LOW (exactly-two-causes overclaim). Governor source-grounded each by grep before accepting. Self-catch on these: 0/4."
|
|
2401
|
+
note: "Recorded per CLAUDE.md §Agent Dispatch invocation-log obligation. Sidecar leg, not an Agent-tool subagent, so the SubagentStop tally does not see it."
|
|
2402
|
+
|
|
2403
|
+
- date: 2026-08-20
|
|
2404
|
+
agent: Explore ×1 (ship_readiness_gate 발췌) · codex/gpt-5.6-terra ×2 (착지계기 REFUTE 레그 · 캘리브레이션 컨트롤)
|
|
2405
|
+
model: opus (orchestrator) / gpt-5.6-terra (codex)
|
|
2406
|
+
caller: FH hub session 08d6fe75 (air node) — peer 로부터 훅 수정 인계받은 축
|
|
2407
|
+
purpose: "정체성 ④ 등급 판정 + 착지 계기 결함 수리. Explore 는 등급표 ④행·§Gate consequence·등급 정의를 **원문 인용으로** 발췌(요약 금지 — 조건을 무르게 만들지 말라고 명시). codex 는 내 수리를 반증"
|
|
2408
|
+
prompt_summary: "① ship_readiness_gate.md 에서 ④행·압도성 절·🟢/🔵 정의·screener 잔여 처리를 파일:줄 붙여 원문 발췌, 못 찾으면 「못 찾음」이라 하고 추정 금지 ② 내 DLC_EXCLUDE_TARGETS 수리를 REFUTE — 비인용 확장/경로매칭/과잉제외/러너배선/레인품질/대안설계 6축, 각 항목에 성립·불성립과 근거 줄"
|
|
2409
|
+
reps: 1
|
|
2410
|
+
outcome: accepted
|
|
2411
|
+
finding: "🟥 **codex HIGH 3 · MED 1 성립, LOW 2 정당 기각 — 자력 적발 0.** HIGH#1 비인용 확장의 글로빙·단어분리로 공백 경로면 제외가 빗나감 = **fail-open, 옛 거짓 양성 재발**. HIGH#2 `./` 접두가 git 경로와 불일치하는데 러너 `-f` 게이트는 실재만 보므로 **무신호로** 실패. LOW#4·#5 는 불성립이나 **하위지적 둘이 유효**했다(조건부 export 가 부모 환경을 상속 · 레인 b 가 `not10` 만 요구해 «안 죽었다» 만 보증). HIGH#6 은 채택 안 함(판정 의미론 변경이라 분리). ★ Explore 는 **카드가 물려준 미결을 무효화**했다 — 「screener 잔여가 명시 잔여냐 보류 사유냐」는 §④ promotion criteria(08-17)가 이미 **재분류**해 뒀다(계기가 재는 것은 «조직 전파»가 아니라 «허브 내부 착지»라 그 조건을 100% 닫아도 ④ 명제는 한 글자도 안 재진다). 즉 결정할 것이 아니라 틀린 이지선다였다."
|
|
2412
|
+
note: "🟥 **사이드카 캘리브레이션이 필요했다.** 첫 두 codex 런이 exit 0 인데 출력이 프롬프트 에코에서 끊겼다(바이트 동일 2,123 = 비결정 아님). 「지적 없음」으로 읽지 않고 known-answer 컨트롤(2+2)을 돌려 **사이드카는 살아 있음**을 확인, 변수 하나씩 갈라 원인이 **긴 프롬프트의 2분 초과**임을 특정했다. exit 0 을 통과로 읽었으면 crossfamily 를 거짓으로 적었을 자리다. ⚠️ Explore 는 요청 밖 항목(P4-1 선행조건 미충족·프로덕션 호출부 0개·§Gate consequence 6행 중 4행 비일관)까지 냈고 그게 판정의 하중이 됐다."
|
|
2413
|
+
cost: "Explore 60,857 tokens (subagent_tokens) · codex ×2 는 CLI 라 미노출 = UNMEASURED · 거버너 = UNMEASURED. 🟥 합계 안 적는다"
|
|
2414
|
+
|
|
2415
|
+
- date: 2026-08-21
|
|
2416
|
+
agent: general-purpose ×4 (L3 drift · L4 connection · L5 pattern · residency ledger)
|
|
2417
|
+
model: opus (orchestrator + all four legs)
|
|
2418
|
+
caller: FH hub session 27a9ba28 — /harness-doctor 정기 진단 (직전 실행 2026-07-20, 캐던스 30일 초과)
|
|
2419
|
+
purpose: "harness-doctor L3~L5 + 메타하네스 residency ledger 를 네 렌즈로 병렬 분해. 거버너는 L1·L1-E·푸터프린트·pointer-illusion·SKILL 크기를 직접 측정하고, 렌즈끼리 파일이 안 겹치게 스코프를 갈랐다."
|
|
2420
|
+
prompt_summary: "각 레그에 계기 규율을 명시 주입 — ① known-positive/known-negative 쌍으로 계기 판별력 먼저 검정하고 못 가르면 UNCALIBRATED ② 카운트마다 손검증 1건 ③ not found = UNMEASURED, 0 아님 ④ tier 확정 금지, 후보만. L5 에는 「기록하지 않는 소스를 grep 해서 0회를 내면 그건 측정이 아니라 생성」을 명시."
|
|
2421
|
+
reps: 1
|
|
2422
|
+
outcome: pending
|
|
2423
|
+
finding: "미착지 — 네 레그 진행 중. 거버너 자체 측정분은 확정: always-loaded 144,205자(M-tier, 임계 80k) · memory-index 25,472자(S-tier, 임계 10k)이며 **로더 하드리밋 초과로 엔트리 5건이 이 세션에 실제로 미로드**(손검증). ★ pointer-illusion 「2건」은 **둘 다 오탐**이었고 원인은 harness-doctor SKILL.md 가 들고 있는 정규식 자신 — `templates/` 접두를 떨어뜨리고 `{project}/{domain}/` 자리표시자를 실경로로 읽는다. 손검증 안 했으면 M-tier 두 건을 지어낼 자리였다."
|
|
2424
|
+
note: "운영자 상시 요청(CLAUDE.local.md, lease→2026-11-09, scope=서브에이전트 디스패치)에 따라 건별 승인 없이 디스패치. 워크플로/deep-research 는 범위 밖이라 안 씀. 단위는 문자(chars)로 통일 — 로더 경고 24.9KB 가 25,472자/1024 와 정확히 일치해 로더도 문자 기준임을 확인(바이트로 쟀으면 40,657 로 과대계상)."
|
|
2425
|
+
cost: "UNMEASURED (완료 알림의 subagent_tokens 로 마감 시 갱신) · 거버너 = UNMEASURED. 합계 안 적는다"
|
|
2426
|
+
|
|
2427
|
+
- date: 2026-08-21
|
|
2428
|
+
agent: general-purpose (×10, consolidated)
|
|
2429
|
+
model: inherited (session default)
|
|
2430
|
+
purpose: "개입규칙 41개 정밀도 프로브 — 블라인드 라벨 2팔 + 창 평가 7배치 + 미판정 1건 재실행"
|
|
2431
|
+
prompt_summary: "라벨팔은 사람 발화만(key), 평가팔은 세션기록만(win) — 블라인드를 프롬프트가 아니라 파일 경계로 걸었다. 서브에이전트가 프로젝트 지시를 상속하므로 프롬프트 블라인드는 안 먹는다"
|
|
2432
|
+
outcome: accepted
|
|
2433
|
+
finding: "라벨 두 팔 완전일치 122/140(87%) — 같은 급 과제의 공개 벤치마크 Cohen κ 0.70~0.74 와 동급. 평가 44창에서 41개 규칙 중 12개만 발화, 판별력(POS−NEG≥3) 2개. 사전등록 중단조건 「무분리」 HIT"
|
|
2434
|
+
note: "🟥 평가팔 하나가 «창 하나를 안 읽고 빈 배열을 냈다»고 자진 신고 — 0으로 안 세고 격리 후 재실행(미판정≠0). 라벨 손검증 6건에서 오탐 1건 확인되어 «애매하면 약한 쪽» 지시가 INTERVENE 을 부풀린 것을 발견, 평가 집합을 «양팔 합치 ∧ 불확실 표시 없음»으로 좁혔다. 위임 자체는 값을 했으나 **판정 축의 결함은 위임이 아니라 운영자가 잡았다**(자력 적발 0)"
|
|
2435
|
+
- date: 2026-08-21
|
|
2436
|
+
agent: 워크플로 5판 (146 에이전트) · fh-meta:{beginner,main-player,challenger} 3 · codex/gpt-5.6-terra REFUTE 2 · Explore 1
|
|
2437
|
+
model: opus (오케스트레이터·페르소나) / gpt-5.6-terra (codex) / sonnet (sim 에이전트)
|
|
2438
|
+
caller: FH hub session 08d6fe75 (air node) — 정체성 ④ 판정 → 재정의 → 챔버 런 → T2 훅
|
|
2439
|
+
purpose: "«세션이 쎄함을 알아채고 확인을 제안하는» 능력의 인큐베이션. 5판 = ① 맥락발화 sim v2(39) ② 되게만들기 3실험(34) ③ tracks 개입 코퍼스 전수(46) ④ 대화원본 개입 전수(9) ⑤ 타이밍 sim(18)"
|
|
2440
|
+
reps: 3 (모든 sim 팔)
|
|
2441
|
+
outcome: accepted
|
|
2442
|
+
finding: "🟥 **자력 적발 0. 열 건 넘게 전부 남이 잡았다** — 운영자 6 · codex(HIGH 7·MED 5, BLOCK 판정) · challenger(S 4, «이미 지어져 배선된 novelty_claim_check.sh 재발명») · beginner(HARD 6, «배치 vs 인터럽트») · 레인 2 · CI 2. ★ 측정으로 확정된 계단: **명시 지시 3/3 · advisory 0/3 · 프레이밍 0/3**(3프로브 18판, 독립 판정자). ★ 개입 코퍼스 전수(대화원본 433발화): 개입 54주장/~43추정, **ⓒ판단결함 32(59%) · ⓑ내부미조회 9 · ⓐ외부미조회 6(11%)** — 초판 설계가 겨눈 «세계에 물어봐» 는 소수 클래스였다. ★ 타이밍 sim 이 **봉인 예측 5중 4를 반증**했고 전부 내 제안에 불리한 방향. 교란(과제가 검색을 이름으로 부름)을 스스로 지목. 교란 안 된 대비 하나가 값을 냈다 — 양팔 다 검색, **힌트 쥔 쪽만 정답 도달(2/3 vs 0/3)** ⇒ 기전은 «찾게 함» 이 아니라 **«이름 붙이게 함»**."
|
|
2443
|
+
note: "🟥 **오케스트레이션의 값은 산출이 아니라 «내가 못 보는 것을 봤다» 였다.** 특히 challenger 가 12일 전 같은 결함·같은 운영자 발화로 지어진 `novelty_claim_check.sh` 를 찾아냈다 — 재발명 방지 능력을 제안하면서 이미 있는 더 엄격한 기계를 못 찾은 것이 이 런의 진짜 산출이다. ⚠️ 「독립 수렴」 을 두 번 썼다가 두 번 다 철회했다(같은 저자·캐논·모델 계열 ⇒ 상호보강이되 상관됨) — 오늘 스스로 인용한 규칙을 몇 시간 뒤 위반한 형태."
|
|
2444
|
+
cost: "39판 4,001,378 · 34판 3,794,737 · 46판 5,660,769 · 9판 1,045,403 · 18판 2,077,707 = **subagent_tokens 16,580,000 (기계 출처)** · codex ×2 = UNMEASURED(CLI) · 거버너 = **UNMEASURED**(세션이 자기 소비를 못 읽는다). 🟥 미측정 칸을 0 으로 접지 않는다."
|
|
@@ -23,7 +23,7 @@ provenance: pmh-dev@cbe3932 knowledge/shared/rules/knowledge_layer_seam.md 역
|
|
|
23
23
|
| 소비자 | 무엇을 읽나 | 언제 | 상태 |
|
|
24
24
|
|---|---|---|---|
|
|
25
25
|
| `phantom-quench` **Step 2-O** | 진입 인덱스(`INDEX.md`/`index.md`/`README.md`/`readme.md` — 판정기와 동일 후보) → 후보 페이지 | **도메인 주장**(조직 고유 사실·용어·정책)이 Step 2 에서 **미선언**일 때 | ✅ **배선됨** (2026-08-10 역수확, FH 1호) |
|
|
26
|
-
| `steel-quench` Step 0.35 | 진입 인덱스(위와 동일 후보) → 관련 정책·용어·도메인 사실 | 공격 각도를 정하기 전 |
|
|
26
|
+
| `steel-quench` Step 0.35 | 진입 인덱스(위와 동일 후보) → 관련 정책·용어·도메인 사실 | 공격 각도를 정하기 전 | ✅ **배선됨** (2026-08-20 역수확, FH 2호) — 🟥 단 **살리언스 층이지 기계 바닥이 아니다.** 훅도 레인도 이 절을 강제하지 않는다. 원 필드 문서에는 강제 취지의 서술이 딸려 있었으나 실측하니 그 훅 레인이 **부재**해서(팬텀 기계-주장) **가져오지 않았다.** 「배선됨」은 «절차가 스킬에 존재한다» 는 뜻이고 «기계가 강제한다» 는 뜻이 아니다 |
|
|
27
27
|
|
|
28
28
|
**1호 소비자 배선의 의미**: 그전까지 도메인 주장은 로컬 선언 파일에도 외부 인용에도 안 걸려
|
|
29
29
|
**전부 🔴 Source-Missing 으로 떨어졌다** — 조직 지식층이 정확히 그 근거인데 탐색 경로에 없었다.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@chrono-meta/fh-gate",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.7.0",
|
|
4
4
|
"description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"keywords": [
|
|
@@ -53,6 +53,7 @@
|
|
|
53
53
|
"CLAUDE.md",
|
|
54
54
|
".claude/registry/agent_cards.json",
|
|
55
55
|
".claude/registry/README.md",
|
|
56
|
+
"docs/USER_GUIDE.md",
|
|
56
57
|
"docs/CONTRIBUTING.md",
|
|
57
58
|
"bin/fh-codex-doctor.js",
|
|
58
59
|
"bin/fh-gate.js",
|
|
@@ -88,12 +89,15 @@
|
|
|
88
89
|
"scripts/adapters/qasp_web_rules.sh",
|
|
89
90
|
"scripts/adapters/fixtures/qasp_web_rules_known_positive.json",
|
|
90
91
|
"scripts/adapters/fixtures/qasp_web_rules_known_negative.json",
|
|
92
|
+
"scripts/adapters/fixtures/mate_agent_boundary_known_positive.md",
|
|
93
|
+
"scripts/adapters/fixtures/mate_agent_boundary_known_negative.md",
|
|
91
94
|
"scripts/test_adapter_lanes.sh",
|
|
92
95
|
"scripts/test_fh_gate_regressions.sh",
|
|
93
96
|
"templates/local_fh_context.md",
|
|
94
97
|
"docs/ETHOS.md",
|
|
95
98
|
"docs/WHY.md",
|
|
96
99
|
"docs/OUTPUT_EVIDENCE.md",
|
|
100
|
+
"docs/platform_sustainability.md",
|
|
97
101
|
"knowledge/shared/GLOSSARY.md",
|
|
98
102
|
"knowledge/shared/patterns",
|
|
99
103
|
"knowledge/shared/plugin-catalog",
|
|
@@ -114,10 +118,15 @@
|
|
|
114
118
|
"scripts/substrate_jump_detector.sh",
|
|
115
119
|
"scripts/tier_census_grep.sh",
|
|
116
120
|
"scripts/test_marker_floor_lanes.sh",
|
|
121
|
+
"scripts/test_marker_defense_lanes.sh",
|
|
117
122
|
"scripts/test_marker_crossfamily_lanes.sh",
|
|
118
123
|
"scripts/test_marker_standpoint_lanes.sh",
|
|
119
124
|
"scripts/test_marker_thirdparty_lanes.sh",
|
|
120
125
|
"scripts/test_marker_axes_run_lanes.sh",
|
|
126
|
+
"scripts/test_marker_soul_check_lanes.sh",
|
|
127
|
+
"scripts/test_heavy_classifier_lanes.sh",
|
|
128
|
+
"scripts/residency_admission_check.sh",
|
|
129
|
+
"scripts/test_residency_admission_lanes.sh",
|
|
121
130
|
"scripts/test_env_purity_lanes.sh",
|
|
122
131
|
"scripts/.env_purity_tokens.defaults",
|
|
123
132
|
"scripts/env_purity_scan.sh",
|
|
@@ -129,7 +138,6 @@
|
|
|
129
138
|
"scripts/chamber_candidate_collect.sh",
|
|
130
139
|
"scripts/chamber_witness.sh",
|
|
131
140
|
"scripts/digest_landing_check.sh",
|
|
132
|
-
"scripts/test_satellite_publish_gate_lanes.sh",
|
|
133
141
|
"scripts/relay_channel.sh",
|
|
134
142
|
"scripts/test_relay_channel_lanes.sh",
|
|
135
143
|
"scripts/fh_session_load.sh",
|
|
@@ -212,6 +220,9 @@
|
|
|
212
220
|
"scripts/judgment_circuit_lint.sh",
|
|
213
221
|
"scripts/novelty_claim_check.sh",
|
|
214
222
|
"templates/settings.Compaction.snippet.json",
|
|
223
|
+
"scripts/fh_hub_identity.sh",
|
|
224
|
+
"scripts/fh_track_resolve.sh",
|
|
225
|
+
"scripts/test_track_resolve_lanes.sh",
|
|
215
226
|
"scripts/field_canon_preload.sh",
|
|
216
227
|
"scripts/test_field_canon_lanes.sh",
|
|
217
228
|
"templates/settings.FieldCanon.snippet.json",
|
|
@@ -222,6 +233,9 @@
|
|
|
222
233
|
"scripts/gate_anchor_check.sh",
|
|
223
234
|
"scripts/portability_lint.sh",
|
|
224
235
|
"scripts/rtk_gross_ablation.sh",
|
|
225
|
-
".claude/capabilities"
|
|
236
|
+
".claude/capabilities",
|
|
237
|
+
"scripts/prior_art_prompt.sh",
|
|
238
|
+
"scripts/test_prior_art_prompt_lanes.sh",
|
|
239
|
+
"templates/settings.PriorArt.snippet.json"
|
|
226
240
|
]
|
|
227
241
|
}
|
|
@@ -10,6 +10,142 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/)
|
|
|
10
10
|
|
|
11
11
|
## Plugin Level
|
|
12
12
|
|
|
13
|
+
### [2.7.0] — 2026-08-21
|
|
14
|
+
|
|
15
|
+
### 🟥 BREAKING (gate): 2026-08-21 이후 날짜의 마커는 `①영혼` 줄이 없으면 커밋이 막힌다 — 없으면 `soul: 없음` 한 줄로 통과한다
|
|
16
|
+
|
|
17
|
+
이 필드는 `CLAUDE.md §자기 대조` 가 **2026-08-09 부터 의무**로 정한 것이고, 이 릴리스는 그것을
|
|
18
|
+
새로 요구하는 게 아니라 **처음으로 읽는다.** 소급 안 한다(그 날짜 이전 마커는 그대로).
|
|
19
|
+
**선언된 부재는 1급 값이다** — `없음`/`none`/`n/a` 는 통과하고, 통과가 아니라 **기록**으로 찍힌다
|
|
20
|
+
(`⚠️ ①영혼: 없음 — declared absent. Recorded, not silent.`). 빈 필드만 막힌다.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
**마커의 `①영혼` 은 6주째 의무였는데, 읽는 코드가 0줄이었다.**
|
|
25
|
+
|
|
26
|
+
컨트롤 동반 실측: 훅에서 `crossfamily` 는 **21곳**에서 검사되는데 `①영혼`/`soul` 을 읽는 코드는
|
|
27
|
+
**0줄**이었고, 게이트 명세(`fh_4axis_gate.md`)는 그 필드 이름조차 안 적는다. 실물 코퍼스에서
|
|
28
|
+
**2026-08-10 이후 마커 98건 중 37건(37.8%)** 이 어떤 표기로도 그 줄을 안 갖고 있었다 — 그중
|
|
29
|
+
손검증한 하나는 **codex+agy 패널 · 28레인 · controls alive · 나머지 필드 전부 채운** 마커다.
|
|
30
|
+
**시킨 축은 다 돌렸고, 아무 기계도 안 읽는 그 한 줄만 없었다.**
|
|
31
|
+
|
|
32
|
+
- **새 축이 아니다 — ⓒ 격리 그라운딩의 확장이다.** `fh_three_layer_canon.md §1-a-2` 의 판별자는
|
|
33
|
+
«무엇을 받았는가»이고, ⓒ 는 이미 «저자가 쓴 문장 + 지금의 트리» 를 받는다 — 사전 선언 검사의
|
|
34
|
+
입력과 한 글자도 다르지 않다. 시제(사전 선언 ↔ 사후 주장)는 적대성과 같은 **자세**지 축이 아니다.
|
|
35
|
+
🟥 초안은 ⓖ 로 뽑으려 했고 **운영자가 기존 카테고리 검토를 지시해 정정했다** — 그대로 갔으면
|
|
36
|
+
정본 6축·`axes-run` enum·덱 표·마커 190건 재해석이 전부 따라왔을 것이고, 그건 정본 안에 이미
|
|
37
|
+
기록된 오류(저자가 ⓓ에 귀속한 5건 중 3건이 다른 축으로 재분류)의 반복이었다.
|
|
38
|
+
- **검사하는 것은 채널뿐이다.** 선언이 있는가(줄머리 키) · 공허하지 않은가 · `soul-check:` 가 닫힌
|
|
39
|
+
enum 인가 · **같은 레코드 안에서 정합적인가**(`reflected` 를 적으려면 그 마커에 실제로 선언이
|
|
40
|
+
있어야 한다). **검사하지 않는 것**: 사전에 썼는가(파일에서 안 갈린다) · 그 판단이 옳은가.
|
|
41
|
+
날조된 `reflected` 는 통과한다 — §4-b(cross-family 가 마커를 읽는다)의 몫이고 여기서 안 닫혔다.
|
|
42
|
+
- **검증**: 레인 44개(BLOCK/PASS 양방향 · 실물 코퍼스 known-pair 2쌍) · 되돌림에서 **정확히 대응
|
|
43
|
+
7레인만** 적색 · 격리 클론 실 커밋 2팔(선언 없음 `rc=1` / 있음 `rc=0 ALL AXES PASSED`).
|
|
44
|
+
- 🟥 **자력 적발 0 이 14건.** cross-family 2계열 분리 발주(codex=diff 축 / agy=주장 축) 10건 +
|
|
45
|
+
첫실사용 4건, **겹친 지적 0**. 그중 하나는 **사이드카의 수리가 새 fail-open 을 만든 것**이다 —
|
|
46
|
+
넓힌 탐지기가 `soul-check:` 줄 자신의 텍스트를 증거로 잡아 BLOCK 레인이 PASS 로 뒤집혔고
|
|
47
|
+
**레인 41개가 전부 초록이었다.** 잡은 것은 격리 클론 첫실사용이다.
|
|
48
|
+
|
|
49
|
+
**유출 스캐너가 zsh 에서 죽었고, 그 죽음이 「깨끗함」과 바이트 동일이었다.** (PR #480)
|
|
50
|
+
|
|
51
|
+
`psa_scan_tagged` 가 `local path` 를 선언하는데 **zsh 에서 `path` 는 `PATH` 와 tied 된 특수
|
|
52
|
+
배열**이라 그 스코프의 `PATH` 가 비고 첫 실행문(`input=$(cat)`)부터 죽는다. 계약이
|
|
53
|
+
「빈 입력 = `return 0`」이라 **출력이 깨끗한 스캔과 구별되지 않았다.** `/public-surface-audit` 가
|
|
54
|
+
한 줄도 안 스캔하고 「유출 없음」을 냈고, 그 스킬은 **Pre-Publish Gate 가 publish 직전에 1번으로
|
|
55
|
+
체이닝하는 렌즈**다.
|
|
56
|
+
|
|
57
|
+
- 훅 경로(`pre-commit`·`pre-push`)는 bash 로 `execve` 되어 **원래 안전했다.** 뚫린 것은
|
|
58
|
+
**에이전트가 Bash 툴로 치는 스킬 본문**이고, macOS 기본 셸이 zsh 라서 난다.
|
|
59
|
+
- 수리는 둘이다: **A** 식별자 개명 · **E** `psa_scan_tagged` 진입에 생존성 가드 배선
|
|
60
|
+
(실패 → `rc=3 NOT SCANNED`). 🟥 **A 만으로는 계열이 안 닫힌다** — 같은 형태가 7곳 더 있고 그중
|
|
61
|
+
6곳은 «오늘 안전한데 그건 호출부가 bash 라서»지 코드가 이식성 있어서가 아니다.
|
|
62
|
+
**A 는 이 실례를, E 는 다음 실례를 막는다.**
|
|
63
|
+
- ⚠️ **소비자에게 보이는 변화**: 패턴 파일이 **부분만 로드되는** 설치본은 이제 `rc=3`
|
|
64
|
+
(NOT SCANNED)를 받는다 — 전에는 `rc=0`(깨끗)이었다. 방향은 옳지만 **막히는 경우가 늘었다**.
|
|
65
|
+
잘못된 패턴 파일을 쓰던 소비자는 그 사실을 이제 통보받는다.
|
|
66
|
+
- **rc 계약 `0 깨끗 · 1 유출 · 3 미측정` 을 1과 3으로 접지 않는다** — 합치면 `PUBLIC_SURFACE_OK=1`
|
|
67
|
+
이 **계기 사망까지 «승인된 유출»로 통과**시킨다. 비가역 표면에서 가장 나쁜 조합이다.
|
|
68
|
+
- **소비자 설치본 실측**(리뷰에서 추가): 깨끗한 클론 → `npm pack` → `node_modules` 경로 추출 →
|
|
69
|
+
**스킬 본문 호출 형태 그대로** known-pair. 수리 전 `zsh` 양성 **rc=0 무보고**(유일한 흔적은
|
|
70
|
+
stderr 한 줄) → 수리 후 **rc=1 보고**. bash 팔은 양쪽 불변(컨트롤).
|
|
71
|
+
|
|
72
|
+
**릴리스 표면 — 두 계보가 한 이름공간을 쓰고 있었다.**
|
|
73
|
+
|
|
74
|
+
GitHub Releases 가 **v0.3.0(8/16)을 「Latest」로** 보여주는 동안 배포물은 **2.6.0** 이었다. 둘 다
|
|
75
|
+
`vX.Y.Z` 를 쓰고 정체성 계보에만 Release 객체가 있었기 때문이다. 🟥 **이건 이 저장소가 자기
|
|
76
|
+
게이트에서 반복해 찾아낸 그 결함(두 층·한 이름)이 자기 버전 번호에 난 것이다.**
|
|
77
|
+
|
|
78
|
+
- **통일하지 않는다** — 두 숫자가 다른 것을 잰다(무엇을 설치하는가 ↔ 얼마나 익었는가). 합치면
|
|
79
|
+
성숙도 신호가 사라지고, **`identity-v1.0.0` = 전정체성 🟢** 라는 마일스톤이 죽는다.
|
|
80
|
+
- 이름공간만 가른다: 정체성 계보는 **`identity-v0.4.0` 부터** 접두어를 갖는다. 기존
|
|
81
|
+
`v0.1.0`·`v0.2.0`·`v0.3.0` 은 **개명하지 않는다** — 공개 ref 재작성은 비가역이라
|
|
82
|
+
Destructive-Op 게이트가 우리 자신의 태그에도 적용된다.
|
|
83
|
+
- `README §Two version numbers` 신설 ·
|
|
84
|
+
🟥 `v2.6.0` Release 객체를 만들었다가 **같은 시각에 되돌렸다**(운영자 지적) — GitHub 의 «Latest»
|
|
85
|
+
배지는 **슬롯이 하나**라 두 계보가 한 페이지에 있으면 경쟁하고, 배지를 쥔 쪽이 «이 저장소가
|
|
86
|
+
뭐라고 말하는가»를 정한다. 패키지 번호가 표제를 가져가면서 **성숙도 주장이 그 아래로 밀렸다.**
|
|
87
|
+
오독은 더 싼 절반(기존 본문 맨 위 한 줄)이 이미 닫았다. ⇒ **Releases 는 정체성 계보만** 나른다 ·
|
|
88
|
+
`v0.3.0` 본문 **맨 위**에 계보 한 줄(본문 4문단 아래엔 이미 있었다 — **gate-locality**: 읽는
|
|
89
|
+
자리에 없으면 없는 것이다).
|
|
90
|
+
- ⚠️ 정본의 stale 하나 같이 정정: `ship_readiness_gate.md` 가 npm 을 *"`1.4.x` range"* 로 적고
|
|
91
|
+
있었다(실제 2.6.0). **산문에 박힌 버전 숫자는 조용히 낡고 사실처럼 읽힌다.**
|
|
92
|
+
|
|
93
|
+
**사이드카는 감사하지 쓰지 않는다.** (`multi_model_sidecar_strategy.md §Runtime Authority`)
|
|
94
|
+
|
|
95
|
+
기존 교리는 «사이드카 finding 은 증거 후보이지 판정이 아니다»를 **판정 축**에서만 말했고 **쓰기
|
|
96
|
+
축이 비어 있었다.** 실측: 적대 감사자로 발주된 사이드카가 워킹트리를 **직접 편집**했고(`mtime`
|
|
97
|
+
으로 적발 — 어떤 게이트도 못 잡았다), 그 수리가 **자기참조 fail-open** 을 새로 만들었다.
|
|
98
|
+
🟥 **금지 근거는 「월권」이 아니라 「고친 쪽과 검사하는 쪽이 같아진다」다.**
|
|
99
|
+
프로세스 판도 같이 적었다 — 폭주 사이드카를 죽일 때 **자기 프로세스 판별자**(모델 핀·PID·
|
|
100
|
+
`pgrep` 선열거)를 갖고 쓰고, **죽인 뒤 무엇이 죽었는지 확인한다**.
|
|
101
|
+
|
|
102
|
+
**규칙은 세 자리 중 하나에 산다.** (`README §Where a rule lives`)
|
|
103
|
+
|
|
104
|
+
상주 층이 무한히 자라야만 하는가라는 질문에 대한 답이다. 🟥 **가운데 자리가 보통 비어 있고,
|
|
105
|
+
그게 공짜다** — **게이트 자신의 오류 메시지**는 행위자가 **행동하는 순간에** 읽는 자리라 상주
|
|
106
|
+
예산을 한 글자도 안 쓴다. 한계도 같이 적었다: **막힐 때만 읽힌다.** 그래서 대체가 아니라 3층이고,
|
|
107
|
+
실측이 그 증거다 — 이번 게이트에서 **기계 +480줄, 상주 산문 ±0**.
|
|
108
|
+
|
|
109
|
+
### 명시 잔여 — 안 닫은 것
|
|
110
|
+
|
|
111
|
+
- **날조된 `reflected(…)` 는 통과한다.** provenance 는 파일에서 안 갈린다. §4-b 의 몫이다.
|
|
112
|
+
- **`①영혼` 소급 안 함** — grace `2026-08-21`, 기존 37건 유지.
|
|
113
|
+
- **형제 축 문법 분열**: `standpoint:` 는 근거를 em-dash 로만 받고 `thirdparty:`·`soul-check:` 는
|
|
114
|
+
괄호로 받는다. 같은 훅 안에서 갈리고, 괄호로 쓰면 *"is not a member of the enum"* 이라는
|
|
115
|
+
**오진**을 낸다(값은 멤버가 맞고 문법만 틀렸다).
|
|
116
|
+
- **psa 계열 7건이 호출부에 의존해 잠들어 있다** — 「남은 7건」이 아니라 **조건부 활성**이다.
|
|
117
|
+
- **`M-1`(상주) · `M-2`(memory) · `M-3`(스킬 활동도)은 안 건드렸다.** 🟥 셋 다 **계기가 미검증**
|
|
118
|
+
이다: 임계 40k·80k·10k **어느 것도 절단점 근거가 없고**(도입 커밋이 축은 20줄 논증하고
|
|
119
|
+
절단점은 한 줄도 안 함, 그리고 도입 당일 대상이 이미 초과였다) · 출하 스캔이 `wc -c`(바이트)를
|
|
120
|
+
**«chars» 라고 라벨**하며(163,456 vs 실제 문자 144,471) · 스킬 활동도 계기가 **언급을 사용으로**
|
|
121
|
+
센다(실행 0인데 «활발» 로 오분류 27종). **재정초 전에 감량하면 근거 없는 목표를 향해 깎는 것**
|
|
122
|
+
이고, 그 절약으로 fail-open 을 산다.
|
|
123
|
+
|
|
124
|
+
### [2.6.0] — 2026-08-20
|
|
125
|
+
|
|
126
|
+
**배포본이 소비자 설치에서 `SELFCHECK: FAIL` 이었다 — 그리고 원인 넷이 전부 «계기가 저자의 머신에 결박» 이었다.** 이 릴리스의 중심은 새 기능이 아니라 그 복구다. 발견 경로는 재출하 준비 중의 손 실행이다: `npm pack` → 추출 → **`node_modules/@chrono-meta/fh-gate` 실경로에서 완주**. 레지스트리에서 받은 **실물 2.5.1** 로도 재현했다(컨트롤).
|
|
127
|
+
|
|
128
|
+
- **위성 레인 둘이 주체 없이 출하됐다.** `test_satellite_publish_gate_lanes.sh`(2.5.1 에 이미 실림) · `test_satellite_profile_schema_lanes.sh`(이번 범위에 추가됨)의 주체는 `frontier_digest_daily.sh` 인데 그건 의도적으로 출하 대상이 아니다(소비자 계정으로 `claude` CLI 를 태운다). 실측: publish_gate **4 passed / 21 failed**, `rc=127`. 🟥 **통과한 쪽이 더 나빴다** — "dispatch 자체가 안 일어남 ✅" 은 러너가 **없어서** 통과한 거짓 초록이다. 선행 사례(`test_frontier_digest_retry.sh`, *"Anchor follows subject"*)대로 `ACCEPTED_ABSENT` 로 내리고, `selfcheck.sh` 는 **주체** 부재를 보고 `_absent_subject_verdict` 로 위임한다 → 이름 있는 SKIP
|
|
129
|
+
- **`digest_landing_check --self-test` 가 폴더 이름에 결박돼 있었다.** 컨트롤이 `basename "$FH"` 로 유도되는데 픽스처 카드가 리터럴 `forge-harness` 를 담고 있어, **디렉터리 이름이 `forge-harness` 일 때만** 컨트롤이 살았다. known-pair(같은 바이트, 이름만 교체): `forge-harness/` **15/15 PASS** · `some-consumer-app/` **5/15 FAIL**. 레인별로 컨트롤을 픽스처 토큰에 고정했다. 🟥 «10 을 기대하는» 레인들도 고정했다 — 고정 전에도 10 을 냈지만 **의도한 사유가 아니라 컨트롤 사망** 때문이었다. 🟥 N1·★N-ctl·★N-ctl-re 는 일부러 유도/사망을 주장하므로 **고정하지 않았다**
|
|
130
|
+
- **`cluster_capability_scan` L13c — 전제만 결박이었다.** `discover` 가 `tracks/`(gitignored, 배포물에 구조적 부재)를 전제하므로, 부재는 **FAIL 이 아니라 이름 있는 SKIP**(미측정 ≠ 0건)으로 낸다. 🟥 **초판은 여기서 하나를 더 «고쳤고», 그게 틀렸다** — 기대 문자열 `^forge-harness\(hub\)` 를 폴더명 결박으로 읽고 `basename` 유도로 바꿨는데, **생산자(`:128`)는 그 리터럴을 낸다.** 유도로 바꾸면 이름이 다른 트리에서 기대와 산출이 갈려 **거짓 FAIL** 이 된다 — 이 파일이 자기 주석에서 경고하는 divergent-normalizer 를 수리가 새로 만든 꼴이다. 되돌렸다. 그 라벨은 디렉터리 이름이 아니라 **허브의 상수 식별자**이고, 양쪽이 같은 상수를 쓰는 한 결박이 아니다. **cross-family(codex/gpt-5.5)가 잡았다 — 자력 적발 0.** known-pair 로 재현: 이름이 다른 트리 + `tracks/` 존재에서 상수판 **20 PASS / 0 FAIL** · 유도판 **19 PASS / 1 FAIL**. 내 소비자 테스트는 그 팔에 **구조적으로 못 닿았다**(거기선 L13c 가 SKIP 이라)
|
|
131
|
+
- **mate 어댑터의 known-pair 픽스처가 안 실렸다.** 게이트는 출하되는데 보정쌍이 소비자 머신에 없어 `HARNESS_ERROR(10)`. `mate_agent_boundary_known_{positive,negative}.md` 를 `files[]` 에 추가 — 이건 **출하하는 쪽**이 맞다(주체가 이미 출하되므로)
|
|
132
|
+
- 검증: 수리 후 소비자 설치 **`SELFCHECK: PASS` (rc=0)**. 🟥 중간에 내 계기가 한 번 틀렸다 — `grep '^FAIL'` 로 세어 «FAIL=0 인데 FAIL» 이라는 가짜 모순을 만들었다. 이 스위트들은 `❌` 로 찍는다
|
|
133
|
+
|
|
134
|
+
**온보딩 메뉴를 세로로 편다.** `·` 로 이어붙인 한 줄 메뉴는 터미널 폭에서 임의로 접혀 문 경계가 안 보인다(운영자 지적). `G-GREET-02`(🐿️+환영문 같은 줄)·`G-GREET-03`(고정 4문)·`G-GREET-05`(문구 리터럴) **셋 다 불변** — 그 프로브들이 박은 것은 문 집합·리터럴·환영문 줄이지 메뉴의 줄 수가 아니다. 플로어 티어 블라인드 sim(레포 밖 cwd·헤드리스·reps=3): **ARM 3/3 세로 · CONTROL 3/3 가로**, 같은 실행에서 G-GREET-02 도 3/3 유지.
|
|
135
|
+
|
|
136
|
+
**BREAKING 없음 — 그리고 그 판정을 적어둔다.** 카드는 «위성 런이 프로필 미선언이면 막힌다» 를 `BREAKING (gate):` 후보로 올려뒀는데, 그 게이트가 사는 `frontier_digest_daily.sh` 가 **출하 대상이 아니라** 소비자의 게이트 수용은 안 바뀐다. 오늘 수리는 FAIL→PASS 라 완화 방향이다.
|
|
137
|
+
|
|
138
|
+
🟥 **이 릴리스가 스스로 낸 교훈**: 결함 넷 중 **셋은 진단이 맞았고 하나는 틀렸는데, 틀린 하나를 자력으로는 못 잡았다.** 소비자-설치 완주라는 계기는 「돌아가나」를 재지 「고친 게 옳은가」를 못 잰다 — SKIP 으로 빠지는 팔은 그 계기가 구조적으로 안 보는 자리다. 잡은 것은 **다른 계열에 diff 를 보낸 것**이다.
|
|
139
|
+
|
|
140
|
+
**같은 릴리스에 함께 나가는 것 — `harness-doctor` 30일 캐던스 수리 (#468).** 진단이 낸 것 중 **지금 출하물에 살아 있던 것**만 골랐다.
|
|
141
|
+
- **팬텀 「500줄 / 16스킬」 임계 제거** (`docs/platform_sustainability.md`, 이번에 출하 대상에 편입). 실제는 1,414줄 / 40스킬이고, `harness-doctor` 는 meta 타깃에서 줄수 행을 «판정 아님»으로 **비활성화**하지 큰 숫자로 갈지 않는다. 🟥 덤이 본체보다 컸다 — 이 팬텀의 사후분석이 *"`500` 은 이 파일 어디에도 없다(grep 0 hits)"* → «런이 지어냈다» 로 결론냈는데 **그 grep 이 자기 파일만 봤다.** 교훈은 «환각했다»가 아니라 **«부재 검사를 틀린 코퍼스에 돌렸다»**
|
|
142
|
+
- **CATALOG 미등재 정본 7건 등재** (컨트롤 `harness_6axis_framework`=2, 대상 7건 전부 0). 그중 `fh_three_layer_canon.md` 는 CLAUDE.md 가 **필독**으로 지정한 문서다 — 색인이 못 찾는 필수 문서는 파일명을 이미 아는 세션이 아닌 한 부재와 구별되지 않는다
|
|
143
|
+
- **`[[wikilink]]` 규약 선언** (`knowledge/shared/GLOSSARY.md` 신설 절). 출하 `.md` 156개에 안 풀리는 타깃 **59 / 출현 98**, 선언이 아무 데도 없었다. 🟥 이건 **깨진 참조가 아니라 출처 표시**다 — 고치려 들지 마라. 메모리 스토어는 운영자별·세션 스코프라 vendoring 은 수리가 아니라 residency 위반이다
|
|
144
|
+
- 🟥 **`CLAUDE.md` 가 자기 기계를 거짓 서술하고 있었다.** *"standpoint 는 Prose-only today — 검증하는 pre-commit 훅도 픽스처도 없다"* → `validate_standpoint_leg()` 는 `pre-commit:798` 정의 · `:1575` 호출이고 `test_marker_standpoint_lanes.sh` 는 `selfcheck.sh:531` 배선이다. **같은 파일이 정반대도 적고 있었다.** 이 계열은 레인이 구조적으로 못 잡는다 — 규칙의 **자기서술**을 그 규칙이 서술하는 **기계**와 대조하는 검사가 없다. 그리고 stale 한 «아직 안 지었다» 는 가장 조용한 드리프트다: 정직한 겸손처럼 읽히면서 **이미 있는 컨트롤의 사용을 억제한다**
|
|
145
|
+
- **M-1(상주 140k)은 안 닫았다.** 상주 원장 32절 전수 결과 **M·S·R 전부 0** — 절 단위 레버가 없다는 것이 실측이다. 유일한 레버는 capability-level merge(후보 6묶음)이고, **그 병합이 실제로 문자를 줄이는지는 UNMEASURED**
|
|
146
|
+
|
|
147
|
+
**명시 잔여**: `package_coverage_check.sh` 는 「참조된 경로가 출하되나」를 보지 「**출하된 레인의 주체가 출하되나**」를 안 본다 — 네 결함 중 셋을 rc=0 으로 통과시켰다. 그 갭은 이번에 안 닫았다.
|
|
148
|
+
|
|
13
149
|
### [2.5.1] — 2026-08-19
|
|
14
150
|
|
|
15
151
|
**진입점 패리티 — 규칙이 출하돼야 발화한다.** 2.5.0 직후 `④-b` 드리프트 검사가 **AGENTS.md 에 공유 체크아웃 규율이 없다**를 냈다. CLAUDE.md 에는 있고 Codex 진입점에는 없는 상태였고, 그 상태에서는 비-Claude 런타임에 그 규칙이 **보이지 않는다**(gate-locality). 이 릴리스는 그 한 항목을 소비자에게 실제로 보내기 위한 것이다.
|
|
@@ -20,6 +20,7 @@ The main agent passes you one of:
|
|
|
20
20
|
- **Mode E (External scan)**: "scan frontier" / "what are people building" / specific topic
|
|
21
21
|
- **Mode F (Full)**: both — default when no mode is specified
|
|
22
22
|
- **Mode T (Technical bridge)**: "can't connect" / "not possible" / "blocked" / "no direct path" / technical constraint hit
|
|
23
|
+
- **Mode X (Intervention cross-check)**: "쎄한데 확인해줘" / "내가 뭘 놓쳤나" / "개입 대조" / a session asking whether it is about to be stopped. Runs Phase 3-b ONLY — no naming, no frontier scan.
|
|
23
24
|
|
|
24
25
|
Optionally: a focus area (e.g., "token efficiency", "agent orchestration", "cascade patterns")
|
|
25
26
|
|
|
@@ -153,6 +154,175 @@ For each gap or absorbed signal:
|
|
|
153
154
|
4. **Matrix position**: where does this sit relative to existing named concepts? (complement / extend / replace)
|
|
154
155
|
5. **Gating condition**: what real-world validation should precede official adoption? (simplicity guard applied)
|
|
155
156
|
|
|
157
|
+
|
|
158
|
+
## Phase 3-b — Intervention algorithm (Mode X, and MANDATORY inside Mode F)
|
|
159
|
+
|
|
160
|
+
Phase 3 above carries the owner's **naming** algorithm. This phase carries the owner's
|
|
161
|
+
**intervention** algorithm — *where the owner has historically stopped a session and turned it.*
|
|
162
|
+
|
|
163
|
+
**Provenance (measured, not asserted)** — census of the conversation corpus itself, not of what
|
|
164
|
+
sessions wrote down afterwards: `~/.claude/projects/…/*.jsonl`, **69 sessions / 2026-07-22–08-21**,
|
|
165
|
+
**433 operator utterances**, semantically classified. **54 interventions claimed · ~43 estimated
|
|
166
|
+
after a 5-sample hand-check (1 false positive) · full hand-verification NOT done.**
|
|
167
|
+
🟥 **CORRECTED 2026-08-21 — that census read 21% of its own corpus and called it 전수.** A full
|
|
168
|
+
re-scan of the same directory with the same discriminator returns **168 sessions · 1,703 operator
|
|
169
|
+
utterances** (uuid-deduplicated; ~1% contamination hand-checked: `<bash-input>` / `<command-message>`,
|
|
170
|
+
17 of 1,703). The census's own note recorded **169MB**, and `du -sh` on that directory is **817MB** —
|
|
171
|
+
the ratio was written down and never compared. **So «54» is a count over a fifth of the corpus, not a
|
|
172
|
+
census.** Do not cite it alone. If the 12% rate holds, the true intervention count is nearer **200**.
|
|
173
|
+
The shapes and the class distribution below are unaffected in *direction* (they were derived from a
|
|
174
|
+
random-in-practice fifth), but every absolute number on this page is a lower bound.
|
|
175
|
+
Detail + the seal comparison: `tracks/_meta/RESULT_2026-08-21_intervention-corpus.md`.
|
|
176
|
+
|
|
177
|
+
🟥 **The dominant class is NOT "you didn't search the world."** Measured distribution:
|
|
178
|
+
`판단결함 32 (59%) · 내부미조회 9 · 외부미조회 6 (11%) · 범위겨냥 6`. A design that treats this
|
|
179
|
+
as a *search* trigger is aiming at an 11% slice — the first draft of this capability did exactly that.
|
|
180
|
+
|
|
181
|
+
**Prior art (2026-08-21) — this task has a published benchmark; price the capability against it.**
|
|
182
|
+
Wu et al., *"User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning
|
|
183
|
+
Signal"*, EMNLP 2025 main (`arXiv:2507.23158`). Numbers read from the PDF text, not from a summary:
|
|
184
|
+
automatic feedback identification with a purpose-built GPT-4o-mini prompt scores **P 61.1 / R 35.9**
|
|
185
|
+
in the *dense* setting (label every turn — the realistic one) and **P 100.0 / R 69.2** in the *sparse*
|
|
186
|
+
setting (the feedback turn is pointed out in advance); inter-annotator agreement **Cohen κ = 0.70
|
|
187
|
+
(binary) / 0.74 (three-way) / 0.60 (fine-grained)** over 54 cross-annotated conversations.
|
|
188
|
+
⇒ Two consequences. **(a)** A weak separation here is the task's difficulty, not this rule set being
|
|
189
|
+
unusually bad — quote precision against **P61/R36**, never against a vacuum. **(b)** Their conclusion
|
|
190
|
+
(*noisy as a learning signal*) converges independently with residual (0) below.
|
|
191
|
+
🟥 **Do not collapse the two tasks.** They read the user's turn and judge post-hoc; this phase predicts
|
|
192
|
+
*before* the user speaks, from the session's own acts, with the next human turn structurally excluded.
|
|
193
|
+
Ours is strictly harder and that difference is the net-new angle — normalizing it away both overstates
|
|
194
|
+
the prior art and invalidates any direct number comparison.
|
|
195
|
+
🟥 Provenance of this paragraph: the world was consulted **only after the operator asked whether it had
|
|
196
|
+
been**; zero external lookups preceded the design. That is identity ④'s measured gap, not a footnote.
|
|
197
|
+
|
|
198
|
+
### The ten shapes — a session in this state is about to be stopped
|
|
199
|
+
|
|
200
|
+
| Shape | The session is… | n |
|
|
201
|
+
|---|---|---|
|
|
202
|
+
| 손안의 기록 미개봉 | reporting «missing / incomplete / impossible» about material it already holds but never opened | 9 |
|
|
203
|
+
| 막힘을 종점으로 접기 | folding at a blocker (hard / absent / over quota / someone else's job / next session) without trying one alternate route | 9 |
|
|
204
|
+
| 산출물 미대조 완료선언 | writing «applied it / fixed it» without re-opening or re-rendering the artifact | 7 |
|
|
205
|
+
| 세계 미조회 설계 | designing from repo-and-session material only, zero external prior art consulted | 6 |
|
|
206
|
+
| 단일 관측 전역 단언 | fixing a cause/state/number from one node or one observation, no second instrument | 6 |
|
|
207
|
+
| 과녁 치환 | working hard on a different target than the one named in the request | 5 |
|
|
208
|
+
| 정적 리뷰로 완료 | closing a wiring/rule change on a static read, never run from the target tier or target harness | 4 |
|
|
209
|
+
| 이상신호 흘려보내기 | passing over a tool error, odd latency, or overwrite side-effect without citing or acting | 4 |
|
|
210
|
+
| 기준 없는 등급 판정 | setting a version digit or identity grade without citing the document that defines it | 2 |
|
|
211
|
+
| 절차 규모 오조준 | defaulting to a heavier procedure than the change needs | 2 |
|
|
212
|
+
|
|
213
|
+
### The self-check list — run these against **this turn's tool calls and output**
|
|
214
|
+
|
|
215
|
+
🟥 **Every line is stated as an observable act, never as a feeling.** That is deliberate: the
|
|
216
|
+
adversarial review that killed the first draft found the fatal flaw was a trigger keyed on the
|
|
217
|
+
session's own sense of uncertainty — *a session that is confident does not report uncertainty, and
|
|
218
|
+
the measured failures were all confident.* These lines are checkable from what the session just did.
|
|
219
|
+
|
|
220
|
+
1. Wrote a paragraph or more of design/definition/gap-diagnosis, and opened **zero** external sources this turn.
|
|
221
|
+
2. About to write «없다 / 미등재 / 미완 / 불가», and never opened the file or history that would hold it.
|
|
222
|
+
3. Wrote «반영 완료» about a document without counting the items in it.
|
|
223
|
+
4. Wrote «applied / fixed / included it all» with **no tool call re-reading that artifact after the edit**.
|
|
224
|
+
5. Was given N items and touched fewer than N, without putting both numbers side by side.
|
|
225
|
+
6. About to write «next session / someone else / later» with **no tool call attempting an alternate route this turn**.
|
|
226
|
+
7. Dropped a verification leg because a sidecar was blocked, with no record of trying another family / local LLM / subagent.
|
|
227
|
+
8. Withdrew its own proposal citing only «hard / side effects», with not one line on how to make it work.
|
|
228
|
+
9. Asked the operator about a peer session's state instead of asking that session via ListAgents/SendMessage.
|
|
229
|
+
10. Routed a candidate to CURATED / drop / hand-off **without one line on how it could become our own capability**.
|
|
230
|
+
11. The file / environment / axis being edited is not the noun the operator named.
|
|
231
|
+
12. Filled a mapping or candidate list only from what exists locally on this machine.
|
|
232
|
+
13. About to write PASS on a rule/wiring change and cannot quote a command run in the target tier or harness with its output.
|
|
233
|
+
14. Ran a «standpoint review» from its own vantage, with no agent dispatched inside the target harness.
|
|
234
|
+
15. Fixed a cause/state/number from one node or one observation, with no second instrument.
|
|
235
|
+
16. Wrote an aggregate count without checking whether already-running or pre-existing items are inside it.
|
|
236
|
+
17. Wrote a time/date/environment fact from memory or inference rather than from a command.
|
|
237
|
+
18. Judged a tool error or warning «non-blocking» and moved on without citing it or acting.
|
|
238
|
+
19. Created or changed a setting and wrote «done» without printing its expiry / default fields.
|
|
239
|
+
20. Regenerated or overwrote a file without a diff showing which prior lines are gone.
|
|
240
|
+
21. Waiting on a run that is taking longer than expected without checking its output or whether a session was created.
|
|
241
|
+
22. Raised a version digit or grade without quoting the document that defines that digit.
|
|
242
|
+
23. Proposed follow-up work larger than the original request without putting a minimal option beside it.
|
|
243
|
+
|
|
244
|
+
### Output for Mode X
|
|
245
|
+
|
|
246
|
+
For each line that fires: quote the session's own act that trips it, and propose **one line** —
|
|
247
|
+
*"확인해볼까?"* — naming the cheapest check that would settle it. **Propose; never decide.**
|
|
248
|
+
Fires nothing → say «걸린 줄 없음» explicitly; silence is not a verdict.
|
|
249
|
+
|
|
250
|
+
### Tier M — signals decidable from the session RECORD (calls + turns + diff), no judgment
|
|
251
|
+
|
|
252
|
+
🟥 **The first draft of this heading said «from the tool-call record alone». That was false**
|
|
253
|
+
(cross-family, 2026-08-21): #9, #10, #12, #15 and #18 require reading the user's turn, the reply, or
|
|
254
|
+
the commit diff — not the call log. The tier's real claim is narrower and is what the heading now
|
|
255
|
+
says: **no judgment is needed**, but more than the call log is read. An evaluator for this tier needs
|
|
256
|
+
a defined input contract (calls · user turns · final reply · staged diff) that **does not exist yet**.
|
|
257
|
+
|
|
258
|
+
A **second census** (same question, different corpus: what sessions *recorded* about
|
|
259
|
+
interventions, `tracks/`+`knowledge/`+memory — 300 scanner hits → 199 claimed → **33 hand-verified,
|
|
260
|
+
18% rejected**) produced signals of a different grade: each one is a **countable fact about this
|
|
261
|
+
session's own calls**, needing no judgment. Both censuses landed on the same class distribution
|
|
262
|
+
(판단결함 dominant · 외부미조회 a minority).
|
|
263
|
+
🟥 **That agreement is CORROBORATING BUT CORRELATED — not independent** (cross-family caught the
|
|
264
|
+
overclaim). Same operator, same canon, same model family; and the `tracks/` records are *derivative
|
|
265
|
+
of the same events* the transcripts hold. Claiming independence would need event-linkage removal, a
|
|
266
|
+
different annotator/model, a pre-registered codebook and blind reclassification — **none were done.**
|
|
267
|
+
|
|
268
|
+
1. An absence/blocked claim (`없다`·`0건`·`not found`·`unavailable`·`막혔`·`overdue`) appears, and the tool call against that subject happened **exactly once** — no second attempt.
|
|
269
|
+
2. A tool output carries a truncation marker (`truncated`·`… N more`·a next/page cursor·line count exactly equal to the limit) and the tool was **never re-called with a different offset/page/cursor**.
|
|
270
|
+
3. A call ended non-zero or errored, and the **same tool with the same arguments was not retried** — the session switched to a different tool instead.
|
|
271
|
+
4. A background handle has produced **0 bytes of stdout for N seconds** and has not exited, and the call carried no timeout.
|
|
272
|
+
5. A freshness/cadence verdict rests on a single glob whose match count is **0** (rendering `not found` as `overdue`).
|
|
273
|
+
6. After session-start `pull`/`fetch`, the newest remote commit is **later than the date field of the card/INDEX that was read**, and **zero** of the files those commits touched were Read.
|
|
274
|
+
7. A staged git-tracked added line contains an absolute home path, a companion-store name, a vendor/product proper noun, or an executable that only `command -v` resolves **on this machine**.
|
|
275
|
+
8. A diff under `package.json files[]` · `templates/` · `plugins/` newly introduces a local-only path or local-only CLI name — an environment dependency entering the shipped set.
|
|
276
|
+
9. A noun phrase or quoted string from the user's turn appears **0 times** in the session's whole commit diff (operator utterance ↔ canon landing).
|
|
277
|
+
10. A quantity token (`N건`·`N자`·`N%`·`HH:MM`) appears in an artifact or final reply, and the session made **no call able to produce it** (`wc`·`grep -c`·`date`·arithmetic).
|
|
278
|
+
11. A time/date predicate (`심야`·`오전`·`어제`·a weekday) was written to a record with **zero `date` calls**.
|
|
279
|
+
12. The first user turn matches the greeting corpus and the first reply carries **neither the 🐿️ literal nor the fixed welcome line**.
|
|
280
|
+
13. A section a rule marks «always include» greps **0 times** in the artifact that rule governs.
|
|
281
|
+
14. A new file is about to be written with **zero** prior Read/Grep against `CATALOG.md` / the skill list / `plugins/**/SKILL.md`, while its name or keywords already match the index.
|
|
282
|
+
15. A skill/agent proper noun the session named as the routing target appears **nowhere in the user's turn** — the session introduced that name.
|
|
283
|
+
16. An external model's or sidecar's **self-report string** is cited as verdict evidence, with **0 calls** running the same probe against a known control.
|
|
284
|
+
17. The diff changes an exit code, a default, or a fail-open/closed direction, and the commit message or 4-axis marker quotes the user's turn **0 times**.
|
|
285
|
+
18. A recommendation to install or use a tool carries **no conditional marker** (`when`·`only if`·`unless`·`~일 때만`) anywhere.
|
|
286
|
+
|
|
287
|
+
### How the two tiers are used
|
|
288
|
+
|
|
289
|
+
```
|
|
290
|
+
Tier M (18) countable from this session's calls → a hook could evaluate these
|
|
291
|
+
Tier J (23) need reading the session's own output → invocation, judged
|
|
292
|
+
```
|
|
293
|
+
**Noise cap (mandatory).** Rank by tier then by how cheap the check is, and surface **at most 2 per
|
|
294
|
+
turn**; hold the rest silently. 🟥 Without this, Mode F makes this phase mandatory and every firing
|
|
295
|
+
emits a proposal — up to **41 «확인해볼까?» in one turn**, which is the nag that trains dismissal and
|
|
296
|
+
kills the capability (cross-family MED). Repeat suggestions dedupe by shape, not by wording.
|
|
297
|
+
|
|
298
|
+
🟥 **Neither tier decides.** Both produce the same one line — *"확인해볼까?"* — naming the cheapest
|
|
299
|
+
check. The operator's bar for this capability is exactly that: *"'쎄함'을 감지하고 사람에게
|
|
300
|
+
「한번 확인해볼까?」 라고 제안하는 것만 가능해도 성공이다."*
|
|
301
|
+
|
|
302
|
+
⚠️ **Named residuals.** (0) **MEASURED 2026-08-21 — and it did not pass its pre-registered bar.**
|
|
303
|
+
(The 18% figure below is the rejection rate of one census's *claims*; it never was these rules'
|
|
304
|
+
precision.) A sealed pre-registration (`sha256 8414a44d…`, re-verified at opening) ran the 41 rules
|
|
305
|
+
against **44 windows** — 22 immediately preceding an operator intervention, 22 preceding an approval —
|
|
306
|
+
with blinding enforced by file boundary (the rule-evaluator never receives the next human turn; the
|
|
307
|
+
labeler never receives the rules). Result: **12 of 41 rules ever fired · 2 discriminate**
|
|
308
|
+
(`J2` +4, `M10` +3) · **10 fired below the +3 line** (J1 · J17 · M2 · J4 · J12 · M1 · J6 · M4 · J18 ·
|
|
309
|
+
M3 — «did not clear the bar in this sample», **not** the same bucket as dead) · **29 never fired at
|
|
310
|
+
all**, 14 of them Tier M. The pre-registered stop condition **«no separation» HIT** (median difference
|
|
311
|
+
+0.5 vs a required +2.0).
|
|
312
|
+
🟥 **The pre-registration itself was defective and that is recorded, not hidden**: a median is a poor
|
|
313
|
+
statistic on a 0–1 sparse count. A post-hoc binary comparison (any rule fires: POS 82% vs NEG 50%,
|
|
314
|
+
Fisher p=0.055) suggests signal — **it does not overturn the verdict**, and claiming it would require
|
|
315
|
+
replication under a *new* pre-registration naming that statistic in advance. Current status:
|
|
316
|
+
**signal possible, not established.**
|
|
317
|
+
⚠️ Scope of that run: **in-sample** (no out-of-sample corpus exists — sessions in other repos are
|
|
318
|
+
one-shot dispatches with no preceding window), which makes it a *favourable* arm: failing it is
|
|
319
|
+
conclusive, passing it would not have been. Detail: `tracks/_meta/RESULT_2026-08-21_intervention-precision.md`. (0-b) Tier M has **no executable definitions** — «N seconds», «same subject»,
|
|
320
|
+
«alternate route», «conditional marker» are undefined; an evaluator schema, window and no-data
|
|
321
|
+
verdict must precede any wiring. (a) Tier M is written but **not wired** — no hook evaluates it yet; both
|
|
322
|
+
tiers currently run on invocation. (b) Neither census hand-verified in full: transcripts 5/54
|
|
323
|
+
checked, tracks 33 of 199 claims verified at an 18% rejection rate. (c) The transcript corpus is
|
|
324
|
+
**one month deep** (2026-07-22 onward); earlier interventions are structurally absent, not zero.
|
|
325
|
+
|
|
156
326
|
## Self-floor discipline (FH floors, applied to the innovator itself)
|
|
157
327
|
|
|
158
328
|
These are FH's own governance floors turned reflexively on this agent's process — an ideation tool
|