@chrono-meta/fh-gate 2.0.1 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/rules/fh_4axis_gate.md +77 -1
- package/.claude-plugin/marketplace.json +3 -3
- package/AGENTS.md +19 -3
- package/CLAUDE.md +240 -4
- package/knowledge/shared/harness-core/capability_composition_contract.md +68 -0
- package/knowledge/shared/harness-core/fh_three_layer_canon.md +38 -12
- package/knowledge/shared/harness-core/field_verdict_crossfamily_gate.md +115 -3
- package/knowledge/shared/harness-core/harness_incubator_doctrine.md +82 -0
- package/knowledge/shared/harness-core/harness_terminal_correlation_and_recommendations.md +49 -8
- package/knowledge/shared/harness-core/ship_readiness_gate.md +205 -1
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +66 -0
- package/knowledge/shared/rules/sister_asset_protocol.md +12 -0
- package/package.json +3 -2
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-commons/agents/quench-challenger.md +6 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +43 -0
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +14 -2
- package/plugins/fh-meta/skills/steel-quench/SKILL_detail.md +13 -0
- package/plugins/fh-meta/skills/verify-bidirectional/SKILL.md +34 -1
- package/scripts/capability_registry_check.sh +75 -12
- package/scripts/chamber_run.sh +24 -3
- package/scripts/chamber_witness.sh +25 -0
- package/scripts/fh-gate.sh +16 -1
- package/scripts/package_coverage_check.sh +10 -0
- package/scripts/prepublish_scope_note.sh +139 -0
- package/scripts/selfcheck.sh +30 -5
- package/scripts/test_fh_gate_regressions.sh +21 -9
- package/scripts/test_marker_axes_run_lanes.sh +87 -3
- package/scripts/test_marker_crossfamily_lanes.sh +31 -1
- package/templates/.git-hooks/pre-commit +225 -10
- package/templates/subagent-tally-hook.json +2 -2
|
@@ -2080,3 +2080,69 @@
|
|
|
2080
2080
|
outcome: accepted
|
|
2081
2081
|
evidence: "1M/4S/2R. 챌린저와 중첩 2건뿐(rc=10 fail-open · 사전순 정렬) — 렌즈가 계열보다 갈랐다. 단독 적발 3건: package.json files[] 누락 · 비UTF8 바이트에서 grep 무매치하는 로케일 fail-open(양쪽 로케일 직접 실행) · 파일스코프 set -e 판별. 그리고 의심 하나를 기각(§ 구분자는 두 로케일에서 실측 통과)"
|
|
2082
2082
|
cost: 105k tokens
|
|
2083
|
+
- date: 2026-08-17
|
|
2084
|
+
agent: fh-meta:beginner (isolated) — 챔버 런 #11 step-4 블라인드 페르소나 1/3
|
|
2085
|
+
purpose: "인물 시뮬레이터 하네스 후보에 대한 냉담 첫 접촉 — 첫 실패 지점·믿을 이유·이해 불가 지점"
|
|
2086
|
+
outcome: accepted
|
|
2087
|
+
evidence: "tool_uses 8(실독 확인 — 어제 sim 8회 tool_uses:0 죽은 컨트롤의 반대). 최대 소득 = 첫 기계적 실패가 크래시가 아니라 무음 통과라는 지목: grounding_gate_v3.py:88-93 _KNOWN_BOOKS 가 'Book Ch:Vs' 주소 문법에 하드코딩이라 임의 인물 코퍼스에선 빈 셋 → 오귀속 검출 b1/b2 가 초록으로 통째 우회. 🟥 최강 발견 1건은 VOID(stale 체크아웃에 조준한 내 계기 결함 — personas_dialogue.json 이 반증). 부재를 0으로 안 렌더한 덕에 그 결함이 드러났다"
|
|
2088
|
+
cost: 102k tokens
|
|
2089
|
+
- date: 2026-08-17
|
|
2090
|
+
agent: fh-meta:main-player (isolated) — 챔버 런 #11 step-4 블라인드 페르소나 2/3
|
|
2091
|
+
purpose: "실사용자 일상 가치 — 티어 선정 후 매일 쓸 값어치가 있는지, 게이트가 값인지 방해인지"
|
|
2092
|
+
outcome: accepted
|
|
2093
|
+
evidence: "tool_uses 9. 핵심 숫자 = 사용 장면 0/3 생존(2 기존수단 대체 · 1 자기 게이트가 금지). 그리고 Light↔Heavy 구조적 상충(게이트 ON→Light 사망 / OFF→Heavy 유일 값 소멸)이 튜닝으로 못 푸는 것임을 코드 주석(v3:233-234 'PARAPHRASE over-blocks … err-safe direction')으로 근거화. Midcore 스킵을 자백하고 '이 스킵이 판정의 최약 고리'라고 스스로 적음. 결론 'KILL the simulator · EMIT the attribution gate'"
|
|
2094
|
+
cost: 102k tokens
|
|
2095
|
+
- date: 2026-08-17
|
|
2096
|
+
agent: fh-meta:challenger (isolated) — 챔버 런 #11 step-4 블라인드 페르소나 3/3
|
|
2097
|
+
purpose: "배출 후보 적대 심사 — 핵심 술어 성립성·재발명·배출가치·실패비용·안 보이는 것"
|
|
2098
|
+
outcome: accepted
|
|
2099
|
+
evidence: "tool_uses 22, 자체 컨트롤 부착(같은 Glob 이 battery4.py·normalization.py 는 잡음 → 계기 생존 증명). 🟥 거버너 사각 2건 적발: EMIT 이 이미 내려진 운영자 결정(출하 없음)과 정면 충돌 · Delphi 가 미승인 Peter Attia 클론을 실제 테이크다운(카파시와 동일 프로파일). 최대 기여 = 술어가 '더 어렵다'가 아니라 '목적과 상충'임을 코드로(grounding_gate.py:317-332). 팬텀 4건은 VOID(stale 체크아웃)이나 본인이 '원인 MED, stale 가능성 배제 못 함'이라고 선제 자백 — 등급 자기하향이 판정 신뢰도를 올렸다"
|
|
2100
|
+
cost: 117k tokens
|
|
2101
|
+
- date: 2026-08-17
|
|
2102
|
+
agent: general-purpose (isolated) — 챔버 런 #11 조건 1(net-new) measured 스캔
|
|
2103
|
+
purpose: "외부 생태계 + FH 내부 자산 전수 스캔으로 재발명 여부를 측정. '없는 것 같다'는 결과가 아니라고 명시 지시"
|
|
2104
|
+
outcome: accepted
|
|
2105
|
+
evidence: "tool_uses 28(WebSearch 12 · WebFetch 3 · 로컬 SKILL 4종 정독). 🟥 이 런의 KILL 을 가른 계기이고, 유일하게 stale-체크아웃 결함의 영향 밖이다(the-bible 은 '안 열었다'고 자백). ⓐ=Delphi.ai 상용 + verbatimeter(--fail 비영종료 CI 게이트) + 내부 corpus-grounding-expander · ⓒ=NeMo Guardrails + 내부 persona_container_schema.md 4그룹 tier-floor · ⓑ만 빈칸인데 그건 배출물이 아니라 업계 미해결 난제. 선행연구 PersonaCite(CHI 2026 EA) 가 후보 루프와 동형. 미탐 자백 6종(영어 전량·Delphi 실계정 미검증·GitHub code search 미실시 등)을 스스로 열거"
|
|
2106
|
+
cost: 157k tokens
|
|
2107
|
+
- date: 2026-08-17
|
|
2108
|
+
agent: codex/gpt-5.5 (cross-family, headless) — 챔버 P1 증인 수리 델타 Axis 2 탈상관
|
|
2109
|
+
purpose: "chamber_run.sh/chamber_witness.sh/test_chamber_run_lanes.sh 197줄 diff 를 다른 계열로 공격. crossfamily 마커 값 확보(하중 변경 — 게이트 exit 동작 + 증인 판정)"
|
|
2110
|
+
outcome: accepted
|
|
2111
|
+
evidence: "7건(A급3·B급4) · **자력 적발 0/7**. A급 전부 손 재현 후 수리 — ⓐ 멱등 가드 `substr($0,9)` 오프셋(`- run: `는 7자, 값은 8부터): 슬러그 첫 글자가 잘려 **어떤 정상 엔트리와도 매칭 안 됨** = 가드가 조용히 무력인데 30레인 전부 초록이었다(재현: `printf -- '- run: g2\\n' | awk '{print substr($0,9)}'` → `2`) ⓑ L11 이 «멱등»을 잰다면서 증인 원장이 아니라 G4 INDEX.md 를 셌다 → L14 신설(증인 실제 복사 후 2회 기록) ⓒ L13-b 가 `step 6 BLOCKED` 한 줄만 grep 해 뒤에 한 줄 더 붙이면 게임 가능 → 차단 블록 전체로 확대. B급 4 중 3 채택(인접성 상태기계 · 리터럴→호출 인자 **집합 대조** · 거짓 주석 정정), 1건 수용(레거시 런이 시끄럽게 막히는 건 안전 방향). 🟥 그리고 **내 주석 하나가 거짓임을 지적**했다(레인이 증인을 복사한다고 적었는데 `grep -c 'cp .*chamber_witness'`=0). 미탐 자백: bash 3.2/BSD·GNU 이식성 구체 결함 0건 — 공백 아티팩트명은 상류에서 이미 거부되어 도달 불가라고 근거까지 댔다"
|
|
2112
|
+
cost: 50k tokens
|
|
2113
|
+
- date: 2026-08-17
|
|
2114
|
+
agent: general-purpose (isolated) — axes-run 4축→6축 조사·설계안
|
|
2115
|
+
purpose: "마커 스펙 정본 위치 · 훅이 실제로 강제하는 것 · 6축 대응표 · 기존 마커 실측 · 마이그레이션 설계안 3개. 파일 수정 금지(read-only)"
|
|
2116
|
+
outcome: accepted
|
|
2117
|
+
evidence: "tool_uses 21. 🟥 승인의 전제를 반증한 것이 최대 기여 — 「기존 마커 전량 무효화」가 거짓이고 훅은 `.axes_23_passed_{branch}_{TODAY}.marker` 한 개만 검증한다(pre-commit:906 인용). 실측 190건 중 차단 0 · 재해석 51. 그리고 스펙 정본이 **어느 rule 파일에도 없다**는 것(grep 0건)을 지목해 §Marker axis fields 신설로 이어졌다. 자기 미확인 3건을 스스로 열거(08-10 자 기호 마커 2건이 어떻게 통과했는지 · 기호 grep 이식성 · 다른 소비자 영향)"
|
|
2118
|
+
cost: 168k tokens
|
|
2119
|
+
- date: 2026-08-17
|
|
2120
|
+
agent: general-purpose (isolated) — deepteam 업스트림 이슈 준비
|
|
2121
|
+
purpose: "전달본 §DELIVERABLE 정리 + 채널 기계 확인 + 이슈 템플릿/라벨 조사 + 사생활 스캔. 게시 금지(read-only)"
|
|
2122
|
+
outcome: partial
|
|
2123
|
+
evidence: "tool_uses 14. 채널을 기계로 확정(hasDiscussionsEnabled=false / hasIssuesEnabled=true, 출력 인용) — 운영자 서술을 맞춰준 게 아니라 확인했다. 전달본이 `sentry_sdk` 누락을 「the one genuinely actionable thing」으로 단정한 것을 코드검색 hit 0 근거로 격하 권고. 🟥 **그 권고가 뒤집혔다** — 검색 대상이 `main` 이었고 설치되는 건 릴리스 1.0.9 다. 깨끗한 venv 실측에서 `import deepteam` 이 실제로 죽었고(telemetry.py:7), 이슈는 사용노트가 아니라 **버그 리포트**로 게시됐다(#263). 정적 검색이 「없다」고 하고 실행이 「있다」고 한 형태 — partial 로 기록"
|
|
2124
|
+
cost: 107k tokens
|
|
2125
|
+
- date: 2026-08-17
|
|
2126
|
+
agent: general-purpose (isolated) — ⓓ 3자대면(날짜 컷오프 마이그레이션 선례 대조)
|
|
2127
|
+
purpose: "「날짜 컷오프 + 표기법이 스키마 버전을 나른다」가 남의 코드베이스에 알려진 패턴인가 안티패턴인가. 반증 기회로 돌릴 것을 명시 지시"
|
|
2128
|
+
outcome: accepted
|
|
2129
|
+
evidence: "tool_uses 2(WebSearch/WebFetch 중심). 🟥 **내 명제 2건을 반증했다** — ⓐ 「표기법이 배열을 선언한다」가 거짓(코퍼스 손 카운트: axes-run 53건 중 기호 4·혼용 1, 기호 4 중 2건이 08-10 자 옛 4축 의미) ⓑ 날짜 컷오프가 프로덕션 도달 불가 분기(호출부가 ${TODAY} 로 경로 구성). 둘 다 내가 재현 확인. 선례 3계열을 URL 근거로 제시(SonarQube new-code-period · PNG 청크 암묵버전 + 그 전제 문장 · git repositoryFormatVersion 3단 롤아웃 · protobuf 필드번호 재사용금지). 「찾지 못함 ≠ 없다」를 스스로 구분해 적음"
|
|
2130
|
+
cost: 136k tokens
|
|
2131
|
+
- date: 2026-08-17
|
|
2132
|
+
agent: general-purpose (isolated) — qasp-dev 입장리뷰(정적+동적)
|
|
2133
|
+
purpose: "전파 자산(pre-commit) 변경에 대한 standpoint 축. 훅 부재 가설을 실행으로 검증 + qasp 자기 정본 독해. 실물 레포 불변"
|
|
2134
|
+
outcome: accepted
|
|
2135
|
+
evidence: "tool_uses 19. 임시 클론에서 git commit 8회, **양방향 컨트롤**(스텁 훅 rc=1 / 실구성 rc=0 · hooksPath unset 시 폴백 rc=1)로 pre-commit 부재 확정. 🟥 **내 프레이밍을 규범적으로 반증** — 「게이트 부재 = S급 결함」이 아니다(qasp 정본 «4축 비적용» + FH 정본 «NEVER installed into field projects», 양쪽 일치). 내가 세운 「.new.NNNNN = 설치 중단 흔적」 가설도 반증(rename 단계 부재, mv 는 폐기된 옛 설계). 부산물로 FH 자기 기록 2건 오진 발견(08-03 철회 후 08-10 재발). 등급 tier2(qasp-dev)"
|
|
2136
|
+
cost: 146k tokens
|
|
2137
|
+
- date: 2026-08-17
|
|
2138
|
+
agent: general-purpose (isolated) — pmh-dev 동적 입장리뷰(tier2)
|
|
2139
|
+
purpose: "6축 전파가 pmh-dev 에서 무엇을 깨는가. 임시 클론에 이식 후 실제 커밋 양팔. 실물 레포 불변"
|
|
2140
|
+
outcome: accepted
|
|
2141
|
+
evidence: "tool_uses 43. 🟥 **자기 방법론 사고를 자백하고 전량 재실행** — 1차가 움직이는 FH 워킹트리를 소스로 써서 중간 측정이 설명 불가능한 값을 냈다(훅 1414줄). `git show a15e384:<path>` 로 동결 후 재실행하고 매 실행마다 계기 적재를 diff 로 대조(IDENTICAL). 커밋 4팔 exit 1/1/0/0 + 되돌림 컨트롤로 귀속 확인. 신규 FAIL 0(기준선 35=전파 후 35) · 기존 마커 42건 재판정 0 · LC_ALL=C sed remap 유니코드 무손상(바이트 IDENTICAL). **내 델타 안의 결함 2건 적발** — 픽스처 헤더가 기호 표기로 옛 4축 의미를 가르침 + selfcheck 주석 stale. 범위 밖 관측 2건을 미확인으로 분리 기재"
|
|
2142
|
+
cost: 180k tokens
|
|
2143
|
+
- date: 2026-08-17
|
|
2144
|
+
agent: codex/gpt-5.5 (cross-family sidecar, headless)
|
|
2145
|
+
purpose: "출하 배선 + axes-run 6축 델타 309줄에 대한 적대 리뷰. 공격 지점 6개를 명시 지정"
|
|
2146
|
+
outcome: accepted
|
|
2147
|
+
evidence: "8건 지적(S2·A3·B3). 그중 2건(혼용 오탐 · 화살표 fail-open)은 내가 투입 전 자력 수리했고 codex 가 독립으로 같은 지점을 지목해 수렴 확인 — 그 둘은 「둘 다 봤다」. **나머지 5건 자력 적발 0**: grep 이 주석 처리된 호출에 매칭(S) · if:false 로 꺼진 잡(S) · 트리거 미검사(A) · 잡 신원 미검사(A) · 마커 파일명 날짜 추출 fail-open(B, 선재). 각 지적에 깨뜨리는 입력을 구체적으로 제시해 손 재현이 즉시 가능했다"
|
|
2148
|
+
cost: 85k tokens
|
|
@@ -12,6 +12,18 @@ When a **sister asset** (another team in the organization · external frontier
|
|
|
12
12
|
- An external resource with **similar scope but different resolution** from a `knowledge/shared/` asset is found
|
|
13
13
|
- An external reference URL repeatedly appears in the weekly audit scanner aggregation
|
|
14
14
|
- User mentions "that other project/team also did this"
|
|
15
|
+
- 🟥 **ACTIVE ADOPTION — you went looking for an external asset, installed it, and RAN it** against
|
|
16
|
+
something this hub owns. Added 2026-08-16 because every condition above is **passive discovery**
|
|
17
|
+
("a resource *is found*"), and the strongest sister signal there is does not look like discovery
|
|
18
|
+
at all: deliberately reaching for an outside tool *because ours did not cover something* is a
|
|
19
|
+
measured resolution difference, not a hunch about one. **Measured miss**: an external red-team
|
|
20
|
+
framework was installed, run against a field harness's safety gate, and found a real bypass — and
|
|
21
|
+
none of it fired this protocol. It was filed in memory as `type: reference`, a **tool pointer**,
|
|
22
|
+
because "a tool I used" was an available and correct-looking category while "sister asset" was a
|
|
23
|
+
category no listed condition named. That is the normalization reflex `CLAUDE.md
|
|
24
|
+
§Envelope-Boundary Discipline` exists to counter, reproduced inside the protocol meant to catch it.
|
|
25
|
+
**The tell**: if you can answer *"what does it do that ours doesn't?"* — you have already done
|
|
26
|
+
step 1's resolution-difference analysis, and owe steps 2–3.
|
|
15
27
|
|
|
16
28
|
## Lightweight path (C-tier — cheap debt entry)
|
|
17
29
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@chrono-meta/fh-gate",
|
|
3
|
-
"version": "2.0
|
|
3
|
+
"version": "2.1.0",
|
|
4
4
|
"description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"keywords": [
|
|
@@ -27,7 +27,7 @@
|
|
|
27
27
|
"prepare": "chmod +x bin/fh-gate.js bin/fh-run.js bin/fh-goal.js bin/fh-codex-doctor.js scripts/fh-gate.sh scripts/fh-run.sh scripts/fh-goal.sh",
|
|
28
28
|
"postinstall": "node scripts/postinstall_notice.js",
|
|
29
29
|
"test": "bash scripts/selfcheck.sh",
|
|
30
|
-
"prepublishOnly": "bash scripts/
|
|
30
|
+
"prepublishOnly": "bash scripts/prepublish_scope_note.sh && bash scripts/publish_freshness_check.sh && bash scripts/version_lockstep_check.sh && bash scripts/package_coverage_check.sh --vs-tarball && bash scripts/public_surface_scan_files.sh",
|
|
31
31
|
"release": "bash scripts/public_surface_scan_files.sh && npm publish"
|
|
32
32
|
},
|
|
33
33
|
"engines": {
|
|
@@ -73,6 +73,7 @@
|
|
|
73
73
|
"scripts/test_version_lockstep_lanes.sh",
|
|
74
74
|
"scripts/package_coverage_check.sh",
|
|
75
75
|
"scripts/publish_freshness_check.sh",
|
|
76
|
+
"scripts/prepublish_scope_note.sh",
|
|
76
77
|
"scripts/lane_runner_check.sh",
|
|
77
78
|
"scripts/test_package_coverage_lanes.sh",
|
|
78
79
|
"scripts/adapters/peer_resolve.sh",
|
|
@@ -53,7 +53,12 @@ This agent repeatedly analyzes harness structure failure patterns. Even when a S
|
|
|
53
53
|
|
|
54
54
|
"This is weak" is not enough. **"Fix it like this" must accompany every valid attack.**
|
|
55
55
|
|
|
56
|
-
> **Role separation from steel-quench Wave 1**: Wave 1 attacks from
|
|
56
|
+
> **Role separation from steel-quench Wave 1**: Wave 1 attacks from **6** angles (reason for existence · real-world validation · bus factor · obsolescence · self-reference · **gate-locality**). quench-challenger attacks **orthogonal** harness-structure-specific 6 axes in an isolated independent instance. It does not replace Wave 1 — it covers the structural layer Wave 1 misses.
|
|
57
|
+
> ⚠️ Corrected 2026-08-16 — this line said "5 angles" and omitted gate-locality after `SKILL.md`
|
|
58
|
+
> added it, and `SKILL_detail.md`'s output template had no row for it either. Three copies of one
|
|
59
|
+
> list, drifted. **Adding a Wave 1 angle means editing all three** (`SKILL.md` table · this
|
|
60
|
+
> paragraph · `SKILL_detail.md` template); the two counts here are deliberately different numbers
|
|
61
|
+
> (Wave 1 = 6 angles, this agent = 6 axes, unrelated sets) so do not "reconcile" them.
|
|
57
62
|
|
|
58
63
|
---
|
|
59
64
|
|
|
@@ -10,6 +10,49 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/)
|
|
|
10
10
|
|
|
11
11
|
## Plugin Level
|
|
12
12
|
|
|
13
|
+
### [2.1.0] — 2026-08-17
|
|
14
|
+
|
|
15
|
+
🟥 **BREAKING (gate)**: `crossfamily: declined` in an Axes 2-3 marker now requires grounds naming a
|
|
16
|
+
record path that **resolves on disk**. Bare `declined`, and `declined` justified by author judgment,
|
|
17
|
+
are **blocked at commit**.
|
|
18
|
+
- **Remedy**: cite where the operator decision lives — e.g.
|
|
19
|
+
`crossfamily: declined — operator declined sidecars, per knowledge/shared/rules/operational_adaptation.md`
|
|
20
|
+
- **Or use the value that describes what actually happened**: if a panel was reachable and you chose
|
|
21
|
+
not to recruit it, that is `DEGRADED_PANEL_UNUSED`, not `declined`.
|
|
22
|
+
- **Why**: `declined` was the only value in the enum with no grounds requirement, so an author
|
|
23
|
+
judgment flowed into the nearest permissive token and passed clean. A cross-family review then
|
|
24
|
+
broke the first (vocabulary-grep) fix three ways — self-validating on the value's own token,
|
|
25
|
+
vacuous keyword passes, and over-blocking genuine declinations phrased in natural prose — so the
|
|
26
|
+
check asserts a **resolvable record** rather than words. It proves a cited record EXISTS, not that
|
|
27
|
+
it says what is claimed; that residual is the marker's own declared scope, stated in the code.
|
|
28
|
+
|
|
29
|
+
**Added**
|
|
30
|
+
- `standpoint:` gains **`tier1b`** (a STATIC read of a target repo — executed nothing) plus a
|
|
31
|
+
decide-in-order procedure, after blind floor-tier sims graded pure cold-reads as `tier2` for three
|
|
32
|
+
rounds, defeating three separate rewordings via the enum's own internal logic.
|
|
33
|
+
- steel-quench Wave 1's sixth angle (**gate-locality**) gains the output-template row it never had —
|
|
34
|
+
a mandatory angle that was structurally unreportable. Gate-locality failing its own check.
|
|
35
|
+
- `verify-bidirectional` category 5 — **prescriptive doctrine statement**, the operator-correction
|
|
36
|
+
shape that fell between its doubt-shaped triggers and its "simple correction, no review" exception.
|
|
37
|
+
- Sister Asset Protocol gains an **active-adoption** condition — all four prior conditions were
|
|
38
|
+
passive discovery, so installing an external framework and running it against our own asset
|
|
39
|
+
matched none of them.
|
|
40
|
+
- Resident doctrine: **Mechanization Boundary** (machinery at irreversible edges and channels;
|
|
41
|
+
judgment left to evolution) · **Local Execution First** (CI is a backstop, never the discovery
|
|
42
|
+
mechanism) · **Skeleton-not-Muscle** (a wiring change is done when the floor tier executes it) ·
|
|
43
|
+
**Expedition** track · **this package's versioning policy** (major reserved for built-anew /
|
|
44
|
+
identity-established / capability-class change — never for tightening an existing gate).
|
|
45
|
+
|
|
46
|
+
**Fixed**
|
|
47
|
+
- `fh-gate.sh` survives a missing or unreadable `package.json`; `selfcheck.sh` decides npm-surface
|
|
48
|
+
applicability at the call site.
|
|
49
|
+
- `capability_registry_check.sh` M3 basename-token matching (a path containing "claude" was silently
|
|
50
|
+
rejected) + a stdin channel for calibration arms.
|
|
51
|
+
- `subagent-tally-hook.json` root resolution is 3-tier, matching the convention the sibling
|
|
52
|
+
`SessionStart` hook already used.
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
13
56
|
### [2.0.1] — 2026-08-16
|
|
14
57
|
|
|
15
58
|
- **feat(cadence)**: `harness-doctor` 30일 캐던스가 산문 제안에서 SessionStart 훅으로 승격 —
|
|
@@ -264,8 +264,20 @@ enforceable layer) from the **declined** case above (chosen floor, first-class,
|
|
|
264
264
|
**shared-body / cross-harness-boundary** change — scoped by *effect* (alters another harness's
|
|
265
265
|
behavior, gate outcome, or interaction contract), not merely by touching a synced file path — also
|
|
266
266
|
emit `standpoint:` alongside `crossfamily:` in the same marker — recruiting family diversity here
|
|
267
|
-
does not substitute for it. Values: `tier1` (content-only, the default) ·
|
|
268
|
-
(
|
|
267
|
+
does not substitute for it. Values: `tier1` (content-only, the default) · **`tier1b(<harness>)`
|
|
268
|
+
(STATIC read of the target's own files — executed nothing)** · `tier2(<harness>)`
|
|
269
|
+
(peer-simulated — **EXECUTED CODE in** the target's own repo and observed the result.
|
|
270
|
+
🟥 **Reading the target's real files, however cold, is `tier1b`, not this.** The discriminator is
|
|
271
|
+
mechanical: *name the command you ran and the output you saw*; cannot name one → `tier1b`, always.
|
|
272
|
+
This line said **"ran the target's own repo, content only"** until 2026-08-17 — actively teaching
|
|
273
|
+
the opposite of the canon it summarizes. 🟥 **RETRACTED (2026-08-17)**: this passage used to add
|
|
274
|
+
that blind Sonnet sims *"STILL graded a pure cold-read `tier2`, both quoting this phrasing"* after
|
|
275
|
+
the other two copies were fixed. Those runs had **`tool_uses: 0`** — nothing was read, so nothing
|
|
276
|
+
was measured, and the claim that the sims "reached for this copy" was itself unfounded
|
|
277
|
+
(`tracks/_meta/fh_completed_2026-08-16.md:690`; the live re-run inverted the grade at reps=1, below
|
|
278
|
+
bar). **The reason to fix this copy needs no sim**: a summary that states the opposite of its canon
|
|
279
|
+
teaches the opposite to whoever reads only the summary, and **two of three copies fixed is a fix
|
|
280
|
+
that does not exist** — that is gate-locality, not a measurement. See `§7` for the canon) · `tier2b(<harness>)` (same operator,
|
|
269
281
|
target's real runtime — local wiring visible, but not an independent reviewer) · `tier3(<harness>)`
|
|
270
282
|
(a *different* operator of the target harness ran it — the only fully independent + local-wiring
|
|
271
283
|
rung) · `not-applicable` (no target-harness standpoint exists — most same-repo dispatches) ·
|
|
@@ -60,9 +60,22 @@ Added Wave 1 attack angles: N items
|
|
|
60
60
|
| Bus factor | S/A/B | [single-person dependency area] | ○/△/× |
|
|
61
61
|
| Platform obsolescence | S/A/B | [vulnerability point] | ○/△/× |
|
|
62
62
|
| Self-referential structure | S/A/B | [closed circuit detection result] | ○/△/× |
|
|
63
|
+
| Gate-locality | S/A/B | [any gate defined only where the enforcing actor never reads it] | ○/△/× |
|
|
63
64
|
|
|
64
65
|
S-grade blockers: N / A-grade: N / B-grade: N
|
|
65
66
|
|
|
67
|
+
> 🟥 **The 6th row was missing until 2026-08-16, and its absence was itself a gate-locality defect.**
|
|
68
|
+
> `SKILL.md` Wave 1 has defined **six** angles since Gate-locality was added; this template shipped
|
|
69
|
+
> **five rows**, and `quench-challenger.md` still said "5 angles" and omitted it from the
|
|
70
|
+
> enumeration. So a mandatory angle had **no slot to be reported in** — an agent filling this table
|
|
71
|
+
> correctly would report the other five and structurally never surface the sixth. That is the exact
|
|
72
|
+
> shape Gate-locality itself is the check for: *a requirement placed where the actor who must
|
|
73
|
+
> satisfy it does not read it.* Found by a sister-asset audit (an external adversarial framework's
|
|
74
|
+
> **typed** attack registry vs. this repo's seven overlapping prose lists) — the fragmentation is
|
|
75
|
+
> what let the three copies drift apart silently. If you add a Wave 1 angle, it lands in **three**
|
|
76
|
+
> places or it does not land: `SKILL.md`'s table, this template, and `quench-challenger.md`'s
|
|
77
|
+
> role-separation paragraph.
|
|
78
|
+
|
|
66
79
|
Optional numeric score (0.0–1.0):
|
|
67
80
|
overall_score: {score}
|
|
68
81
|
[0.0–0.3] S-grade present → immediate blocker, do not proceed
|
|
@@ -39,8 +39,41 @@ Also triggered in external user environments by these natural language phrases:
|
|
|
39
39
|
| "what's your basis?", "why do you think that?" | Baseline grep trigger |
|
|
40
40
|
| "check that one more time" | Self-validation request |
|
|
41
41
|
|
|
42
|
+
### 🟥 5. Prescriptive doctrine statement — the category that fell through (added 2026-08-16)
|
|
43
|
+
|
|
44
|
+
Every trigger above is shaped as **doubt about a claim**. A large class of operator correction is not
|
|
45
|
+
doubt at all — it is a **standing rule being handed over**, stated as fact or instruction:
|
|
46
|
+
|
|
47
|
+
| Shape | Real examples (2026-08-16 session) |
|
|
48
|
+
|---|---|
|
|
49
|
+
| *"X should also include Y"* | *"입장리뷰에는 정적리뷰뿐만 아니라 **동적리뷰도 포함되어야 할 거야**"* |
|
|
50
|
+
| *"the real point of X is Z"* (redefining, not doubting) | *"부스팅보다 인큐베이터의 장점은 그 레포 전체를 감싸서 …**모든 것을 조작할 수 있는 권한**이 있는 거야"* |
|
|
51
|
+
| *"isn't this too late / wrong-ordered?"* (rhetorical, expects agreement) | *"이 실패가 CI 확인 단계에서야 발견되는 건 매우 **늦은 게 아닐까**"* |
|
|
52
|
+
| *"from now on, do W at Z"* | *"앞으로도 마감할 때 그 갈래로 이어갈 수 있게 **알아서 정리해줘**"* |
|
|
53
|
+
|
|
54
|
+
**Why these were missed 4 times out of 4** — they sit in the blind spot between the triggers (which
|
|
55
|
+
expect a *challenge*) and the Exceptions below (which release a *simple correction* with "no
|
|
56
|
+
review"). A doctrine statement is neither: it does not dispute a claim, and it is not a one-off fix
|
|
57
|
+
to apply and forget. Treated as the exception, it gets a verbal acknowledgment and evaporates at
|
|
58
|
+
session end. **Measured**: four such statements in one session, all acknowledged in conversation,
|
|
59
|
+
**zero landed in any file** until a later review grepped for them and found nothing.
|
|
60
|
+
|
|
61
|
+
**The tell is grammatical, and it is language-independent** — the listed phrases above are all
|
|
62
|
+
English interrogatives, which is why a Korean declarative (`~해야 할 거야` · `~인 거야` · `~아닐까`)
|
|
63
|
+
matched none of them. Do not fix this by appending more literals (Grep-Collision Treadmill, P10);
|
|
64
|
+
the discriminator is: **does this utterance describe how I should behave from now on, rather than
|
|
65
|
+
what is wrong with this one output?** If yes, it is this category regardless of language or phrasing.
|
|
66
|
+
|
|
67
|
+
**Required action** — an acknowledgment is not compliance. The statement must land **in a file** in
|
|
68
|
+
the same session: canon (`CLAUDE.md` / `knowledge/`) if it governs future behavior, memory if it is
|
|
69
|
+
about this operator, a signal file if it needs a decision first. Then say WHERE it landed, so the
|
|
70
|
+
operator can see it did.
|
|
71
|
+
|
|
42
72
|
**Exceptions** (this skill does NOT apply):
|
|
43
|
-
- Simple user correction ("this is wrong, redo it") = direct negation → immediate correction (no review)
|
|
73
|
+
- Simple user correction ("this is wrong, redo it") = direct negation → immediate correction (no review).
|
|
74
|
+
⚠️ **Not the same as category 5 above** — "redo this" is scoped to one output; "from now on do X"
|
|
75
|
+
is a standing rule. When both readings fit, take it as category 5: over-landing costs a file edit,
|
|
76
|
+
under-landing loses the rule entirely.
|
|
44
77
|
- This harness AI self-catch (no external counter-argument) = `fact-checker` rule (narrow 1 / broad N+1)
|
|
45
78
|
|
|
46
79
|
## Execution Steps
|
|
@@ -54,7 +54,7 @@ RC_OK=0; RC_REJECT=1; RC_HARNESS=10
|
|
|
54
54
|
# `summary`/`tags` 는 **추천 전용**(2026-08-16, cluster-wizard). 판정에는 안 쓰이고
|
|
55
55
|
# `cluster_capability_scan.sh recommend` 의 어휘 매칭에만 쓰인다. 닫힌 목록에 넣는 이유는
|
|
56
56
|
# 이 목록의 목적이 «오타 축 무음 드롭 방지» 이기 때문이다 — 안 넣으면 정당한 키가 SCHEMA 로 막힌다.
|
|
57
|
-
CLOSED_KEYS="id entry requires_cwd summary tags verdict_channel verdict_enum verdict_stdout_key upstream_argv echoes_upstream approval reversibility residency degrade tier_floor writes judge verdict_binding calibration_positive_args calibration_positive_expect calibration_negative_args calibration_negative_expect"
|
|
57
|
+
CLOSED_KEYS="id entry requires_cwd summary tags verdict_channel verdict_enum verdict_stdout_key upstream_argv echoes_upstream approval reversibility residency degrade tier_floor writes judge verdict_binding calibration_positive_args calibration_positive_expect calibration_negative_args calibration_negative_expect calibration_positive_stdin calibration_negative_stdin"
|
|
58
58
|
|
|
59
59
|
# 「안 돌았다」를 뜻하는 이름들 — 추가조항(§ⓑ.4 B1)이 요구하는 구분항
|
|
60
60
|
DIDNOTRUN_NAMES="DID_NOT_RUN DIDNOTRUN NOT_RUN NO_TARGET SKIPPED UNMEASURED NOT_CONFIGURED HARNESS_ERROR"
|
|
@@ -70,6 +70,9 @@ _parse() {
|
|
|
70
70
|
CAP_id=""; CAP_entry=""; CAP_requires_cwd=""; CAP_verdict_channel=""
|
|
71
71
|
CAP_verdict_enum=""; CAP_verdict_stdout_key=""; CAP_judge=""; CAP_writes=""
|
|
72
72
|
CAP_cal_pos_args=""; CAP_cal_pos_expect=""; CAP_cal_neg_args=""; CAP_cal_neg_expect=""
|
|
73
|
+
# 다중 capfile 실행에서 앞 파일의 선언이 뒤 파일로 새지 않게 한다 — 초기화 누락은
|
|
74
|
+
# "뒤 파일이 선언하지 않은 stdin 으로 돌았다" 를 만들고, 그건 조용히 통과한다.
|
|
75
|
+
CAP_cal_pos_stdin=""; CAP_cal_neg_stdin=""
|
|
73
76
|
UNKNOWN_KEYS=""; CAP_requires_cwd_was_self=0
|
|
74
77
|
# 「선언했는데 값이 비었다」와 「선언 자체가 없다」는 다른 사실이다 — 관측 범위 블록이
|
|
75
78
|
# 둘을 구분해 적으려면 파서가 그 구분을 보존해야 한다(빈 문자열 하나로는 못 나눈다).
|
|
@@ -93,6 +96,8 @@ _parse() {
|
|
|
93
96
|
calibration_positive_args) CAP_cal_pos_args="$val"; CAP_cal_pos_args_seen=1 ;;
|
|
94
97
|
calibration_positive_expect) CAP_cal_pos_expect="$val" ;;
|
|
95
98
|
calibration_negative_args) CAP_cal_neg_args="$val"; CAP_cal_neg_args_seen=1 ;;
|
|
99
|
+
calibration_positive_stdin) CAP_cal_pos_stdin="$val" ;;
|
|
100
|
+
calibration_negative_stdin) CAP_cal_neg_stdin="$val" ;;
|
|
96
101
|
calibration_negative_expect) CAP_cal_neg_expect="$val" ;;
|
|
97
102
|
esac
|
|
98
103
|
done < "$1"
|
|
@@ -119,10 +124,43 @@ _validate_arm_args() { # $1=args → 셸 메타문자/상위경로 탈출을
|
|
|
119
124
|
return 0
|
|
120
125
|
}
|
|
121
126
|
|
|
122
|
-
|
|
127
|
+
# stdin 파일 검증 — args 와 같은 규율(메타문자·상위경로 탈출 거부) + 실재 확인.
|
|
128
|
+
_validate_arm_stdin() { # $1=선언된 경로(빈 값 허용)
|
|
129
|
+
[ -n "$1" ] || return 0
|
|
130
|
+
case "$1" in
|
|
131
|
+
*'|'*|*';'*|*'&'*|*'>'*|*'<'*|*'`'*|*'$('*|*$'\n'*)
|
|
132
|
+
_fail "M4" "캘리브레이션 stdin 경로에 셸 메타문자가 있다: $1"; return 1 ;;
|
|
133
|
+
*'../'*|/*)
|
|
134
|
+
_fail "M4" "캘리브레이션 stdin 경로가 레포 밖을 가리킨다(상대경로만 허용): $1"; return 1 ;;
|
|
135
|
+
esac
|
|
136
|
+
[ -f "$CAP_requires_cwd/$1" ] || {
|
|
137
|
+
_fail "M4" "캘리브레이션 stdin 파일이 없다: $1 (선언은 실재하는 픽스처를 가리켜야 한다)"; return 1; }
|
|
138
|
+
return 0
|
|
139
|
+
}
|
|
140
|
+
|
|
141
|
+
_run_arm() { # $1=extra args · $2=stdin 파일(선택) → ARM_RC / ARM_NAME. 파이프로 읽지 않는다(PIPE-VERDICT).
|
|
142
|
+
# ── STDIN 은 항상 정의된다. 상속하지 않는다. ────────────────────────────────
|
|
143
|
+
# 실측 2026-08-16 (자기선언 캠페인, the-bible arm): 이 함수는 `> /dev/null 2>&1` 로
|
|
144
|
+
# stdout/stderr 만 막고 **stdin 은 검사기의 것을 물려받았다.** 결과가 두 가지로 나빴다.
|
|
145
|
+
# ⓐ **비결정**: 같은 선언이 터미널·파이프·CI 어디서 돌리느냐에 따라 arm 의 입력이 달라진다.
|
|
146
|
+
# 판정 계기 안의 비결정은 그 자체로 결함이다 — 초록의 이유가 실행 환경에 달리게 된다.
|
|
147
|
+
# ⓑ **표현 불가**: 대상을 stdin 으로 받는 진입점(JSON 브리지 등)은 양·음 arm 이 **둘 다 빈
|
|
148
|
+
# 입력**으로 돌아 같은 코드를 내고, 판별 실패로 REJECT 된다. 능력이 없어서가 아니라
|
|
149
|
+
# **스키마에 채널이 없어서** 막힌 것이다. 실제로 4방향으로 갈리는 게이트가 그렇게 막혔다.
|
|
150
|
+
# ⇒ 선언된 파일이 있으면 그걸 먹이고, 없으면 **명시적으로 /dev/null** 을 먹인다. 후자가
|
|
151
|
+
# 중요하다 — "선언 안 했으니 상속"이 아니라 "선언 안 했으면 빈 입력"이라고 **정해야** ⓐ 가 닫힌다.
|
|
152
|
+
# 🟥 `set --` 는 위치 인자를 **덮어쓴다** — 그 뒤의 `$2` 는 stdin 파일이 아니라 entry 의 두 번째
|
|
153
|
+
# 토큰이다. 초판이 그걸 밟았고, 그 결과 두 arm 이 **진입점 스크립트 자신을 stdin 으로 먹었다.**
|
|
154
|
+
# 발견 경위가 이 수정의 요점이다: 그때 **음성 arm 만 빨개졌고 양성 arm 은 통과했다** — 스크립트
|
|
155
|
+
# 소스에 마침 양성 픽스처와 같은 토큰이 들어 있어서다. 즉 **둘 다 틀린 입력을 먹었는데 하나가
|
|
156
|
+
# 우연히 기대값과 맞았다.** 한쪽만 보고 있었으면 «양성은 되는데 음성이 이상하다»로 오진했다.
|
|
157
|
+
# ⇒ set -- 앞에서 이름 있는 변수로 잡는다.
|
|
158
|
+
local _stdin="${2:-}"
|
|
123
159
|
( cd "$CAP_requires_cwd" 2>/dev/null || exit 127
|
|
124
160
|
# shellcheck disable=SC2086 # argv 토큰 분리는 의도 (noglob 로 확장은 막혀 있다)
|
|
125
|
-
set -- $CAP_entry $1
|
|
161
|
+
set -- $CAP_entry $1
|
|
162
|
+
if [ -n "$_stdin" ]; then "$@" < "$_stdin"; else "$@" < /dev/null; fi
|
|
163
|
+
) > /dev/null 2>&1
|
|
126
164
|
ARM_RC=$?
|
|
127
165
|
ARM_NAME="$(_enum_name_of "$ARM_RC" "$CAP_verdict_enum")"
|
|
128
166
|
}
|
|
@@ -256,11 +294,35 @@ _check_one() {
|
|
|
256
294
|
# ── M3 모델 독립성 (선언 검사 + 명백한 모순만) ─────────────────────────────
|
|
257
295
|
case "$CAP_judge" in
|
|
258
296
|
mechanical)
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
297
|
+
# 토큰의 BASENAME 으로 본다 — 문자열 전체 부분매칭이 아니다.
|
|
298
|
+
#
|
|
299
|
+
# 실측 2026-08-16 (자기선언 캠페인, qasp-dev arm 이 지목): 이 검사는 `*claude*` 처럼
|
|
300
|
+
# entry **문자열 전체**에 부분매칭했다. 그래서 진입점이 모델을 전혀 안 부르는데도
|
|
301
|
+
# **경로에 그 단어가 있다는 이유만으로** REJECT 났다 — 실제 사례는 스크래치패드 경로
|
|
302
|
+
# `/private/tmp/claude-501/...` 였고, 경로만 바꾸니 같은 선언이 PASS 했다.
|
|
303
|
+
# 구조적으로 더 나쁜 경우가 둘 있다: ⓐ 진입점이 `.claude/` 아래 사는 하네스는 **전부**
|
|
304
|
+
# 막힌다(FH 자기 선언이 `.claude/capabilities/` 에 사는 걸 생각하면 남 얘기가 아니다)
|
|
305
|
+
# ⓑ `scripts/claude_md_lint.sh` 처럼 이름에 그 단어가 든 순수 기계 스크립트도 막힌다.
|
|
306
|
+
#
|
|
307
|
+
# 과차단은 «안전한 방향»이 아니다 — 정본이 명시하듯 override 를 습관화시켜 같은 훅의
|
|
308
|
+
# 다른 게이트까지 무장해제시킨다. 그리고 이 오탐은 **무음**이었다: 거부 사유가
|
|
309
|
+
# "모델 CLI 를 부른다" 로 찍히므로, 읽는 사람은 자기 진입점을 의심하지 경로를 의심하지 않는다.
|
|
310
|
+
#
|
|
311
|
+
# 이 검사가 실제로 답할 수 있는 질문은 «entry 줄이 모델 CLI 를 **직접 이름으로** 부르는가»
|
|
312
|
+
# 뿐이다. 래퍼 안에서 부르는 건 선언 층에서 원래 못 본다(M6 효과 프로브의 몫).
|
|
313
|
+
# 그러니 각 토큰의 basename 을 보고, 확장자를 떼고, **정확히 그 이름일 때만** 막는다.
|
|
314
|
+
_m3_bad=""
|
|
315
|
+
for _tok in $CAP_entry; do
|
|
316
|
+
_base="${_tok##*/}"; _base="${_base%.sh}"; _base="${_base%.js}"; _base="${_base%.py}"
|
|
317
|
+
case "$_base" in
|
|
318
|
+
claude|codex|gemini|copilot|ollama|llm) _m3_bad="$_base" ;;
|
|
319
|
+
esac
|
|
320
|
+
done
|
|
321
|
+
if [ -n "$_m3_bad" ]; then
|
|
322
|
+
_fail "M3" "judge: mechanical 선언인데 entry 가 모델 CLI 를 직접 부른다: $_m3_bad ($CAP_entry)"
|
|
323
|
+
else
|
|
324
|
+
_ok "M3" "judge: mechanical (entry 토큰에 모델 CLI 없음 — basename 기준)"
|
|
325
|
+
fi ;;
|
|
264
326
|
model) _ok "M3" "judge: model — 선언됨(합법). 조합에서 이 PASS 는 NON_CLEARING 이다" ;;
|
|
265
327
|
'') _fail "M3" "judge 축 미선언 — 모델 개입 여부가 불명이면 조합이 계산될 수 없다" ;;
|
|
266
328
|
*) _fail "M3" "judge 값이 {mechanical|model} 밖: '$CAP_judge'" ;;
|
|
@@ -314,15 +376,16 @@ _check_one() {
|
|
|
314
376
|
elif [ "$FILE_FAILED" -eq 1 ] && [ -z "${CRC_FORCE_M4:-}" ]; then
|
|
315
377
|
printf ' ⏭ M4 — 앞선 축이 실패해 실행 생략(SKIPPED, PASS 아님)\n'
|
|
316
378
|
OBS_WHY="앞선 축 실패로 M4 생략"
|
|
317
|
-
elif ! _validate_arm_args "$CAP_cal_pos_args" || ! _validate_arm_args "$CAP_cal_neg_args"
|
|
379
|
+
elif ! _validate_arm_args "$CAP_cal_pos_args" || ! _validate_arm_args "$CAP_cal_neg_args" \
|
|
380
|
+
|| ! _validate_arm_stdin "${CAP_cal_pos_stdin:-}" || ! _validate_arm_stdin "${CAP_cal_neg_stdin:-}"; then
|
|
318
381
|
printf ' ⏭ M4 — args 검문 실패로 arm 을 실행하지 않았다(SKIPPED, PASS 아님)\n'
|
|
319
382
|
OBS_WHY="args 검문 실패"
|
|
320
383
|
else
|
|
321
384
|
local pos_name neg_name pos_rc neg_rc pos_name2 pos_rc2
|
|
322
|
-
_run_arm "$CAP_cal_pos_args"; pos_rc="$ARM_RC"; pos_name="$ARM_NAME"
|
|
385
|
+
_run_arm "$CAP_cal_pos_args" "${CAP_cal_pos_stdin:-}"; pos_rc="$ARM_RC"; pos_name="$ARM_NAME"
|
|
323
386
|
[ "$pos_rc" = "127" ] && _fail "M4" "requires_cwd 로 진입 실패 — arm 을 돌릴 수 없다"
|
|
324
|
-
_run_arm "$CAP_cal_pos_args"; pos_name2="$ARM_NAME"; pos_rc2="$ARM_RC" # reps=2 (M3 부분 방어)
|
|
325
|
-
_run_arm "$CAP_cal_neg_args"; neg_rc="$ARM_RC"; neg_name="$ARM_NAME"
|
|
387
|
+
_run_arm "$CAP_cal_pos_args" "${CAP_cal_pos_stdin:-}"; pos_name2="$ARM_NAME"; pos_rc2="$ARM_RC" # reps=2 (M3 부분 방어)
|
|
388
|
+
_run_arm "$CAP_cal_neg_args" "${CAP_cal_neg_stdin:-}"; neg_rc="$ARM_RC"; neg_name="$ARM_NAME"
|
|
326
389
|
# 관측 원장에 남긴다 — 판정과 무관하게, 이 실행이 **실제로 밟은 exit** 만 적는다.
|
|
327
390
|
OBS_STATE=run; OBS_POS_RC="$pos_rc"; OBS_POS_RC2="$pos_rc2"; OBS_NEG_RC="$neg_rc"
|
|
328
391
|
OBS_POS_NAME="$pos_name"; OBS_NEG_NAME="$neg_name"
|
package/scripts/chamber_run.sh
CHANGED
|
@@ -125,8 +125,11 @@ if [ ! -f "$WS/BUDGET.md" ]; then
|
|
|
125
125
|
|
|
126
126
|
# Route through goal-quench's budget gate for an expensive run, then record here.
|
|
127
127
|
# Demo-scale runs may self-cap — but an ESTIMATE line is mandatory (this is the entry cap).
|
|
128
|
+
#
|
|
129
|
+
# 🟥 ACTUAL 은 이 파일에 적지 마라 — step 6 이 별도 ACTUAL.md 를 쓴다.
|
|
130
|
+
# 이유: 이 파일의 판정-전 해시가 순서 증인(ship_readiness_gate §② P1)이다. 사후에
|
|
131
|
+
# 이 파일을 고치면 증인이 TAMPERED 로 죽는다 (2026-08-17 런 #11 실측).
|
|
128
132
|
ESTIMATE: <e.g. ~3 persona dispatches, demo-scale, self-capped | or a token budget>
|
|
129
|
-
ACTUAL: <filled at step 6>
|
|
130
133
|
EOF
|
|
131
134
|
echo " ⛔ step 3 BLOCKED: record an ESTIMATE in $WS/BUDGET.md (budget-entry gate G2), then re-run."; exit 1
|
|
132
135
|
fi
|
|
@@ -203,8 +206,26 @@ if [ -f "$_WITNESS" ]; then
|
|
|
203
206
|
fi
|
|
204
207
|
|
|
205
208
|
# STEP 6 — actual cost / carry-forward record.
|
|
206
|
-
|
|
207
|
-
|
|
209
|
+
# 🟥 별도 파일이다. BUDGET.md 가 아니다 — 그리고 그 분리가 이 게이트의 핵심이다.
|
|
210
|
+
# 2026-08-17 런 #11(첫 형식 완주)이 실측한 결함: step 3 이 BUDGET.md 의 해시를 순서 증인으로
|
|
211
|
+
# 기록하는데 step 6 이 같은 파일에 ACTUAL 을 요구하며 하드 차단했다 → 완주하면 반드시
|
|
212
|
+
# 증인 아티팩트가 사후 변경되고 verify 가 TAMPERED 를 낸다 → **완주한 런은 구조적으로
|
|
213
|
+
# ② 승급 근거가 될 수 없었다.** 대상 선정 실수가 아니라 한 파일에 두 역할(①불변 요구 =
|
|
214
|
+
# 증인 · ②변경 요구 = 사후 캘리브레이션)이 겹친 것이고, 개별로는 둘 다 옳아서 어느 쪽
|
|
215
|
+
# 코드를 읽어도 안 보였다. 그래서 스펙(P1 이 BUDGET 을 증인으로 지목)은 안 건드리고
|
|
216
|
+
# **변경 요구만** 빼낸다. ACTUAL.md 는 정의상 판정 후 산물이므로 **증인에 기록하지 않는다.**
|
|
217
|
+
if [ ! -f "$WS/ACTUAL.md" ]; then
|
|
218
|
+
cat > "$WS/ACTUAL.md" <<EOF
|
|
219
|
+
# ACTUAL — $SLUG (chamber run, 판정 후 기록)
|
|
220
|
+
#
|
|
221
|
+
# 이 파일은 **순서 증인 대상이 아니다** — 판정 후에 쓰이는 것이 정상이므로 해시를 걸지 않는다.
|
|
222
|
+
# ESTIMATE 는 BUDGET.md 에 있고 그 파일은 판정 전에 얼어 있다. 대조는 사람이 한다.
|
|
223
|
+
ACTUAL: <number/summary — dispatch 수 · 토큰 · 예산 대비 편차>
|
|
224
|
+
EOF
|
|
225
|
+
fi
|
|
226
|
+
if grep -qE '^ACTUAL:[[:space:]]*<' "$WS/ACTUAL.md" || ! grep -qE '^ACTUAL:[[:space:]]*\S' "$WS/ACTUAL.md"; then
|
|
227
|
+
echo " ⛔ step 6 BLOCKED: record ACTUAL cost in $WS/ACTUAL.md (actual-vs-estimate calibration)."
|
|
228
|
+
echo " 🟥 BUDGET.md 가 아니다 — 그 파일은 판정 전 증인이라 사후 변경하면 증인이 죽는다."
|
|
208
229
|
echo " Expected — a LINE PREFIX, not a heading:"
|
|
209
230
|
echo " ACTUAL: <number/summary> ← this form is checked"
|
|
210
231
|
echo " ## ACTUAL ← a heading does NOT satisfy it"
|
|
@@ -140,6 +140,31 @@ HDR
|
|
|
140
140
|
echo "❌ 원장 헤더 쓰기 실패: $LEDGER" >&2; return 10
|
|
141
141
|
fi
|
|
142
142
|
fi
|
|
143
|
+
# ── 멱등: 같은 (run, artifact, sha) 삼중항이 이미 있으면 다시 안 적는다 ──────────
|
|
144
|
+
# 기원 2026-08-17 런 #11(첫 형식 완주): 이 함수가 **무조건 append** 라서 러너를 advance
|
|
145
|
+
# 할 때마다 step 2~5 가 재실행되며 같은 해시가 다시 쌓였다 — 4스텝 런에 **엔트리 11개**.
|
|
146
|
+
# 증인 판정 자체는 안 틀리지만(같은 해시는 같은 결론) 공개 tracked 파일이 부풀고,
|
|
147
|
+
# verify 출력이 같은 줄을 여러 번 뱉어 **읽는 사람이 「몇 건이 문제인가」를 오독**한다.
|
|
148
|
+
# 🟥 내용이 바뀐 경우는 **여전히 새 엔트리로 남는다** — 그게 TAMPERED 를 성립시키는
|
|
149
|
+
# 증거이므로 여기서 접으면 안 된다. 접는 것은 «완전히 동일한 재기록»뿐이다.
|
|
150
|
+
# 🟥 오프셋 주의 — `- run: ` 는 **7자**라 값은 8부터다. 초판이 9로 썼고(cross-family
|
|
151
|
+
# gpt-5.5 적발, A급) 그러면 슬러그 첫 글자가 잘려 **어떤 정상 엔트리와도 매칭되지 않는다**
|
|
152
|
+
# → 가드가 조용히 무력화되고 append 가 계속된다. 재현: `printf -- '- run: g2\n' |
|
|
153
|
+
# awk '{print substr($0,9)}'` → `2`. **레인이 이걸 못 잡았다**(아래 L14 신설로 닫음).
|
|
154
|
+
# 인접성도 강제한다: run → artifact → sha256 이 **연속**일 때만 인정. 안 그러면 깨진
|
|
155
|
+
# YAML 에서 stale run/art 가 무관한 sha 와 짝지어 **필요한 append 를 눌러버린다**(같은 리뷰 B급).
|
|
156
|
+
if awk -v r="$slug" -v a="$base" -v s="$h" '
|
|
157
|
+
/^- run: / { run=substr($0,8); art=""; prev="run"; next }
|
|
158
|
+
/^ artifact: / { if (prev=="run") { art=substr($0,13); prev="art" }
|
|
159
|
+
else { art=""; prev="" } ; next }
|
|
160
|
+
/^ sha256: / { if (prev=="art" && run==r && art==a && substr($0,11)==s) { found=1; exit }
|
|
161
|
+
prev=""; next }
|
|
162
|
+
{ prev="" }
|
|
163
|
+
END { exit(found?0:1) }
|
|
164
|
+
' "$LEDGER" 2>/dev/null; then
|
|
165
|
+
echo "witnessed(already): $slug/$base — 동일 해시가 이미 원장에 있다(재기록 안 함)"
|
|
166
|
+
return 0
|
|
167
|
+
fi
|
|
143
168
|
if ! printf -- '- run: %s\n artifact: %s\n sha256: %s\n recorded: %s\n' \
|
|
144
169
|
"$slug" "$base" "$h" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" >> "$LEDGER" 2>/dev/null; then
|
|
145
170
|
echo "❌ 원장 append 실패: $LEDGER" >&2; return 10
|
package/scripts/fh-gate.sh
CHANGED
|
@@ -34,7 +34,22 @@ set -euo pipefail
|
|
|
34
34
|
FH_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
|
|
35
35
|
# Single source of truth: read version from the package.json shipped alongside this script.
|
|
36
36
|
# No jq dependency (users may not have it); fall back to "unknown" if unreadable.
|
|
37
|
-
VERSION
|
|
37
|
+
# VERSION — the `|| true` is load-bearing under `set -euo pipefail`, and it is the ONLY part that
|
|
38
|
+
# is. Without it, a package.json that is missing OR unreadable makes this pipeline non-zero, and
|
|
39
|
+
# because the whole thing is an ASSIGNMENT, `set -e` aborts the script right here — so the
|
|
40
|
+
# `:-unknown` fallback on the next line never runs and the gate exits 1 before doing any work.
|
|
41
|
+
# MEASURED 2026-08-16: a sibling harness carries this file verbatim but ships no package.json (it
|
|
42
|
+
# is not an npm package), so its gate was dead on every invocation and its 30-lane regression
|
|
43
|
+
# suite reported 30 failures that were all this one line. One lane appeared to PASS only because
|
|
44
|
+
# its expected exit code happened to be 1 — the textbook signature of a dead control. Nobody had
|
|
45
|
+
# looked, because that repo had no CI. Latent HERE too, not merely a sibling's problem: a copy of
|
|
46
|
+
# this script run from a directory with no package.json died the same way.
|
|
47
|
+
# A `[ -f "$FH_ROOT/package.json" ]` guard was written first and then REMOVED, because a 4-way
|
|
48
|
+
# mutant matrix refuted it: with the file merely absent either mechanism alone suffices, but with
|
|
49
|
+
# the file present-and-unreadable the guard passes and only `|| true` survives. A guard that
|
|
50
|
+
# covers nothing `|| true` does not already cover is decoration, and the first version of this
|
|
51
|
+
# comment claimed both were load-bearing — that claim was false.
|
|
52
|
+
VERSION="$(sed -n 's/.*"version"[[:space:]]*:[[:space:]]*"\([^"]*\)".*/\1/p' "$FH_ROOT/package.json" 2>/dev/null | head -1 || true)"
|
|
38
53
|
VERSION="${VERSION:-unknown}"
|
|
39
54
|
CALLER_CWD="$(pwd -P)"
|
|
40
55
|
_TMPDIR="${TMPDIR:-/tmp}"
|
|
@@ -51,6 +51,16 @@ ACCEPTED_ABSENT=(
|
|
|
51
51
|
# read as claims about THEIR file. The shipped docs name it as the place verdicts live in the
|
|
52
52
|
# harness repo, which is a pointer for contributors, not a promise of a shipped artifact.
|
|
53
53
|
".claude/regression/ablation_verdicts.md"
|
|
54
|
+
# 챔버 **순서 증인 원장**(ship_readiness_gate §② P1). 바로 위 ablation_verdicts 와 **같은
|
|
55
|
+
# 논거**다: 이건 THIS repo 의 챔버 런이 언제 무엇을 고정했는지에 대한 기록이고, 소비자의
|
|
56
|
+
# 챔버 런은 그들 것이다 — 출하 문서에서 인용된 채로 딸려가면 소비자가 **자기 런에 대한
|
|
57
|
+
# 주장**으로 읽는다. 게다가 이 원장은 우리 런 slug·시각을 담은 공개 표면이라 소비자에게
|
|
58
|
+
# 보내는 것은 정보 유출 방향으로도 틀렸다.
|
|
59
|
+
# 🟥 **부재가 소비자 쪽 기능을 깨지 않는다** — `chamber_witness.sh do_record` 는 원장이
|
|
60
|
+
# 없으면 헤더를 만들어 생성한다(같은 파일의 `[ ! -f "$LEDGER" ]` 분기). 즉 스크립트는
|
|
61
|
+
# 출하되고 원장은 소비자 머신에서 처음 쓸 때 생긴다. 「없으면 죽는다」가 아니라
|
|
62
|
+
# 「없는 게 정상 초기 상태」다.
|
|
63
|
+
"knowledge/shared/learnings/chamber_ordering_witness.yaml"
|
|
54
64
|
"scripts/sync-to-be.sh"
|
|
55
65
|
"scripts/sync_guard_check.sh"
|
|
56
66
|
# Return path (companion store → hub) and its anchor. Same reason as the forward path above: the
|