@chrono-meta/fh-gate 2.1.0 → 2.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +3 -3
- package/AGENTS.md +10 -3
- package/knowledge/shared/harness-core/harness_terminal_correlation_and_recommendations.md +36 -2
- package/knowledge/shared/harness-core/ship_readiness_gate.md +36 -0
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +36 -0
- package/package.json +5 -1
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-commons/skills/ko-tech-writer/fixtures/known_negative.md +19 -0
- package/plugins/fh-commons/skills/ko-tech-writer/fixtures/known_positive.md +20 -0
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +45 -0
- package/scripts/chamber_run.sh +15 -2
- package/scripts/ko_tech_writer_calibrate.py +127 -0
- package/scripts/publish_freshness_check.sh +21 -2
- package/scripts/selfcheck.sh +2 -1
- package/scripts/test_ko_tech_writer_lanes.sh +61 -0
|
@@ -11,13 +11,13 @@
|
|
|
11
11
|
"plugins": [
|
|
12
12
|
{
|
|
13
13
|
"name": "fh-meta",
|
|
14
|
-
"version": "2.
|
|
15
|
-
"description": "New in 2.1.0: BREAKING (gate): `crossfamily: declined` in an Axes 2-3 marker now requires grounds naming a record path that RESOLVES on disk — bare `declined`, and `declined` justified by author judgment, are blocked at commit. Remedy: cite where the operator decision lives (e.g. `.. — operator declined sidecars, per knowledge/shared/rules/operational_adaptation.md`), or use `DEGRADED_PANEL_UNUSED` if a panel was reachable and you chose not to recruit it — which is what author judgment actually is. `declined` was the only enum value with no grounds requirement; a cross-family review then broke the first (vocabulary-grep) fix three ways — self-validating on the value's own token, vacuous keyword passes, and over-blocking real declinations in natural prose — so the check asserts a resolvable record instead of words. Also: standpoint axis gains `tier1b` (a STATIC read of a target repo, executed nothing) plus a decide-in-order procedure, after blind floor-tier sims graded pure cold-reads as `tier2` three rounds running; steel-quench Wave 1's sixth angle (gate-locality) gains the output-template row it never had, so a mandatory angle stops being structurally unreportable; verify-bidirectional gains category 5 (prescriptive doctrine statement); Sister Asset Protocol gains an active-adoption trigger; new resident doctrine — Mechanization Boundary, Local Execution First, Skeleton-not-Muscle, Expedition track, and this package's versioning policy. Hub meta-operations toolkit — 35 skills + 7 agents. New in 2.0.1: harness-doctor cadence hook, portability lint wired into pre-commit, branch_claim.sh claim-count-vs-tree-count warning, louder confidentiality-scan fail-open notice, fh-gate.sh missing-package.json survival, identity ① reclassified 🟢 (cross-harness adapters + relay argument channel). New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
14
|
+
"version": "2.2.0",
|
|
15
|
+
"description": "New in 2.2.0: BREAKING (gate): chamber step 6 now reads ACTUAL.md, not BUDGET.md — an in-flight chamber run whose actual cost sits in BUDGET.md blocks until the ACTUAL: line moves to tracks/_chamber/<slug>/ACTUAL.md (the runner prints the path). Why: BUDGET.md's pre-verdict hash IS the ordering witness, and step 6 hard-blocked until that same file changed, so every run that reached COMPLETE necessarily mutated a witnessed artifact and verify returned TAMPERED — the chamber's promotion condition was unsatisfiable by construction, not by strictness. Two roles (immutable witness / post-verdict calibration sink) had collided in one file; each was correct alone, so neither side's code showed the conflict. Also: ko-tech-writer Step 2/4-b scans are now calibration-backed (known-pair fixtures + reproducible command, shipped) — discrimination is proven, 'zero residue' is explicitly NOT; chamber lane suite 12 -> 33 including the runner x witness seam no test covered; chamber_run.sh now teaches the two-commit discipline (gate hashes and verdict hash must land in separate commits/PRs — it previously advised the opposite). New in 2.1.0: BREAKING (gate): `crossfamily: declined` in an Axes 2-3 marker now requires grounds naming a record path that RESOLVES on disk — bare `declined`, and `declined` justified by author judgment, are blocked at commit. Remedy: cite where the operator decision lives (e.g. `.. — operator declined sidecars, per knowledge/shared/rules/operational_adaptation.md`), or use `DEGRADED_PANEL_UNUSED` if a panel was reachable and you chose not to recruit it — which is what author judgment actually is. `declined` was the only enum value with no grounds requirement; a cross-family review then broke the first (vocabulary-grep) fix three ways — self-validating on the value's own token, vacuous keyword passes, and over-blocking real declinations in natural prose — so the check asserts a resolvable record instead of words. Also: standpoint axis gains `tier1b` (a STATIC read of a target repo, executed nothing) plus a decide-in-order procedure, after blind floor-tier sims graded pure cold-reads as `tier2` three rounds running; steel-quench Wave 1's sixth angle (gate-locality) gains the output-template row it never had, so a mandatory angle stops being structurally unreportable; verify-bidirectional gains category 5 (prescriptive doctrine statement); Sister Asset Protocol gains an active-adoption trigger; new resident doctrine — Mechanization Boundary, Local Execution First, Skeleton-not-Muscle, Expedition track, and this package's versioning policy. Hub meta-operations toolkit — 35 skills + 7 agents. New in 2.0.1: harness-doctor cadence hook, portability lint wired into pre-commit, branch_claim.sh claim-count-vs-tree-count warning, louder confidentiality-scan fail-open notice, fh-gate.sh missing-package.json survival, identity ① reclassified 🟢 (cross-harness adapters + relay argument channel). New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
16
16
|
"source": "./plugins/fh-meta"
|
|
17
17
|
},
|
|
18
18
|
{
|
|
19
19
|
"name": "fh-commons",
|
|
20
|
-
"version": "2.
|
|
20
|
+
"version": "2.2.0",
|
|
21
21
|
"description": "Project-agnostic utility skills — 5 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate · ko-tech-writer) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
|
|
22
22
|
"source": "./plugins/fh-commons"
|
|
23
23
|
}
|
package/AGENTS.md
CHANGED
|
@@ -121,15 +121,22 @@ Because non-Claude runtimes do not auto-load Claude path rules, apply these rule
|
|
|
121
121
|
`ⓐ=… ⓑ=→standpoint ⓒ=… ⓓ=… ⓔ=… ⓕ=…`; markers dated earlier keep the old **ASCII four**
|
|
122
122
|
(`a b c d`) and are not retroactively blocked. 🟥 **The two arrays are not the same letters —
|
|
123
123
|
old `b` (first real use) is now `ⓔ`, old `d` (revert probe) is now `ⓕ`.** Copying an old line
|
|
124
|
-
forward silently swaps two axes and raises no error
|
|
125
|
-
|
|
124
|
+
forward silently swaps two axes and raises no error. **Which array a marker used is decided by
|
|
125
|
+
the date in its filename** (`< 2026-08-17` = old four). ⚠️ The notation is NOT the discriminator
|
|
126
|
+
— that claim stood in this file for part of 2026-08-17 and a hand-count of the corpus refuted it:
|
|
127
|
+
2 of the 4 circled-key markers on disk are dated 2026-08-10 and carry the OLD meanings. Aligning
|
|
128
|
+
the notation still helps going forward; it does not work backwards.
|
|
129
|
+
`standpoint:` remains the canonical field for ⓑ (the `axes-run` entry
|
|
126
130
|
is only a pointer to it, and a pointer at an empty field is blocked); **its value enum is still
|
|
127
131
|
validated by nothing** — that is the one remaining gap, and it is not the same thing as the axis
|
|
128
132
|
being unmechanized. Format spec: `.claude/rules/fh_4axis_gate.md §Marker axis fields`.
|
|
129
133
|
(Two drift corrections landed here on 2026-08-17: first this sentence said "four" while its own
|
|
130
134
|
next clause described the +1 — caught by the session-close ④-b CC↔Codex parity check — and then
|
|
131
135
|
the machine layer moved to six the same day.)
|
|
132
|
-
**ⓓ has no field at all** — record it in prose
|
|
136
|
+
⚠️ Until 2026-08-17 this entry point added "**ⓓ has no field at all** — record it in prose".
|
|
137
|
+
That is now **false**: `ⓓ=` is a required key like the rest. The retraction is kept visible
|
|
138
|
+
rather than deleted, because a Codex-side reader who memorised the old line would otherwise
|
|
139
|
+
keep writing markers without ⓓ and see them blocked with no idea why.
|
|
133
140
|
🟥 ⓑ **standpoint is itself split** (2026-08-17): a STATIC read of the target's own files is
|
|
134
141
|
`tier1b` and **executes nothing**; `tier2`+ asserts that something was RUN — the discriminator is
|
|
135
142
|
mechanical, *name the command you ran and the output you saw*. Measured on one delta: the static
|
|
@@ -232,13 +232,47 @@ residency 가 걸린 환경(조직 내부·규제·고객 데이터)에서는 **
|
|
|
232
232
|
|---|---|---|
|
|
233
233
|
| **Step 0 레지스터** | 독자용 아키텍처 기술문서 (문어 존댓말/명료체) | mandatory-pass |
|
|
234
234
|
| **Step 1 캘리브레이션** | FH Knowledge Core 정본 표본(`knowledge/shared/harness-core/fh_ecosystem_positioning.md`) 서식 및 톤 대조 완료 | mandatory-pass |
|
|
235
|
-
| **Step 2 문체 규율** | 번역투·조각문 5종 스캔 «잔여 0건 (양성 컨트롤 동반)» 주장 |
|
|
235
|
+
| **Step 2 문체 규율** | 번역투·조각문 5종 스캔 «잔여 0건 (양성 컨트롤 동반)» 주장 | 🟡 **부분 해소 (2026-08-17)** — 계기의 **판별력**은 이제 재현 가능하다: `bash scripts/test_ko_tech_writer_lanes.sh`(9레인, 픽스처 `plugins/fh-commons/skills/ko-tech-writer/fixtures/known_{positive,negative}.md`). 5클래스 전부 «양성 ≥1 · 음성 0» 으로 갈린다. 🟥 **그러나 이 행의 원 주장(«이 문서에서 잔여 0건»)은 여전히 미검증**이다 — 판별력과 잔여는 다른 명제이고, 후자는 문서마다 따로 재야 한다. 아래 §Step2-calibration 참조 |
|
|
236
236
|
| **Step 3 정직 수위** | 저자 내부 집계 규율 서술 제거 및 독자 의사결정 기반 정보 보존 | judged |
|
|
237
|
-
| **Step 4 수치·주장 게이트** | 수치 전수 추출 + «전칭 단정 스캔 잔여 0건» 주장 | 🟥 `UNCALIBRATED — 자기반증`: 그 «0건» 시점에 본문에 전칭 단정이 **3건 살아 있었다**(«완벽 보호» ×2, «가장 뛰어난 하네스») — 임포트 심사가 손으로 잡아 정정했다. 계기는 초록인데 대상을 안 쟀다 |
|
|
237
|
+
| **Step 4 수치·주장 게이트** | 수치 전수 추출 + «전칭 단정 스캔 잔여 0건» 주장 | 🟡 **부분 해소 (2026-08-17)** — 전칭 단정 후보 검출 2패턴(어휘형·부정형)의 판별력이 같은 스위트로 재현된다(양성 2·2건 / 음성 0·0건). 🟥 **자기반증 사실 자체는 그대로 남는다** — 아래 원 판정 유지. 그리고 «수치 전수 추출» 쪽은 **미보정**이다(이 스위트가 안 다룬다). ↓ 원 판정: 🟥 `UNCALIBRATED — 자기반증`: 그 «0건» 시점에 본문에 전칭 단정이 **3건 살아 있었다**(«완벽 보호» ×2, «가장 뛰어난 하네스») — 임포트 심사가 손으로 잡아 정정했다. 계기는 초록인데 대상을 안 쟀다 |
|
|
238
238
|
| **Step 5 지각 QA** | Mermaid 다이어그램 노드 레이블 및 Decision Tree ASCII 렌더링 시각 확인 완료 | judged |
|
|
239
239
|
| **글로벌 인프라 심사** | 논리 격리 vs 보안 샌드박싱 분리, 자원 경합 및 Context Budgeting 피드백 반영 | 🟥 `LOCAL-ONLY ATTESTATION — UNVERIFIED`: 저자 런타임(Antigravity) 측 심사이고, 짝으로 적혀 있던 `research` 서브에이전트는 **이 레포에 존재하지 않는다**(등록 에이전트 8종 중 없음). FH 안에서 재현 불가 |
|
|
240
240
|
| **적대적 공격 심사** | 직교적 3계층 모델 재정립, $N_{human}$ vs $M_{subagent}$ 인지 분리, 샌드박싱 조건 개고 | 🟥 `LOCAL-ONLY ATTESTATION — UNVERIFIED`: `challenger` 는 실재하는 FH 에이전트지만(`plugins/fh-meta/agents/challenger.md`), 이 행이 가리키는 실행의 마커·로그가 없다. **이름의 실재는 실행의 증거가 아니다** |
|
|
241
241
|
|
|
242
|
+
<a name="step2-calibration"></a>
|
|
243
|
+
### §Step2-calibration — 2026-08-17, 무엇이 해소됐고 무엇이 안 됐나
|
|
244
|
+
|
|
245
|
+
**기원**: 챔버 런 #12(`prosody-lens`, KILL)가 이 두 행을 **배출 판정의 결정적 근거**로 인용했다 —
|
|
246
|
+
*"같은 계열 계기가 미보정인데 하나 더 짓는 것은 재발명이자 미보정 계기의 증식이다."*
|
|
247
|
+
운영자 결정(2026-08-17): **새 계기보다 이 부채가 먼저.**
|
|
248
|
+
|
|
249
|
+
```
|
|
250
|
+
재현 커맨드 bash scripts/test_ko_tech_writer_lanes.sh
|
|
251
|
+
픽스처 plugins/fh-commons/skills/ko-tech-writer/fixtures/known_positive.md
|
|
252
|
+
plugins/fh-commons/skills/ko-tech-writer/fixtures/known_negative.md
|
|
253
|
+
결과 9 레인 전건 통과 — 5클래스 + 전칭 2패턴 + META 컨트롤 2
|
|
254
|
+
엔진 ripgrep 15.1.0 (SKILL.md 가 rg 로 고정한 그 엔진)
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
**해소된 것**: 「컨트롤이 무엇이었는지·재현 커맨드·출력이 하나도 없다」 — 셋 다 생겼다.
|
|
258
|
+
각 클래스가 **양성 ≥1 · 음성 0** 으로 갈린다(존재 확인이 아니라 판별 확인).
|
|
259
|
+
|
|
260
|
+
🟥 **해소되지 않은 것 — 축소하지 않는다**
|
|
261
|
+
1. **원 주장(«이 문서에서 잔여 0건»)은 여전히 미검증.** 판별력과 잔여는 다른 명제다.
|
|
262
|
+
이 스위트를 근거로 «잔여 0» 을 주장하면 강등 사유가 그대로 재발한다 — 스위트 자신이
|
|
263
|
+
출력 말미에 그렇게 인쇄한다.
|
|
264
|
+
2. **Step 4 의 자기반증 사실은 그대로 남는다**(«0건» 시점에 전칭 단정 3건 생존).
|
|
265
|
+
3. **«수치 전수 추출»은 이 스위트가 안 다룬다** — Step 4 의 절반은 여전히 미보정.
|
|
266
|
+
4. 🟥 **부수 발견 — 정본의 분류가 과장이다.** SKILL.md 는 *"앞 다섯 줄은 기계 검출 가능"*
|
|
267
|
+
이라 적었는데, **실제로 grep 을 싣고 있는 것은 C1(줄표)·C5(소유 직역) 둘뿐**이고
|
|
268
|
+
C2·C3·C4 는 산문 힌트다. 특히 **C4 조각문은 패턴이 없다**(«서술어 없는 마침»).
|
|
269
|
+
첫 캘리브레이션이 그걸 드러냈다 — 내가 임의로 지은 C4 패턴이 **양성을 0건으로 놓쳤다.**
|
|
270
|
+
지금 실린 C4 는 **닫힌 명사 어휘 목록**이라 recall 이 낮고, 넓히려면 어휘 추가가 아니라
|
|
271
|
+
형태소 분석이 필요하다. **낮다는 사실을 스크립트 주석에 적었다.**
|
|
272
|
+
5. 이 스위트는 `rg` 부재 시 **SKIP 이 아니라 rc=10** 으로 끝난다(미측정을 통과로 렌더 금지).
|
|
273
|
+
⚠️ 실측 계기: 이 개발 머신의 `grep` 은 **ugrep 7.5.0**(GNU 아님)이라 한글 word 경계가
|
|
274
|
+
갈릴 수 있다 — SKILL.md 가 엔진을 `rg` 로 못 박은 이유가 여기서 실증된다.
|
|
275
|
+
|
|
242
276
|
> 🟥 **이 부록 전체의 지위**: 저자 런타임의 **자기신고**이며, 위 «잔여 0건»·«PASS» 는 아티팩트로
|
|
243
277
|
> 뒷받침되지 않는다. FH 자기 규율상 이것은 증거가 아니라 저자의 주장이다
|
|
244
278
|
> (`fh_4axis_gate.md §Reviewer-visible evidence` 의 degrade 라벨을 그대로 적용).
|
|
@@ -515,6 +515,42 @@ thing being counted"* 라고 경고한 그 형태가, **그 경고를 적은 채
|
|
|
515
515
|
잴 것은 파일이 아니라 **지시**였다. 되돌림 실측: `git show HEAD:` 판으로 되돌리면 L13-b 가
|
|
516
516
|
적색(`step6 이 증인 아티팩트를 고치라고 시킨다`), 컨트롤은 초록 유지.
|
|
517
517
|
|
|
518
|
+
### 🟥 §P1-2026-08-17-b — 그리고 그 수리로도 아직 부족했다. **커밋을 합치는 모든 경로가 증인을 죽인다**
|
|
519
|
+
|
|
520
|
+
런 #12(`prosody-lens`, KILL)가 수리 후 첫 런이었고, TAMPERED 는 사라졌으나 **`UNORDERED`** 가 나왔다.
|
|
521
|
+
판정 코드가 이유를 명시한다 — `chamber_witness.sh do_verify`:
|
|
522
|
+
|
|
523
|
+
```bash
|
|
524
|
+
# 주석 원문: "같은 커밋(또는 같은 초)에 들어온 verdict 는 «먼저» 를 증명하지 못한다"
|
|
525
|
+
if [ "$verdict_ts" -le "$pre_max_ts" ]; then # ← -le. 같은 «초» 도 실패
|
|
526
|
+
echo "UNORDERED — verdict 가 pre-verdict 아티팩트보다 먼저이거나 같은 시점에 커밋됐다."
|
|
527
|
+
return 1
|
|
528
|
+
fi
|
|
529
|
+
```
|
|
530
|
+
이건 의도된 설계다(초판이 `-lt` 라 same-second 역순을 통과시켰고 cross-family 가 잡았다).
|
|
531
|
+
문제는 **그 조건을 만드는 경로가 이 레포의 표준 절차 안에 셋이나 있다**는 것이다:
|
|
532
|
+
|
|
533
|
+
| 경로 | 실측 |
|
|
534
|
+
|---|---|
|
|
535
|
+
| **P-a `--squash` 머지** | 런 #11 은 브랜치에서 게이트/verdict 를 **따로** 커밋해 순서가 성립했다(`e86f796`→`f79cc05`→`852064b`). PR #414 를 squash 하자 **main 에서 원장을 건드린 커밋 = 1개** — 게이트 4 + verdict 가 전부 그 하나에 접혔다 |
|
|
536
|
+
| **P-b 저자가 한 커밋에 배치** | 런 #12 에서 4해시를 한 번에 커밋(`50ef761`) → 브랜치에서부터 UNORDERED. 실측 `verdict_commit=1786935355 latest_pre=1786935355` |
|
|
537
|
+
| **P-c `--delete-branch`** | 순서를 담은 커밋이 도달 불가가 된다(런 #12 의 첫 증인 커밋이 그렇게 됐다) |
|
|
538
|
+
|
|
539
|
+
🟥 **정정**: 위 §P1-2026-08-17 을 *"P1 이 만족 가능해졌다"* 로 읽지 마라. 정확히는
|
|
540
|
+
**피처 브랜치에서 만족 가능하고, 이 레포가 의무화한 squash 머지가 main 에서 그것을 파괴한다.**
|
|
541
|
+
|
|
542
|
+
**채택된 처방 ⓐ (운영자 결정 2026-08-17) — 정책 변경 0**:
|
|
543
|
+
> **게이트 해시와 verdict 해시를 두 PR 로 분리한다.** 각각 squash 해도 main 에 **2커밋**이 남아
|
|
544
|
+
> 순서가 보존된다. 저자 규율이 따라온다 — **게이트 커밋과 verdict 커밋을 절대 합치지 않는다**(P-b 차단).
|
|
545
|
+
|
|
546
|
+
**배선**(산문으로 안 남긴다): `chamber_run.sh` 가 **행위자가 읽는 자리에서** 이 규율을 인쇄한다 —
|
|
547
|
+
게이트 기록 시 «verdict 보다 먼저, 별도 커밋으로», verdict 기록 시 «게이트 해시들과 **다른 커밋**»,
|
|
548
|
+
그리고 UNORDERED 일 때 세 경로를 이름으로 열거. 앵커는 `test_chamber_run_lanes.sh` **L15/L15b/L15c**
|
|
549
|
+
(33레인) — 🟥 **그 레인이 생긴 이유는 러너가 틀린 처방을 가르치고 있었기 때문이다**:
|
|
550
|
+
*"원장 해시를 커밋한 뒤 재실행하면 증인이 성립한다"* — 한 커밋에 넣으면 성립하지 않는다.
|
|
551
|
+
**게이트 문구는 «막는가»만이 아니라 «옳은 것을 가르치는가»도 재야 한다**
|
|
552
|
+
(`[[feedback_gate_prescription_is_unverified]]`, 2026-08-17 3번째 사례).
|
|
553
|
+
|
|
518
554
|
⚠️ **런 #11 자신은 TAMPERED 로 남는다. 소급 수리하지 않는다** — 과거 9런을 back-fill 하지 않는 것과
|
|
519
555
|
같은 이유다(사후에 쓴 기록은 증명하지 않는다). **P1 은 이 수리 이후의 EMIT 런에서 처음 만족 가능하다.**
|
|
520
556
|
🟥 그리고 이 수리는 **P1 을 만족시키지 않는다** — 만족을 *가능하게* 만들 뿐이다. ②는 🔵 유지.
|
|
@@ -2146,3 +2146,39 @@
|
|
|
2146
2146
|
outcome: accepted
|
|
2147
2147
|
evidence: "8건 지적(S2·A3·B3). 그중 2건(혼용 오탐 · 화살표 fail-open)은 내가 투입 전 자력 수리했고 codex 가 독립으로 같은 지점을 지목해 수렴 확인 — 그 둘은 「둘 다 봤다」. **나머지 5건 자력 적발 0**: grep 이 주석 처리된 호출에 매칭(S) · if:false 로 꺼진 잡(S) · 트리거 미검사(A) · 잡 신원 미검사(A) · 마커 파일명 날짜 추출 fail-open(B, 선재). 각 지적에 깨뜨리는 입력을 구체적으로 제시해 손 재현이 즉시 가능했다"
|
|
2148
2148
|
cost: 85k tokens
|
|
2149
|
+
- date: 2026-08-17
|
|
2150
|
+
agent: fh-meta:beginner (isolated) — 챔버 런 #12 step-4 블라인드 1/3 (낭독 렌즈)
|
|
2151
|
+
purpose: "「소리 내어 읽히는가」 계기 후보의 냉담 첫 접촉 — 정의가 서는가, 첫 실패 지점, ko-tech-writer 와의 겹침"
|
|
2152
|
+
outcome: accepted
|
|
2153
|
+
evidence: "tool_uses 3(실독). HARD 4/SOFT 2 · **첫 산출물 도달 실패**. 최대 소득 = *「소리 내어 읽힌다」의 예시 셋이 서로 다른 층이라 정의를 역추정하는 것조차 안 된다* — 번역투=어법층(소리 무관) · 호흡=글자수 · 운율만 소리. 🟥 **그 결함은 후보가 아니라 내 한 줄 요약이 갖고 있었다**(정본은 A~F 로 층을 갈라놨다). ko-tech-writer 대조로 Step 2(번역투 7클래스)·Step 2-b(리듬 수치 대조)를 줄번호로 짚음"
|
|
2154
|
+
cost: 112k tokens
|
|
2155
|
+
- date: 2026-08-17
|
|
2156
|
+
agent: fh-meta:main-player (isolated) — 챔버 런 #12 step-4 블라인드 2/3
|
|
2157
|
+
purpose: "실사용자 일상 가치 — 매일 쓸 층이 있는가, 별도 계기인가 한 스텝인가"
|
|
2158
|
+
outcome: accepted
|
|
2159
|
+
evidence: "tool_uses 4. 판정 **NO** — 장면 3/3 대체, 증명된 사용자 **n=1·월 3회**. 🟥 결정적 한 줄: *「후보의 known-positive 가 이미 남의 집에 있다」*(ko-tech-writer:84-90 의 124자 문장·어간 4연속 실측) — **픽스처가 남의 집에 있는 계기는 별도 계기가 아니다.** 그리고 한국어 기술문서 인구를 *「못 잡는다 — UNMEASURED 이지 크다도 작다도 아니다」* 로 **스스로 자백**해 판정 신뢰도를 올렸다. 진짜 구멍 지목: 낭독 대본 레지스터의 Step 5 N/A"
|
|
2160
|
+
cost: 106k tokens
|
|
2161
|
+
- date: 2026-08-17
|
|
2162
|
+
agent: fh-meta:challenger (isolated) — 챔버 런 #12 step-4 블라인드 3/3
|
|
2163
|
+
purpose: "배출 후보 적대 심사 — 술어 성립성·재발명·정답라벨·보이지 않는 것"
|
|
2164
|
+
outcome: partial
|
|
2165
|
+
evidence: "tool_uses 12, S 4건. **VALID**: S3(기계화 가능분은 이미 있는 스텝) · S4(표면지표 환원 시 **양방향 오탐**, known-negative 3건을 손으로 제작해 시연) · A3(정답 라벨 주체 부재, SKILL.md:235 «생성자=평가자 금지»). 🟥 **S2 는 거버너가 반증**: *「산출물 회수 불가, rg 낭독 tracks/ → no matches」* 의 repro 가 **재현 안 된다**(실측 12파일, 정본 신호 272줄 실재). S1(«양식 미도달»)은 과장 — 음절수·동음이의는 텍스트에서 계산되는 소리 속성. ⇒ **사이드카 발견은 소스로 닫히기 전엔 판정이 아니다**(오늘 2번째). 본인이 *「나도 스캐너를 안 돌렸고 나 역시 소리를 못 낸다」* 로 등급을 스스로 깎은 것은 신뢰도를 올렸다"
|
|
2166
|
+
cost: 131k tokens
|
|
2167
|
+
- date: 2026-08-17
|
|
2168
|
+
agent: general-purpose (isolated) — 챔버 런 #12 조건 1(net-new) measured 스캔
|
|
2169
|
+
purpose: "외부 생태계(한국어권 1순위) + 로컬 전수 스캔으로 재발명 여부 측정. 런 #11 의 자백된 사각(영어 전량)을 INTENT 가 필수 조건으로 박았다"
|
|
2170
|
+
outcome: accepted
|
|
2171
|
+
evidence: "tool_uses 27(WebSearch 12 중 **한국어 7** · GitHub API 15 · WebFetch 4 · 로컬 정독). 이 런의 KILL 을 가른 계기이고 **거버너의 프레이밍 오류 영향을 가장 덜 받았다**. ⓑ번역투 **포화** — 2026-06~08 3개월간 한국어 번역투 제거 스킬 **10+ 신설**(humanizer-ko 35패턴 등) · ⓒ처방형식은 이 분야 **기본값** · ⓐ운율만 외부 공백(GitHub 전체 `낭독` SKILL.md **1건=우리 미러**). 🟥 **결정적 근거를 찾았다** — harness_terminal_correlation:235-236 이 ko-tech-writer Step 2·4 를 이미 `UNCALIBRATED` 로 강등(Step 4 는 자기반증까지). 미탐 자백 7종을 스스로 열거(GitHub API 403 4쿼리 미실행 · 국립국어원 규칙목록 미개봉 · description 기준 판정 등)"
|
|
2172
|
+
cost: 152k tokens
|
|
2173
|
+
- date: 2026-08-17
|
|
2174
|
+
agent: codex/gpt-5.5 (cross-family sidecar) — qasp-dev #174 load-bearing 게이트 리뷰
|
|
2175
|
+
purpose: "게이트 exit enum·면제 경로 변경의 degrade-direction / bypass / 테스트 판별력 적대검증"
|
|
2176
|
+
outcome: partial
|
|
2177
|
+
evidence: "5건 지목, 전부 diff 인용 동반. **실행으로 갈랐다 — 2 반증 · 3 확인.** 🟥 반증 둘(깊은 `src/**/*.py`·중첩 `scripts/ci/*.sh` 가 샌다)의 근거 오류가 같다: **bash `case` 의 `*` 는 `/` 를 먹는다**(경로 글롭 semantics 로 읽음) — 실측 rc=1 로 둘 다 정상 차단. 확인 셋 중 **이 PR 이 새로 만든 것은 S1-c 하나**(sandbox 제외가 침묵, 종전 rc=4→rc=0 무흔적)이고 그것만 수리. 나머지 둘은 기존 결함·재발 0건이라 미착수. 테스트 판별력에서 되돌림 에이전트와 **갈렸고 코덱스가 옳았다**(파일명 substring 은 그 이름이 위반 목록에 실려도 참 → 보고 형태로 조임). ⇒ 사이드카 발견은 소스로 닫히기 전엔 판정이 아니다 — 이번엔 40% 가 틀렸다"
|
|
2178
|
+
cost: 68k tokens
|
|
2179
|
+
- date: 2026-08-17
|
|
2180
|
+
agent: general-purpose (isolated, sonnet) — qasp-dev #174 되돌림 프로브 ⓕ축
|
|
2181
|
+
purpose: "PR 이 추가한 테스트가 장식인지 실측 — 수리 4건 개별 되돌림 3단(적용확인→실행→복원) + substring 8행 판별력"
|
|
2182
|
+
outcome: accepted
|
|
2183
|
+
evidence: "tool_uses 25(계기 생존 — 카드가 경고한 `tool_uses: 0` 아님). **4/4 앵커 생존**, 매번 대응 레인만 적색(S1→2건·S2→1건·S3→2건·B1→known-pair 2건), 무관 25~28개는 통과 → 과결합 0. 복원 후 29 passed 재확인, 트리 clean. 🟥 substring 8행은 **전부 ✅ 로 판정했는데 코덱스와 갈렸고 이쪽이 졌다** — 「뚫리는 입력을 제시 못 함」을 판별력 있음으로 읽었으나, 그 단언이 *주장*하는 것(«범위 밖 보고가 났다»)은 파일명만으로 증명되지 않는다. **없음을 증명 못 함 ≠ 판별력 있음**"
|
|
2184
|
+
cost: 135k tokens
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@chrono-meta/fh-gate",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.2.0",
|
|
4
4
|
"description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"keywords": [
|
|
@@ -72,6 +72,10 @@
|
|
|
72
72
|
"scripts/test_lane_runner_lanes.sh",
|
|
73
73
|
"scripts/test_version_lockstep_lanes.sh",
|
|
74
74
|
"scripts/package_coverage_check.sh",
|
|
75
|
+
"scripts/test_ko_tech_writer_lanes.sh",
|
|
76
|
+
"scripts/ko_tech_writer_calibrate.py",
|
|
77
|
+
"plugins/fh-commons/skills/ko-tech-writer/fixtures/known_positive.md",
|
|
78
|
+
"plugins/fh-commons/skills/ko-tech-writer/fixtures/known_negative.md",
|
|
75
79
|
"scripts/publish_freshness_check.sh",
|
|
76
80
|
"scripts/prepublish_scope_note.sh",
|
|
77
81
|
"scripts/lane_runner_check.sh",
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
<!-- KNOWN-NEGATIVE 픽스처 — Step 2/4-b 스캔이 «잡으면 안 되는» 표본.
|
|
2
|
+
같은 주제·비슷한 길이의 정상 산문. 히트가 나오면 그 스캔은 과차단이다. -->
|
|
3
|
+
|
|
4
|
+
# 표본 — 같은 내용을 정상 문체로
|
|
5
|
+
|
|
6
|
+
파이프라인은 세 단계로 나뉘고, 각 단계가 서로 다른 결함을 잡습니다.
|
|
7
|
+
|
|
8
|
+
- **격리**: 에이전트를 별도 컨텍스트에서 돌리는 방식입니다.
|
|
9
|
+
|
|
10
|
+
관측된 분업은 검출을 기계가 맡고 판정을 사람이 맡는 형태입니다.
|
|
11
|
+
|
|
12
|
+
여기서 한 층이 걸립니다.
|
|
13
|
+
|
|
14
|
+
이 스캐너에는 이전 상태가 없습니다.
|
|
15
|
+
그 판정에는 재현 경로가 없습니다.
|
|
16
|
+
|
|
17
|
+
이 방식에서는 오탐이 세 건 나왔고, 그 셋을 손으로 확인했습니다.
|
|
18
|
+
후보 목록에서 두 건이 빠졌으며, 남은 항목은 개별로 판정했습니다.
|
|
19
|
+
그 경로는 두 차례 관측됐습니다.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
<!-- KNOWN-POSITIVE 픽스처 — Step 2/4-b 스캔이 «잡아야 하는» 표본.
|
|
2
|
+
🟥 이 파일은 일부러 번역투를 담고 있다. 문체 교정 대상이 아니다.
|
|
3
|
+
각 줄 끝 주석이 어느 클래스를 심었는지 밝힌다. 클래스당 최소 1건. -->
|
|
4
|
+
|
|
5
|
+
# 표본 — 일부러 심은 번역투
|
|
6
|
+
|
|
7
|
+
파이프라인은 세 단계로 나뉩니다 — 그리고 각 단계가 서로 다른 결함을 잡습니다. <!-- C1 줄표 이어붙임 -->
|
|
8
|
+
|
|
9
|
+
- **격리** — 에이전트를 별도 컨텍스트에서 돌리는 방식입니다. <!-- C2 용어-머리 -->
|
|
10
|
+
|
|
11
|
+
관측된 분업은: 검출은 기계가, 판정은 사람이. <!-- C3 콜론 나열투 -->
|
|
12
|
+
|
|
13
|
+
여기서 걸리는 층. <!-- C4 조각문 -->
|
|
14
|
+
|
|
15
|
+
이 스캐너는 이전 상태를 갖지 않습니다. <!-- C5 소유 직역 -->
|
|
16
|
+
그 판정은 재현 경로를 가지고 있지 않습니다. <!-- C5 소유 직역(변형) -->
|
|
17
|
+
|
|
18
|
+
이 방식으로는 오탐이 전혀 나오지 않았습니다. <!-- C6 전칭 단정(Step 4-b) -->
|
|
19
|
+
후보 목록에서 하나도 빠지지 않았고, 예외 없이 모두 통과했습니다. <!-- C6 전칭 단정 -->
|
|
20
|
+
그 경로는 한 번도 잡히지 않았다. <!-- C6 부정형 전칭 -->
|
|
@@ -10,6 +10,51 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/)
|
|
|
10
10
|
|
|
11
11
|
## Plugin Level
|
|
12
12
|
|
|
13
|
+
### [2.2.0] — 2026-08-17
|
|
14
|
+
|
|
15
|
+
🟥 **BREAKING (gate)**: chamber step 6 now reads `ACTUAL.md`, **not `BUDGET.md`**. An in-flight
|
|
16
|
+
chamber run whose actual cost was written into `BUDGET.md` will **block at step 6** until the value
|
|
17
|
+
moves to a new `ACTUAL.md` in the same workspace.
|
|
18
|
+
- **Remedy**: the runner prints the exact path when it blocks — move the `ACTUAL:` line to
|
|
19
|
+
`tracks/_chamber/<slug>/ACTUAL.md`. `BUDGET.md` keeps `ESTIMATE:` only.
|
|
20
|
+
- **Why**: `BUDGET.md`'s pre-verdict hash IS the ordering witness (`ship_readiness_gate §② P1`), and
|
|
21
|
+
step 6 was hard-blocking until that same file changed. So **every run that reached COMPLETE
|
|
22
|
+
necessarily mutated a witnessed artifact** and `verify` returned `TAMPERED` — identity ②'s only
|
|
23
|
+
promotion condition was unsatisfiable by construction, not by strictness. Measured on chamber run
|
|
24
|
+
#11, the first run ever taken through step 7. Two roles (immutable witness / post-verdict
|
|
25
|
+
calibration sink) had collided in one file; each was correct alone, so neither side's code showed
|
|
26
|
+
the conflict.
|
|
27
|
+
|
|
28
|
+
**Added**
|
|
29
|
+
- `scripts/ko_tech_writer_calibrate.py` + `scripts/test_ko_tech_writer_lanes.sh` + two known-pair
|
|
30
|
+
fixtures — the discrimination of `ko-tech-writer` Step 2 (five translationese classes) and
|
|
31
|
+
Step 4-b (universal-claim candidates) is now **reproducible**: positive ≥1 / negative 0 per class,
|
|
32
|
+
plus two META controls. Wired into `selfcheck.sh` and shipped in `files[]`.
|
|
33
|
+
🟥 **What this does NOT prove**: "zero residue" in any real document. Discrimination and residue
|
|
34
|
+
are different propositions; the suite prints that warning itself.
|
|
35
|
+
- Chamber lane suite **12 → 33 lanes**, including the **runner × witness seam** that no test covered
|
|
36
|
+
(the runner's lanes excluded the witness by design; the witness self-test ran it standalone).
|
|
37
|
+
|
|
38
|
+
**Changed**
|
|
39
|
+
- `chamber_run.sh` now teaches the **two-commit discipline** where the actor reads it: gate hashes
|
|
40
|
+
and the verdict hash must land in **separate commits** (and, for a squash-merge repo, **separate
|
|
41
|
+
PRs**). Committing them together yields `UNORDERED` — the runner previously advised the opposite.
|
|
42
|
+
- `chamber_witness.sh do_record` skips a byte-identical re-record of the same
|
|
43
|
+
`(run, artifact, sha)` triple. A **changed** artifact still appends — that is the tamper evidence.
|
|
44
|
+
|
|
45
|
+
**Fixed**
|
|
46
|
+
- `harness_terminal_correlation_and_recommendations.md` Appendix: Step 2 / Step 4 rows move from
|
|
47
|
+
🟥 `UNCALIBRATED` to 🟡 **partially resolved**, with the reproduction command recorded. The
|
|
48
|
+
original claims ("zero residue" in that document; Step 4's numeric extraction half) remain
|
|
49
|
+
**unverified and are labelled as such**.
|
|
50
|
+
|
|
51
|
+
**Known residual (not closed)**
|
|
52
|
+
- Identity ② stays **RC**. No run has yet been `WITNESSED` on `main`; the two-PR prescription is
|
|
53
|
+
reasoned from the verify logic, and its efficacy is provable only by the next EMIT run.
|
|
54
|
+
- `--delete-branch` still orphans ordering evidence — warned in prose, not blocked.
|
|
55
|
+
- Of Step 2's five "machine-detectable" classes, only two ship an actual grep in `SKILL.md`; the
|
|
56
|
+
calibration surfaced that the doc over-claims. C4's pattern is a closed noun list with **low recall**.
|
|
57
|
+
|
|
13
58
|
### [2.1.0] — 2026-08-17
|
|
14
59
|
|
|
15
60
|
🟥 **BREAKING (gate)**: `crossfamily: declined` in an Axes 2-3 marker now requires grounds naming a
|
package/scripts/chamber_run.sh
CHANGED
|
@@ -56,7 +56,15 @@ _witness_record() { # $1 = artifact path
|
|
|
56
56
|
return 0
|
|
57
57
|
fi
|
|
58
58
|
if bash "$_WITNESS" record "$SLUG" "$1" >/dev/null 2>&1; then
|
|
59
|
-
|
|
59
|
+
# 🟥 «커밋해야 된다» 로만 적으면 **틀린 처방을 가르친다**(2026-08-17 런 #12 실측:
|
|
60
|
+
# 네 해시를 한 커밋에 배치했더니 브랜치에서부터 UNORDERED). verify 는
|
|
61
|
+
# `[ verdict_ts -le pre_max_ts ]` 로 판정하므로 **같은 커밋 = 같은 초 = 실패**다.
|
|
62
|
+
if [ "$(basename "$1")" = "EMISSION_VERDICT.md" ]; then
|
|
63
|
+
echo " ↳ 순서 증인: EMISSION_VERDICT.md 해시 기록됨"
|
|
64
|
+
echo " 🟥 이 해시는 **게이트 해시들과 다른 커밋**으로 올려야 한다. 같이 커밋하면 UNORDERED 다."
|
|
65
|
+
else
|
|
66
|
+
echo " ↳ 순서 증인: $(basename "$1") 해시 기록됨 — **verdict 보다 먼저, 별도 커밋으로** 올려라"
|
|
67
|
+
fi
|
|
60
68
|
else
|
|
61
69
|
echo " ⚠️ 순서 증인 기록 실패: $(basename "$1") (증인 없음 — 미측정이지 0이 아니다)"
|
|
62
70
|
fi
|
|
@@ -197,7 +205,12 @@ if [ -f "$_WITNESS" ]; then
|
|
|
197
205
|
echo " ✅ 이 EMIT 은 순서 증인을 가진다 — identity ② 승격 근거로 사용 가능"
|
|
198
206
|
else
|
|
199
207
|
echo " 🟥 이 EMIT 에는 순서 증인이 없다(rc=${_wrc:-?}) — identity ② 승격 근거로 쓰지 마라."
|
|
200
|
-
echo "
|
|
208
|
+
echo " 처방 — 커밋을 **합치지 마라**. 세 경로가 각각 증인을 죽인다(2026-08-17 실측):"
|
|
209
|
+
echo " ① 한 커밋에 배치 → 같은 초라 UNORDERED (런 #12 에서 재현)"
|
|
210
|
+
echo " ② --squash 머지 → 게이트+verdict 가 main 에서 한 커밋으로 접힌다 (런 #11)"
|
|
211
|
+
echo " ③ --delete-branch → 순서를 담은 커밋이 도달 불가가 된다"
|
|
212
|
+
echo " 올바른 형태: 게이트 해시 커밋 → (별도) verdict 해시 커밋, **두 PR 로 분리해 각각 머지**."
|
|
213
|
+
echo " 그러면 squash 해도 main 에 2커밋이 남아 순서가 보존된다."
|
|
201
214
|
# ★ 출력만 하고 exit 0 으로 끝내면, stdout 을 안 읽는 호출자에게는 증인 없는 EMIT 도
|
|
202
215
|
# **성공**이다(cross-family F8). EMIT 은 승격 주장이므로 기계적으로 구분되어야 한다.
|
|
203
216
|
WITNESS_EMIT_FAIL=1
|
|
@@ -0,0 +1,127 @@
|
|
|
1
|
+
#!/usr/bin/env python3
|
|
2
|
+
# -*- coding: utf-8 -*-
|
|
3
|
+
"""ko_tech_writer_calibrate.py — `ko-tech-writer` Step 2 / Step 4-b 스캔의 **판별력** known-pair.
|
|
4
|
+
|
|
5
|
+
왜 파이썬인가 (엔진 선택은 취향이 아니라 실측 결과다)
|
|
6
|
+
─────────────────────────────────────────────────────────
|
|
7
|
+
`SKILL.md` Step 4 가 허용 엔진을 명시한다: *"유니코드 인식 엔진(ripgrep · GNU grep UTF-8
|
|
8
|
+
로케일 · Python `re` — **`grep -P` 는 예외**)"*. 초판은 `rg` 로 고정했고 **CI 에서 죽었다** —
|
|
9
|
+
ubuntu-latest 에 ripgrep 이 없다(2026-08-17 실측, `HARNESS ERROR ... exited 10`).
|
|
10
|
+
같은 실행에서 이식성 린트도 걸렸다: 셸 파일 안의 `[가-힣]` 문자 범위는
|
|
11
|
+
**GNU grep C 로케일에서 exit 2 로 죽고 그 결과가 «무매치» 로 읽힌다** — 즉 fail-open.
|
|
12
|
+
|
|
13
|
+
Python `re` 는 ⓐ 세 허용 엔진 중 **어디에나 있는 유일한 것**이고 ⓑ 유니코드를 기본으로
|
|
14
|
+
인식하며 ⓒ 패턴이 셸 밖으로 나가 로케일 문제를 구조적으로 없앤다. **격하가 아니라
|
|
15
|
+
정본이 이미 허용한 등가 엔진으로의 이동**이다.
|
|
16
|
+
|
|
17
|
+
무엇을 주고 무엇을 안 주는가 — 섞지 마라
|
|
18
|
+
─────────────────────────────────────────────────────────
|
|
19
|
+
준다 각 스캔이 **양성을 잡고 음성을 안 잡는가**(=판별력). 재현 커맨드와 실행 출력.
|
|
20
|
+
안 준다 그 스캔이 **임의 문서에서 잔여 0건인가**. 그건 문서마다 따로 재는 것이다.
|
|
21
|
+
⇒ 이걸 통과했다고 «Step 2 잔여 0» 을 주장하면
|
|
22
|
+
`harness_terminal_correlation_and_recommendations.md` 의 강등 사유가 그대로 재발한다.
|
|
23
|
+
|
|
24
|
+
exit: 0 전 레인 판별 · 1 판별 실패 · 10 harness error(픽스처 부재)
|
|
25
|
+
"""
|
|
26
|
+
import os
|
|
27
|
+
import re
|
|
28
|
+
import sys
|
|
29
|
+
|
|
30
|
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
|
31
|
+
FIX = os.path.join(HERE, "..", "plugins", "fh-commons", "skills", "ko-tech-writer", "fixtures")
|
|
32
|
+
POS = os.path.join(FIX, "known_positive.md")
|
|
33
|
+
NEG = os.path.join(FIX, "known_negative.md")
|
|
34
|
+
|
|
35
|
+
# (라벨, 패턴) — Step 2 의 «기계 검출 가능» 5클래스 + Step 4-b 의 2패턴.
|
|
36
|
+
#
|
|
37
|
+
# 🟥 정본의 분류가 과장이라는 것이 첫 캘리브레이션에서 드러났다. SKILL.md 는
|
|
38
|
+
# "앞 다섯 줄은 기계 검출 가능" 이라 적었는데, **실제로 grep 을 싣고 있는 것은
|
|
39
|
+
# C1(줄표)·C5(소유 직역) 둘뿐**이고 C2·C3·C4 의 «검출 힌트» 칸은 산문 서술이다.
|
|
40
|
+
# 특히 C4 는 «서술어 없는 마침» 이라는 서술뿐 패턴이 없다.
|
|
41
|
+
LANES = [
|
|
42
|
+
("C1 줄표 이어붙임", r"(다|것|음) — "),
|
|
43
|
+
("C2 용어-머리", r"^\s*[-*]\s+\*\*[^*]+\*\* — "),
|
|
44
|
+
("C3 콜론 나열투", r"[가-힣]은: "),
|
|
45
|
+
# C4: SKILL.md 에 패턴이 없다. 아래는 **후보 검출**이지 판정이 아니다
|
|
46
|
+
# (SKILL.md 자신의 규율: "grep은 후보를 표시할 뿐, 남길지 판정은 문맥으로").
|
|
47
|
+
# 닫힌 명사 어휘라 **recall 이 낮다** — 낮은 쪽이 과차단보다 안전한 방향이고,
|
|
48
|
+
# 넓히려면 어휘를 늘리는 게 아니라 형태소 분석이 필요하다.
|
|
49
|
+
("C4 조각문(후보·닫힌 어휘)", r"(층|것|점|축|건|뿐|바|셈)\.\s*(<!--|$)"),
|
|
50
|
+
("C5 소유 직역", r"(을|를) (갖|가지)"),
|
|
51
|
+
("C6 전칭 어휘", r"전부|모두|하나도|전혀|일절|예외 없이"),
|
|
52
|
+
("C6b 부정형 전칭", r"(안|못) ?(했|만들|나오|잡히)|지 않았(다|습니다)"),
|
|
53
|
+
]
|
|
54
|
+
|
|
55
|
+
# 절대 잡히면 안 되는 토큰 — 「늘 잡는다」를 내는 죽은 계기 배제용
|
|
56
|
+
NEVER = r"ZZZ_absent_token_ZZZ"
|
|
57
|
+
|
|
58
|
+
|
|
59
|
+
def count(path, pattern):
|
|
60
|
+
"""패턴에 매칭되는 **줄 수**. 파일 부재는 예외로 올린다(무음 0 금지)."""
|
|
61
|
+
rx = re.compile(pattern, re.MULTILINE)
|
|
62
|
+
n = 0
|
|
63
|
+
with open(path, encoding="utf-8") as fh:
|
|
64
|
+
for line in fh:
|
|
65
|
+
if rx.search(line.rstrip("\n")):
|
|
66
|
+
n += 1
|
|
67
|
+
return n
|
|
68
|
+
|
|
69
|
+
|
|
70
|
+
def main():
|
|
71
|
+
for p in (POS, NEG):
|
|
72
|
+
if not os.path.isfile(p):
|
|
73
|
+
sys.stderr.write("harness error: fixture missing %s\n" % p)
|
|
74
|
+
return 10
|
|
75
|
+
|
|
76
|
+
print("engine: python%d.%d re (unicode-aware)"
|
|
77
|
+
% (sys.version_info[0], sys.version_info[1]))
|
|
78
|
+
print("known-positive: %s" % os.path.relpath(POS, os.path.join(HERE, "..")))
|
|
79
|
+
print("known-negative: %s" % os.path.relpath(NEG, os.path.join(HERE, "..")))
|
|
80
|
+
print("")
|
|
81
|
+
|
|
82
|
+
npass = nfail = 0
|
|
83
|
+
for label, pat in LANES:
|
|
84
|
+
p = count(POS, pat)
|
|
85
|
+
n = count(NEG, pat)
|
|
86
|
+
if p >= 1 and n == 0:
|
|
87
|
+
print("OK %s -- positive %d / negative 0 (discriminates)" % (label, p))
|
|
88
|
+
npass += 1
|
|
89
|
+
elif p < 1:
|
|
90
|
+
print("FAIL %s -- MISSES the positive (positive %d). A blind instrument." % (label, p))
|
|
91
|
+
nfail += 1
|
|
92
|
+
else:
|
|
93
|
+
print("FAIL %s -- HITS the negative (negative %d). Over-blocks; no discrimination."
|
|
94
|
+
% (label, n))
|
|
95
|
+
nfail += 1
|
|
96
|
+
|
|
97
|
+
# META 컨트롤 1 — 없는 토큰이 양쪽 다 0인가 (계기가 «늘 잡는다» 를 내지 않는다)
|
|
98
|
+
if count(POS, NEVER) == 0 and count(NEG, NEVER) == 0:
|
|
99
|
+
print("OK META control -- an absent token matches nothing in either fixture")
|
|
100
|
+
npass += 1
|
|
101
|
+
else:
|
|
102
|
+
print("FAIL META control -- an absent token matched. Do not trust this suite.")
|
|
103
|
+
nfail += 1
|
|
104
|
+
|
|
105
|
+
# META 컨트롤 2 — 두 픽스처가 실제로 다른 파일인가
|
|
106
|
+
with open(POS, encoding="utf-8") as a, open(NEG, encoding="utf-8") as b:
|
|
107
|
+
if a.read() != b.read():
|
|
108
|
+
print("OK META control -- positive/negative fixtures differ")
|
|
109
|
+
npass += 1
|
|
110
|
+
else:
|
|
111
|
+
print("FAIL META control -- the two fixtures are identical; results are meaningless")
|
|
112
|
+
nfail += 1
|
|
113
|
+
|
|
114
|
+
print("")
|
|
115
|
+
if nfail == 0:
|
|
116
|
+
print("CALIBRATED (%d lanes) -- Step 2 five classes + Step 4-b two patterns discriminate."
|
|
117
|
+
% npass)
|
|
118
|
+
print(" NOT proven: 'zero residue' in any real document. That is measured per document.")
|
|
119
|
+
print(" Citing this suite as evidence of 'zero residue' reproduces the very downgrade")
|
|
120
|
+
print(" recorded in harness_terminal_correlation_and_recommendations.md.")
|
|
121
|
+
return 0
|
|
122
|
+
print("NOT CALIBRATED (%d/%d failed)" % (nfail, npass + nfail))
|
|
123
|
+
return 1
|
|
124
|
+
|
|
125
|
+
|
|
126
|
+
if __name__ == "__main__":
|
|
127
|
+
sys.exit(main())
|
|
@@ -46,8 +46,27 @@ run_checks() {
|
|
|
46
46
|
# ── ① 워킹트리가 깨끗한가 ────────────────────────────────────────────────
|
|
47
47
|
# `--porcelain` 은 추적 파일의 수정·스테이징·미추적을 전부 낸다. 미추적까지 세는 것은
|
|
48
48
|
# 의도적이다 — 2026-08-13 사고는 **남의 미커밋 파일**이 팩된 것이었다.
|
|
49
|
-
|
|
50
|
-
|
|
49
|
+
# 🟥 cross-family(codex/gpt-5.5, 2026-08-17) 가 이 한 줄에서 S급 3건을 냈고 S1 은 재현됐다.
|
|
50
|
+
# S1 초판 `git status --porcelain 2>/dev/null` 는 stderr 를 버리고 rc 를 안 봐서,
|
|
51
|
+
# **git 이 실패하면 빈 문자열이 「깨끗함」으로 읽혔다.** 비가역 게이트의 fail-open.
|
|
52
|
+
# 재현: printf x >/tmp/badindex; GIT_INDEX_FILE=/tmp/badindex bash <이 파일>
|
|
53
|
+
# S3 `status.showUntrackedFiles=no` 로 untracked 이 은폐된다 → -c 로 덮고 -uall 강제
|
|
54
|
+
# A4 `assume-unchanged`/`skip-worktree` 비트로 tracked 변경이 은폐된다 → ls-files -v 로 검사
|
|
55
|
+
local dirty status_out status_rc
|
|
56
|
+
status_out=$(git -c status.showUntrackedFiles=normal status --porcelain --untracked-files=all 2>&1)
|
|
57
|
+
status_rc=$?
|
|
58
|
+
if [ "$status_rc" -ne 0 ]; then
|
|
59
|
+
bad "git status 가 rc=$status_rc 로 실패했다 — **미측정은 통과가 아니다**:"
|
|
60
|
+
printf '%s\n' "$status_out" | head -5 | sed 's/^/ /'
|
|
61
|
+
return 1
|
|
62
|
+
fi
|
|
63
|
+
local hidden
|
|
64
|
+
hidden=$(git ls-files -v 2>/dev/null | grep -E '^[hS]' || true)
|
|
65
|
+
if [ -n "$hidden" ]; then
|
|
66
|
+
bad "assume-unchanged / skip-worktree 비트가 걸린 tracked 파일 — status 엔 안 보이지만 npm 은 팩한다:"
|
|
67
|
+
printf '%s\n' "$hidden" | head -5 | sed 's/^/ /'
|
|
68
|
+
fi
|
|
69
|
+
dirty=$(printf '%s\n' "$status_out" | grep -vE '^\?\? tracks/' || true)
|
|
51
70
|
if [ -n "$dirty" ]; then
|
|
52
71
|
bad "워킹트리가 깨끗하지 않다 — publish 는 커밋이 아니라 이 트리를 팩한다:"
|
|
53
72
|
printf '%s\n' "$dirty" | head -10 | sed 's/^/ /'
|
package/scripts/selfcheck.sh
CHANGED
|
@@ -530,7 +530,8 @@ for _pair in \
|
|
|
530
530
|
"scripts/residency_closure_scan.py|scripts/test_residency_closure_lanes.sh" \
|
|
531
531
|
"scripts/reviewer_capability_corpus.tsv|scripts/test_reviewer_capability_conformance.sh" \
|
|
532
532
|
"scripts/field_canon_preload.sh|scripts/test_field_canon_lanes.sh" \
|
|
533
|
-
"scripts/stale_clone_guard.sh|scripts/test_stale_clone_guard_lanes.sh"
|
|
533
|
+
"scripts/stale_clone_guard.sh|scripts/test_stale_clone_guard_lanes.sh" \
|
|
534
|
+
"plugins/fh-commons/skills/ko-tech-writer/SKILL.md|scripts/test_ko_tech_writer_lanes.sh"
|
|
534
535
|
do
|
|
535
536
|
_subj="${_pair%%|*}"; _anc="${_pair##*|}"; _lbl="${_anc##*/}"
|
|
536
537
|
if [ ! -f "$_subj" ]; then
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# test_ko_tech_writer_lanes.sh — `ko-tech-writer` Step 2 / Step 4-b 판별력 known-pair (얇은 래퍼).
|
|
3
|
+
#
|
|
4
|
+
# ─────────────────────────────────────────────────────────────────────────────
|
|
5
|
+
# 왜 필요한가 — 이 레포가 자기 스캔을 UNCALIBRATED 로 강등해 놨다
|
|
6
|
+
# ─────────────────────────────────────────────────────────────────────────────
|
|
7
|
+
# `knowledge/shared/harness-core/harness_terminal_correlation_and_recommendations.md`:
|
|
8
|
+
# Step 2 "잔여 0건(양성 컨트롤 동반)" 주장 -> UNCALIBRATED: 컨트롤·재현커맨드·출력이
|
|
9
|
+
# 하나도 없다. 재현 불가한 0은 0의 증거가 아니다
|
|
10
|
+
# Step 4 "전칭 단정 잔여 0건" 주장 -> UNCALIBRATED, 자기반증(그 시점에 3건 생존)
|
|
11
|
+
#
|
|
12
|
+
# 챔버 런 #12(2026-08-17, prosody-lens KILL)가 이 부채를 배출 판정의 결정적 근거로
|
|
13
|
+
# 인용했다 — "같은 계열 계기가 미보정인데 하나 더 짓는 것은 미보정 계기의 증식".
|
|
14
|
+
# 운영자 결정: 새 계기보다 이 부채가 먼저. 이 파일이 그 부채를 갚는다.
|
|
15
|
+
#
|
|
16
|
+
# ─────────────────────────────────────────────────────────────────────────────
|
|
17
|
+
# 왜 래퍼인가 — 패턴이 셸 밖에 있어야 하는 실측 이유가 둘 있다
|
|
18
|
+
# ─────────────────────────────────────────────────────────────────────────────
|
|
19
|
+
# 초판은 패턴을 이 파일 안에 두고 엔진을 ripgrep 으로 고정했다. CI 가 둘 다 반증했다
|
|
20
|
+
# (2026-08-17, ubuntu-latest):
|
|
21
|
+
# 1) ripgrep 부재 -> 스위트가 rc=10 으로 죽었다(HARNESS ERROR). 설계는 옳게 작동했다 —
|
|
22
|
+
# SKIP 이 아니라 시끄럽게 죽었다 — 그러나 CI 에서 못 도는 앵커는 앵커가 아니다.
|
|
23
|
+
# 2) 이식성 린트 적중 — 셸 파일 안의 한글 문자 범위는 GNU grep C 로케일에서 exit 2 로
|
|
24
|
+
# 죽고 그 결과가 "무매치"로 읽힌다. 즉 fail-open.
|
|
25
|
+
# 두 결함이 같은 처방으로 닫힌다: **패턴을 Python 으로 옮긴다.** SKILL.md 가 허용 엔진으로
|
|
26
|
+
# Python `re` 를 이미 명시했으므로 격하가 아니라 등가 엔진으로의 이동이고, python3 는
|
|
27
|
+
# 세 허용 엔진 중 어디에나 있는 유일한 것이다.
|
|
28
|
+
#
|
|
29
|
+
# ─────────────────────────────────────────────────────────────────────────────
|
|
30
|
+
# 무엇이 재어지는가 — Step 2 다섯 클래스 + Step 4-b 두 패턴
|
|
31
|
+
# ─────────────────────────────────────────────────────────────────────────────
|
|
32
|
+
# Step 2 C1 줄표 이어붙임 · C2 용어-머리 · C3 콜론 나열투 · C4 조각문 · C5 소유 직역
|
|
33
|
+
# Step 4-b C6 전칭 어휘 · C6b 부정형 전칭
|
|
34
|
+
# 각 레인의 PASS 조건은 **양성 >= 1 그리고 음성 == 0** 이다. 존재 확인이 아니라 판별
|
|
35
|
+
# 확인이라는 점이 핵심 — known-negative 가 없으면 "오탐이 있나" 조차 못 재고, 그러면
|
|
36
|
+
# PASS 는 "늘 잡는 계기" 와 구분되지 않는다.
|
|
37
|
+
#
|
|
38
|
+
# 🟥 첫 캘리브레이션이 정본의 과장을 드러냈다: SKILL.md 는 Step 2 의 "앞 다섯 줄은 기계
|
|
39
|
+
# 검출 가능" 이라 적었는데, 실제로 grep 을 싣고 있는 것은 C1·C5 둘뿐이고 C2·C3·C4 의
|
|
40
|
+
# 검출 힌트 칸은 산문 서술이다. 특히 C4 는 패턴이 없어서, 내가 임의로 지은 패턴이
|
|
41
|
+
# 양성을 0건으로 놓쳤다 — 계기를 세우자마자 대상의 결함이 나온 형태다.
|
|
42
|
+
#
|
|
43
|
+
# 🟥 이 스위트가 PASS 를 내도 주장하면 안 되는 것: "Step 2 잔여 0건" · "Step 4 잔여 0건".
|
|
44
|
+
# 판별력과 잔여는 다른 명제이고, 후자는 문서마다 따로 잰다. 그 혼동이 애초에 두 행을
|
|
45
|
+
# UNCALIBRATED 로 강등시킨 사유다.
|
|
46
|
+
#
|
|
47
|
+
# 사용: bash scripts/test_ko_tech_writer_lanes.sh
|
|
48
|
+
# exit: 0 판별 확인(전 레인 PASS) · 1 판별 실패 · 10 harness error(python3/픽스처 부재)
|
|
49
|
+
|
|
50
|
+
set -uo pipefail
|
|
51
|
+
|
|
52
|
+
SRC="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
53
|
+
PY="$SRC/ko_tech_writer_calibrate.py"
|
|
54
|
+
|
|
55
|
+
command -v python3 >/dev/null 2>&1 || {
|
|
56
|
+
echo "harness error: python3 absent -- cannot measure. Not a pass." >&2; exit 10; }
|
|
57
|
+
[ -f "$PY" ] || { echo "harness error: missing $PY" >&2; exit 10; }
|
|
58
|
+
|
|
59
|
+
echo "-- ko-tech-writer Step2/4-b discrimination known-pair --"
|
|
60
|
+
python3 "$PY"
|
|
61
|
+
exit $?
|