@chrono-meta/fh-gate 2.6.0 → 2.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/.claude/rules/fh_4axis_gate.md +26 -3
  2. package/.claude-plugin/marketplace.json +2 -2
  3. package/AGENTS.md +24 -3
  4. package/CLAUDE.md +39 -12
  5. package/README.ja.md +32 -8
  6. package/README.ko.md +32 -8
  7. package/README.md +112 -15
  8. package/README.zh.md +28 -8
  9. package/knowledge/shared/harness-core/field_verdict_crossfamily_gate.md +19 -5
  10. package/knowledge/shared/harness-core/harness_incubator_doctrine.md +12 -2
  11. package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +32 -0
  12. package/knowledge/shared/harness-core/ship_readiness_gate.md +77 -8
  13. package/knowledge/shared/learnings/subagent_invocations_log.yaml +90 -0
  14. package/knowledge/shared/rules/knowledge_layer_seam.md +1 -1
  15. package/package.json +13 -2
  16. package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
  17. package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
  18. package/plugins/fh-meta/CHANGELOG.md +111 -0
  19. package/plugins/fh-meta/agents/persona-innovator.md +170 -0
  20. package/plugins/fh-meta/skills/public-surface-audit/SKILL_detail.md +16 -1
  21. package/plugins/fh-meta/skills/steel-quench/SKILL.md +97 -0
  22. package/scripts/adapters/peer_resolve.sh +58 -6
  23. package/scripts/cluster_capability_scan.sh +42 -13
  24. package/scripts/digest_landing_check.sh +142 -2
  25. package/scripts/fh_hub_identity.sh +83 -0
  26. package/scripts/fh_session_load.sh +53 -5
  27. package/scripts/fh_track_resolve.sh +114 -0
  28. package/scripts/field_canon_preload.sh +50 -5
  29. package/scripts/package_coverage_check.sh +8 -0
  30. package/scripts/prior_art_prompt.sh +168 -0
  31. package/scripts/psa_scan_lib.sh +201 -15
  32. package/scripts/residency_admission_check.sh +204 -0
  33. package/scripts/selfcheck.sh +88 -0
  34. package/scripts/test_adapter_lanes.sh +67 -2
  35. package/scripts/test_heavy_classifier_lanes.sh +144 -0
  36. package/scripts/test_marker_defense_lanes.sh +152 -0
  37. package/scripts/test_marker_soul_check_lanes.sh +211 -0
  38. package/scripts/test_prior_art_prompt_lanes.sh +128 -0
  39. package/scripts/test_psa_singlefile_lanes.sh +351 -1
  40. package/scripts/test_residency_admission_lanes.sh +60 -0
  41. package/scripts/test_track_resolve_lanes.sh +158 -0
  42. package/templates/.git-hooks/pre-commit +400 -4
  43. package/templates/.git-hooks/pre-push +17 -2
  44. package/templates/settings.PriorArt.snippet.json +15 -0
@@ -20,6 +20,7 @@ The main agent passes you one of:
20
20
  - **Mode E (External scan)**: "scan frontier" / "what are people building" / specific topic
21
21
  - **Mode F (Full)**: both — default when no mode is specified
22
22
  - **Mode T (Technical bridge)**: "can't connect" / "not possible" / "blocked" / "no direct path" / technical constraint hit
23
+ - **Mode X (Intervention cross-check)**: "쎄한데 확인해줘" / "내가 뭘 놓쳤나" / "개입 대조" / a session asking whether it is about to be stopped. Runs Phase 3-b ONLY — no naming, no frontier scan.
23
24
 
24
25
  Optionally: a focus area (e.g., "token efficiency", "agent orchestration", "cascade patterns")
25
26
 
@@ -153,6 +154,175 @@ For each gap or absorbed signal:
153
154
  4. **Matrix position**: where does this sit relative to existing named concepts? (complement / extend / replace)
154
155
  5. **Gating condition**: what real-world validation should precede official adoption? (simplicity guard applied)
155
156
 
157
+
158
+ ## Phase 3-b — Intervention algorithm (Mode X, and MANDATORY inside Mode F)
159
+
160
+ Phase 3 above carries the owner's **naming** algorithm. This phase carries the owner's
161
+ **intervention** algorithm — *where the owner has historically stopped a session and turned it.*
162
+
163
+ **Provenance (measured, not asserted)** — census of the conversation corpus itself, not of what
164
+ sessions wrote down afterwards: `~/.claude/projects/…/*.jsonl`, **69 sessions / 2026-07-22–08-21**,
165
+ **433 operator utterances**, semantically classified. **54 interventions claimed · ~43 estimated
166
+ after a 5-sample hand-check (1 false positive) · full hand-verification NOT done.**
167
+ 🟥 **CORRECTED 2026-08-21 — that census read 21% of its own corpus and called it 전수.** A full
168
+ re-scan of the same directory with the same discriminator returns **168 sessions · 1,703 operator
169
+ utterances** (uuid-deduplicated; ~1% contamination hand-checked: `<bash-input>` / `<command-message>`,
170
+ 17 of 1,703). The census's own note recorded **169MB**, and `du -sh` on that directory is **817MB** —
171
+ the ratio was written down and never compared. **So «54» is a count over a fifth of the corpus, not a
172
+ census.** Do not cite it alone. If the 12% rate holds, the true intervention count is nearer **200**.
173
+ The shapes and the class distribution below are unaffected in *direction* (they were derived from a
174
+ random-in-practice fifth), but every absolute number on this page is a lower bound.
175
+ Detail + the seal comparison: `tracks/_meta/RESULT_2026-08-21_intervention-corpus.md`.
176
+
177
+ 🟥 **The dominant class is NOT "you didn't search the world."** Measured distribution:
178
+ `판단결함 32 (59%) · 내부미조회 9 · 외부미조회 6 (11%) · 범위겨냥 6`. A design that treats this
179
+ as a *search* trigger is aiming at an 11% slice — the first draft of this capability did exactly that.
180
+
181
+ **Prior art (2026-08-21) — this task has a published benchmark; price the capability against it.**
182
+ Wu et al., *"User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning
183
+ Signal"*, EMNLP 2025 main (`arXiv:2507.23158`). Numbers read from the PDF text, not from a summary:
184
+ automatic feedback identification with a purpose-built GPT-4o-mini prompt scores **P 61.1 / R 35.9**
185
+ in the *dense* setting (label every turn — the realistic one) and **P 100.0 / R 69.2** in the *sparse*
186
+ setting (the feedback turn is pointed out in advance); inter-annotator agreement **Cohen κ = 0.70
187
+ (binary) / 0.74 (three-way) / 0.60 (fine-grained)** over 54 cross-annotated conversations.
188
+ ⇒ Two consequences. **(a)** A weak separation here is the task's difficulty, not this rule set being
189
+ unusually bad — quote precision against **P61/R36**, never against a vacuum. **(b)** Their conclusion
190
+ (*noisy as a learning signal*) converges independently with residual (0) below.
191
+ 🟥 **Do not collapse the two tasks.** They read the user's turn and judge post-hoc; this phase predicts
192
+ *before* the user speaks, from the session's own acts, with the next human turn structurally excluded.
193
+ Ours is strictly harder and that difference is the net-new angle — normalizing it away both overstates
194
+ the prior art and invalidates any direct number comparison.
195
+ 🟥 Provenance of this paragraph: the world was consulted **only after the operator asked whether it had
196
+ been**; zero external lookups preceded the design. That is identity ④'s measured gap, not a footnote.
197
+
198
+ ### The ten shapes — a session in this state is about to be stopped
199
+
200
+ | Shape | The session is… | n |
201
+ |---|---|---|
202
+ | 손안의 기록 미개봉 | reporting «missing / incomplete / impossible» about material it already holds but never opened | 9 |
203
+ | 막힘을 종점으로 접기 | folding at a blocker (hard / absent / over quota / someone else's job / next session) without trying one alternate route | 9 |
204
+ | 산출물 미대조 완료선언 | writing «applied it / fixed it» without re-opening or re-rendering the artifact | 7 |
205
+ | 세계 미조회 설계 | designing from repo-and-session material only, zero external prior art consulted | 6 |
206
+ | 단일 관측 전역 단언 | fixing a cause/state/number from one node or one observation, no second instrument | 6 |
207
+ | 과녁 치환 | working hard on a different target than the one named in the request | 5 |
208
+ | 정적 리뷰로 완료 | closing a wiring/rule change on a static read, never run from the target tier or target harness | 4 |
209
+ | 이상신호 흘려보내기 | passing over a tool error, odd latency, or overwrite side-effect without citing or acting | 4 |
210
+ | 기준 없는 등급 판정 | setting a version digit or identity grade without citing the document that defines it | 2 |
211
+ | 절차 규모 오조준 | defaulting to a heavier procedure than the change needs | 2 |
212
+
213
+ ### The self-check list — run these against **this turn's tool calls and output**
214
+
215
+ 🟥 **Every line is stated as an observable act, never as a feeling.** That is deliberate: the
216
+ adversarial review that killed the first draft found the fatal flaw was a trigger keyed on the
217
+ session's own sense of uncertainty — *a session that is confident does not report uncertainty, and
218
+ the measured failures were all confident.* These lines are checkable from what the session just did.
219
+
220
+ 1. Wrote a paragraph or more of design/definition/gap-diagnosis, and opened **zero** external sources this turn.
221
+ 2. About to write «없다 / 미등재 / 미완 / 불가», and never opened the file or history that would hold it.
222
+ 3. Wrote «반영 완료» about a document without counting the items in it.
223
+ 4. Wrote «applied / fixed / included it all» with **no tool call re-reading that artifact after the edit**.
224
+ 5. Was given N items and touched fewer than N, without putting both numbers side by side.
225
+ 6. About to write «next session / someone else / later» with **no tool call attempting an alternate route this turn**.
226
+ 7. Dropped a verification leg because a sidecar was blocked, with no record of trying another family / local LLM / subagent.
227
+ 8. Withdrew its own proposal citing only «hard / side effects», with not one line on how to make it work.
228
+ 9. Asked the operator about a peer session's state instead of asking that session via ListAgents/SendMessage.
229
+ 10. Routed a candidate to CURATED / drop / hand-off **without one line on how it could become our own capability**.
230
+ 11. The file / environment / axis being edited is not the noun the operator named.
231
+ 12. Filled a mapping or candidate list only from what exists locally on this machine.
232
+ 13. About to write PASS on a rule/wiring change and cannot quote a command run in the target tier or harness with its output.
233
+ 14. Ran a «standpoint review» from its own vantage, with no agent dispatched inside the target harness.
234
+ 15. Fixed a cause/state/number from one node or one observation, with no second instrument.
235
+ 16. Wrote an aggregate count without checking whether already-running or pre-existing items are inside it.
236
+ 17. Wrote a time/date/environment fact from memory or inference rather than from a command.
237
+ 18. Judged a tool error or warning «non-blocking» and moved on without citing it or acting.
238
+ 19. Created or changed a setting and wrote «done» without printing its expiry / default fields.
239
+ 20. Regenerated or overwrote a file without a diff showing which prior lines are gone.
240
+ 21. Waiting on a run that is taking longer than expected without checking its output or whether a session was created.
241
+ 22. Raised a version digit or grade without quoting the document that defines that digit.
242
+ 23. Proposed follow-up work larger than the original request without putting a minimal option beside it.
243
+
244
+ ### Output for Mode X
245
+
246
+ For each line that fires: quote the session's own act that trips it, and propose **one line** —
247
+ *"확인해볼까?"* — naming the cheapest check that would settle it. **Propose; never decide.**
248
+ Fires nothing → say «걸린 줄 없음» explicitly; silence is not a verdict.
249
+
250
+ ### Tier M — signals decidable from the session RECORD (calls + turns + diff), no judgment
251
+
252
+ 🟥 **The first draft of this heading said «from the tool-call record alone». That was false**
253
+ (cross-family, 2026-08-21): #9, #10, #12, #15 and #18 require reading the user's turn, the reply, or
254
+ the commit diff — not the call log. The tier's real claim is narrower and is what the heading now
255
+ says: **no judgment is needed**, but more than the call log is read. An evaluator for this tier needs
256
+ a defined input contract (calls · user turns · final reply · staged diff) that **does not exist yet**.
257
+
258
+ A **second census** (same question, different corpus: what sessions *recorded* about
259
+ interventions, `tracks/`+`knowledge/`+memory — 300 scanner hits → 199 claimed → **33 hand-verified,
260
+ 18% rejected**) produced signals of a different grade: each one is a **countable fact about this
261
+ session's own calls**, needing no judgment. Both censuses landed on the same class distribution
262
+ (판단결함 dominant · 외부미조회 a minority).
263
+ 🟥 **That agreement is CORROBORATING BUT CORRELATED — not independent** (cross-family caught the
264
+ overclaim). Same operator, same canon, same model family; and the `tracks/` records are *derivative
265
+ of the same events* the transcripts hold. Claiming independence would need event-linkage removal, a
266
+ different annotator/model, a pre-registered codebook and blind reclassification — **none were done.**
267
+
268
+ 1. An absence/blocked claim (`없다`·`0건`·`not found`·`unavailable`·`막혔`·`overdue`) appears, and the tool call against that subject happened **exactly once** — no second attempt.
269
+ 2. A tool output carries a truncation marker (`truncated`·`… N more`·a next/page cursor·line count exactly equal to the limit) and the tool was **never re-called with a different offset/page/cursor**.
270
+ 3. A call ended non-zero or errored, and the **same tool with the same arguments was not retried** — the session switched to a different tool instead.
271
+ 4. A background handle has produced **0 bytes of stdout for N seconds** and has not exited, and the call carried no timeout.
272
+ 5. A freshness/cadence verdict rests on a single glob whose match count is **0** (rendering `not found` as `overdue`).
273
+ 6. After session-start `pull`/`fetch`, the newest remote commit is **later than the date field of the card/INDEX that was read**, and **zero** of the files those commits touched were Read.
274
+ 7. A staged git-tracked added line contains an absolute home path, a companion-store name, a vendor/product proper noun, or an executable that only `command -v` resolves **on this machine**.
275
+ 8. A diff under `package.json files[]` · `templates/` · `plugins/` newly introduces a local-only path or local-only CLI name — an environment dependency entering the shipped set.
276
+ 9. A noun phrase or quoted string from the user's turn appears **0 times** in the session's whole commit diff (operator utterance ↔ canon landing).
277
+ 10. A quantity token (`N건`·`N자`·`N%`·`HH:MM`) appears in an artifact or final reply, and the session made **no call able to produce it** (`wc`·`grep -c`·`date`·arithmetic).
278
+ 11. A time/date predicate (`심야`·`오전`·`어제`·a weekday) was written to a record with **zero `date` calls**.
279
+ 12. The first user turn matches the greeting corpus and the first reply carries **neither the 🐿️ literal nor the fixed welcome line**.
280
+ 13. A section a rule marks «always include» greps **0 times** in the artifact that rule governs.
281
+ 14. A new file is about to be written with **zero** prior Read/Grep against `CATALOG.md` / the skill list / `plugins/**/SKILL.md`, while its name or keywords already match the index.
282
+ 15. A skill/agent proper noun the session named as the routing target appears **nowhere in the user's turn** — the session introduced that name.
283
+ 16. An external model's or sidecar's **self-report string** is cited as verdict evidence, with **0 calls** running the same probe against a known control.
284
+ 17. The diff changes an exit code, a default, or a fail-open/closed direction, and the commit message or 4-axis marker quotes the user's turn **0 times**.
285
+ 18. A recommendation to install or use a tool carries **no conditional marker** (`when`·`only if`·`unless`·`~일 때만`) anywhere.
286
+
287
+ ### How the two tiers are used
288
+
289
+ ```
290
+ Tier M (18) countable from this session's calls → a hook could evaluate these
291
+ Tier J (23) need reading the session's own output → invocation, judged
292
+ ```
293
+ **Noise cap (mandatory).** Rank by tier then by how cheap the check is, and surface **at most 2 per
294
+ turn**; hold the rest silently. 🟥 Without this, Mode F makes this phase mandatory and every firing
295
+ emits a proposal — up to **41 «확인해볼까?» in one turn**, which is the nag that trains dismissal and
296
+ kills the capability (cross-family MED). Repeat suggestions dedupe by shape, not by wording.
297
+
298
+ 🟥 **Neither tier decides.** Both produce the same one line — *"확인해볼까?"* — naming the cheapest
299
+ check. The operator's bar for this capability is exactly that: *"'쎄함'을 감지하고 사람에게
300
+ 「한번 확인해볼까?」 라고 제안하는 것만 가능해도 성공이다."*
301
+
302
+ ⚠️ **Named residuals.** (0) **MEASURED 2026-08-21 — and it did not pass its pre-registered bar.**
303
+ (The 18% figure below is the rejection rate of one census's *claims*; it never was these rules'
304
+ precision.) A sealed pre-registration (`sha256 8414a44d…`, re-verified at opening) ran the 41 rules
305
+ against **44 windows** — 22 immediately preceding an operator intervention, 22 preceding an approval —
306
+ with blinding enforced by file boundary (the rule-evaluator never receives the next human turn; the
307
+ labeler never receives the rules). Result: **12 of 41 rules ever fired · 2 discriminate**
308
+ (`J2` +4, `M10` +3) · **10 fired below the +3 line** (J1 · J17 · M2 · J4 · J12 · M1 · J6 · M4 · J18 ·
309
+ M3 — «did not clear the bar in this sample», **not** the same bucket as dead) · **29 never fired at
310
+ all**, 14 of them Tier M. The pre-registered stop condition **«no separation» HIT** (median difference
311
+ +0.5 vs a required +2.0).
312
+ 🟥 **The pre-registration itself was defective and that is recorded, not hidden**: a median is a poor
313
+ statistic on a 0–1 sparse count. A post-hoc binary comparison (any rule fires: POS 82% vs NEG 50%,
314
+ Fisher p=0.055) suggests signal — **it does not overturn the verdict**, and claiming it would require
315
+ replication under a *new* pre-registration naming that statistic in advance. Current status:
316
+ **signal possible, not established.**
317
+ ⚠️ Scope of that run: **in-sample** (no out-of-sample corpus exists — sessions in other repos are
318
+ one-shot dispatches with no preceding window), which makes it a *favourable* arm: failing it is
319
+ conclusive, passing it would not have been. Detail: `tracks/_meta/RESULT_2026-08-21_intervention-precision.md`. (0-b) Tier M has **no executable definitions** — «N seconds», «same subject»,
320
+ «alternate route», «conditional marker» are undefined; an evaluator schema, window and no-data
321
+ verdict must precede any wiring. (a) Tier M is written but **not wired** — no hook evaluates it yet; both
322
+ tiers currently run on invocation. (b) Neither census hand-verified in full: transcripts 5/54
323
+ checked, tracks 33 of 199 claims verified at an 18% rejection rate. (c) The transcript corpus is
324
+ **one month deep** (2026-07-22 onward); earlier interventions are structurally absent, not zero.
325
+
156
326
  ## Self-floor discipline (FH floors, applied to the innovator itself)
157
327
 
158
328
  These are FH's own governance floors turned reflexively on this agent's process — an ideation tool
@@ -25,9 +25,24 @@ if [ -r "scripts/psa_scan_lib.sh" ]; then
25
25
  # Feed every tracked file as path<TAB>line, the stream psa_scan_tagged consumes. Sourcing the lib
26
26
  # without these calls is a no-op scan — measured on this repo (PSA_STREAM stayed unset), so the
27
27
  # calls are spelled out here rather than pointed at.
28
+ # 🟥 rc 를 반드시 받아라 (2026-08-21 배선 리뷰 S-1). 이 파이프는 **원래 뚫렸던 바로 그 진입점**이고,
29
+ # 바로 아래 coverage 조건이 `$?` 를 덮으므로 여기서 안 받으면 계약이 소실된다.
30
+ # 계약: 0=신고할 것 없음 · 1=유출(이미 인쇄됨) · 3=NOT SCANNED(계기 사망)
31
+ # ⚠️ `_psa_rc=0` 은 반드시 `while` **밖**에 둔다 — 안에 두면 루프 본문이라 매 줄 초기화되고
32
+ # 파이프 서브셸에 갇힌다(초판이 그렇게 넣었고 `bash -n` 은 통과했다. 문법은 맞고 의미가 틀린다).
33
+ _psa_rc=0
28
34
  while IFS= read -r f; do
29
35
  awk -v p="$f" '{printf "%s\t%s\n", p, $0}' "$f" 2>/dev/null
30
- done < /tmp/_psa_tracked.txt | psa_scan_tagged
36
+ done < /tmp/_psa_tracked.txt | psa_scan_tagged || _psa_rc=$?
37
+ # 🟥 3 은 «깨끗» 이 아니라 «안 쟀다» 다. 여기서 멈춰야 한다 — 이 스킬은 publish 직전에
38
+ # Pre-Publish Gate 가 1번으로 체이닝하는 렌즈이고, 그 자리에서 미측정을 통과시키면
39
+ # 아래 coverage 줄이 «defaults-only 로는 스캔했다» 는 인상까지 얹는다.
40
+ if [ "$_psa_rc" -eq 3 ]; then
41
+ echo "⛔ INSTRUMENT DEAD: the scanner did not run (rc=3). NOT SCANNED is not clean."
42
+ echo " Fix first — run under bash (zsh special vars can blank PATH inside the matcher),"
43
+ echo " and confirm psa_load ran (PSA_STREAM non-empty). Do NOT report a verdict from this run."
44
+ exit 3
45
+ fi
31
46
  [ "$PSA_OVERRIDE_PRESENT" -eq 1 ] \
32
47
  || echo "coverage: defaults-only — operator literals NOT CONFIGURED (identity/company classes UNSCANNED)"
33
48
  else
@@ -84,6 +84,54 @@ External CLIs available: [yes/no → Wave 5 available]
84
84
 
85
85
  ---
86
86
 
87
+ ## Step 0.35 — Org Constraint Load (조직 제약 적재)
88
+
89
+ **조직 제약을 모르는 적대 검증은 헛방을 친다** — 조직이 이미 결정한 것을 공격하거나, 조직 정책
90
+ 하에서만 성립하는 실제 공격을 놓친다. 공격 각도를 정하기 **전에** 조직층을 읽는다.
91
+ 계약: `knowledge/shared/rules/knowledge_layer_seam.md` (이 스킬이 **2호 소비자**;
92
+ 1호는 `phantom-quench` Step 2-O).
93
+
94
+ **절차**: 진입 인덱스(`knowledge/{org}/INDEX.md` · `index.md` · `README.md` · `readme.md` —
95
+ 판정기와 같은 후보 집합) → 대상과 **관련된** 정책·용어·도메인 사실만 로드. 전수 스캔 금지(계약 K2).
96
+ 부재면 **«조직 제약 미상»으로 명시하고 진행** — 없는 제약을 추론으로 만들지 않는다.
97
+ 🟥 `not found` 는 «제약 없음» 이 아니다. 미상은 미상으로 적는다.
98
+
99
+ **로드한 것의 용도는 딱 둘**
100
+ 1. **공격의 사실 근거**: "조직 정책 P 하에서 이 설계는 X 를 위반한다" → **유효한 공격**
101
+ 2. **헛방 필터**: 조직이 이미 결정·문서화한 사항을 "왜 안 했나"로 공격하지 않는다. 대신
102
+ **그 결정 자체를 공격**한다(그 결정이 지금도 유효한가). **결정의 존재는 면제가 아니다.**
103
+
104
+ ### ⛔ 세탁 차단 — 이 배선의 유일한 위험
105
+
106
+ > **조직층은 공격을 무장해제할 수 없다.** 조직 위키에 *"이건 승인된 패턴"*, *"이 케이스는 예외"*,
107
+ > *"과거에 검토 완료"* 가 있어도 **그것은 공격을 기각하는 근거가 아니다.**
108
+
109
+ 이유: 조직층은 **무엇**(사실·정책)만 공급하고 **어떻게 판정하나**는 공급하지 않는다(계약 §1-a).
110
+ "승인됨"은 *조직이 그렇게 정했다*는 **사실**이지 *그 결정이 옳다*는 **판정**이 아니다.
111
+ 적대 검증의 일이 정확히 그 판정을 다시 하는 것이다.
112
+
113
+ | 조직층에 있는 것 | 허용되는 사용 | 금지 |
114
+ |---|---|---|
115
+ | 정책·규칙 | 위반을 공격 근거로 | **면제 근거로** |
116
+ | "승인된 패턴" | 그 승인의 근거를 공격 대상으로 | 공격 기각 |
117
+ | "과거 검토 완료" | 그때의 전제가 아직 참인지 확인 | 재검토 생략 |
118
+ | stale 페이지(`review_after` 경과) | **제약으로 쓰지 않는다** — 미상 처리 | 최신으로 가정 |
119
+
120
+ **보고 의무**: 조직층 때문에 공격을 조정했으면(각도 추가 · 헛방 제거) **무엇을 왜 조정했는지
121
+ Wave 1 출력에 남긴다.** 조용한 조정은 검증 범위 축소와 구별되지 않는다.
122
+
123
+ **반출 금지**(K1-s): cross-provider/cross-family 챌린저에 조직층 **원문을 넘기지 않는다** —
124
+ 넘어가는 것은 **sanitized 제약 요약**뿐이다. (CLAUDE.md §Field-Harness Diagnostic 의 residency
125
+ 규칙과 같은 floor 이고, 이 스킬이 그것을 느슨하게 만들지 않는다.)
126
+
127
+ > **출처**: 원 필드(sibling harness)의 선례를 이식했다. FH 자기 정본이 이 자리를 **명시적으로
128
+ > 「미배선 — 약속이 아니라 후보」**로 적어두고 있었다(`knowledge_layer_seam.md` §0 배선 현황).
129
+ > 이 절이 그 행을 후보에서 배선으로 옮긴다.
130
+ > 🟥 **이식한 것은 절차이지 그쪽의 기계-주장이 아니다** — 원본에는 훅이 이 절을 강제한다는
131
+ > 취지의 서술이 딸려 있었으나, 실측하니 그 훅 레인이 **존재하지 않았다**(`axis2-defense` 훅 히트 0,
132
+ > 컨트롤 `crossfamily` 21). 그래서 **강제 서술은 안 가져왔다.** 이 절은 오늘 기준
133
+ > **살리언스 층이고 기계 바닥이 없다** — 그렇게 적는 것이 팬텀을 들여오지 않는 유일한 방법이다.
134
+
87
135
  ## Step 0.4 — Specialized Reviewer Discovery
88
136
 
89
137
  For the target artifact, scan installed agents for a domain-specific adversarial reviewer:
@@ -211,6 +259,55 @@ A finding here is a real-code attack (Wave 1 execution principle) — cite the e
211
259
 
212
260
  ---
213
261
 
262
+ ## Wave 1-D — Defense Questions (floor tiers · mechanically required)
263
+
264
+ Three questions, asked of **your own findings and numbers**, before Wave 1 is done. They are written
265
+ out rather than left to judgment because that is exactly what makes them portable: measured n=6 in
266
+ the origin field, a floor-tier pass executes the attack angles above without defect but **does not
267
+ spontaneously ask these three**, while a stronger tier does. That is a **checklist gap, not a
268
+ capability gap** — and by `sonnet_floor_doctrine.md` a harness whose behaviour depends on which model
269
+ is driving is *defective*, not merely limited.
270
+
271
+ | # | Question | What a real answer looks like |
272
+ |:---:|---|---|
273
+ | **재현성** | Can another session reproduce this verdict from the same inputs? | The exact command, or `file:line`, another session runs. "It's reproducible" is not an answer. |
274
+ | **비교공정성** | Were the two arms measured under the same conditions? | reps, inputs and environment named for **both** arms. An asymmetry you found and left in place counts — say so. |
275
+ | **추정층위** | Is each number a measurement, an estimate, or a quotation? | Which, per number — and for a measurement, what showed the instrument works **on this target** (known-pair). |
276
+
277
+ **Where the answers go** — one line in the Axes 2+3 marker:
278
+
279
+ ```
280
+ axis2-defense: reproducibility=<…> fairness=<…> estimation-layer=<…>
281
+ ```
282
+
283
+ **Enforcement, stated exactly.** `templates/.git-hooks/pre-commit` → `validate_defense_leg()` runs
284
+ this **at the floor tiers only** (`floor-status: sonnet-floor` or `below-floor`), and checks
285
+ **presence · completeness · non-vacuity**: all three sub-answers must exist and `ok`/`yes`/`n/a` is
286
+ rejected as a filled form rather than an answer. Fixtures: `scripts/test_marker_defense_lanes.sh`
287
+ (17 lanes: known-pairs both directions, two over-block controls, four prescription assertions, and a
288
+ call-site pin — because a suite that extracts the function and calls it directly stays green when
289
+ the hook stops calling it, which is precisely the failure being imported against).
290
+ 🟥 **It cannot check whether the answers are TRUE.** That stays with the operator and the weekly
291
+ audit, exactly like every other marker field — do not read the hook's PASS as verification.
292
+
293
+ **Why the trigger is narrow.** `below-floor` occurs **0 times** across the existing marker corpus, so
294
+ gating on it alone would be a decoration that never fires; `sonnet-floor` occurs 11. The wide reading
295
+ ("any marker carrying numbers") is deliberately **not** taken — pricing this axis at a near-universal
296
+ rate is the over-trigger `field_verdict_crossfamily_gate.md §7` rejects, and a field required
297
+ everywhere becomes a rubber stamp.
298
+
299
+ > **Provenance, and what was deliberately NOT imported.** Absorbed from a sibling harness
300
+ > (2026-08-20). That document additionally amends its floor rule so a below-floor pass carrying this
301
+ > line **plus** a crossfamily record passes **without the operator ack**. 🟥 That is a *loosening* of
302
+ > an existing FH gate and **was not adopted** — here the leg is purely additive and `below-floor`
303
+ > still requires `below-floor-ack`, unchanged.
304
+ > 🟥 The same document asserted its own hook read this field and its own fixture suite pinned it.
305
+ > Measured twice with a control (`crossfamily` → 21 hits in the same files): **both were 0**. The
306
+ > prose was portable; the machine was not. Everything claimed in this section's *Enforcement*
307
+ > paragraph was built here, and the fixture file named there is the receipt.
308
+
309
+ ---
310
+
214
311
  ## Wave 2 — Defense Principles
215
312
 
216
313
  **3 Defense Principles**: (1) Reinforce with external cases via WebSearch — "unique to us" or "structural pattern"?
@@ -35,6 +35,38 @@
35
35
  # rc=3 NO_ROOTS — 루트 목록 자체가 비었다(전제 파손. 「없다」가 아니다)
36
36
  # stderr : rc!=0 일 때 한 줄 사유
37
37
 
38
+ # ── 단일 소스 배선 (2026-08-21 F-1) ──────────────────────────────────────────
39
+ # 🟥 위 주석이 «스캐너와 같은 순서여야 한다» 고 **관례로** 적어 두었던 그 중복을 여기서 닫는다.
40
+ # `scripts/fh_track_resolve.sh` 가 track 이름 → 레포 루트 해석의 단일 소스이고,
41
+ # 아래 `FH_PROJECTS_HOME` 팔은 그 라이브러리를 부른다.
42
+ #
43
+ # ⚠️ **`FH_CLUSTER_ROOTS` 팔은 닫히지 않는다 — 조회 모델이 다르기 때문이다.**
44
+ # 라이브러리는 «루트 하나 밑의 `$root/$c`» 를 묻는다. 이 팔은 «임의 절대경로 목록의
45
+ # `basename` 이 별칭인가» 를 묻는다. 이름만 같고 의미가 다른 둘을 억지로 합치면 그건
46
+ # 이 수리가 막으려는 결함과 같은 종류다. 그래서 그 팔은 자기 열거를 유지한다
47
+ # (⇒ `fh_peer_alias_candidates` 는 남는다. 부재 메시지와 강등 폴백도 이걸 쓴다).
48
+ #
49
+ # 라이브러리가 없으면 **수리 이전 동작**으로 강등한다 — 다른 세 소비자와 같은 형태다.
50
+ # (폴백은 의도적으로 옛 구현 그대로다. 공백 결함까지 포함해서 «수리 이전» 이다.)
51
+ _FH_TRLIB="$(cd -P "$(dirname "${BASH_SOURCE[0]:-$0}")/.." 2>/dev/null && pwd)/fh_track_resolve.sh"
52
+ # shellcheck source=scripts/fh_track_resolve.sh
53
+ [ -f "$_FH_TRLIB" ] && . "$_FH_TRLIB"
54
+ type fh_resolve_track_root >/dev/null 2>&1 || fh_resolve_track_root() {
55
+ local n="$1" root="$2" c hits="" first="" nh
56
+ for c in "$n" "${n}-dev" "$(printf '%s' "$n" | tr '_' '-')"; do
57
+ [ -d "$root/$c" ] || continue
58
+ case " $hits " in *" $c "*) continue ;; esac
59
+ hits="$hits $c"; [ -n "$first" ] || first="$c"
60
+ done
61
+ nh=$(printf '%s' "$hits" | wc -w | tr -d ' ')
62
+ if [ "${nh:-0}" -gt 1 ]; then
63
+ printf '%s|AMBIGUOUS:%s' "$root/$n" "$(printf '%s' "$hits" | sed 's/^ //; s/ /,/g')"; return 0
64
+ fi
65
+ [ -n "$first" ] || { printf '%s|UNRESOLVED' "$root/$n"; return 0; }
66
+ [ "$first" = "$n" ] && { printf '%s|' "$root/$n"; return 0; }
67
+ printf '%s|alias:%s' "$root/$first" "$first"
68
+ }
69
+
38
70
  # 별칭 후보 — 닫힌 목록. 스캐너의 `_resolve_root` 와 **같은 순서**여야 한다.
39
71
  fh_peer_alias_candidates() { # $1=name → 후보를 줄 단위로
40
72
  printf '%s\n' "$1" "$1-dev" "$(printf '%s' "$1" | tr '_' '-')"
@@ -74,13 +106,33 @@ EOF
74
106
  printf 'peer-resolve: 프로젝트 홈이 없다: %s (전제 파손이지 «peer 없음» 이 아니다)\n' "$home" >&2
75
107
  return 3
76
108
  fi
77
- while IFS= read -r c; do
78
- [ -d "$home/$c" ] || continue
79
- case " $hits " in *" $home/$c "*) continue ;; esac
80
- hits="$hits $home/$c"; [ -n "$first" ] || first="$home/$c"
81
- done <<EOF
82
- $(fh_peer_alias_candidates "$n")
109
+ # 🟥 여기가 단일 소스로 닫힌 자리다. 옛 인라인 루프(별칭 3종 + 공백 구분자 dedup)를
110
+ # `fh_resolve_track_root` 대체한다. 출력 계약은 **아래에서 옛 표기로 복원**한다 —
111
+ # 소비자 4종과 레인은 여전히 rc(0/1/2/3) + stderr 문구를 읽는다.
112
+ local rr note
113
+ rr="$(fh_resolve_track_root "$n" "$home" dir)"
114
+ note="${rr##*|}"
115
+ case "$note" in
116
+ AMBIGUOUS:*)
117
+ # 옛 표기 복원: 라이브러리는 **후보 이름**을 콤마로 주고, 옛 메시지는 **절대경로**를
118
+ # «, » 로 이어 붙였다. 공백 없는 이름에서는 바이트 동일하다.
119
+ local _l="${note#AMBIGUOUS:}" _c _paths=""
120
+ while IFS= read -r _c; do
121
+ [ -n "$_c" ] || continue
122
+ if [ -n "$_paths" ]; then _paths="$_paths, $home/$_c"; else _paths="$home/$_c"; fi
123
+ done <<EOF
124
+ $(printf '%s' "$_l" | tr ',' '\n')
83
125
  EOF
126
+ printf 'peer-resolve: 한 이름에 레포 여럿 — 고르지 않는다: %s\n' "$_paths" >&2
127
+ return 2
128
+ ;;
129
+ UNRESOLVED)
130
+ first="" # 아래 공통 «부재» 분기로 떨어진다 (rc=1 · 옛 문구 그대로)
131
+ ;;
132
+ *)
133
+ first="${rr%|*}"
134
+ ;;
135
+ esac
84
136
  fi
85
137
 
86
138
  nh=$(printf '%s' "$hits" | wc -w | tr -d ' ')
@@ -76,6 +76,31 @@ CAP_SUBDIR=".claude/capabilities"
76
76
  # 테스트·다른 배치를 위해 `FH_CLUSTER_ROOTS`(콜론 구분 절대경로 목록)로 통째 대체할 수 있다.
77
77
  PROJECTS_HOME="${FH_PROJECTS_HOME:-$HOME/projects}"
78
78
 
79
+ # track 이름 → 레포 루트 해석은 scripts/fh_track_resolve.sh 가 단일 소스다 (2026-08-21 F-1).
80
+ # 🟥 **이 파일이 그 라이브러리의 원본이다** — 닫힌 별칭 3종 + 별칭 표면화 + 모호거부는 여기서
81
+ # 뽑아낸 것이고, 갈라져 있던 나머지 두 소비자(fh_session_load.sh · field_canon_preload.sh)가
82
+ # 같은 벌을 쓰게 되는 것이 이 배선의 전부다. 이 파일의 동작은 바뀌지 않는다.
83
+ # 라이브러리가 없으면 **수리 이전 동작**(= 아래 인라인 구현)으로 강등한다 — 다른 두 소비자와
84
+ # 같은 형태다. 새 실패 모드를 만들지 않는다.
85
+ _FH_TRLIB="$(cd "$(dirname "${BASH_SOURCE[0]:-$0}")" && pwd)/fh_track_resolve.sh"
86
+ # shellcheck source=scripts/fh_track_resolve.sh
87
+ [ -f "$_FH_TRLIB" ] && . "$_FH_TRLIB"
88
+ type fh_resolve_track_root >/dev/null 2>&1 || fh_resolve_track_root() {
89
+ local n="$1" root="$2" c hits="" first=""
90
+ for c in "$n" "$n-dev" "$(printf '%s' "$n" | tr '_' '-')"; do
91
+ [ -d "$root/$c" ] || continue
92
+ case " $hits " in *" $c "*) continue ;; esac
93
+ hits="$hits $c"; [ -n "$first" ] || first="$c"
94
+ done
95
+ local nh; nh=$(printf '%s' "$hits" | wc -w | tr -d ' ')
96
+ if [ "${nh:-0}" -gt 1 ]; then
97
+ printf '%s|AMBIGUOUS:%s' "$root/$n" "$(printf '%s' "$hits" | sed 's/^ //;s/ /,/g')"; return 0
98
+ fi
99
+ [ -n "$first" ] || { printf '%s|UNRESOLVED' "$root/$n"; return 0; }
100
+ [ "$first" = "$n" ] && { printf '%s|' "$root/$n"; return 0; }
101
+ printf '%s|alias:%s' "$root/$first" "$first"
102
+ }
103
+
79
104
  _die() { printf '❌ %s\n' "$1" >&2; exit "$RC_HARNESS"; }
80
105
 
81
106
  # ── `.cap` 파일 열거 — **단 하나의 소스** ────────────────────────────────────
@@ -150,20 +175,24 @@ _enumerate_harnesses() {
150
175
  # 🟥 그래도 «조용히 추측» 하지는 않는다. 별칭은 **닫힌 목록**이고, 맞은 별칭은
151
176
  # `_ALIAS_USED` 로 표면화된다. 여러 개가 동시에 맞으면 **고르지 않고 모호로 낸다** —
152
177
  # 둘 중 하나를 조용히 고르는 것이 이 파일이 반대하는 그 접힘이다.
178
+ # 🟥 이 함수는 이제 **어댑터**다 — 해석은 `fh_resolve_track_root`(단일 소스)가 하고, 여기서는
179
+ # 이 파일의 **표시 계약**으로 옮긴다. 두 계약이 미묘하게 다르기 때문에 조용히 통과시키지
180
+ # 않는다:
181
+ # 라이브러리 `alias:<c>` ↔ 여기 `경로 별칭 — alias:<c>` (OK 행의 DETAIL 로 직접 출력됨)
182
+ # 라이브러리 `UNRESOLVED` ↔ 여기 빈 문자열 (부재는 `_harness_status` 가
183
+ # UNREACHABLE 로 판정한다)
184
+ # `AMBIGUOUS:<...>` 는 양쪽이 같은 표기이고 `_harness_status` 의 `AMBIGUOUS:*` 분기가
185
+ # 그대로 받는다. 술어는 `dir` — 이 스캐너는 «디렉토리가 있나» 를 묻는다(git 레포 여부가
186
+ # 아니다). 그 차이는 의도이지 갈라짐이 아니다(라이브러리 헤더 참조).
153
187
  _resolve_root() { # $1=track 이름 → "<root>|<alias-note>"
154
- local n="$1" c hits="" first="" note=""
155
- for c in "$n" "$n-dev" "$(printf '%s' "$n" | tr '_' '-')"; do
156
- [ -d "$PROJECTS_HOME/$c" ] || continue
157
- case " $hits " in *" $c "*) continue ;; esac
158
- hits="$hits $c"; [ -n "$first" ] || first="$c"
159
- done
160
- local nh; nh=$(printf '%s' "$hits" | wc -w | tr -d ' ')
161
- if [ "${nh:-0}" -gt 1 ]; then
162
- printf '%s|AMBIGUOUS:%s' "$PROJECTS_HOME/$n" "$(printf '%s' "$hits" | sed 's/^ //;s/ /,/g')"
163
- return
164
- fi
165
- [ -n "$first" ] && [ "$first" != "$n" ] && note="경로 별칭 — alias:$first"
166
- printf '%s|%s' "$PROJECTS_HOME/${first:-$n}" "$note"
188
+ local rr note
189
+ rr="$(fh_resolve_track_root "$1" "$PROJECTS_HOME" dir)"
190
+ note="${rr#*|}"
191
+ case "$note" in
192
+ alias:*) printf '%s|경로 별칭 %s' "${rr%%|*}" "$note" ;;
193
+ UNRESOLVED) printf '%s|' "${rr%%|*}" ;;
194
+ *) printf '%s' "$rr" ;;
195
+ esac
167
196
  }
168
197
 
169
198
  # 한 하네스의 상태를 판정한다. 출력 한 줄: "<name>\t<STATUS>\t<n>\t<detail>"