@chrono-meta/fh-gate 1.4.66 → 1.4.67
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +5 -0
- package/CATALOG.md +6 -0
- package/knowledge/shared/harness-core/self_evolution_routine.md +8 -0
- package/package.json +1 -1
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
|
@@ -11,13 +11,13 @@
|
|
|
11
11
|
"plugins": [
|
|
12
12
|
{
|
|
13
13
|
"name": "fh-meta",
|
|
14
|
-
"version": "1.4.
|
|
14
|
+
"version": "1.4.67",
|
|
15
15
|
"description": "Hub meta-operations toolkit — 34 skills + 7 agents. New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
16
16
|
"source": "./plugins/fh-meta"
|
|
17
17
|
},
|
|
18
18
|
{
|
|
19
19
|
"name": "fh-commons",
|
|
20
|
-
"version": "1.4.
|
|
20
|
+
"version": "1.4.67",
|
|
21
21
|
"description": "Project-agnostic utility skills — 4 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
|
|
22
22
|
"source": "./plugins/fh-commons"
|
|
23
23
|
}
|
package/AGENTS.md
CHANGED
|
@@ -115,6 +115,11 @@ So four things that govern behavior are not going to reach you on their own. Rea
|
|
|
115
115
|
`.claude/rules/fh_4axis_gate.md` — **open it directly**; nothing will load it for you. The commit is
|
|
116
116
|
hard-blocked by `templates/.git-hooks/pre-commit` regardless of runtime, so skipping the read does not
|
|
117
117
|
skip the gate — it just means you meet the block without knowing what it wants.
|
|
118
|
+
⚠️ If you run `templates/regression_guard.sh` (Axis 1) yourself, **`exit 0` means PASS *or* SKIP** —
|
|
119
|
+
SKIP being "no staged file matched the gate's pathspec", which is *not checked*, not *checked and
|
|
120
|
+
clean*. Read stdout for `REGRESSION_GUARD_RESULT=skip` to tell them apart. Judging by exit code
|
|
121
|
+
alone lets an unexamined change report as a pass (measured 2026-07-22: that is exactly what the
|
|
122
|
+
commit hook did until it was fixed).
|
|
118
123
|
2. **Company residency is absolute** (CLAUDE.md §Field-Harness Diagnostic): raw company source, secrets,
|
|
119
124
|
hostnames, internal repo/asset names, stack traces, and unredacted findings **never leave the local
|
|
120
125
|
machine** — not to an external *or same-family* cloud model, not through a browser/API tool, not into a
|
package/CATALOG.md
CHANGED
|
@@ -8,6 +8,12 @@ AI reads this file first when searching past work. Open individual files for det
|
|
|
8
8
|
|
|
9
9
|
<!-- Add entries in reverse date order (newest at top) -->
|
|
10
10
|
|
|
11
|
+
### 2026-07-22 (2) | forge-harness · qasp · pmh | #weekly-audit, #entrypoint-drift, #gate-fail-open, #union-silent-drop, #instrument-attribution, #npm-release
|
|
12
|
+
**File:** tracks/_audit/weekly_audit_2026-07-22.md · AGENTS.md §Non-Claude runtimes item 4 · templates/regression_guard.sh · (qasp) src/api/ensemble.py · (pmh) AGENTS.md §Orchestration Gates
|
|
13
|
+
Weekly audit (07-15~07-22) plus the three cross-repo fixes it surfaced. **Audit's largest finding was a card claim that was false**: the session card's red-flag "frontier-digest job not running — zero logs, zero output" did not survive a hand check (14/14 launchd fires, 12/14 outputs, that day's digest present at 09:03). The real defect is a 14.3% output-miss whose two instances both die on `Connection closed mid-response`, and whose 07-18 retry+watchdog fix engaged **neither mechanism** on its first failure day — recorded as unfixed, root cause not isolated. Card-vs-reality drift reached the N=3 recurrence threshold, so the prescription is a mechanical probe rather than another habit rule. **Entry-point drift** closed in both harnesses (FH PR #163, field meta-harness PR #25): a runtime-default governor rule had landed only in the Claude-native entry point, invisible to every other runtime. The target-tier blind sim rejected the first port — the *wording*, carried over verbatim, read as coercive to a cold third-party reader and was reproduced 3/3, once escalating to "I would flag this to the repo owner". Rewritten as a scope statement; converged 2/2. **Gate fail-open** (PR #165): Axis 1's pathspec omitted three asset classes the canonical rule declares covered, and the resulting not-checked state rendered as a green PASS; adversarial review then caught the fix's own over-blocking (a one-word prose edit produced a hard block) before it could train `--no-verify`. **Field harness UNION** shipped a silent-drop path: two divergent fence-unwrap implementations meant a response accepted by the single backend was discarded whole by the ensemble — in a component whose entire justification is not discarding findings. Published `@chrono-meta/fh-gate@1.4.66`.
|
|
14
|
+
- Decision: port constraints across entry points, never wording — the audience's trust relationship changes with the location. Reversible surfaces report SKIP distinctly from PASS rather than switching to a blocking exit code. Company-derived doc conflicts resolve toward the operator-approved integration branch, never toward a feature branch's older state.
|
|
15
|
+
- Open: the failed single-vs-UNION detection ratio stays UNCALIBRATED — a demonstrated drop path and the "anchor absorbs the gain" hypothesis produce the same number, and no counter existed to separate them; re-measure with the new instrumentation. Skip-vs-pass distinction is wired into the commit hook only; three prose-level consumers still judge by exit code. Card-drift probe and the retry/watchdog reproduction harness remain unbuilt.
|
|
16
|
+
|
|
11
17
|
### 2026-07-22 | forge-harness | #intent-marshaling, #doctrine, #purpose-organization, #leader-briefing, #conference, #pre-registration
|
|
12
18
|
**File:** knowledge/shared/harness-core/intent_marshaling_general_work.md · CLAUDE.md §Intent Marshaling · (companion store) leader/TF briefing pair · handoff §5-§7
|
|
13
19
|
Two operator insights forged into doctrine. **Intent Marshaling** (PR #161, mirrored to the field meta-harness as its PR #23): the runtime twin of intent-machinization — a work-shaped request in plain language triggers a mechanical capability scan (trust tiers carried in scan output), a one-line compose proposal, then run-first execution; gap declarable only by citing the scan; no new gates (install/persist/outward actions route to existing ones). Verified by Sonnet known-pair sim (2/2 separated) + codex cross-family R1 4/4-confirmed findings (2 HIGH: non-FH ask-tier auto-run hole, per-action reversibility fail-open) fixed to R2 CONVERGED. **Leader-judgment briefing pair** (companion store, operator-approved): QA-team edition (4 judgment axes, measured/pending boundary, act-by-act glossary from operator definitions) + org-TF edition (domain-agnostic layers as protagonist, field harness as the n=1 evidence case, method-stack 8/8 as the domain-agnostic quantitative anchor) — both persona-audited (SHIP_AFTER_M, all findings applied; the audits caught the docs' own optimism twice, which became a self-evidencing section). Conference talk submitted (title A) with **metric pre-registration** pinned in the handoff (calibration / field / org metric sets + before-baseline warning + freeze-timeline insurance).
|
|
@@ -106,6 +106,14 @@ If `steel-quench`/`phantom-quench` are unavailable in the routine session, note
|
|
|
106
106
|
`Axis N: skipped (skill unavailable)` — Axis 1 PASS alone unblocks a *draft* PR (Axes 2–3 are the
|
|
107
107
|
operator's residual at merge review).
|
|
108
108
|
|
|
109
|
+
> ⚠️ **"Axis 1 PASS" 는 종료코드로 판정하지 마라.** `regression_guard.sh` 의 `exit 0` 은
|
|
110
|
+
> **PASS 와 SKIP 을 둘 다** 뜻한다 — SKIP 은 "게이트 pathspec 에 걸린 파일이 없었다"이지
|
|
111
|
+
> "검사했고 괜찮다"가 아니다. 무인 루틴이 종료코드만 보면 **미검사가 draft PR 을 unblock 한다.**
|
|
112
|
+
> 구분: stdout 에 `REGRESSION_GUARD_RESULT=skip` 이 있으면 SKIP 이다 — 그 경우 Axis 1 은
|
|
113
|
+
> *통과가 아니라 미실행*이므로 Axes 2–3 과 같은 운영자 잔여로 올려라.
|
|
114
|
+
> (2026-07-22 실측: `pre-commit` 이 정확히 이 혼동을 일으켜 AGENTS.md 변경이 `✅ PASS` 를
|
|
115
|
+
> 받고 지나갔다. `pre-commit` 은 배선 완료 · 이 루틴을 포함한 나머지 소비자는 **미배선 잔여**.)
|
|
116
|
+
|
|
109
117
|
**Skip-visibility at the handoff (load-bearing for honest HITL escalation).** When Axis 2
|
|
110
118
|
(challenger / steel-quench) is skipped, the adversarial check did not run autonomously — it becomes
|
|
111
119
|
the *merger's* responsibility. An intelligent hand-off must make that visible at the hand-off point,
|
package/package.json
CHANGED