@chrono-meta/fh-gate 1.4.47 → 1.4.48
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md
CHANGED
|
@@ -276,12 +276,28 @@ dispatch's own `model` parameter; the session model/plan-mode does **not** propa
|
|
|
276
276
|
> *self-development* is where tier depth measurably pays (design-increment finding), while operation
|
|
277
277
|
> does not. Sub-agent token costs are CC-visible in the session jsonl under `message.model`.
|
|
278
278
|
|
|
279
|
-
**Measured, not asserted** (
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
above-rubric *design* increments (developing the harness, not running it) — which is
|
|
283
|
-
Sonnet with **tier-floored dispatch** covering the depth-sensitive turns, and a
|
|
284
|
-
recommended only for harness-editing sessions.
|
|
279
|
+
**Measured, not asserted** (worked examples): on a blind rule-application battery, *operating* FH is
|
|
280
|
+
near model-flat — **every Claude tier measured scores 94–100%** (Fable, Opus 4.8, Sonnet 4.6 and 5,
|
|
281
|
+
Haiku 4.5); the few lost points are format discipline, never a trap or gate-class miss. The tiers
|
|
282
|
+
separate only on above-rubric *design* increments (developing the harness, not running it) — which is
|
|
283
|
+
why the default is Sonnet with **tier-floored dispatch** covering the depth-sensitive turns, and a
|
|
284
|
+
pinned stronger model is recommended only for harness-editing sessions.
|
|
285
|
+
|
|
286
|
+
This is stated as an **invariant, not a per-model leaderboard**. Two structural laws, neither of which a
|
|
287
|
+
new release overturns:
|
|
288
|
+
|
|
289
|
+
1. **Operation (구동) flattens across tiers** — the rules-in-context do the work, so every tier ceilings
|
|
290
|
+
on rule-application (Sonnet 5 tied Opus 4.8 at the battery ceiling in a 2026-07-03 replication).
|
|
291
|
+
2. **Depth (design increments) is tier-ordered, and the order is fixed *within a generation*** — a lower
|
|
292
|
+
tier never overtakes the higher of the **same** generation (tiers are priced to be worth it, so the
|
|
293
|
+
vendor keeps them ordered). *Across* generations a newer lower-tier model can surpass an older
|
|
294
|
+
higher-tier one (Sonnet 5 ≥ Opus 4.8 on operation is exactly this cross-generation case) — but the
|
|
295
|
+
current top tier of any generation still wins its own depth turns.
|
|
296
|
+
|
|
297
|
+
So the doctrine is permanent, not perishable: **default to the mid tier for operation; escalate to the
|
|
298
|
+
current top tier for depth.** Re-measurement is warranted only when a new model becomes a field-main
|
|
299
|
+
*candidate* (a one-time cross-generation threshold check), never to re-confirm same-generation tier order
|
|
300
|
+
— that is guaranteed by design. Details + dated runs: `docs/OUTPUT_EVIDENCE.md` §Validation signals.
|
|
285
301
|
|
|
286
302
|
If you use external CLIs (Gemini, Codex, `gh copilot`) as sidecars, their costs are billed to their own quota and not visible in CC's token display.
|
|
287
303
|
|
|
@@ -292,7 +308,7 @@ the canary / cheap-breadth rungs only:
|
|
|
292
308
|
|
|
293
309
|
| Tier | Spec | Runs locally | What it buys |
|
|
294
310
|
|---|---|---|---|
|
|
295
|
-
| **Minimum** | anything that runs Claude Code | nothing | full methodology + gates; operating FH is ~model-flat (
|
|
311
|
+
| **Minimum** | anything that runs Claude Code | nothing | full methodology + gates; operating FH is ~model-flat across every tier tested (94–100%) |
|
|
296
312
|
| **Recommended** | laptop-class, ~16GB RAM | one 8B-class quantized model (e.g. an 8B / small Gemma) | a token-free **floor canary** (pre-screen before a metered sim) · offline triage · a cheap-breadth panel arm |
|
|
297
313
|
| **Optional (heavy)** | ~24GB VRAM GPU | a 27–32B model | a *stronger* decorrelation canary |
|
|
298
314
|
|
package/package.json
CHANGED
|
@@ -83,6 +83,13 @@ When AI generates artifacts without reading the source, those artifacts look lik
|
|
|
83
83
|
| **Partial reading** | Source partially read, rest filled in with inference | A |
|
|
84
84
|
| **Reconstruction contamination** | Source was read but LLM modified values/conditions during paraphrase | A |
|
|
85
85
|
|
|
86
|
+
> **External anchor**: *package hallucination* — an LLM emitting a non-existent or invalid
|
|
87
|
+
> package name, which an attacker can then register under the hallucinated name as a supply-chain
|
|
88
|
+
> vector — is a documented instance of the Phantom Claim class (arXiv:2607.02052, *Mitigating
|
|
89
|
+
> Package Hallucinations in Large Language Models via Model Editing*). That work mitigates the failure inside the
|
|
90
|
+
> model; phantom-quench catches the same class at the artifact surface by back-tracing each
|
|
91
|
+
> referenced name to a declared source.
|
|
92
|
+
|
|
86
93
|
---
|
|
87
94
|
|
|
88
95
|
## Execution Steps
|
|
@@ -112,6 +112,12 @@ Treat the adapter output as the isolated challenger result for Wave 1. This pres
|
|
|
112
112
|
|
|
113
113
|
> **Import origin** (sister-asset cross-audit 2026-06-14, `tracks/_audit/session_2026_06_14_official-plugins-cross-audit.md`): skill-creator + plugin-dev/skill-reviewer measure trigger accuracy **empirically**; FH's skill gate ("3+ NL triggers") and steel-quench's trigger-collision attack are **judged**, not measured. This probe converts that one verdict to **measured** — the mechanical-anchor discipline (a terminal trigger verdict should rest on a count, not an inference: the W4-4 question applied to the skill's own description).
|
|
114
114
|
|
|
115
|
+
> **External frame**: treating a trigger/prompt phrase set as a first-class artifact that needs a
|
|
116
|
+
> coverage criterion analog to code coverage is the position argued in arXiv:2607.02057, *Prompt
|
|
117
|
+
> Coverage Adequacy*. FH's instrument here is a fire-count over should-fire / near-miss phrases,
|
|
118
|
+
> not that paper's attention-based test-suite coverage — the anchor grounds the coverage-adequacy
|
|
119
|
+
> framing, not a drop-in metric.
|
|
120
|
+
|
|
115
121
|
**Fires only when** `artifact_type = skill_md` (Step 0.3 canonical enum, `tpa_schema.md` — i.e. a SKILL.md) AND the **trigger surface** changed. *Trigger surface* = exactly the `description:` YAML field **plus** the `## Triggers` / `## Trigger Phrases` section (and nothing else — a body-wave or procedure edit with both of those untouched does **not** fire it). Any other artifact has no trigger surface — note `Step 0.5: skipped (not a skill_md trigger change)` and proceed.
|
|
116
122
|
|
|
117
123
|
**Procedure** (prose-scale — FH routes + governs, it does **not** rebuild skill-creator's eval engine; no Python harness, no vector store):
|