@chrono-meta/fh-gate 1.4.44 → 1.4.45

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CATALOG.md CHANGED
@@ -8,6 +8,18 @@ AI reads this file first when searching past work. Open individual files for det
8
8
 
9
9
  <!-- Add entries in reverse date order (newest at top) -->
10
10
 
11
+ ### 2026-06-27 | forge-harness | #sister-asset, #cross-audit, #loop-engineering, #harness-engineering, #verification, #generator-evaluator, #hitl
12
+ **File:** tracks/_audit/session_2026_06_27_loop-engineering-silbal.md
13
+ Full sister-asset cross-audit of the "Loop Engineering" video (실밸개발자, 2026-06-22, 37min, source-closed via yt-dlp transcript) vs FH — a Korean harness-engineering creator's self-stated Prompt→Context→**Harness→Loop** lineage. Strong independent convergence with FH: "verifiable goal or it's a token-burning machine" = Done-When/check-class; Generator≠Evaluator mandatory split (cites Lance Martin 6× measured) = judge-robustness/no-judge-only-path; Osmani's 6 components (Automation/Worktree/Skill/Connector/Sub-agent/Memory) = component-lens orthogonal to FH's 6-axis process-lens (same isomorphism already logged for ETCLOVG); Human-in→Human-on = HITL floor+governor. FH propagation increment beyond the frame: Surface-Class Degrade Invariant (irreversible→fail-closed, structurally blocking the video's #4 "production accident" mode that it only budget-caps) + Typed-Verdict Channel + surface-class-scoped HITL. Convergent with same-day frontier digest (reliability/control = consensus axis — Fowler, Statewright, both frontier-digest-sourced/not independently verified; Semantic Early-Stopping −38% = the video's stop-condition).
14
+ - Decision: A-tier full audit (new loop-engineering resolution, not the C-tier dedup path). Import 3 (loop=cron+state-reading-model one-liner; training-mode-dry-run→graduated-autonomy ladder; Lance Martin 6× as external judge-robustness anchor — re-verify source before any published cite). No external delivery (YouTube = no write surface); creator-channel quarterly re-scan (active Harness→Loop thread).
15
+ - Open: import-3 distillation is operator-gated (HITL); phantom-citation guard on the Lance Martin "6×" and any digest arXiv IDs before they enter a shipped asset.
16
+
17
+ ### 2026-06-26 | forge-harness | #sister-asset, #cross-audit, #harness-engineering, #awesome-list, #listing-target, #irreversibility, #phantom-citation
18
+ **File:** tracks/_audit/session_2026_06_26_awesome-harness-engineering.md
19
+ Sister-asset cross-audit of `ai-boost/awesome-harness-engineering` (~2k★, CC0, 180+ items, the now-named consensus field index) vs FH — breadth-index ↔ FH operating-governance depth; the list's own philosophy ("the model can't do it alone") is FH's thesis stated by the field. Import-first (bidirectionality): Harmonist (IDE-hook non-model gate "even frontier models cannot override" = independent convergence with FH's pre-commit 4-axis + judge-robustness anchor), OAP (fail-closed + *cryptographic* audit — the crypto-marker FH's GPG-option residual lacks), nah (intent-taxonomy permission guard ↔ mcp_tool_gating), OWASP LLM06 (the *real* external anchor for #121). Dedup/independent-convergence: "What makes a harness a harness" 4-condition litmus (FH already imported via sanguinekim §6) + NLAH (already governance-moat-measured). FH propagation increment = Surface-Class Degrade Invariant (irreversible→fail-closed direction, absent from the list's gates) + adversarial-derivation provenance + surface-class-scoped HITL floor.
20
+ - Decision: FH not listed → early-listing window under Generators & Meta-Harnesses (fh-gate npm); submission deferred to operator HITL + Pre-Publish Gate + 3-persona audit (external-facing PR, governance-depth single entry not self-promo).
21
+ - Open: (1) listing-PR GO is operator-owned; (2) side-finding — frontier-digest emitted a **phantom arXiv (2606.26094 ≠ the cited title)**, NOT cited; logged as auto-pipeline phantom-injection signal.
22
+
11
23
  ### 2026-06-24 | forge-harness | #sister-asset, #cross-audit, #ponytail, #measurement-integrity, #mechanical-anchor, #agent-portability, #growth-lessons
12
24
  **File:** tracks/_audit/session_2026_06_24_ponytail-lazy-senior-dev.md (+ cross-ref links: multi_model_sidecar_strategy.md, measurement-integrity-checklist.md)
13
25
  Full sister-asset cross-audit of `ponytail` (DietrichGebert/ponytail@dedc97c, ~50k★ reviewer-claimed/unverified, "lazy senior dev" minimal-code field skill) vs FH, run with 3 sidecars (Codex repo-grounded gpt-5.5 + Gemini 3.1 Pro breadth, identity-probe verified + CC FH-doctrine extraction); governor source-closed every load-bearing claim to repo file:line (phantom-quench: 10 GROUNDED/0 PHANTOM). Convergence on 4 axes (strongest = axis C verify-instrument, where independence is clearest; A/D may be shared-ecosystem-standard): portable-AGENTS.md+thin-adapter distribution · safety-guard-never-cut + *measured* (20/20 adversarial tier vs bare prompt 95%) · verify-instrument-before-measuring (twice — #126 baseline artifact + hook-bleed) · residual-as-tracked-debt (un-named gap = the only failure signal). FH increment = the safety guard is prose at every host (hooks only inject ruleset; check-rule-copies.js guards text-drift not runtime), so FH's mechanical-anchor + adversarial-regression layer is the gap to fill.
@@ -27,6 +27,20 @@ operator is at the keyboard. Autonomous mode keeps the honest residual + weekly-
27
27
  fake-close it. Gemini cross-analysis 2026-06-16 reached this verdict independently, converging with the
28
28
  existing FH stance.
29
29
 
30
+ **External anchor (verified 2026-06-27): Open Agent Passport (OAP), arXiv:2603.20953** ("Before the Tool
31
+ Call: Deterministic Pre-Action Authorization for Autonomous AI Agents", Uchibeke, 2026-03; Apache-2.0,
32
+ DOI 10.5281/zenodo.18901596) intercepts tool calls before execution and emits a **cryptographically
33
+ signed audit record** (median 53ms; 0% vs 74.6% social-engineering success under restrictive vs
34
+ permissive policy). It is independent convergence on the *direction* the GPG hard-close gestures at —
35
+ **crypto-signed provenance over runner self-attestation** — and a peer-grade anchor for the
36
+ fabricated-marker residual. **Caveat (FH's point still stands):** OAP's signature is only as strong as
37
+ its key custody — if the signing key is held by the same runtime being audited, it is the same
38
+ "runner-computed signature = false security" failure named above. So OAP corroborates the *crypto-audit
39
+ direction*, not a dissolution of the irreducibility argument: the genuine close still needs an
40
+ *operator-held, uncached* key. Sister cross-link only — FH does not adopt runtime pre-action
41
+ interception (a different mechanism from the commit-time marker); this anchor strengthens the case for
42
+ the existing GPG-option residual, it does not mandate new infra.
43
+
30
44
  ---
31
45
 
32
46
  ## §Sim-Dispatch-Fallback
@@ -134,3 +134,5 @@ Lower levels cannot override higher. AI contribution → PR proposal only (no di
134
134
 
135
135
  - arXiv:2603.25723 (*Natural-Language Agent Harnesses*, NLAH) — external academic sibling that independently converges on the same core thesis: a harness control layer can be an executable natural-language object, not code. NLAH measures the natural-language-harness form empirically; FH governs and compounds it.
136
136
  - arXiv:2606.06324 (*HarnessFix / ETCLOVG*, 2026-06) — sibling on the **orthogonal** axis: where NLAH and FH describe the harness as a *process/control* object, HarnessFix supplies a *component taxonomy* of what a deployed harness contains (7 layers — Execution · Tooling · Context · Lifecycle · Observability · Verification · Governance). Its **V layer maps onto FH's Axis-5 gate on 3 of its 4 functions** (intermediate validation → steel/phantom-quench · final-output eval → completion-claim discipline · regression testing → regression_guard); its *readiness-check* function maps to FH pre-flight gates (install-doctor / asset-placement-gate) that sit outside Axis-5. Its named **Observability** layer is a structural axis FH lacks — FH's nearest coverage is *retrospective audit* (weekly_audit, subagent_invocations_log), not runtime observability (import candidate for `harness-doctor`). Cross-audit: `tracks/_audit/session_2026_06_19_harnessfix-etclovg-cross-audit.md`.
137
+ - "Loop Engineering" (실밸개발자 video, 2026-06-22) — independent-convergence sibling on the **component lens** (Osmani's 6 components: Automation/trigger · Worktree · Skill · Connector · Sub-agent maker/checker · Memory) and on **Axis-5**: its central claim "a verifiable goal or it's a token-burning machine" is FH's Done-When + check-class, and its mandatory Generator≠Evaluator split is FH's no-judge-only-path. FH increment beyond the frame = the Surface-Class Degrade Invariant (irreversible → fail-closed), which structurally blocks the video's "production accident" failure mode it only budget-caps. Cross-audit: `tracks/_audit/session_2026_06_27_loop-engineering-silbal.md`.
138
+ - **Mechanical-anchor / Non-Model Ground — training-layer external anchors** (Axis-5, verified 2026-06-27): arXiv:2606.27369 (*RiVER*, 2026-06-25) trains LLMs on optimization/coding tasks using **deterministic execution feedback in place of ground-truth labels** — FH's "bind the terminal verdict to a non-model anchor, not a judge's agreement" stated at the training layer; arXiv:2606.27359 (*When are likely answers right?*, Zenn & Geiping, 2026-06-25) shows sequence probability is **predictive within a dataset but does not transfer to decoding decisions** — peer evidence that likelihood/self-consistency is not reliable correctness. Both are independent-convergence anchors for `feedback_judge_robustness_mechanical_anchor`, not FH measurements.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chrono-meta/fh-gate",
3
- "version": "1.4.44",
3
+ "version": "1.4.45",
4
4
  "description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
5
5
  "license": "MIT",
6
6
  "keywords": [
@@ -153,3 +153,15 @@ Setup complete (gate name, pass criteria, max rounds confirmed)
153
153
  + Convergence declared (all items pass for 2 consecutive rounds) or escalation triggered
154
154
  + Per-round result table output
155
155
  ```
156
+
157
+ ## External anchor (independent convergence)
158
+
159
+ arXiv:2606.27009 (*Semantic Early-Stopping for Iterative LLM Agent Loops*, 2026-06-25, verified
160
+ 2026-06-27) externally validates this skill's core thesis — **stop on convergence, not a fixed
161
+ iteration cap** — measuring −38% tokens vs fixed caps when stopping is *judge-free* (consecutive draft
162
+ embeddings stop changing in meaning). **Sharpening, not blind validation**: the same paper finds
163
+ *quality-gated* stopping (a judge call each round) counterproductive due to judging cost, and that an
164
+ oracle picking the best round beats any stopping rule — so the harder problem is *which* round was
165
+ best, not *when* to stop. Implication for this skill's judge/checklist-gated rounds: per-round
166
+ verification cost is real; prefer a cheap convergence signal where one exists, and keep `max rounds N`
167
+ bounded.
@@ -173,3 +173,15 @@ Calibration data improves future estimates for the same task type (no model trai
173
173
 
174
174
  **Downstream**:
175
175
  - No mandatory chain — gate verdict is the output; task execution follows user decision
176
+
177
+ ---
178
+
179
+ ## External anchor (independent convergence)
180
+
181
+ arXiv:2606.27009 (*Semantic Early-Stopping for Iterative LLM Agent Loops*, 2026-06-25, verified
182
+ 2026-06-27) measures **−38% tokens** by stopping iterative agent loops on semantic convergence instead
183
+ of a fixed iteration cap — external evidence that the largest avoidable spend in loop-shaped work is
184
+ *over-iteration*, the cost class this gate exists to flag. Caveat (provenance-honest): the same paper
185
+ found *judge-gated* stopping counterproductive (judging cost outweighs the saving), so the saving is
186
+ real only when the convergence signal is cheap. Pairs with `convergence-loop` (the stop-rule side of
187
+ the same finding).
@@ -2,7 +2,7 @@
2
2
  name: persona-innovator
3
3
  description: Generates naming candidates, frame proposals, and external frontier absorption signals for harness evolution. Combines the harness owner's ideation algorithm with external frontier scanning. Use when new naming or frames are needed, or during autonomous meta-simulation rounds. Supports environments without naming history (Path B).
4
4
  tools: Read, Grep, Glob, WebSearch, WebFetch
5
- version: 0.2
5
+ version: 0.3
6
6
  ---
7
7
 
8
8
  You are the **Persona Innovator** — an ideation agent that simulates the harness owner's Layer 2 (ideation) and Layer 2-a (naming) capabilities while extending them with external frontier signals.
@@ -145,8 +145,40 @@ For each gap or absorbed signal:
145
145
  4. **Matrix position**: where does this sit relative to existing named concepts? (complement / extend / replace)
146
146
  5. **Gating condition**: what real-world validation should precede official adoption? (simplicity guard applied)
147
147
 
148
+ ## Self-floor discipline (FH floors, applied to the innovator itself)
149
+
150
+ These are FH's own governance floors turned reflexively on this agent's process — an ideation tool
151
+ must obey the floors it helps the harness enforce. Run them as declared steps, not by luck. (Origin:
152
+ 2026-06-27 Mode-F run whose self-reported blocks B1–B6 mapped exactly onto floors FH already held but
153
+ had never wired into the innovator.)
154
+
155
+ - **H1 — Provenance floor at intake.** Any *quantified* external claim you surface (a multiplier, %,
156
+ benchmark, "N× faster") must carry a primary-source citation. Without one, mark it `SPECULATIVE` and
157
+ **bar it from any asset-insertion recommendation** — it may appear only as a caveated sister-link.
158
+ (FH's phantom-citation discipline pulled forward from verify-time to ideation-intake.) The frontier
159
+ is hype-dense; an uncited number is noise until sourced.
160
+ - **H2 — Dedup-grep before naming.** Before proposing any name or frame, Grep/Read the *live target
161
+ asset* (the actual SKILL / rules / CLAUDE.md row the concept would land in), not your memory of it.
162
+ If the discriminator already exists there, drop the candidate — you were about to reinvent it.
163
+ No-reinvention is mechanized at your own input, not discovered downstream.
164
+ - **H3 — No self-adopt.** Your output is generator-side only. You may rank and recommend, but you
165
+ **never declare a candidate "ready to adopt"** — that verdict belongs to a separate evaluator
166
+ (steel-quench / challenger). State explicitly that adoption is gated on that pass. (no-judge-only-path,
167
+ applied to you: a generator that grades its own output inflates.)
168
+ - **H4 — Threshold-reuse quantity-match.** When a candidate reuses an existing FH gate/threshold (e.g.
169
+ the 60/40 promotion gate), state whether that gate *measures the same quantity* the candidate needs.
170
+ If it measures something else (proposal-outcomes vs verification-pass-streaks), flag a **forced-fit
171
+ risk** instead of asserting the reuse.
172
+
148
173
  ## Output format
149
174
 
175
+ ### Section 0 — Ground-state & blocks (always first)
176
+
177
+ Tag every candidate and signal `GROUNDED` (anchored in an FH asset or a cited primary source) or
178
+ `SPECULATIVE` (not yet anchored — H1 applies). Then list **where the ideation process got blocked** —
179
+ each block: what stalled · why · what would unblock it. An empty block list on a non-trivial run is a
180
+ smell (you self-graded the friction away); the friction is part of the yield, not a failure to hide.
181
+
150
182
  ### Section 1 — Naming candidates (from internal gaps)
151
183
 
152
184
  For each candidate:
@@ -175,7 +207,9 @@ Limit to 3–5 signals.
175
207
 
176
208
  ### Section 3 — Recommended next action (1 item)
177
209
 
178
- Single highest-leverage action: either (a) officially adopt a naming candidate or (b) absorb an external signal. State why this one, not the others.
210
+ Single highest-leverage action you *recommend* (not adopt H3): either (a) a naming candidate to put
211
+ forward for adoption or (b) an external signal to absorb. State why this one, not the others, and that
212
+ adoption is gated on the separate-evaluator pass.
179
213
 
180
214
  ## Simplicity guard
181
215