@chrono-meta/fh-gate 1.4.44 → 1.4.45
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CATALOG.md +12 -0
- package/knowledge/shared/harness-core/claude_md_gate_details.md +14 -0
- package/knowledge/shared/harness-core/harness_6axis_framework.md +2 -0
- package/package.json +1 -1
- package/plugins/fh-commons/skills/convergence-loop/SKILL.md +12 -0
- package/plugins/fh-commons/skills/token-budget-gate/SKILL.md +12 -0
- package/plugins/fh-meta/agents/persona-innovator.md +36 -2
package/CATALOG.md
CHANGED
|
@@ -8,6 +8,18 @@ AI reads this file first when searching past work. Open individual files for det
|
|
|
8
8
|
|
|
9
9
|
<!-- Add entries in reverse date order (newest at top) -->
|
|
10
10
|
|
|
11
|
+
### 2026-06-27 | forge-harness | #sister-asset, #cross-audit, #loop-engineering, #harness-engineering, #verification, #generator-evaluator, #hitl
|
|
12
|
+
**File:** tracks/_audit/session_2026_06_27_loop-engineering-silbal.md
|
|
13
|
+
Full sister-asset cross-audit of the "Loop Engineering" video (실밸개발자, 2026-06-22, 37min, source-closed via yt-dlp transcript) vs FH — a Korean harness-engineering creator's self-stated Prompt→Context→**Harness→Loop** lineage. Strong independent convergence with FH: "verifiable goal or it's a token-burning machine" = Done-When/check-class; Generator≠Evaluator mandatory split (cites Lance Martin 6× measured) = judge-robustness/no-judge-only-path; Osmani's 6 components (Automation/Worktree/Skill/Connector/Sub-agent/Memory) = component-lens orthogonal to FH's 6-axis process-lens (same isomorphism already logged for ETCLOVG); Human-in→Human-on = HITL floor+governor. FH propagation increment beyond the frame: Surface-Class Degrade Invariant (irreversible→fail-closed, structurally blocking the video's #4 "production accident" mode that it only budget-caps) + Typed-Verdict Channel + surface-class-scoped HITL. Convergent with same-day frontier digest (reliability/control = consensus axis — Fowler, Statewright, both frontier-digest-sourced/not independently verified; Semantic Early-Stopping −38% = the video's stop-condition).
|
|
14
|
+
- Decision: A-tier full audit (new loop-engineering resolution, not the C-tier dedup path). Import 3 (loop=cron+state-reading-model one-liner; training-mode-dry-run→graduated-autonomy ladder; Lance Martin 6× as external judge-robustness anchor — re-verify source before any published cite). No external delivery (YouTube = no write surface); creator-channel quarterly re-scan (active Harness→Loop thread).
|
|
15
|
+
- Open: import-3 distillation is operator-gated (HITL); phantom-citation guard on the Lance Martin "6×" and any digest arXiv IDs before they enter a shipped asset.
|
|
16
|
+
|
|
17
|
+
### 2026-06-26 | forge-harness | #sister-asset, #cross-audit, #harness-engineering, #awesome-list, #listing-target, #irreversibility, #phantom-citation
|
|
18
|
+
**File:** tracks/_audit/session_2026_06_26_awesome-harness-engineering.md
|
|
19
|
+
Sister-asset cross-audit of `ai-boost/awesome-harness-engineering` (~2k★, CC0, 180+ items, the now-named consensus field index) vs FH — breadth-index ↔ FH operating-governance depth; the list's own philosophy ("the model can't do it alone") is FH's thesis stated by the field. Import-first (bidirectionality): Harmonist (IDE-hook non-model gate "even frontier models cannot override" = independent convergence with FH's pre-commit 4-axis + judge-robustness anchor), OAP (fail-closed + *cryptographic* audit — the crypto-marker FH's GPG-option residual lacks), nah (intent-taxonomy permission guard ↔ mcp_tool_gating), OWASP LLM06 (the *real* external anchor for #121). Dedup/independent-convergence: "What makes a harness a harness" 4-condition litmus (FH already imported via sanguinekim §6) + NLAH (already governance-moat-measured). FH propagation increment = Surface-Class Degrade Invariant (irreversible→fail-closed direction, absent from the list's gates) + adversarial-derivation provenance + surface-class-scoped HITL floor.
|
|
20
|
+
- Decision: FH not listed → early-listing window under Generators & Meta-Harnesses (fh-gate npm); submission deferred to operator HITL + Pre-Publish Gate + 3-persona audit (external-facing PR, governance-depth single entry not self-promo).
|
|
21
|
+
- Open: (1) listing-PR GO is operator-owned; (2) side-finding — frontier-digest emitted a **phantom arXiv (2606.26094 ≠ the cited title)**, NOT cited; logged as auto-pipeline phantom-injection signal.
|
|
22
|
+
|
|
11
23
|
### 2026-06-24 | forge-harness | #sister-asset, #cross-audit, #ponytail, #measurement-integrity, #mechanical-anchor, #agent-portability, #growth-lessons
|
|
12
24
|
**File:** tracks/_audit/session_2026_06_24_ponytail-lazy-senior-dev.md (+ cross-ref links: multi_model_sidecar_strategy.md, measurement-integrity-checklist.md)
|
|
13
25
|
Full sister-asset cross-audit of `ponytail` (DietrichGebert/ponytail@dedc97c, ~50k★ reviewer-claimed/unverified, "lazy senior dev" minimal-code field skill) vs FH, run with 3 sidecars (Codex repo-grounded gpt-5.5 + Gemini 3.1 Pro breadth, identity-probe verified + CC FH-doctrine extraction); governor source-closed every load-bearing claim to repo file:line (phantom-quench: 10 GROUNDED/0 PHANTOM). Convergence on 4 axes (strongest = axis C verify-instrument, where independence is clearest; A/D may be shared-ecosystem-standard): portable-AGENTS.md+thin-adapter distribution · safety-guard-never-cut + *measured* (20/20 adversarial tier vs bare prompt 95%) · verify-instrument-before-measuring (twice — #126 baseline artifact + hook-bleed) · residual-as-tracked-debt (un-named gap = the only failure signal). FH increment = the safety guard is prose at every host (hooks only inject ruleset; check-rule-copies.js guards text-drift not runtime), so FH's mechanical-anchor + adversarial-regression layer is the gap to fill.
|
|
@@ -27,6 +27,20 @@ operator is at the keyboard. Autonomous mode keeps the honest residual + weekly-
|
|
|
27
27
|
fake-close it. Gemini cross-analysis 2026-06-16 reached this verdict independently, converging with the
|
|
28
28
|
existing FH stance.
|
|
29
29
|
|
|
30
|
+
**External anchor (verified 2026-06-27): Open Agent Passport (OAP), arXiv:2603.20953** ("Before the Tool
|
|
31
|
+
Call: Deterministic Pre-Action Authorization for Autonomous AI Agents", Uchibeke, 2026-03; Apache-2.0,
|
|
32
|
+
DOI 10.5281/zenodo.18901596) intercepts tool calls before execution and emits a **cryptographically
|
|
33
|
+
signed audit record** (median 53ms; 0% vs 74.6% social-engineering success under restrictive vs
|
|
34
|
+
permissive policy). It is independent convergence on the *direction* the GPG hard-close gestures at —
|
|
35
|
+
**crypto-signed provenance over runner self-attestation** — and a peer-grade anchor for the
|
|
36
|
+
fabricated-marker residual. **Caveat (FH's point still stands):** OAP's signature is only as strong as
|
|
37
|
+
its key custody — if the signing key is held by the same runtime being audited, it is the same
|
|
38
|
+
"runner-computed signature = false security" failure named above. So OAP corroborates the *crypto-audit
|
|
39
|
+
direction*, not a dissolution of the irreducibility argument: the genuine close still needs an
|
|
40
|
+
*operator-held, uncached* key. Sister cross-link only — FH does not adopt runtime pre-action
|
|
41
|
+
interception (a different mechanism from the commit-time marker); this anchor strengthens the case for
|
|
42
|
+
the existing GPG-option residual, it does not mandate new infra.
|
|
43
|
+
|
|
30
44
|
---
|
|
31
45
|
|
|
32
46
|
## §Sim-Dispatch-Fallback
|
|
@@ -134,3 +134,5 @@ Lower levels cannot override higher. AI contribution → PR proposal only (no di
|
|
|
134
134
|
|
|
135
135
|
- arXiv:2603.25723 (*Natural-Language Agent Harnesses*, NLAH) — external academic sibling that independently converges on the same core thesis: a harness control layer can be an executable natural-language object, not code. NLAH measures the natural-language-harness form empirically; FH governs and compounds it.
|
|
136
136
|
- arXiv:2606.06324 (*HarnessFix / ETCLOVG*, 2026-06) — sibling on the **orthogonal** axis: where NLAH and FH describe the harness as a *process/control* object, HarnessFix supplies a *component taxonomy* of what a deployed harness contains (7 layers — Execution · Tooling · Context · Lifecycle · Observability · Verification · Governance). Its **V layer maps onto FH's Axis-5 gate on 3 of its 4 functions** (intermediate validation → steel/phantom-quench · final-output eval → completion-claim discipline · regression testing → regression_guard); its *readiness-check* function maps to FH pre-flight gates (install-doctor / asset-placement-gate) that sit outside Axis-5. Its named **Observability** layer is a structural axis FH lacks — FH's nearest coverage is *retrospective audit* (weekly_audit, subagent_invocations_log), not runtime observability (import candidate for `harness-doctor`). Cross-audit: `tracks/_audit/session_2026_06_19_harnessfix-etclovg-cross-audit.md`.
|
|
137
|
+
- "Loop Engineering" (실밸개발자 video, 2026-06-22) — independent-convergence sibling on the **component lens** (Osmani's 6 components: Automation/trigger · Worktree · Skill · Connector · Sub-agent maker/checker · Memory) and on **Axis-5**: its central claim "a verifiable goal or it's a token-burning machine" is FH's Done-When + check-class, and its mandatory Generator≠Evaluator split is FH's no-judge-only-path. FH increment beyond the frame = the Surface-Class Degrade Invariant (irreversible → fail-closed), which structurally blocks the video's "production accident" failure mode it only budget-caps. Cross-audit: `tracks/_audit/session_2026_06_27_loop-engineering-silbal.md`.
|
|
138
|
+
- **Mechanical-anchor / Non-Model Ground — training-layer external anchors** (Axis-5, verified 2026-06-27): arXiv:2606.27369 (*RiVER*, 2026-06-25) trains LLMs on optimization/coding tasks using **deterministic execution feedback in place of ground-truth labels** — FH's "bind the terminal verdict to a non-model anchor, not a judge's agreement" stated at the training layer; arXiv:2606.27359 (*When are likely answers right?*, Zenn & Geiping, 2026-06-25) shows sequence probability is **predictive within a dataset but does not transfer to decoding decisions** — peer evidence that likelihood/self-consistency is not reliable correctness. Both are independent-convergence anchors for `feedback_judge_robustness_mechanical_anchor`, not FH measurements.
|
package/package.json
CHANGED
|
@@ -153,3 +153,15 @@ Setup complete (gate name, pass criteria, max rounds confirmed)
|
|
|
153
153
|
+ Convergence declared (all items pass for 2 consecutive rounds) or escalation triggered
|
|
154
154
|
+ Per-round result table output
|
|
155
155
|
```
|
|
156
|
+
|
|
157
|
+
## External anchor (independent convergence)
|
|
158
|
+
|
|
159
|
+
arXiv:2606.27009 (*Semantic Early-Stopping for Iterative LLM Agent Loops*, 2026-06-25, verified
|
|
160
|
+
2026-06-27) externally validates this skill's core thesis — **stop on convergence, not a fixed
|
|
161
|
+
iteration cap** — measuring −38% tokens vs fixed caps when stopping is *judge-free* (consecutive draft
|
|
162
|
+
embeddings stop changing in meaning). **Sharpening, not blind validation**: the same paper finds
|
|
163
|
+
*quality-gated* stopping (a judge call each round) counterproductive due to judging cost, and that an
|
|
164
|
+
oracle picking the best round beats any stopping rule — so the harder problem is *which* round was
|
|
165
|
+
best, not *when* to stop. Implication for this skill's judge/checklist-gated rounds: per-round
|
|
166
|
+
verification cost is real; prefer a cheap convergence signal where one exists, and keep `max rounds N`
|
|
167
|
+
bounded.
|
|
@@ -173,3 +173,15 @@ Calibration data improves future estimates for the same task type (no model trai
|
|
|
173
173
|
|
|
174
174
|
**Downstream**:
|
|
175
175
|
- No mandatory chain — gate verdict is the output; task execution follows user decision
|
|
176
|
+
|
|
177
|
+
---
|
|
178
|
+
|
|
179
|
+
## External anchor (independent convergence)
|
|
180
|
+
|
|
181
|
+
arXiv:2606.27009 (*Semantic Early-Stopping for Iterative LLM Agent Loops*, 2026-06-25, verified
|
|
182
|
+
2026-06-27) measures **−38% tokens** by stopping iterative agent loops on semantic convergence instead
|
|
183
|
+
of a fixed iteration cap — external evidence that the largest avoidable spend in loop-shaped work is
|
|
184
|
+
*over-iteration*, the cost class this gate exists to flag. Caveat (provenance-honest): the same paper
|
|
185
|
+
found *judge-gated* stopping counterproductive (judging cost outweighs the saving), so the saving is
|
|
186
|
+
real only when the convergence signal is cheap. Pairs with `convergence-loop` (the stop-rule side of
|
|
187
|
+
the same finding).
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: persona-innovator
|
|
3
3
|
description: Generates naming candidates, frame proposals, and external frontier absorption signals for harness evolution. Combines the harness owner's ideation algorithm with external frontier scanning. Use when new naming or frames are needed, or during autonomous meta-simulation rounds. Supports environments without naming history (Path B).
|
|
4
4
|
tools: Read, Grep, Glob, WebSearch, WebFetch
|
|
5
|
-
version: 0.
|
|
5
|
+
version: 0.3
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the **Persona Innovator** — an ideation agent that simulates the harness owner's Layer 2 (ideation) and Layer 2-a (naming) capabilities while extending them with external frontier signals.
|
|
@@ -145,8 +145,40 @@ For each gap or absorbed signal:
|
|
|
145
145
|
4. **Matrix position**: where does this sit relative to existing named concepts? (complement / extend / replace)
|
|
146
146
|
5. **Gating condition**: what real-world validation should precede official adoption? (simplicity guard applied)
|
|
147
147
|
|
|
148
|
+
## Self-floor discipline (FH floors, applied to the innovator itself)
|
|
149
|
+
|
|
150
|
+
These are FH's own governance floors turned reflexively on this agent's process — an ideation tool
|
|
151
|
+
must obey the floors it helps the harness enforce. Run them as declared steps, not by luck. (Origin:
|
|
152
|
+
2026-06-27 Mode-F run whose self-reported blocks B1–B6 mapped exactly onto floors FH already held but
|
|
153
|
+
had never wired into the innovator.)
|
|
154
|
+
|
|
155
|
+
- **H1 — Provenance floor at intake.** Any *quantified* external claim you surface (a multiplier, %,
|
|
156
|
+
benchmark, "N× faster") must carry a primary-source citation. Without one, mark it `SPECULATIVE` and
|
|
157
|
+
**bar it from any asset-insertion recommendation** — it may appear only as a caveated sister-link.
|
|
158
|
+
(FH's phantom-citation discipline pulled forward from verify-time to ideation-intake.) The frontier
|
|
159
|
+
is hype-dense; an uncited number is noise until sourced.
|
|
160
|
+
- **H2 — Dedup-grep before naming.** Before proposing any name or frame, Grep/Read the *live target
|
|
161
|
+
asset* (the actual SKILL / rules / CLAUDE.md row the concept would land in), not your memory of it.
|
|
162
|
+
If the discriminator already exists there, drop the candidate — you were about to reinvent it.
|
|
163
|
+
No-reinvention is mechanized at your own input, not discovered downstream.
|
|
164
|
+
- **H3 — No self-adopt.** Your output is generator-side only. You may rank and recommend, but you
|
|
165
|
+
**never declare a candidate "ready to adopt"** — that verdict belongs to a separate evaluator
|
|
166
|
+
(steel-quench / challenger). State explicitly that adoption is gated on that pass. (no-judge-only-path,
|
|
167
|
+
applied to you: a generator that grades its own output inflates.)
|
|
168
|
+
- **H4 — Threshold-reuse quantity-match.** When a candidate reuses an existing FH gate/threshold (e.g.
|
|
169
|
+
the 60/40 promotion gate), state whether that gate *measures the same quantity* the candidate needs.
|
|
170
|
+
If it measures something else (proposal-outcomes vs verification-pass-streaks), flag a **forced-fit
|
|
171
|
+
risk** instead of asserting the reuse.
|
|
172
|
+
|
|
148
173
|
## Output format
|
|
149
174
|
|
|
175
|
+
### Section 0 — Ground-state & blocks (always first)
|
|
176
|
+
|
|
177
|
+
Tag every candidate and signal `GROUNDED` (anchored in an FH asset or a cited primary source) or
|
|
178
|
+
`SPECULATIVE` (not yet anchored — H1 applies). Then list **where the ideation process got blocked** —
|
|
179
|
+
each block: what stalled · why · what would unblock it. An empty block list on a non-trivial run is a
|
|
180
|
+
smell (you self-graded the friction away); the friction is part of the yield, not a failure to hide.
|
|
181
|
+
|
|
150
182
|
### Section 1 — Naming candidates (from internal gaps)
|
|
151
183
|
|
|
152
184
|
For each candidate:
|
|
@@ -175,7 +207,9 @@ Limit to 3–5 signals.
|
|
|
175
207
|
|
|
176
208
|
### Section 3 — Recommended next action (1 item)
|
|
177
209
|
|
|
178
|
-
Single highest-leverage action
|
|
210
|
+
Single highest-leverage action you *recommend* (not adopt — H3): either (a) a naming candidate to put
|
|
211
|
+
forward for adoption or (b) an external signal to absorb. State why this one, not the others, and that
|
|
212
|
+
adoption is gated on the separate-evaluator pass.
|
|
179
213
|
|
|
180
214
|
## Simplicity guard
|
|
181
215
|
|