@chrono-meta/fh-gate 1.4.42 → 1.4.43
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md
CHANGED
|
@@ -68,7 +68,9 @@ Agents in this registry belong to the **Automation layer**. Skills (in `plugins/
|
|
|
68
68
|
>
|
|
69
69
|
> When unsure, treat raw / observational / operator-specific material as **private-first** and promote only the polished result to public. (Concrete per-operator bindings — exact companion-store path, sync mechanism — live in the operator's local config, not here.)
|
|
70
70
|
|
|
71
|
-
> **Multi-model sidecar (validated)**: Any FH user can delegate to other models via sidecar — Gemini CLI, OpenAI/Codex CLI, or Copilot CLI's model catalog — invoked with `Bash` from within the Claude Code session. FH is the orchestrating harness; the sidecar is a routing/access layer (not a second harness — different layer entirely). Validated empirically: `echo "prompt" | gemini` works inside a CC session and produces usable output. Sidecar calls are Bash invocations, not agent dispatches — they bypass this registry and are coordinated inline by the skill. Capability routing matters too: Gemini/Antigravity is the natural multimodal sidecar, while a Codex
|
|
71
|
+
> **Multi-model sidecar (validated)**: Any FH user can delegate to other models via sidecar — Gemini CLI, OpenAI/Codex CLI, or Copilot CLI's model catalog — invoked with `Bash` from within the Claude Code session. FH is the orchestrating harness; the sidecar is a routing/access layer (not a second harness — different layer entirely). Validated empirically: `echo "prompt" | gemini` works inside a CC session and produces usable output. Sidecar calls are Bash invocations, not agent dispatches — they bypass this registry and are coordinated inline by the skill. Capability routing matters too: Gemini/Antigravity is the natural breadth/multimodal sidecar, while Codex's primary cast is the **repo-grounded audit** sidecar (file reads · grep/source-close · diff & patch · gate execution · phantom/backtrace) — **not** discovery/design-depth; a Codex session with Browser/Chrome connectors mounted can additionally take live web-flow automation as a capability-routed handoff. In a local FH workspace that pairs the public methodology mirror with a private companion store (the `*-be` pattern), route by workspace capability while preserving each repository's ownership boundary. See `knowledge/shared/harness-core/multi_model_sidecar_strategy.md §Runtime Authority` for the authority model and the full pattern.
|
|
72
|
+
|
|
73
|
+
> **Runtime authority — hard stop line (Codex / non-Claude runtimes):** your findings are **evidence candidates, not terminal verdicts**. They are not final until the governor source-closes them against a **mechanical anchor** (a local file hit · a literal source span · a passing check) — **never governor agreement alone**. You are a capability-routed **sidecar**, not a co-governor: there is one explicit governor per context. Full doctrine: `knowledge/shared/harness-core/multi_model_sidecar_strategy.md §Runtime Authority`.
|
|
72
74
|
|
|
73
75
|
---
|
|
74
76
|
|
package/CATALOG.md
CHANGED
|
@@ -8,6 +8,12 @@ AI reads this file first when searching past work. Open individual files for det
|
|
|
8
8
|
|
|
9
9
|
<!-- Add entries in reverse date order (newest at top) -->
|
|
10
10
|
|
|
11
|
+
### 2026-06-24 | forge-harness | #sister-asset, #cross-audit, #ponytail, #measurement-integrity, #mechanical-anchor, #agent-portability, #growth-lessons
|
|
12
|
+
**File:** tracks/_audit/session_2026_06_24_ponytail-lazy-senior-dev.md (+ cross-ref links: multi_model_sidecar_strategy.md, measurement-integrity-checklist.md)
|
|
13
|
+
Full sister-asset cross-audit of `ponytail` (DietrichGebert/ponytail@dedc97c, ~50k★ reviewer-claimed/unverified, "lazy senior dev" minimal-code field skill) vs FH, run with 3 sidecars (Codex repo-grounded gpt-5.5 + Gemini 3.1 Pro breadth, identity-probe verified + CC FH-doctrine extraction); governor source-closed every load-bearing claim to repo file:line (phantom-quench: 10 GROUNDED/0 PHANTOM). Convergence on 4 axes (strongest = axis C verify-instrument, where independence is clearest; A/D may be shared-ecosystem-standard): portable-AGENTS.md+thin-adapter distribution · safety-guard-never-cut + *measured* (20/20 adversarial tier vs bare prompt 95%) · verify-instrument-before-measuring (twice — #126 baseline artifact + hook-bleed) · residual-as-tracked-debt (un-named gap = the only failure signal). FH increment = the safety guard is prose at every host (hooks only inject ruleset; check-rule-copies.js guards text-drift not runtime), so FH's mechanical-anchor + adversarial-regression layer is the gap to fill.
|
|
14
|
+
- Decision: import 3 (platform-native table, --selftest dogfood example, behavior-grader sharpening for prompt-regression); propagate 3 to ponytail (mechanical-anchor option, adversarial regression on minimized diffs, reps≥3 on safety) via humble issue after persona audit; growth = a **field-skill spin-out that feeds the hub**, NOT re-pointing the meta-harness toward virality (reference-asset identity held; missing lever = a visible before/after).
|
|
15
|
+
- Open: external #3 delivery gated on 3+ persona × 4-axis audit + operator GO.
|
|
16
|
+
|
|
11
17
|
### 2026-06-14 | forge-harness | #crucible-mode, #total-immersion-absorption, #design-decision-lens, #completion-claim-discipline, #self-forge, #sister-asset
|
|
12
18
|
**File:** knowledge/shared/harness-core/crucible_mode.md + harness_design_decision_lens.md + harness_6axis_framework.md (Completion-claim discipline) + tracks/_audit/session_2026_06_14_wikidocs-deep-sweep.md
|
|
13
19
|
Content-level deep cross-audit of two wikidocs sister books (19689 백과사전 / 19736 Allen 멀티에이전트) via live-surface Playwright ingest + Gemini/Codex debate-loop + governor source-close, then **absorbed every candidate that passed the identity gate** (FH-identity-preserving + positively-expandable). Three assets: (1) `harness_design_decision_lens.md` — the 7 architectural-bet decisions as an orthogonal companion to the 6-axis lifecycle (only net-new = the framing; rest ALREADY-HAVE, honestly marked); (2) 6-axis **Completion-claim discipline** — a "done" claim must carry evidence + failure-checks-run + residual risk, non-vacuous; (3) **`crucible_mode.md`** — names the total-immersion absorption *stance* (throw the whole corpus in, melt under adversarial heat, keep only what bonds to an **unmeltable adamantium core**; rejections are boundary-defining). Each absorption was itself put through the crucible (quench-challenger + persona-auditor + Sonnet blind sim) — the crucible doc's own quench caught 3 of its defects (incl. a phantom worked-instance claim) before commit.
|
|
@@ -45,6 +45,11 @@ mechanical assertion**: the measurement harness records the *verified* model ide
|
|
|
45
45
|
the requested slug. Prose discipline is sufficient for internal dogfooding; a published claim earns the
|
|
46
46
|
mechanical log.
|
|
47
47
|
|
|
48
|
+
> **External dogfood (a second, field-layer instance — n=1 external, a signal not a settled frontier):**
|
|
49
|
+
> the sister skill `ponytail` ships a runnable instance of the precondition behind all three modes —
|
|
50
|
+
> a `--selftest` that proves each instrument (`good===true && bad===false`) before any API spend, and
|
|
51
|
+
> two caught instrument contaminations. Detail + pinned citations: `tracks/_audit/session_2026_06_24_ponytail-lazy-senior-dev.md` §2-C (single source).
|
|
52
|
+
|
|
48
53
|
---
|
|
49
54
|
|
|
50
55
|
**Origin** (2026-06-22 harvest-loop): three failure modes observed across the-bible L2 model panel
|
|
@@ -648,3 +648,4 @@ Missing any layer = compression risk. (Path conventions adapt per project — se
|
|
|
648
648
|
- A sister-harness `sidecar-orchestrator` SKILL.md (2026-06-01) — gh copilot + corporate endpoint + 3-tier fallback + 3-layer persistence
|
|
649
649
|
- arXiv:2605.26302 AgingBench — compression aging defense rationale
|
|
650
650
|
- `hybrid_orchestration_architecture_roadmap.md` — proposed (not-yet-implemented) architecture direction that would generalize this sidecar strategy into a hybrid orchestration engine
|
|
651
|
+
- **Sister asset** — `ponytail` (github DietrichGebert/ponytail@dedc97c; "lazy senior dev" minimal-code field skill, 14-host portable `AGENTS.md` + thin adapters) converges on the portable-`AGENTS.md`-as-entrypoint + thin-adapter distribution this doc codifies (portable-AGENTS.md is itself a recognized 2026 standard — convergence, not provably independent derivation). Cross-audit: `tracks/_audit/session_2026_06_24_ponytail-lazy-senior-dev.md`
|
package/package.json
CHANGED