@chrono-meta/fh-gate 1.4.68 → 1.4.69

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -11,13 +11,13 @@
11
11
  "plugins": [
12
12
  {
13
13
  "name": "fh-meta",
14
- "version": "1.4.68",
14
+ "version": "1.4.69",
15
15
  "description": "Hub meta-operations toolkit — 34 skills + 7 agents. New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
16
16
  "source": "./plugins/fh-meta"
17
17
  },
18
18
  {
19
19
  "name": "fh-commons",
20
- "version": "1.4.68",
20
+ "version": "1.4.69",
21
21
  "description": "Project-agnostic utility skills — 4 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
22
22
  "source": "./plugins/fh-commons"
23
23
  }
package/CATALOG.md CHANGED
@@ -8,6 +8,18 @@ AI reads this file first when searching past work. Open individual files for det
8
8
 
9
9
  <!-- Add entries in reverse date order (newest at top) -->
10
10
 
11
+ ### 2026-07-24 (2) | forge-harness · forge-wiki · llmwiki-template | #sister-links, #full-gate, #frontier-digest-angle-rule, #query-refresh, #mapping, #light-harness, #graph-engineering
12
+ **File:** knowledge/shared/dialogue/memory_intent_recall.md · knowledge/shared/harness-core/field_harness_diagnostic.md · knowledge/shared/harness-core/multi_model_sidecar_strategy.md · plugins/fh-meta/skills/harness-doctor/SKILL.md · plugins/fh-meta/skills/frontier-digest/SKILL_detail.md · tracks/forge-wiki/ · tracks/llmwiki-template/
13
+ Second half of the 07-24 hub session (PR #174 + two field repos). **Sister anchors executed under the full 4-axis gate** (the four links #173 deferred): GRACE into memory_intent_recall (instruction-graph boundary spelled out) and harness-doctor Step 2-E (C-layer maintenance-path ledger, ETCLOVG-style anchor-only), the 7-criteria rubric as a considered-and-held note with two re-check triggers in field_harness_diagnostic, and the wsff RL-incentive causal anchor ("maintainability has no fast oracle") beside the harness-ceiling principle in multi_model_sidecar_strategy. **frontier-digest repaired on two axes**: the stale arXiv canonical query replaced ("AI software testing"→"LLM agent evaluation", refresh criterion written down) and an Angle rule added to the synthesis prompt closing the methodology-angle silent drop (fh_signal 07-24 #3) — verified by a blind Sonnet known-pair sim (dual-angle item surfaced, single-angle item not forced) plus a quench-challenger pass (HIGH 0 · MED 2 · LOW 3; MED+2 LOW repaired in-branch). **Two sibling repos mapped and light-harnessed** (forge-wiki = chamber run #9's first EMIT, public; llmwiki-template = its company-side private twin): tracks/ dirs, session rules, .claudeignore — env card and MCP gating skipped by judgment (residency / no external mount); both merged (forge-wiki #4, llmwiki-template #2), llmwiki's CLAUDE.md kept operator-local. Positioning decision recorded in both tracks: **a wiki that sits under the harness** (Harness ⊃ context/knowledge component) — "wiki" is the outward word, "harness knowledge layer" the architecture coordinate.
14
+ - Decision: first live measurement of the Angle rule and the new query is the next 09:00 digest run — manifest predictions pinned; npm republish proposed at close rather than auto-published; forge-wiki merged its hub-linked CLAUDE.md publicly by operator approval.
15
+ - Open: ETCLOVG citation's operator-local label retrofit; wsff follow-up (design FH's own deep-clarify before/after measurement) unstarted; Graph-layer chamber candidate maturing with multi-source anchors; company handoff (qasp-preflight→mate PR workflow, 07-25) carried to next session.
16
+
17
+ ### 2026-07-24 | forge-harness | #sister-asset, #grace, #context-quality-rubric, #graph-engineering, #judged-vs-measured, #frontier-digest
18
+ **File:** tracks/_audit/session_2026_07_24_grace-context-rubric.md · tracks/_meta/fh_signal_2026-07-24_frontier-digest.md
19
+ Bundled sister-asset triage closing the 10-day GRACE execution lag, plus the 07-24 digest chain. **GRACE (arXiv:2607.09175)** registered as an A-tier sister: typed-semantic-graph context maintenance with scoped verification (validate only the local typed neighborhood of modified nodes). Import candidates: neighborhood-first consistency checking for memory-hygiene/verify-bidirectional (a cost ceiling on the "re-grep everything" propagation rule), and the checkpoint-incremental-reconstruction shape as an external anchor for delta-update discipline. Category boundary held explicitly: GRACE's instruction graph ≠ the execution-orchestration Graph layer (chamber candidate) ≠ the memory recall graph — three different categories, not one "graph" asset. **Context-quality 7-criteria rubric (arXiv:2607.14275)** verdict: HOLD — scoring runs on ProofAgent-Harness multi-juror consensus, i.e. judged-not-mechanical, failing the digest's stated adoption precondition (mechanically scoreable on a known pair); re-check triggers named (deterministic rubric ships, or token-efficiency/tool-schema subset proves mechanizable). **Graph engineering** upgraded from n=1 video to multi-source convergence (AI Builder Club 4-question graph-vs-loop discriminator · TrueFoundry 7 governance requirements — delta for FH is typed run-identifier propagation only). **wsff.md (humanlayer, Dex) triaged same-pass** after full read: the digest's claimed absorbable unit — a "review burden hours→minutes measurement design" — does **not exist in the source** (its quantified content is the Faros AI correlation dataset, self-caveated by the author as "correlation signal, not smoking gun"; the front-loading benefit is an unmeasured personal assertion), so the import downgrades to motivation for FH to design its own deep-clarify before/after measurement. Its floor-vs-ceiling thesis dedups against the existing harness-ceiling principle; the genuinely new increments are the RL-incentive causal argument ("maintainability has no fast oracle, so RL cannot reward it" — the sharpest external anchor for why the HITL floor is structural, not transitional) and the Faros correlation numbers with caveat inherited.
20
+ - Decision: rubric lens NOT adopted (judge-only as shipped); wsff absorbed as two anchors, not doctrine (digest overclaim corrected in the audit record); knowledge/-file sister links deferred to a full-gate Mode D session (citation tokens trigger the substantive carve-out); effort A/B Window-2 resumption deferred to post-Saturday by operator.
21
+ - Open: 4 sister cross-links pending (memory_intent_recall · field_harness_diagnostic · harness-doctor 2-E · harness-ceiling anchor); digest→execution return-path gap still open as signal; GRACE ships no code — mechanism anchor only.
22
+
11
23
  ### 2026-07-22 (2) | forge-harness · qasp · pmh | #weekly-audit, #entrypoint-drift, #gate-fail-open, #union-silent-drop, #instrument-attribution, #npm-release
12
24
  **File:** tracks/_audit/weekly_audit_2026-07-22.md · AGENTS.md §Non-Claude runtimes item 4 · templates/regression_guard.sh · (qasp) src/api/ensemble.py · (pmh) AGENTS.md §Orchestration Gates
13
25
  Weekly audit (07-15~07-22) plus the three cross-repo fixes it surfaced. **Audit's largest finding was a card claim that was false**: the session card's red-flag "frontier-digest job not running — zero logs, zero output" did not survive a hand check (14/14 launchd fires, 12/14 outputs, that day's digest present at 09:03). The real defect is a 14.3% output-miss whose two instances both die on `Connection closed mid-response`, and whose 07-18 retry+watchdog fix engaged **neither mechanism** on its first failure day — recorded as unfixed, root cause not isolated. Card-vs-reality drift reached the N=3 recurrence threshold, so the prescription is a mechanical probe rather than another habit rule. **Entry-point drift** closed in both harnesses (FH PR #163, field meta-harness PR #25): a runtime-default governor rule had landed only in the Claude-native entry point, invisible to every other runtime. The target-tier blind sim rejected the first port — the *wording*, carried over verbatim, read as coercive to a cold third-party reader and was reproduced 3/3, once escalating to "I would flag this to the repo owner". Rewritten as a scope statement; converged 2/2. **Gate fail-open** (PR #165): Axis 1's pathspec omitted three asset classes the canonical rule declares covered, and the resulting not-checked state rendered as a green PASS; adversarial review then caught the fix's own over-blocking (a one-word prose edit produced a hard block) before it could train `--no-verify`. **Field harness UNION** shipped a silent-drop path: two divergent fence-unwrap implementations meant a response accepted by the single backend was discarded whole by the ensemble — in a component whose entire justification is not discarding findings. Published `@chrono-meta/fh-gate@1.4.66`.
@@ -207,3 +207,4 @@ constellations + the convergent/divergent split — all at prose scale, no new s
207
207
  - `memory-hygiene` (skill) — staleness pass; link-evolution (A-MEM) and the §D.8 store-side injection filter live here
208
208
  - `plugins/fh-meta/agents/persona-innovator` (the fh-meta agent) — the divergent-mode consumer this substrate unlocks
209
209
  - `knowledge/shared/rules/operational_adaptation.md` — the UAP is itself intent-recalled; shared tier-dependence rationale
210
+ - GRACE (arXiv:2607.09175) — external sister, *update-side* counterpart to this doc's recall-side 1-hop traversal: maintains agent instructions as a typed semantic graph and verifies edits **only within the modified node's local typed neighborhood** (instruction-graph substrate — mechanism-level analogy, not this memory graph). Mechanism anchor only (no code shipped); cross-audit: `tracks/_audit/session_2026_07_24_grace-context-rubric.md` (operator-local record)
@@ -23,6 +23,14 @@ is HITL — the diagnostic **proposes**, never auto-edits.
23
23
  | **Verdict/gate degrade** | `scripts/degrade_direction_scan.sh` | a field verdict/gate helper that degrades toward permissive (advisory pre-screen) |
24
24
  | **Loop-readiness** (황민호 loop-eng 5-question lens, 2026-07-10 — detail home: `loop_engineering.md`, incl. the FH loop inventory + design-time discipline) | *Loop-runtime axis — net-new vs Structure* (harness-doctor scans static form; this scans whether the path closes a loop). **Mechanical grep**: `/goal-quench`·`/loop` wiring present · check-class token declared. **Judged**: is the persisted state (card/handoff/memory) actually reloaded · is the declared check-class anchored, not judged-only · does the path halt. Done-When *presence* → see Structure row (no double-grep). **Adversarial pair** (for the judged sub-checks — decorrelated, behavior-vs-checklist): a target-tier blind sim that *runs* the path and observes whether it halts + persists, rather than re-checklisting it (the harness litmus shares this lens's axis, so it is a co-lens, not the adversary). | an agent path that *runs but doesn't loop*: no completion criterion (Done-When absent), judged-only validation with no anchor, no halt/budget guard (runaway/cost), or no state carried to the next run — the 5 questions (initiate · complete · validate · halt · persist) with 0 answers |
25
25
 
26
+ > **Considered-and-held, 7th-lens candidate (2026-07-24)**: the context-quality 7-criteria rubric
27
+ > (arXiv:2607.14275 — role clarity · guardrail coverage · instruction consistency · tool schema ·
28
+ > grounding sufficiency · injection hardening · token efficiency) is **not** adopted as a lens: its
29
+ > shipped scoring is ProofAgent-Harness multi-juror consensus — judge-only, failing the measured bar
30
+ > this table holds. Re-check triggers: ⓐ a deterministic scoring rubric ships upstream, or ⓑ the
31
+ > token-efficiency / tool-schema subset proves mechanically scoreable on a known pair. Do not
32
+ > re-propose without one of the two. Cross-audit: `tracks/_audit/session_2026_07_24_grace-context-rubric.md` (operator-local record).
33
+
26
34
  ## Output
27
35
 
28
36
  One ranked list, `M` (must-fix) / `S` (should-fix) / `R` (recommended) — same tiering as
@@ -136,6 +136,11 @@ FH is a **multi-runtime harness with explicit runtime authority**, not Claude-on
136
136
 
137
137
  A sidecar is recruited where it adds *decorrelated* value; its ceiling is still set by the governor — the
138
138
  harness lifts a model to its own ceiling, it does not move it ([[feedback_harness_ceiling_principle]]).
139
+ External causal anchor for *why* the ceiling sits in the weights (wsff.md, HumanLayer
140
+ `advanced-context-engineering-for-coding-agents` repo, Dex 2026, triaged 2026-07-24): RL rewards are pass/fail in seconds while architectural-decay costs surface over
141
+ weeks — *"maintainability has no fast oracle, so we can't reward for it during RL"* — hence the
142
+ human-review floor is structural, not a transitional patch. (Its Faros AI numbers are correlation-only;
143
+ the author's own caveat travels with any citation.)
139
144
  "Aggressive" Codex/Gemini use is bounded by the fit task-class above, never a blanket main-seat swap.
140
145
 
141
146
  ### Vendor-native harness — the main layer stays multi-CLI, never Copilot-consolidated
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@chrono-meta/fh-gate",
3
- "version": "1.4.68",
3
+ "version": "1.4.69",
4
4
  "description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
5
5
  "license": "MIT",
6
6
  "keywords": [
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "fh-commons",
3
- "version": "1.4.68",
3
+ "version": "1.4.69",
4
4
  "engines": {
5
5
  "claudeCode": ">=1.0.0"
6
6
  },
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "fh-meta",
3
- "version": "1.4.68",
3
+ "version": "1.4.69",
4
4
  "engines": {
5
5
  "claudeCode": ">=1.0.0"
6
6
  },
@@ -26,7 +26,10 @@ Collection criteria: score > 10, keyword-relevant items only. Max 15 items.
26
26
  ### arxiv
27
27
 
28
28
  ```bash
29
- for Q in "multi-agent LLM" "AI software testing" "context engineering agents"; do
29
+ # query refresh 2026-07-24: "AI software testing" (exact-phrase) went stale — newest hit was Sept 2024
30
+ # (07-24 run instrument note). Replaced with "LLM agent evaluation"; refresh again when a query's
31
+ # newest hit is >6 months old two runs in a row.
32
+ for Q in "multi-agent LLM" "LLM agent evaluation" "context engineering agents"; do
30
33
  curl -s --max-time 8 \
31
34
  "https://export.arxiv.org/api/query?search_query=all:${Q// /+}&max_results=2&sortBy=submittedDate&sortOrder=descending"
32
35
  done
@@ -75,6 +78,10 @@ Resolve a *video-harvest* capability via the Sidecar Engine Resolution Protocol
75
78
 
76
79
  ## §Synthesis-Prompt
77
80
 
81
+ > Angle-rule provenance: added 2026-07-24 (fh_signal 07-24 #3 — Bun-Rust methodology-angle silent
82
+ > drop; scope limited to collection-filtered items per Axis-2 challenger cost finding). Provenance
83
+ > lives here, outside the fenced prompt — the synthesis model gets only the bare rule.
84
+
78
85
  ### With Anthropic API
79
86
 
80
87
  ```
@@ -91,6 +98,12 @@ FH Context:
91
98
 
92
99
  [Insert collected data]
93
100
 
101
+ Angle rule: an item you reject on its primary angle (e.g. substrate/model/runtime news),
102
+ if it passed collection filtering, gets ONE explicit second look before discarding: does it
103
+ carry a separate methodology/harness angle (orchestration pattern, workflow scale,
104
+ operational practice)? If yes, judge that angle on its own merits; if no, discard silently —
105
+ no output line about the discard, and never force an angle that isn't there.
106
+
94
107
  Output format:
95
108
  ## This Week's Frontier Highlights (max 3)
96
109
  **[Title]** — FH connection point in one sentence
@@ -49,6 +49,12 @@ coverage lens** — not the paper's scoring model. A **coverage checklist, not a
49
49
  harness needs all seven (a read-only doc harness needs no Execution or Governance). For each layer ask
50
50
  "is there an asset covering it?"; surface gaps, let the human judge if each is real for *this* harness.
51
51
 
52
+ > **Orthogonal-sister ledger (same adoption rule — anchor only, no scoring model imported)**:
53
+ > GRACE (arXiv:2607.09175, cross-audit `tracks/_audit/session_2026_07_24_grace-context-rubric.md`,
54
+ > operator-local record) — typed-semantic-graph **scoped verification** for the **C (Context) layer's
55
+ > maintenance path**: edits validated only within the modified node's local typed neighborhood. Cite
56
+ > when a C-layer gap involves *how context updates are verified*, not just whether context assets exist.
57
+
52
58
  | Layer | Covers | Typical FH asset | If the gap is judged real → priority hint |
53
59
  |---|---|---|---|
54
60
  | **E** Execution | isolated/reproducible env, bounded autonomy | Agent-dispatch isolation · goal-quench budget | advisory |