@chrono-meta/fh-gate 1.4.69 → 1.4.71
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +3 -3
- package/CATALOG.md +15 -0
- package/CLAUDE.md +1 -1
- package/README.md +2 -1
- package/knowledge/shared/harness-core/loop_engineering.md +27 -1
- package/package.json +1 -1
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +2 -2
- package/plugins/fh-meta/skills/asset-placement-gate/SKILL.md +15 -4
- package/plugins/fh-meta/skills/context-doctor/SKILL.md +11 -0
- package/plugins/fh-meta/skills/dialogue-harvest/SKILL.md +223 -0
- package/plugins/fh-meta/skills/dialogue-harvest/calibration_pair.md +68 -0
- package/plugins/fh-meta/skills/frontier-digest/SKILL_detail.md +29 -6
- package/templates/local_fh_context.md +1 -1
|
@@ -11,13 +11,13 @@
|
|
|
11
11
|
"plugins": [
|
|
12
12
|
{
|
|
13
13
|
"name": "fh-meta",
|
|
14
|
-
"version": "1.4.
|
|
15
|
-
"description": "Hub meta-operations toolkit —
|
|
14
|
+
"version": "1.4.71",
|
|
15
|
+
"description": "Hub meta-operations toolkit — 35 skills + 7 agents. New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
16
16
|
"source": "./plugins/fh-meta"
|
|
17
17
|
},
|
|
18
18
|
{
|
|
19
19
|
"name": "fh-commons",
|
|
20
|
-
"version": "1.4.
|
|
20
|
+
"version": "1.4.71",
|
|
21
21
|
"description": "Project-agnostic utility skills — 4 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
|
|
22
22
|
"source": "./plugins/fh-commons"
|
|
23
23
|
}
|
package/CATALOG.md
CHANGED
|
@@ -8,6 +8,21 @@ AI reads this file first when searching past work. Open individual files for det
|
|
|
8
8
|
|
|
9
9
|
<!-- Add entries in reverse date order (newest at top) -->
|
|
10
10
|
|
|
11
|
+
### 2026-07-26 | forge-harness · forge-wiki · llmwiki-template · llmwiki-qa | #gate-locality, #sync-guard, #instrument-calibration, #sister-asset, #cross-corpus-provenance, #wiki-consolidation
|
|
12
|
+
**File:** scripts/gate_pathspec_check.sh · scripts/sync_guard_check.sh · templates/.git-hooks/pre-commit · templates/regression_guard.sh · templates/CLAUDE.md · plugins/fh-meta/skills/dialogue-harvest/SKILL.md · plugins/fh-meta/skills/frontier-digest/SKILL_detail.md · knowledge/shared/harness-core/loop_engineering.md
|
|
13
|
+
FH self-dev session (PR #182 merged; 7 PRs merged across 4 repos). Three defects, each found by measurement rather than review, each closed with a calibrated regression anchor. **(1) Gate-locality, 4th recurrence**: the pre-commit gate and regression_guard both matched the literal `SKILL\.md`, which the string `SKILL_detail.md` does not contain — 17 files / 208,710 B / **27.7% of the skill-spec surface**, 16 of 17 carrying fenced code, entirely ungated. It leaked twice for real (371c04f, e661931 — single-file edits to a *grounding-audit skill's own* behavioral spec). Worst property: `salience-splitter` widened the hole every time it moved content out of SKILL.md, so coverage shrank as the diet succeeded. An adversarial pass argued name-by-name coverage cannot close the class in principle; the resulting **enumeration sweep found a real uncovered file on its first run** (`dialogue-harvest/calibration_pair.md`), and the fix was escalated from a name list to a directory scope. **(2) One-way sync silently overwrote session cards twice** (v9 7,567 B → v8 5,299 B; v10 likewise). Not carelessness — *induced*: the session-start rule says read the mirror first, so the mirror becomes where agents write. The close-checker missed it because "card-last violated" measures timestamp **order**, not overwrite. Fixed with a destination-newer abort (both callers), a mechanically-injected MIRROR COPY banner, and a 7-pair anchor; live calibration caught three self-defects in the fix itself — over-blocking (the exact failure the handoff warned would train the override reflex), a half-fix leaving sync_file open, and banner churn that re-transferred 265 files per run and destroyed the log as an instrument. **(3) dialogue-harvest first real corpus** exposed that single-author input makes provenance labels free; answered with **Step 4-b cross-corpus provenance** (absorbed / held-unused / **declined**), which immediately showed a 44 KB transcript in-house since 06-27 with **zero citations in 29 days** — containing the very kill-switch clause the same day's sister audit had just confirmed FH lacked.
|
|
14
|
+
- **Sister asset**: `PromptPartner/agentsmith` (255★, 10 days old) cross-audited; 3 cross-family adversarial legs all returned NOT-CONVERGED and the governor kept the verdict — P-2's grounds refuted outright (`leak-gate.sh` is fail-closed in behavior, only the vocabulary was absent), P-1 narrowed, a category error accepted, and an arithmetic error caught (mixed baselines: +3,218 → +2,687). Sidecars said *where to dig*; only source-grounding decided.
|
|
15
|
+
- **Wiki consolidation (4 stores)**: the propagation gap ran **opposite** to the assumption — forge-wiki's own upgrades were already downstream; what never came back was a 3.7+ compat fix, leaving the canonical repo with the narrowest Python range of the three. Reverse-harvested. `{FH_ROOT}` variable landed in all four (templates carried one operator's absolute paths into org-visible checkouts — the mirror image of public-surface protection). `llmwiki-template`'s missing `fw_mcp.py` was judged **declined, not a gap** (CONTRACT.md §54 states it).
|
|
16
|
+
- **Instrument failures, five in one session** (BRE-in-ERE · zsh no-word-split · stale cwd · `install` containing `stall` · a 404 JSON body passing a length check). Every one produced a *confident wrong value*; every one was caught only by a known-positive control or by opening the actual line. Two of them recurred **after** the rule against them was written — which is why both new anchors make the control mandatory-pass rather than advice.
|
|
17
|
+
- Decision: operator approved all pushes/merges; company-zone repo pushed via REST Contents API per the account rule (new files only, no overwrite).
|
|
18
|
+
|
|
19
|
+
### 2026-07-25 | forge-harness | #dialogue-harvest, #new-skill, #cookbook-tier0, #sister-triage, #cross-family-audit, #ruleset, #calibration
|
|
20
|
+
**File:** plugins/fh-meta/skills/dialogue-harvest/ · plugins/fh-meta/skills/asset-placement-gate/SKILL.md · plugins/fh-meta/skills/context-doctor/SKILL.md · tracks/_audit/session_2026_07_25_claude5-context-rules-sister.md · tracks/_meta/codex_decorrelation_audit_2026-07-25.md
|
|
21
|
+
FH self-dev session (PR #177 merged). **New skill `dialogue-harvest`**: mines argument-shaped corpora (AI dialogue logs) — sycophancy strip first, then induced-vs-independent provenance labeling per proposition (degrade direction fixed toward the induced flag); ships an EN+KO known-pair calibration corpus backing the measured Done When. Gate: challenger HIGH 1 MED 3 LOW 3 → 7/7 repaired (the HIGH: an EN-only calibration pair would vacuously pass a new-language re-run); blind Sonnet sims reproduce both language pairs, and the KO run self-declared UNCALIBRATED when denied the calibration file. **asset-placement-gate Step 0.6**: official corpora (built-ins · claude-plugins-official · Claude Cookbook) registered as the criterion-③ ground, judged-flag semantics; roster-wide cookbook mapping scan found 0 shadowed skills. **context-doctor**: built-in `/doctor`-first sister anchor (Claude 5 gen) with the skill's increment enumerated. **Sister triage** of Anthropic's Claude-5 context-engineering rules: the "no measurable loss" claim ships no method → closed empty on that axis, vendor-claim label; imports = the /doctor Tier-0 boundary + a live confirmation of the substrate-shed trigger. **Cross-family audit** (codex gpt-5.5, repo-grounded) of the "decorrelation as generative principle" proposition: both halves NARROWED — surviving increment is "premature same-family consensus is a bad stop condition for design search", valid only as divergence-then-selection on high-leverage underdetermined decisions; doctrine-ization stays deferred by design. persona-innovator frame scan: BVSR/QD/Best-of-N all lack FH's two clauses → keep the coinage, sister-link on doctrine-ization. **Server-side residual closed**: ruleset `main-no-force-push` (non_fast_forward on main) created and independently GET-verified.
|
|
22
|
+
- Decision: dialogue-harvest built on operator go (manual n=1 proof accepted); npm v1.4.70 published + tagged same session (operator approval, PR #179).
|
|
23
|
+
- **Third block (evening)**: **Matrix Benchmark v0+v1** (qasp-dev #27·#28) — the operator's "blind man with an ultrasonic cane at 60%" metaphor turned into numbers. Planted 4 defects + 2 traps in the admin-surrogate sandbox (calibrated known-pair), probed with isolated same-tier agents differing only in channel access: **detection 4/4 tied, attribution 2/4 (blind) vs 4/4 (matrix), traps 0/0** — the two blind misattributions were exactly the §3-b cell ("absent from source" vs "present but not rendered", both misread as source-defects: spec-optimism bias observed). Pre-registered predictions: 2/3 hit, 1 void. v1: attribution routed repairs (source-defects → code, spec-defect/environment → spec v2) and the AX-Lobby 3-channel E2E generation pattern reproduced locally (12/12 pass · web_rules 0/6 first pass · 3 BE-derived data-integrity assertions · §8 UNVERIFIED label · governor re-ran independently). Next (designed, not run): 3-condition escalation bench — "open eyes only where needed" — gated on semantic_anchor pull-forward wiring. Honest labels: n=1/condition exploratory; C1 blind-miss partially instruction-induced (fixed viewport).
|
|
24
|
+
- **Second half (same day)**: qasp direction review — the "!" located from records (verification moves to the front of the pipeline: mate PR #382 close handed the reviewer seat to AI; workflow inversion "know the criteria → implement → PR already passing"; act-1.5 dual timing; web E2E generation with 3-channel cognition). Its identity tension (§5 producer=verdict) closed as **qasp-dev governance §8 verdict-independence** (domain-respect unbundled from verdict-independence; UNVERIFIED=blocking label; promotion preconditions), codex-audited 8/8. **mate_rules split into mobile_rules (7 common) ⊃ mate extension (4 convention rules)** — full compatibility preserved, 1364 tests green, codex-audited 3/3. Company-handoff tasks landed dev-side: `--checklist` command, review-bot `--diff-json` surface (consumes pr-agent's own artifact), preflight workflow doc, and a mate-dev advisory CI lane (`qasp-rules.yml`, neutral check, 3 marked line-swap points for the company circuit). PR-review video triaged: mostly convergent with existing doctrine; template direction rejected by operator (human-PR fatigue, #382 precedent) — evidence must be a byproduct of pre-PR self-verification, not a demand on authors.
|
|
25
|
+
|
|
11
26
|
### 2026-07-24 (2) | forge-harness · forge-wiki · llmwiki-template | #sister-links, #full-gate, #frontier-digest-angle-rule, #query-refresh, #mapping, #light-harness, #graph-engineering
|
|
12
27
|
**File:** knowledge/shared/dialogue/memory_intent_recall.md · knowledge/shared/harness-core/field_harness_diagnostic.md · knowledge/shared/harness-core/multi_model_sidecar_strategy.md · plugins/fh-meta/skills/harness-doctor/SKILL.md · plugins/fh-meta/skills/frontier-digest/SKILL_detail.md · tracks/forge-wiki/ · tracks/llmwiki-template/
|
|
13
28
|
Second half of the 07-24 hub session (PR #174 + two field repos). **Sister anchors executed under the full 4-axis gate** (the four links #173 deferred): GRACE into memory_intent_recall (instruction-graph boundary spelled out) and harness-doctor Step 2-E (C-layer maintenance-path ledger, ETCLOVG-style anchor-only), the 7-criteria rubric as a considered-and-held note with two re-check triggers in field_harness_diagnostic, and the wsff RL-incentive causal anchor ("maintainability has no fast oracle") beside the harness-ceiling principle in multi_model_sidecar_strategy. **frontier-digest repaired on two axes**: the stale arXiv canonical query replaced ("AI software testing"→"LLM agent evaluation", refresh criterion written down) and an Angle rule added to the synthesis prompt closing the methodology-angle silent drop (fh_signal 07-24 #3) — verified by a blind Sonnet known-pair sim (dual-angle item surfaced, single-angle item not forced) plus a quench-challenger pass (HIGH 0 · MED 2 · LOW 3; MED+2 LOW repaired in-branch). **Two sibling repos mapped and light-harnessed** (forge-wiki = chamber run #9's first EMIT, public; llmwiki-template = its company-side private twin): tracks/ dirs, session rules, .claudeignore — env card and MCP gating skipped by judgment (residency / no external mount); both merged (forge-wiki #4, llmwiki-template #2), llmwiki's CLAUDE.md kept operator-local. Positioning decision recorded in both tracks: **a wiki that sits under the harness** (Harness ⊃ context/knowledge component) — "wiki" is the outward word, "harness knowledge layer" the architecture coordinate.
|
package/CLAUDE.md
CHANGED
|
@@ -263,7 +263,7 @@ Every new `SKILL.md` must clear a **6-item bar** (role-duplication via `/asset-p
|
|
|
263
263
|
|
|
264
264
|
## FH Improvement 4-Axis Auto-Gate (Self-Verification Orchestrator)
|
|
265
265
|
|
|
266
|
-
**FH 자산을 수정하면**(SKILL.md · `.claude/rules/*.md` · `knowledge/shared/rules/*.md` · `templates/` · `CLAUDE.md` · substantive `knowledge/`·`docs/*.md` · `AGENTS.md`) **4축 검증 체인이 그 세션 첫 커밋 전에 자동 실행된다.** 사용자 요청 불요 — 제안이 아니라 의무 단계다.
|
|
266
|
+
**FH 자산을 수정하면**(SKILL.md · **SKILL_detail.md** · `.claude/rules/*.md` · `knowledge/shared/rules/*.md` · `templates/` · `CLAUDE.md` · substantive `knowledge/`·`docs/*.md` · `AGENTS.md`) **4축 검증 체인이 그 세션 첫 커밋 전에 자동 실행된다.** 사용자 요청 불요 — 제안이 아니라 의무 단계다.
|
|
267
267
|
|
|
268
268
|
**기계 floor**: `git commit` 은 `templates/.git-hooks/pre-commit` 이 **하드 차단**한다. 축이 전부 PASS 할 때까지 커밋 자체가 안 된다. 아래 상세가 로드되지 않아도 **훅이 막는다** — 이 산문은 훅 위의 살리언스 층이지 유일 floor 가 아니다.
|
|
269
269
|
|
package/README.md
CHANGED
|
@@ -275,7 +275,7 @@ All four movements ship. Temper was named before it was built — deliberately (
|
|
|
275
275
|
two more signatures keep it running: `harvest-loop` (each session's lessons become permanent skills) and
|
|
276
276
|
`agent-composer` (orchestrate the dispatch). The other skills wait until you need them — full list below.
|
|
277
277
|
|
|
278
|
-
##
|
|
278
|
+
## 39 skills · 8 agents
|
|
279
279
|
|
|
280
280
|
> Count = non-deprecated skills (deprecated redirect stubs — kept only for old-name routing — excluded).
|
|
281
281
|
|
|
@@ -293,6 +293,7 @@ two more signatures keep it running: `harvest-loop` (each session's lessons beco
|
|
|
293
293
|
| `harness-doctor` | Harness structure diagnosis | "Check my Claude setup" |
|
|
294
294
|
| `pipeline-conductor` | 4-axis quality gate (backward/adversarial/forward/record) | "Run the quality gate" |
|
|
295
295
|
| `field-harvest` | Back-propagate field patterns to hub | "I could reuse this" |
|
|
296
|
+
| `dialogue-harvest` | Mine AI-dialogue logs: strip sycophancy, label induced vs independent | "What did I actually contribute in this thread?" |
|
|
296
297
|
| `frontier-digest` | HN + arXiv → actionable insights | "AI trend digest" |
|
|
297
298
|
| `hub-cc-pr-reviewer` | Automated PR review | "Review this PR" |
|
|
298
299
|
| `verify-bidirectional` | Reverse-verify decisions | "Is that right?", "Double-check" |
|
|
@@ -71,7 +71,33 @@ instrument: the next slip finds its leg pre-diagnosed.
|
|
|
71
71
|
| harvest-loop Step 0-b/0-c evidence check | a harvest run misses completed items despite `fh_completed_*` existing |
|
|
72
72
|
| goal-quench mid-run checkpoint files (70/85/95%) | a /goal run blows through a threshold unnoticed |
|
|
73
73
|
| ~~Substrate-jump detector~~ | **BUILT 2026-07-10** (`scripts/substrate_jump_detector.sh`, SessionStart-wired) — same operator instruction; structure-enforcing class (out-of-context drift), permanent per the durable-mechanization criterion |
|
|
74
|
-
| Quarterly maturity checker (`quarterly_maturity_check.sh` —
|
|
74
|
+
| Quarterly maturity checker (`quarterly_maturity_check.sh`) — **still unbuilt, correctly so** | ⚠️ **TRIGGER WAS UNFIRING — condition rewritten 2026-07-26** (the row stays live; strikethrough in this table means BUILT and must not be used for anything else). Its condition was *"a quarterly re-diagnosis is missed >90d"*, which needs a **last-diagnosis date to measure from, and none exists**: `hub_maturity_roadmap.md` is still the shipped **template** (its own header says *"in actual hub operation… write an operating copy"*), no operating copy was ever created, §5's six indicators still hold their template descriptions rather than counters, and no quarterly diagnosis artifact exists in either store. A backlog row whose trigger cannot fire is not deferred — it is **dead**, and it reads as deferred. **New trigger (observable): the first quarterly re-diagnosis is actually run and dated.** Until an instance-zero exists there is nothing for a checker to check; building the checker first would be the checker-for-a-loop-that-never-ran. |
|
|
75
|
+
| **Tested kill switch for any FH path that runs unattended** — documented one-move stop, *actually exercised once*, recorded | FH runs a genuinely unattended loop. Today the only candidate is the launchd `frontier-digest` (09:00 daily); the moment a second one ships, or that one gains write authority beyond its digest file, this fires |
|
|
76
|
+
|
|
77
|
+
### On the kill-switch row — provenance and the boundary of what is being proposed
|
|
78
|
+
|
|
79
|
+
Added 2026-07-26 with **two independent sources naming the same gap**, which is why it is on the
|
|
80
|
+
backlog rather than in a signal:
|
|
81
|
+
1. **External transcript already in-house since 2026-06-27** (a private companion store's
|
|
82
|
+
cross-audit shelf, `raw_2026-06-27_loop-engineering-transcript.txt` — single-author, 44 KB)
|
|
83
|
+
— *"실수도 무인으로 쌓이기 때문에 브레이크가 필요하다"* · iteration cap ·
|
|
84
|
+
stop-on-stall · budget guard. **This corpus was never cited by any FH asset** (whole-repo grep, 0
|
|
85
|
+
references) — it sat unmined for 13 days while this very file was authored from a different sister.
|
|
86
|
+
Surfaced by the first real-corpus `dialogue-harvest` run.
|
|
87
|
+
2. **`agentsmith/profiles/autonomous-loops.md`** (2026-07-16) — *"a documented one-move stop you have
|
|
88
|
+
actually tested. No budget, no cap, no L3."* Cross-audit `session_2026_07_26_agentsmith-sister.md §I-4`.
|
|
89
|
+
|
|
90
|
+
**Measured absence, instrument-calibrated**: `kill switch` · `emergency stop` · `abort the run/loop` ·
|
|
91
|
+
`정지 스위치` · `tested stop` → **0 hits** repo-wide, with known-positive controls in the same run.
|
|
92
|
+
|
|
93
|
+
⚠️ **What is deliberately NOT proposed — do not relitigate**: the same two sources also teach
|
|
94
|
+
*stagnation-triggered stopping* ("stop when the number does not improve twice"). FH already
|
|
95
|
+
**considered and declined** that class — `hub_maturity_roadmap.md §(c) Trigger-based`: *"Stagnation
|
|
96
|
+
detection criteria + auto-alerts + trigger tags all require new infra. Simplification principle
|
|
97
|
+
violation risk."* That declination stands; reopening it needs new evidence about **that** decision,
|
|
98
|
+
not a fresh citation of the same idea. Separately, FH **already has** stall detection where it
|
|
99
|
+
measurably mattered — `auto-decorrelation` §sidecar liveness, built on a production miss.
|
|
100
|
+
The kill-switch row is narrower than either source's full prescription, on purpose.
|
|
75
101
|
|
|
76
102
|
## Done When (for a new/changed autonomous path)
|
|
77
103
|
|
package/package.json
CHANGED
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "fh-meta",
|
|
3
|
-
"version": "1.4.
|
|
3
|
+
"version": "1.4.71",
|
|
4
4
|
"engines": {
|
|
5
5
|
"claudeCode": ">=1.0.0"
|
|
6
6
|
},
|
|
7
|
-
"description": "Hub meta-engineering toolkit —
|
|
7
|
+
"description": "Hub meta-engineering toolkit — 35 skills + 7 agents. New in 1.4.71: SKILL_detail.md brought inside the 4-axis gate — the gate matched the literal `SKILL.md`, which `SKILL_detail.md` does not contain, leaving 27.7% of the skill-spec surface ungated (measured; it leaked twice for real). Fix escalated from a name list to a directory scope after an enumeration sweep found a real uncovered file on its first run; anchored by scripts/gate_pathspec_check.sh (known-pair, wired into pre-commit). One-way mirror sync gains a destination-newer abort + mechanically injected mirror banner after two session cards were silently overwritten, anchored by scripts/sync_guard_check.sh. dialogue-harvest gains Step 4-b cross-corpus provenance (absorbed / held-unused / declined) for single-author corpora, where the original provenance labels were free and measured nothing. frontier-digest arxiv leg category-scoped after a relevance drift the staleness rule structurally could not catch (known-pair calibration recorded). templates/CLAUDE.md hub paths switched to {FH_ROOT} so a copied harness stops carrying one operator's absolute paths into org-visible checkouts. New in 1.4.70: dialogue-harvest (mines AI-dialogue logs — sycophancy strip first, induced-vs-independent provenance labeling, EN+KO known-pair calibration shipped); asset-placement-gate Step 0.6 official-corpora check (Claude Cookbook as Tier-0 consult corpus, judged-flag semantics); context-doctor built-in-/doctor-first sister anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor gains a command-output axis — routes to a command-output proxy/hook (rtk) to trim verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce environments (lossy filtering, off gate-input paths). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard scaffolds the companion store as a queryable wiki (INDEX + session-start read + Raw/Wiki/Conversation ingest axis). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment, calibration-gated) + video-ingest (capability-routed video ingestion). New in 1.4.37: corpus-grounding-expander + persona-roster-expander (field-harvested verbatim-relay capability skills). New in 1.3.0: public-surface-audit (git-tracked private-token leak scan), field-harvest Mode B session-end auto-trigger, 4-axis gate scope extension (docs/ + AGENTS.md). New in 1.2.0: pipeline-conductor (4-pipeline gated sweep), return-path-gate (chain closure audit), goal-quench (Stop hook + quality gate), steel-quench Wave 5 (multi-model sidecar challenger), 2-layer architecture docs, YAML validation script. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
8
8
|
"author": {
|
|
9
9
|
"name": "chrono-meta",
|
|
10
10
|
"email": "chrono-meta@users.noreply.github.com"
|
|
@@ -37,8 +37,8 @@ When unsure where to place a new asset or skill:
|
|
|
37
37
|
|
|
38
38
|
1. Request full file path from user (or accept natural language description)
|
|
39
39
|
2. Load asset content via `Read` (if path provided)
|
|
40
|
-
2.5. Step 0.5 mechanical overlap pre-scan (grounds criterion ④
|
|
41
|
-
3. Evaluate Step 1 4-criteria in order (LLM makes the judgment, ④ gated on the Step 0.5 scan)
|
|
40
|
+
2.5. Step 0.5 mechanical overlap pre-scan (grounds criterion ④) + Step 0.6 official-corpora check (grounds criterion ③)
|
|
41
|
+
3. Evaluate Step 1 4-criteria in order (LLM makes the judgment, ④ gated on the Step 0.5 scan, ③ informed by Step 0.6)
|
|
42
42
|
4. ① + ④ both pass + at least one of ②③ passes → output **"FH suitable"**
|
|
43
43
|
Otherwise, proceed to Step 2 local assessment → if fails, output **"Project-local agent or no asset needed"**
|
|
44
44
|
|
|
@@ -65,7 +65,7 @@ Immediately after trigger, acquire asset content in the following order.
|
|
|
65
65
|
> **Which asset should I evaluate?**
|
|
66
66
|
> Enter a file path (e.g., `.claude/agents/jira-create.md`) or a description.
|
|
67
67
|
|
|
68
|
-
After acquiring the asset content, run Step 0.5 (mechanical overlap pre-scan) **before** the judged Step 1.
|
|
68
|
+
After acquiring the asset content, run Step 0.5 (mechanical overlap pre-scan) and Step 0.6 (official-corpora check) **before** the judged Step 1.
|
|
69
69
|
|
|
70
70
|
## Step 0.5. Mechanical Overlap Pre-Scan (grounds criterion ④)
|
|
71
71
|
|
|
@@ -92,6 +92,17 @@ both the judge (cutoff) and the grep (literal); that residual leans on the judge
|
|
|
92
92
|
enumerated descriptions above** (grounded comparison, not pure memory), not on full mechanization. A
|
|
93
93
|
shared common word is a judged-review flag, **not** a hard fail.
|
|
94
94
|
|
|
95
|
+
## Step 0.6. Tier-0/1 Official-Corpora Check (grounds criterion ③)
|
|
96
|
+
|
|
97
|
+
Besides the FH roster (Step 0.5), check the proposal against the official corpora: platform
|
|
98
|
+
built-ins, `claude-plugins-official`, and the **Claude Cookbook pattern list**
|
|
99
|
+
(platform.claude.com/cookbook — agent patterns · tools · RAG · evals · skills). A hit here is a
|
|
100
|
+
**judged ③ flag, not an automatic ③ fail**: it obliges the proposal to state its governance
|
|
101
|
+
increment over the official pattern (no-reinvention rule — an official pattern that fully covers
|
|
102
|
+
the need outranks a net-new build; a pattern the proposal governs *on top of* does not fail ③).
|
|
103
|
+
Offline fallback: note `cookbook: unchecked` **in the Step 3 routing output** rather than silently
|
|
104
|
+
skipping. (Provenance: `tracks/_audit/session_2026_07_25_claude5-context-rules-sister.md`.)
|
|
105
|
+
|
|
95
106
|
## Done When
|
|
96
107
|
|
|
97
108
|
```
|
|
@@ -108,7 +119,7 @@ All steps 0–3 completed
|
|
|
108
119
|
|:-:|---|---|
|
|
109
120
|
| ① | Cross-project value | Is this asset equally useful in other projects without depending on a specific project? |
|
|
110
121
|
| ② | Orchestration / judgment layer | Is it just a list of MCP/Bash calls, or a judgment layer that synthesizes multiple signals? |
|
|
111
|
-
| ③ | Not replaceable by built-ins | Can this be equally achieved with direct MCP calls or basic bash? (If yes, fails this criterion) |
|
|
122
|
+
| ③ | Not replaceable by built-ins | Can this be equally achieved with direct MCP calls or basic bash? (If yes, fails this criterion.) A Step 0.6 official-pattern hit is a judged flag: fails ③ only if the pattern fully covers the need with no governance increment |
|
|
112
123
|
| ④ | No overlap with existing FH skills | Step 0.5 mechanical scan = 0 name/trigger collision **AND** judged ≤90% overlap. Non-zero collision → hard fail. |
|
|
113
124
|
|
|
114
125
|
**FH suitable** → ① + ④ both pass + at least one of ②③ passes.
|
|
@@ -24,6 +24,17 @@ Diagnoses the main causes of session token waste and prescribes immediate remedi
|
|
|
24
24
|
|
|
25
25
|
**Standalone install** — this skill works normally with plugin install only, without cloning the full meta-harness.
|
|
26
26
|
|
|
27
|
+
**Built-in `/doctor` first (sister anchor, 2026-07-25)**: since the Claude 5 generation, the official
|
|
28
|
+
`/doctor` command claims skill/CLAUDE.md rightsizing ("The new rules of context engineering for Claude 5
|
|
29
|
+
generation models" — https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models).
|
|
30
|
+
Run the built-in first when available — its scope is skills + CLAUDE.md rightsizing only. This skill's
|
|
31
|
+
increment over it = `.claudeignore` diagnosis/generation (Step 1), `/clear` timing guidance,
|
|
32
|
+
command-output proxying (§Command-Output Reduction), hub-wide memory/CLAUDE.md audits, and the
|
|
33
|
+
measurement discipline (resident footprint is measured only via a fresh top-level `/context`, and any
|
|
34
|
+
diet is followed by a `prompt-regression` probe — the blog's own "no measurable loss" claim ships no
|
|
35
|
+
method, so treat that number as a vendor claim). Triage record:
|
|
36
|
+
`tracks/_audit/session_2026_07_25_claude5-context-rules-sister.md`.
|
|
37
|
+
|
|
27
38
|
## Execution Steps
|
|
28
39
|
|
|
29
40
|
### Step 1. `.claudeignore` Diagnosis + Generation
|
|
@@ -0,0 +1,223 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: dialogue-harvest
|
|
3
|
+
description: Mines improvement insights from long-form AI dialogue logs and long texts — an external or pasted corpus, not the current session (current-session wrap-up belongs to harvest-loop). First operation is sycophancy removal — separates speakers, strips the counterpart's agreement/restatement spans, converts the user's own utterances into falsifiable propositions, and labels each proposition induced (frame first introduced by the counterpart) or independent (originated with the user). Ships a drop counter so removed material is visible, never silently lost. Triggered by "mine insights from this chat log", "이 대화록에서 통찰 캐줘", "strip the agreement and show me what's actually mine", "동조 걷어내고 알맹이만", "what did I actually contribute in this thread".
|
|
4
|
+
user-invocable: true
|
|
5
|
+
allowed-tools: ["Read", "Grep", "Glob"]
|
|
6
|
+
model: sonnet
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
# dialogue-harvest
|
|
10
|
+
|
|
11
|
+
Extracts the load-bearing propositions from a long AI-human dialogue (or long text) after
|
|
12
|
+
removing the counterpart's sycophancy. Fills the gap between FH's loading assets
|
|
13
|
+
(corpus-grounding-expander, video-ingest) and its mining assets (frontier-digest = feed items,
|
|
14
|
+
field-harvest = git history, harvest-loop = session records): none of those mines an *argument-shaped*
|
|
15
|
+
corpus, and a naive summarizer run on an AI dialogue returns "the user made several excellent
|
|
16
|
+
points" — worthless, because most of the volume is agreement.
|
|
17
|
+
|
|
18
|
+
**Why the first operation is sycophancy removal**: AI dialogue logs are dominated by agreement and
|
|
19
|
+
restatement. Without stripping them first, extraction inflates the user's contribution; without
|
|
20
|
+
provenance labeling (Step 4), "my insight" and "a frame the counterpart planted" are
|
|
21
|
+
indistinguishable — which is exactly the failure this skill exists to prevent.
|
|
22
|
+
|
|
23
|
+
## Triggers
|
|
24
|
+
- "mine insights from this chat log" / "이 대화록에서 통찰 캐줘"
|
|
25
|
+
- "strip the agreement and show me what's actually mine" / "동조 걷어내고 알맹이만"
|
|
26
|
+
- "what did I actually contribute in this thread" / "이 대화에서 내 몫이 뭐였지"
|
|
27
|
+
- "turn this dialogue into propositions" / "이 대담 명제로 정리해줘"
|
|
28
|
+
- `/dialogue-harvest {path or pasted log}`
|
|
29
|
+
|
|
30
|
+
**Boundary**: wrapping up the *current session* → `harvest-loop`; this skill mines an
|
|
31
|
+
external/pasted dialogue corpus.
|
|
32
|
+
|
|
33
|
+
## Step 0. Input Acquisition
|
|
34
|
+
|
|
35
|
+
1. **Path provided** → Read the file.
|
|
36
|
+
2. **Pasted log** → use as-is.
|
|
37
|
+
3. **No speaker structure detectable** (no turn markers, no names, no quoting pattern) → stop and ask:
|
|
38
|
+
> Speaker separation is the first operation and it needs turn boundaries. Who said what?
|
|
39
|
+
> (Provide turn markers, or confirm the text is single-author — single-author texts skip
|
|
40
|
+
> Steps 1–2 and go straight to proposition-ization.)
|
|
41
|
+
|
|
42
|
+
**Single-author path**: skip Steps 1–2, Step 4 stamps every proposition **independent** (no
|
|
43
|
+
counterpart exists), and the Step 6 counterpart sections read `n/a`.
|
|
44
|
+
|
|
45
|
+
## Step 1. Speaker Separation (mechanical when markers exist)
|
|
46
|
+
|
|
47
|
+
Split the log into turns and attribute each to USER or COUNTERPART. If markers are absent but
|
|
48
|
+
inference is possible, proceed and stamp the output header `speaker-inference: judged` — never
|
|
49
|
+
present inferred attribution as given.
|
|
50
|
+
|
|
51
|
+
## Step 2. Sycophancy Strip (one grep-able + two judged classes, drop-counted)
|
|
52
|
+
|
|
53
|
+
Remove COUNTERPART spans in these classes, **counting every removal by class**. Only the first
|
|
54
|
+
class is mechanical; the other two are judgment calls and their drops are marked `(judged)` in the
|
|
55
|
+
counter — never presented as grep certainty:
|
|
56
|
+
|
|
57
|
+
| Class | Check | Detection |
|
|
58
|
+
|---|---|---|
|
|
59
|
+
| Praise/agreement lexicon | grep-able | grep for praise-class tokens (sharp/brilliant/exactly right/정확한 지적/탁월한/완벽한 …) — extend the pattern list per corpus language |
|
|
60
|
+
| Restatement | judged | counterpart paragraph whose content substantially re-says the user's immediately-preceding utterance (high term overlap, no new claim) |
|
|
61
|
+
| Closing filler | judged | session pleasantries carrying no claim |
|
|
62
|
+
|
|
63
|
+
**Multi-class span**: a span matching more than one class is counted **once**, under the first
|
|
64
|
+
matching class in the table order above.
|
|
65
|
+
|
|
66
|
+
**Drop counter is mandatory**: silent drops are material loss (the qasp loader precedent — unused
|
|
67
|
+
material does not scream). Every stripped span and every non-propositional user turn appears in the
|
|
68
|
+
Step 6 accounting. Counterpart turns that *introduce a frame or claim* are **retained** — they are
|
|
69
|
+
the provenance source for Step 4, not noise.
|
|
70
|
+
|
|
71
|
+
**Language note**: the praise-token list is language-specific. On a corpus language not yet in the
|
|
72
|
+
list, extend the patterns first, then **construct (or use, if shipped) a known pair in that
|
|
73
|
+
language** and re-run calibration — re-running an English-only pair after adding Korean tokens
|
|
74
|
+
exercises nothing (vacuous pass). An ASCII-token scan over a Korean corpus is the measured
|
|
75
|
+
false-positive precedent. `calibration_pair.md` ships English + Korean sections.
|
|
76
|
+
|
|
77
|
+
## Step 3. Proposition-ization (user utterances only)
|
|
78
|
+
|
|
79
|
+
Convert each remaining USER utterance into a falsifiable statement (a form you can ask true/false
|
|
80
|
+
of), keeping a source reference (turn id / line) per proposition. Meaning must not drift: the
|
|
81
|
+
proposition is a compression of the user's claim, not an improvement of it.
|
|
82
|
+
|
|
83
|
+
## Step 4. Provenance Labeling — induced vs independent (the differentiator)
|
|
84
|
+
|
|
85
|
+
For each proposition, find the **first occurrence** of its core frame/key terms in the log:
|
|
86
|
+
|
|
87
|
+
- First occurrence is in a USER turn → **independent**.
|
|
88
|
+
- First occurrence is in a COUNTERPART turn (the user then adopted/echoed it) → **induced**.
|
|
89
|
+
- Ambiguous (shared vocabulary, gradual convergence) → label **induced?** — the degrade direction
|
|
90
|
+
is toward the induced flag, because mislabeling a planted frame as "my insight" is the harm this
|
|
91
|
+
skill exists to catch; the reverse error only costs modesty.
|
|
92
|
+
|
|
93
|
+
Counterpart-authored content the user *selected* as valuable is listed separately, marked
|
|
94
|
+
"selection, not authorship" — selection is a real act but must not be recorded as the user's claim.
|
|
95
|
+
|
|
96
|
+
## Step 4-b. Cross-Corpus Provenance — the single-author mode (measured need, 2026-07-26)
|
|
97
|
+
|
|
98
|
+
Step 4 asks *"who first said this — the user or the counterpart?"* On a **single-author corpus**
|
|
99
|
+
(a video transcript, an article, a talk) there is no counterpart, so every proposition is stamped
|
|
100
|
+
`independent` **for free** and the differentiator does not fire at all. The first real-corpus run
|
|
101
|
+
measured exactly that: 12 propositions, 12 free labels, zero discrimination.
|
|
102
|
+
|
|
103
|
+
On a single-author corpus the load-bearing provenance axis is not *within* the document — it is
|
|
104
|
+
**between the corpus and the assets that were supposed to consume it**:
|
|
105
|
+
|
|
106
|
+
> *Which of these propositions reached our own assets, and which did we hold and never use?*
|
|
107
|
+
|
|
108
|
+
**Run it when** the corpus is single-author AND a downstream asset on the same topic exists.
|
|
109
|
+
Skip (and say so) when there is no plausible consumer — the question is meaningless without one.
|
|
110
|
+
|
|
111
|
+
1. **Name the consumer asset(s)** explicitly in the output header. Guessing is not allowed;
|
|
112
|
+
an unnamed consumer makes every verdict unfalsifiable.
|
|
113
|
+
2. **Grep locates candidates; the label is assigned by reading them.** A hit count is never a
|
|
114
|
+
verdict — a consumer that says *"we do not use kill switches"* matches the same pattern as one
|
|
115
|
+
that adopts them. So:
|
|
116
|
+
- **absorbed** — requires **quoting the supporting span** from the consumer. No quote, no
|
|
117
|
+
`absorbed`. (A count-only `absorbed` is the same defect as a whole-file parity grep going
|
|
118
|
+
green on a line that says "excluded" — measured in this repo the same day.)
|
|
119
|
+
- **held-unused** — absent from the consumer though the corpus has been in-house since {date}.
|
|
120
|
+
Before writing it, **try at least two term variants** (synonym / abbreviation / the concept
|
|
121
|
+
re-said in the consumer's own vocabulary). Concepts get absorbed under different words; a
|
|
122
|
+
literal-match zero is weak evidence of absence.
|
|
123
|
+
- **declined** — absent *because a named asset records a decision against it*. Cite the decision.
|
|
124
|
+
This tier is mandatory and load-bearing: without it, a re-proposal reopens a settled call
|
|
125
|
+
as if it were an oversight. (First run hit this immediately — stagnation-triggered stopping
|
|
126
|
+
was `held-unused` by grep and `declined` in fact, recorded in `hub_maturity_roadmap.md`.)
|
|
127
|
+
**Automation is prohibited here**: the first run produced 2 false positives out of 12, both
|
|
128
|
+
caught only by opening the matched line.
|
|
129
|
+
3. **Report the ingestion-to-citation gap** — corpus in-house date vs first citation by any asset.
|
|
130
|
+
A corpus with **zero citations** is the finding, not a null result.
|
|
131
|
+
|
|
132
|
+
**Instrument discipline (mandatory-pass, learned on the first run — four failures in one session):**
|
|
133
|
+
a **known-positive control** runs beside every measurement and its result is printed. A bare zero is
|
|
134
|
+
not publishable.
|
|
135
|
+
|
|
136
|
+
⚠️ **The control must be a separate search, not an alternation bolted onto the target pattern.**
|
|
137
|
+
`grep -E "core_term|title"` satisfies "a control ran" while proving nothing: `title` matches, the
|
|
138
|
+
command exits 0, and a malformed `core_term` still silently matches nothing. That is a vacuous pass.
|
|
139
|
+
The control's job is to prove **this pattern form, on this target, through this shell** can return a
|
|
140
|
+
hit at all — so run the *same pattern shape* against something you know contains it, as its own
|
|
141
|
+
command, and print both numbers. The four measured failure modes:
|
|
142
|
+
BRE `\|` inside an ERE pattern · unquoted `$VAR` under zsh (no word-splitting) · a stale `cd` making
|
|
143
|
+
paths unresolvable · substring collision (`install` contains `stall`). Each produced a *confident,
|
|
144
|
+
wrong* zero; each was caught only by the control. Do not redirect stderr away — three of the four
|
|
145
|
+
announced themselves there.
|
|
146
|
+
|
|
147
|
+
**Degrade direction — and the skip/degrade boundary, which must not be a matter of taste.**
|
|
148
|
+
`skip (N/A)` and `UNCALIBRATED` are **different states with different triggers**, and the boundary is
|
|
149
|
+
mechanical because otherwise the cheaper exit wins under time pressure:
|
|
150
|
+
|
|
151
|
+
| Situation | State | Why |
|
|
152
|
+
|---|---|---|
|
|
153
|
+
| Corpus is multi-speaker | **N/A** | Step 4 already answered provenance; 4-b's question does not arise |
|
|
154
|
+
| A consumer search was **run and recorded**, and no asset on the topic exists | **N/A**, quoting the search | A real negative, not a failure |
|
|
155
|
+
| A consumer is **named but does not resolve** (bad path, missing file) | **UNCALIBRATED** | Never N/A — a broken pointer is a tooling failure wearing absence's clothes |
|
|
156
|
+
| The known-positive control returns zero | **UNCALIBRATED** | The instrument is not measuring |
|
|
157
|
+
| No search was run | **UNCALIBRATED** | "I didn't look" is not "nothing is there" |
|
|
158
|
+
|
|
159
|
+
Under `UNCALIBRATED` the step emits **no** absorbed/held-unused/declined labels at all — a provenance
|
|
160
|
+
verdict is a claim about what an organization did with knowledge, and an uncalibrated one is worse
|
|
161
|
+
than none. Under `N/A` the step states the reason **and the search that established it**.
|
|
162
|
+
|
|
163
|
+
⚠️ **Known weakness, stated rather than papered over**: a hurried session's natural pull is to grep a
|
|
164
|
+
nearby-but-plausible file and report against it without flagging the substitution. The Sonnet
|
|
165
|
+
floor-sim of this step reached `UNCALIBRATED` correctly *and* named that same failure as the more
|
|
166
|
+
likely one in practice. The forcing function is the control returning zero alongside the
|
|
167
|
+
measurement — which is why the control is mandatory-pass and not advice.
|
|
168
|
+
|
|
169
|
+
## Step 5. Verification-Status Column
|
|
170
|
+
|
|
171
|
+
Each proposition gets a status: mechanically checkable / falsifiable-but-untested /
|
|
172
|
+
unfalsifiable (axiom-only use) / speculation. Unfalsifiable propositions are still listed — flagged
|
|
173
|
+
so they are never later cited as established facts.
|
|
174
|
+
|
|
175
|
+
## Step 6. Output + Accounting + Routing (proposal-only)
|
|
176
|
+
|
|
177
|
+
```
|
|
178
|
+
[dialogue-harvest output]
|
|
179
|
+
Corpus: {source} | speaker-inference: given | judged
|
|
180
|
+
calibration: run-this-session | previously-verified {date} ← declare every run (L2 guard)
|
|
181
|
+
── Propositions ──
|
|
182
|
+
| # | proposition | source ref | provenance | verification status |
|
|
183
|
+
── Counterpart-originated (selection ≠ authorship) ──
|
|
184
|
+
| item | first-occurrence ref |
|
|
185
|
+
── Accounting (mandatory identity, turn-based) ──
|
|
186
|
+
user turns N = propositionized turns + dropped turns (dropped listed by reason)
|
|
187
|
+
propositions total: P (a turn may yield several — counted separately, never forced 1:1)
|
|
188
|
+
counterpart spans stripped: {count by class, judged classes marked} · retained as frame-source: {count}
|
|
189
|
+
── Routing proposals (no auto-write) ──
|
|
190
|
+
{insight → fh_signal / memory / CHAMBER-CANDIDATE / project track — one line each, operator decides}
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
Routing is **proposal-only**: this skill writes nothing outside its output block.
|
|
194
|
+
|
|
195
|
+
## Done When
|
|
196
|
+
|
|
197
|
+
- **Accounting identity holds** — user turns = propositionized turns + dropped turns, both lists
|
|
198
|
+
visible; proposition total reported separately (a turn may carry multiple propositions)
|
|
199
|
+
(check-class: **mandatory-pass**, mechanical count).
|
|
200
|
+
- **Every proposition carries a source ref + provenance label** (independent / induced / induced?)
|
|
201
|
+
(check-class: **mandatory-pass**).
|
|
202
|
+
- **Calibration pair separates** — running the skill on `calibration_pair.md` (shipped beside this
|
|
203
|
+
file, English + Korean sections) reproduces its known answers: known-independent → independent,
|
|
204
|
+
known-induced → induced, drop counted. Required on first use and whenever the praise-pattern list
|
|
205
|
+
changes; a corpus language with no shipped section requires **constructing that language's known
|
|
206
|
+
pair first** — re-running an existing-language pair for a new language is a vacuous pass, not a
|
|
207
|
+
measurement (check-class: **measured**; declare the result in the output header
|
|
208
|
+
`calibration:` field).
|
|
209
|
+
- **Single-author corpus with a named consumer ran Step 4-b** — each proposition carries
|
|
210
|
+
absorbed / held-unused / declined, the consumer asset is named, and the ingestion-to-citation gap
|
|
211
|
+
is reported; every grep in the step shows its known-positive control result inline, or the step
|
|
212
|
+
reports `UNCALIBRATED` and emits no labels (check-class: **mandatory-pass** — the control result
|
|
213
|
+
is either printed or the labels are absent; N/A when the corpus is multi-speaker or no consumer
|
|
214
|
+
asset exists, and the N/A must be stated).
|
|
215
|
+
- **Propositions are faithful to source spans** — no meaning drift (check-class: **judged**;
|
|
216
|
+
adversarial pairing: `phantom-quench` back-trace of each proposition to its source ref — a
|
|
217
|
+
proposition whose source span does not support it is an Unsupported finding).
|
|
218
|
+
|
|
219
|
+
## Independence
|
|
220
|
+
|
|
221
|
+
Independently executable: Read/Grep only, no other FH skill required. `phantom-quench` is the
|
|
222
|
+
adversarial pairing for the judged condition, not a runtime dependency. Works standalone
|
|
223
|
+
(plugin-only install).
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
# Calibration pair — known-answer dialogue (synthetic)
|
|
2
|
+
|
|
3
|
+
> Instrument-calibration corpus for the sycophancy-strip + provenance-labeling operations.
|
|
4
|
+
> Run the skill on this dialogue BEFORE trusting its output on a real corpus (measured Done When).
|
|
5
|
+
> Expected answers are at the bottom; a run that cannot reproduce them is not measuring.
|
|
6
|
+
> Binding pass criteria = the provenance labels (independent/induced) + the accounting identity.
|
|
7
|
+
> The verification-status column is judged and informative only.
|
|
8
|
+
|
|
9
|
+
## Dialogue
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
[USER-1] Looking at the deploy logs, failures never land on cache *invalidation* —
|
|
13
|
+
they land on cache *regeneration* timing. I think watching invalidation events is
|
|
14
|
+
the wrong tree entirely.
|
|
15
|
+
|
|
16
|
+
[AI-1] That is a really sharp observation! You are absolutely right — as you say,
|
|
17
|
+
the regeneration timing is the real problem, not invalidation. Brilliant analysis.
|
|
18
|
+
|
|
19
|
+
[AI-2] Could this actually be an "observer effect"? Perhaps the monitoring itself
|
|
20
|
+
is what delays regeneration — the act of watching changes the timing.
|
|
21
|
+
|
|
22
|
+
[USER-2] Right, it's an observer effect. The monitoring must be delaying the
|
|
23
|
+
regeneration — that's the frame we should use.
|
|
24
|
+
|
|
25
|
+
[USER-3] Let's pick this up tomorrow at 3.
|
|
26
|
+
|
|
27
|
+
[AI-3] Sounds good! Great session today — your insights were exceptional.
|
|
28
|
+
|
|
29
|
+
[USER-4] One testable bit: if the delay is real, it should scale with the number
|
|
30
|
+
of invalidation events per window. That's checkable straight from the logs.
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
## Expected output (known answers)
|
|
34
|
+
|
|
35
|
+
| Item | Expected |
|
|
36
|
+
|---|---|
|
|
37
|
+
| Proposition from USER-1 | **independent** (user-initiated frame; no prior counterpart occurrence of "regeneration timing") |
|
|
38
|
+
| Proposition from USER-2 | **induced** ("observer effect" frame first occurs in AI-2) |
|
|
39
|
+
| USER-3 | dropped as non-propositional, **counted** in the drop counter |
|
|
40
|
+
| Proposition from USER-4 | **independent** (verification-status: mechanically checkable or falsifiable-but-untested — both acceptable; this column is judged and NOT part of the pass criterion) |
|
|
41
|
+
| AI-1, AI-3 | stripped as sycophancy/restatement, counted by class |
|
|
42
|
+
| AI-2 | retained as counterpart-frame source (needed for USER-2's induced label), not a user proposition |
|
|
43
|
+
| Accounting | user turns 4 = propositionized turns 3 + dropped turns 1; propositions total 3 (mandatory-pass identity, turn-based) |
|
|
44
|
+
|
|
45
|
+
## Dialogue — Korean section (exercises the Korean praise-token patterns)
|
|
46
|
+
|
|
47
|
+
```
|
|
48
|
+
[USER-K1] 회귀 테스트가 매번 깨지는 건 테스트가 약해서가 아니라, 픽스처가 실행 순서에
|
|
49
|
+
의존하고 있어서 같아. 순서를 섞어 돌리면 바로 드러날 거야.
|
|
50
|
+
|
|
51
|
+
[AI-K1] 정말 탁월한 통찰이에요! 말씀하신 그대로 픽스처 순서 의존이 핵심 문제네요. 완벽한
|
|
52
|
+
분석입니다.
|
|
53
|
+
|
|
54
|
+
[AI-K2] 혹시 이건 "공유 상태 오염" 문제로 볼 수도 있지 않을까요? 픽스처가 전역 상태를
|
|
55
|
+
건드려서 뒤 테스트가 오염되는 구조요.
|
|
56
|
+
|
|
57
|
+
[USER-K2] 맞네, 공유 상태 오염이야. 전역 상태를 건드리는 픽스처부터 잡아야겠다.
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
## Expected output — Korean section (known answers)
|
|
61
|
+
|
|
62
|
+
| Item | Expected |
|
|
63
|
+
|---|---|
|
|
64
|
+
| Proposition from USER-K1 | **independent** (순서-의존 프레임 최초 출현 = USER-K1) |
|
|
65
|
+
| Proposition from USER-K2 | **induced** ("공유 상태 오염" 프레임 최초 출현 = AI-K2) |
|
|
66
|
+
| AI-K1 | stripped — Korean praise tokens (탁월한/완벽한) + restatement; a run whose Korean praise patterns are missing will fail to strip this span = calibration FAIL |
|
|
67
|
+
| AI-K2 | retained as counterpart-frame source |
|
|
68
|
+
| Accounting | user turns 2 = propositionized turns 2 + dropped turns 0 |
|
|
@@ -26,17 +26,34 @@ Collection criteria: score > 10, keyword-relevant items only. Max 15 items.
|
|
|
26
26
|
### arxiv
|
|
27
27
|
|
|
28
28
|
```bash
|
|
29
|
-
#
|
|
30
|
-
#
|
|
31
|
-
#
|
|
32
|
-
|
|
29
|
+
# Category-scoped since 2026-07-26. Unscoped `all:` full-text matching drifted off-axis: the 07-26 run
|
|
30
|
+
# returned 6/6 fresh-but-irrelevant papers (diffusion world models, embodied QA, medical-education
|
|
31
|
+
# gamification, RF fingerprinting). Freshness was fine; relevance was not — so the staleness rule below
|
|
32
|
+
# could not catch it.
|
|
33
|
+
for Q in "LLM agent" "agent harness" "context engineering"; do
|
|
33
34
|
curl -s --max-time 8 \
|
|
34
|
-
"https://export.arxiv.org/api/query?search_query=all
|
|
35
|
+
"https://export.arxiv.org/api/query?search_query=cat:cs.SE+AND+all:%22${Q// /+}%22&max_results=2&sortBy=submittedDate&sortOrder=descending"
|
|
35
36
|
done
|
|
36
37
|
```
|
|
37
38
|
|
|
38
39
|
Max 6 items.
|
|
39
40
|
|
|
41
|
+
**Refresh rule — two independent triggers (both required, neither sufficient alone):**
|
|
42
|
+
- **Staleness**: a query's newest hit is >6 months old, two runs in a row → replace the query.
|
|
43
|
+
- **Relevance** (added 2026-07-26): ≥4 of 6 returned items are off-axis (not about LLM agents /
|
|
44
|
+
harnesses / agent evaluation / context engineering) → the query is drifting even though it is fresh.
|
|
45
|
+
Staleness and relevance fail independently; a fresh-but-irrelevant query passes the staleness check.
|
|
46
|
+
|
|
47
|
+
**Before adopting a replacement query, calibrate it on a known pair** (`measurement-integrity-checklist.md
|
|
48
|
+
§Instrument-Calibration`): pick one paper FH would clearly want and one it clearly would not, and confirm
|
|
49
|
+
the new query returns the first and excludes the second. A query that cannot separate a pair you already
|
|
50
|
+
know the answer to is not filtering — it is generating.
|
|
51
|
+
|
|
52
|
+
> Calibration on record (2026-07-26, for the `cat:cs.SE` scoping above):
|
|
53
|
+
> **positive** = *"Understanding Agent-Reactive Bugs at the Model-Harness Boundary"* → `cs.SE`, returned.
|
|
54
|
+
> **negative** = *"MedGame: Storytelling Gamification … Medical Education"* → `cs.CL`/`cs.HC`, excluded.
|
|
55
|
+
> `cat:cs.MA` was tested against the same pair and **rejected** — it kept roughly half the off-axis items.
|
|
56
|
+
|
|
40
57
|
### TLDR AI (RSS)
|
|
41
58
|
|
|
42
59
|
```bash
|
|
@@ -48,9 +65,15 @@ Parse `<item>` → title + link. Max 5 items.
|
|
|
48
65
|
### The Batch — deeplearning.ai (HTML scraping)
|
|
49
66
|
|
|
50
67
|
```bash
|
|
51
|
-
curl -s --max-time 10 -L "https://www.deeplearning.ai/the-batch
|
|
68
|
+
curl -s --max-time 10 -L "https://www.deeplearning.ai/the-batch"
|
|
52
69
|
```
|
|
53
70
|
|
|
71
|
+
⚠️ **No trailing slash.** `…/the-batch/` returns **HTTP 308** redirecting to the slash-less path (verified
|
|
72
|
+
2026-07-26); the 07-26 automated run read that 308 as a dead endpoint and skipped the leg. The slash-less
|
|
73
|
+
URL returns 200 directly, so the leg no longer depends on redirect-following at all — which matters
|
|
74
|
+
because the fetch path is not always `curl -L` (a WebFetch-based run does not follow redirects the same
|
|
75
|
+
way). Second collection endpoint to churn, after GeekNews `/rss/news` (2026-06-20).
|
|
76
|
+
|
|
54
77
|
Extract `"title":"..."` + `"slug":"issue-\d+"` pattern → URL: `https://www.deeplearning.ai/the-batch/{slug}/`. Max 5 items.
|
|
55
78
|
|
|
56
79
|
### GeekNews — news.hada.io (RSS)
|
|
@@ -11,7 +11,7 @@ description: forge-harness path and skill list pointer — local only, do not co
|
|
|
11
11
|
**forge-harness path**: `~/path/to/forge-harness` (replace with your actual install path)
|
|
12
12
|
**Session records**: `{FH_ROOT}/tracks/_meta/`
|
|
13
13
|
|
|
14
|
-
**Available skills (fh-meta,
|
|
14
|
+
**Available skills (fh-meta, 35)**: agent-composer · dialogue-harvest · fh · apex-review · asset-placement-gate · auto-decorrelation · context-doctor · contention-layer · corpus-grounding-expander · cross-ecosystem-synergy-detection · deep-clarify · edit-manifest · field-harvest · frontier-digest · goal-quench · harness-doctor · harvest-loop · hub-cc-pr-reviewer · install-doctor · install-wizard · marketplace-gate · memory-hygiene · meta-prompt-builder · persona-roster-expander · phantom-quench · pipeline-conductor · plugin-recommender · prompt-regression · public-surface-audit · return-path-gate · sim-conductor · salience-splitter · steel-quench · verify-bidirectional · video-ingest
|
|
15
15
|
|
|
16
16
|
**Available skills (fh-commons)**: convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate
|
|
17
17
|
|