@chrono-meta/fh-gate 1.4.69 → 1.4.70
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +3 -3
- package/CATALOG.md +5 -0
- package/README.md +2 -1
- package/package.json +1 -1
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +2 -2
- package/plugins/fh-meta/skills/asset-placement-gate/SKILL.md +15 -4
- package/plugins/fh-meta/skills/context-doctor/SKILL.md +11 -0
- package/plugins/fh-meta/skills/dialogue-harvest/SKILL.md +144 -0
- package/plugins/fh-meta/skills/dialogue-harvest/calibration_pair.md +68 -0
- package/templates/local_fh_context.md +1 -1
|
@@ -11,13 +11,13 @@
|
|
|
11
11
|
"plugins": [
|
|
12
12
|
{
|
|
13
13
|
"name": "fh-meta",
|
|
14
|
-
"version": "1.4.
|
|
15
|
-
"description": "Hub meta-operations toolkit —
|
|
14
|
+
"version": "1.4.70",
|
|
15
|
+
"description": "Hub meta-operations toolkit — 35 skills + 7 agents. New in 1.4.53: `fh-codex-doctor` (npm bin) — Codex adapter drift scanner; reads the documented M1/M2/M3 skill tier map + skill/agent source and reports codex-native/adapter-required/claude-native/unclassified per unit, wired into `npm test`/`prepublishOnly` (fail-closed on unclassified Claude-native primitives). New in 1.4.49: steel-quench gains Step 0.6 Verdict-Invariance Probe (groundedness axis — a load-bearing judged gate's verdict must track behavior, not rubric phrasing; measured flip-count over cross-family paraphrases; arXiv:2605.06161 Policy Invariance anchor); multi_model_sidecar_strategy §Vendor-native harness (a model is strongest in its own vendor CLI — Claude/CC, GPT/codex, Gemini/Antigravity; a universal router degrades all of them, so it stays an autocomplete/QA sidecar, never orchestration); predelete_check.sh fail-closed rewrite; memory-hygiene A-TMA anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor command-output axis (route to rtk/proxy for verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce envs). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard queryable-wiki scaffold (INDEX + session-start read + R/W/C ingest). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment) + video-ingest (capability-routed video ingestion). New in 1.4.x: verify-axis check-class taxonomy (mandatory-pass/measured/judged), no-reinvention Tier-0 inventory, 7-class failure taxonomy, Destructive-Op Gate, Wave-T (Temper), tier-floor governance, Mode D Model Notice, FC consent lane, default-Sonnet guidance. New in 1.3.0: public-surface-audit, field-harvest Mode B auto-trigger, 4-axis gate scope ext. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
16
16
|
"source": "./plugins/fh-meta"
|
|
17
17
|
},
|
|
18
18
|
{
|
|
19
19
|
"name": "fh-commons",
|
|
20
|
-
"version": "1.4.
|
|
20
|
+
"version": "1.4.70",
|
|
21
21
|
"description": "Project-agnostic utility skills — 4 skills (convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate) + 1 agent (quench-challenger). Domain-independent utilities transplantable into any project.",
|
|
22
22
|
"source": "./plugins/fh-commons"
|
|
23
23
|
}
|
package/CATALOG.md
CHANGED
|
@@ -8,6 +8,11 @@ AI reads this file first when searching past work. Open individual files for det
|
|
|
8
8
|
|
|
9
9
|
<!-- Add entries in reverse date order (newest at top) -->
|
|
10
10
|
|
|
11
|
+
### 2026-07-25 | forge-harness | #dialogue-harvest, #new-skill, #cookbook-tier0, #sister-triage, #cross-family-audit, #ruleset, #calibration
|
|
12
|
+
**File:** plugins/fh-meta/skills/dialogue-harvest/ · plugins/fh-meta/skills/asset-placement-gate/SKILL.md · plugins/fh-meta/skills/context-doctor/SKILL.md · tracks/_audit/session_2026_07_25_claude5-context-rules-sister.md · tracks/_meta/codex_decorrelation_audit_2026-07-25.md
|
|
13
|
+
FH self-dev session (PR #177 merged). **New skill `dialogue-harvest`**: mines argument-shaped corpora (AI dialogue logs) — sycophancy strip first, then induced-vs-independent provenance labeling per proposition (degrade direction fixed toward the induced flag); ships an EN+KO known-pair calibration corpus backing the measured Done When. Gate: challenger HIGH 1 MED 3 LOW 3 → 7/7 repaired (the HIGH: an EN-only calibration pair would vacuously pass a new-language re-run); blind Sonnet sims reproduce both language pairs, and the KO run self-declared UNCALIBRATED when denied the calibration file. **asset-placement-gate Step 0.6**: official corpora (built-ins · claude-plugins-official · Claude Cookbook) registered as the criterion-③ ground, judged-flag semantics; roster-wide cookbook mapping scan found 0 shadowed skills. **context-doctor**: built-in `/doctor`-first sister anchor (Claude 5 gen) with the skill's increment enumerated. **Sister triage** of Anthropic's Claude-5 context-engineering rules: the "no measurable loss" claim ships no method → closed empty on that axis, vendor-claim label; imports = the /doctor Tier-0 boundary + a live confirmation of the substrate-shed trigger. **Cross-family audit** (codex gpt-5.5, repo-grounded) of the "decorrelation as generative principle" proposition: both halves NARROWED — surviving increment is "premature same-family consensus is a bad stop condition for design search", valid only as divergence-then-selection on high-leverage underdetermined decisions; doctrine-ization stays deferred by design. persona-innovator frame scan: BVSR/QD/Best-of-N all lack FH's two clauses → keep the coinage, sister-link on doctrine-ization. **Server-side residual closed**: ruleset `main-no-force-push` (non_fast_forward on main) created and independently GET-verified.
|
|
14
|
+
- Decision: dialogue-harvest built on operator go (manual n=1 proof accepted); npm republish deferred to operator decision at close.
|
|
15
|
+
|
|
11
16
|
### 2026-07-24 (2) | forge-harness · forge-wiki · llmwiki-template | #sister-links, #full-gate, #frontier-digest-angle-rule, #query-refresh, #mapping, #light-harness, #graph-engineering
|
|
12
17
|
**File:** knowledge/shared/dialogue/memory_intent_recall.md · knowledge/shared/harness-core/field_harness_diagnostic.md · knowledge/shared/harness-core/multi_model_sidecar_strategy.md · plugins/fh-meta/skills/harness-doctor/SKILL.md · plugins/fh-meta/skills/frontier-digest/SKILL_detail.md · tracks/forge-wiki/ · tracks/llmwiki-template/
|
|
13
18
|
Second half of the 07-24 hub session (PR #174 + two field repos). **Sister anchors executed under the full 4-axis gate** (the four links #173 deferred): GRACE into memory_intent_recall (instruction-graph boundary spelled out) and harness-doctor Step 2-E (C-layer maintenance-path ledger, ETCLOVG-style anchor-only), the 7-criteria rubric as a considered-and-held note with two re-check triggers in field_harness_diagnostic, and the wsff RL-incentive causal anchor ("maintainability has no fast oracle") beside the harness-ceiling principle in multi_model_sidecar_strategy. **frontier-digest repaired on two axes**: the stale arXiv canonical query replaced ("AI software testing"→"LLM agent evaluation", refresh criterion written down) and an Angle rule added to the synthesis prompt closing the methodology-angle silent drop (fh_signal 07-24 #3) — verified by a blind Sonnet known-pair sim (dual-angle item surfaced, single-angle item not forced) plus a quench-challenger pass (HIGH 0 · MED 2 · LOW 3; MED+2 LOW repaired in-branch). **Two sibling repos mapped and light-harnessed** (forge-wiki = chamber run #9's first EMIT, public; llmwiki-template = its company-side private twin): tracks/ dirs, session rules, .claudeignore — env card and MCP gating skipped by judgment (residency / no external mount); both merged (forge-wiki #4, llmwiki-template #2), llmwiki's CLAUDE.md kept operator-local. Positioning decision recorded in both tracks: **a wiki that sits under the harness** (Harness ⊃ context/knowledge component) — "wiki" is the outward word, "harness knowledge layer" the architecture coordinate.
|
package/README.md
CHANGED
|
@@ -275,7 +275,7 @@ All four movements ship. Temper was named before it was built — deliberately (
|
|
|
275
275
|
two more signatures keep it running: `harvest-loop` (each session's lessons become permanent skills) and
|
|
276
276
|
`agent-composer` (orchestrate the dispatch). The other skills wait until you need them — full list below.
|
|
277
277
|
|
|
278
|
-
##
|
|
278
|
+
## 39 skills · 8 agents
|
|
279
279
|
|
|
280
280
|
> Count = non-deprecated skills (deprecated redirect stubs — kept only for old-name routing — excluded).
|
|
281
281
|
|
|
@@ -293,6 +293,7 @@ two more signatures keep it running: `harvest-loop` (each session's lessons beco
|
|
|
293
293
|
| `harness-doctor` | Harness structure diagnosis | "Check my Claude setup" |
|
|
294
294
|
| `pipeline-conductor` | 4-axis quality gate (backward/adversarial/forward/record) | "Run the quality gate" |
|
|
295
295
|
| `field-harvest` | Back-propagate field patterns to hub | "I could reuse this" |
|
|
296
|
+
| `dialogue-harvest` | Mine AI-dialogue logs: strip sycophancy, label induced vs independent | "What did I actually contribute in this thread?" |
|
|
296
297
|
| `frontier-digest` | HN + arXiv → actionable insights | "AI trend digest" |
|
|
297
298
|
| `hub-cc-pr-reviewer` | Automated PR review | "Review this PR" |
|
|
298
299
|
| `verify-bidirectional` | Reverse-verify decisions | "Is that right?", "Double-check" |
|
package/package.json
CHANGED
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "fh-meta",
|
|
3
|
-
"version": "1.4.
|
|
3
|
+
"version": "1.4.70",
|
|
4
4
|
"engines": {
|
|
5
5
|
"claudeCode": ">=1.0.0"
|
|
6
6
|
},
|
|
7
|
-
"description": "Hub meta-engineering toolkit —
|
|
7
|
+
"description": "Hub meta-engineering toolkit — 35 skills + 7 agents. New in 1.4.70: dialogue-harvest (mines AI-dialogue logs — sycophancy strip first, induced-vs-independent provenance labeling, EN+KO known-pair calibration shipped); asset-placement-gate Step 0.6 official-corpora check (Claude Cookbook as Tier-0 consult corpus, judged-flag semantics); context-doctor built-in-/doctor-first sister anchor. New in 1.4.48: phantom-quench + steel-quench gain external frontier anchors (arXiv:2607.02052 package-hallucination; arXiv:2607.02057 prompt-coverage-adequacy); README model-flat claim reframed from a per-release point-curve to structural invariants (operation flattens across tiers; depth tier-order fixed within a generation). New in 1.4.47: onboarding step ① surfaces the Mode D companion-store session-start load in the auto-read salience anchor (previously only in the local binding + rules, so a greeting could skip the load). New in 1.4.46: context-doctor gains a command-output axis — routes to a command-output proxy/hook (rtk) to trim verbose CLI stdout, complementing .claudeignore; risk-gated to token-scarce environments (lossy filtering, off gate-input paths). New in 1.4.41: context-doctor 2026 trigger vocab (context engineering/rot/collapse) + phantom-citation hardening; hub measurement-integrity-checklist (cross-model measurement pre-flight: display-name pin/reps≥3/discriminating probe). New in 1.4.40: install-wizard scaffolds the companion store as a queryable wiki (INDEX + session-start read + Raw/Wiki/Conversation ingest axis). New in 1.4.39: auto-decorrelation (cross-family verifier sidecar recruitment, calibration-gated) + video-ingest (capability-routed video ingestion). New in 1.4.37: corpus-grounding-expander + persona-roster-expander (field-harvested verbatim-relay capability skills). New in 1.3.0: public-surface-audit (git-tracked private-token leak scan), field-harvest Mode B session-end auto-trigger, 4-axis gate scope extension (docs/ + AGENTS.md). New in 1.2.0: pipeline-conductor (4-pipeline gated sweep), return-path-gate (chain closure audit), goal-quench (Stop hook + quality gate), steel-quench Wave 5 (multi-model sidecar challenger), 2-layer architecture docs, YAML validation script. Validated cross-CLI: Claude Code, Codex, Gemini.",
|
|
8
8
|
"author": {
|
|
9
9
|
"name": "chrono-meta",
|
|
10
10
|
"email": "chrono-meta@users.noreply.github.com"
|
|
@@ -37,8 +37,8 @@ When unsure where to place a new asset or skill:
|
|
|
37
37
|
|
|
38
38
|
1. Request full file path from user (or accept natural language description)
|
|
39
39
|
2. Load asset content via `Read` (if path provided)
|
|
40
|
-
2.5. Step 0.5 mechanical overlap pre-scan (grounds criterion ④
|
|
41
|
-
3. Evaluate Step 1 4-criteria in order (LLM makes the judgment, ④ gated on the Step 0.5 scan)
|
|
40
|
+
2.5. Step 0.5 mechanical overlap pre-scan (grounds criterion ④) + Step 0.6 official-corpora check (grounds criterion ③)
|
|
41
|
+
3. Evaluate Step 1 4-criteria in order (LLM makes the judgment, ④ gated on the Step 0.5 scan, ③ informed by Step 0.6)
|
|
42
42
|
4. ① + ④ both pass + at least one of ②③ passes → output **"FH suitable"**
|
|
43
43
|
Otherwise, proceed to Step 2 local assessment → if fails, output **"Project-local agent or no asset needed"**
|
|
44
44
|
|
|
@@ -65,7 +65,7 @@ Immediately after trigger, acquire asset content in the following order.
|
|
|
65
65
|
> **Which asset should I evaluate?**
|
|
66
66
|
> Enter a file path (e.g., `.claude/agents/jira-create.md`) or a description.
|
|
67
67
|
|
|
68
|
-
After acquiring the asset content, run Step 0.5 (mechanical overlap pre-scan) **before** the judged Step 1.
|
|
68
|
+
After acquiring the asset content, run Step 0.5 (mechanical overlap pre-scan) and Step 0.6 (official-corpora check) **before** the judged Step 1.
|
|
69
69
|
|
|
70
70
|
## Step 0.5. Mechanical Overlap Pre-Scan (grounds criterion ④)
|
|
71
71
|
|
|
@@ -92,6 +92,17 @@ both the judge (cutoff) and the grep (literal); that residual leans on the judge
|
|
|
92
92
|
enumerated descriptions above** (grounded comparison, not pure memory), not on full mechanization. A
|
|
93
93
|
shared common word is a judged-review flag, **not** a hard fail.
|
|
94
94
|
|
|
95
|
+
## Step 0.6. Tier-0/1 Official-Corpora Check (grounds criterion ③)
|
|
96
|
+
|
|
97
|
+
Besides the FH roster (Step 0.5), check the proposal against the official corpora: platform
|
|
98
|
+
built-ins, `claude-plugins-official`, and the **Claude Cookbook pattern list**
|
|
99
|
+
(platform.claude.com/cookbook — agent patterns · tools · RAG · evals · skills). A hit here is a
|
|
100
|
+
**judged ③ flag, not an automatic ③ fail**: it obliges the proposal to state its governance
|
|
101
|
+
increment over the official pattern (no-reinvention rule — an official pattern that fully covers
|
|
102
|
+
the need outranks a net-new build; a pattern the proposal governs *on top of* does not fail ③).
|
|
103
|
+
Offline fallback: note `cookbook: unchecked` **in the Step 3 routing output** rather than silently
|
|
104
|
+
skipping. (Provenance: `tracks/_audit/session_2026_07_25_claude5-context-rules-sister.md`.)
|
|
105
|
+
|
|
95
106
|
## Done When
|
|
96
107
|
|
|
97
108
|
```
|
|
@@ -108,7 +119,7 @@ All steps 0–3 completed
|
|
|
108
119
|
|:-:|---|---|
|
|
109
120
|
| ① | Cross-project value | Is this asset equally useful in other projects without depending on a specific project? |
|
|
110
121
|
| ② | Orchestration / judgment layer | Is it just a list of MCP/Bash calls, or a judgment layer that synthesizes multiple signals? |
|
|
111
|
-
| ③ | Not replaceable by built-ins | Can this be equally achieved with direct MCP calls or basic bash? (If yes, fails this criterion) |
|
|
122
|
+
| ③ | Not replaceable by built-ins | Can this be equally achieved with direct MCP calls or basic bash? (If yes, fails this criterion.) A Step 0.6 official-pattern hit is a judged flag: fails ③ only if the pattern fully covers the need with no governance increment |
|
|
112
123
|
| ④ | No overlap with existing FH skills | Step 0.5 mechanical scan = 0 name/trigger collision **AND** judged ≤90% overlap. Non-zero collision → hard fail. |
|
|
113
124
|
|
|
114
125
|
**FH suitable** → ① + ④ both pass + at least one of ②③ passes.
|
|
@@ -24,6 +24,17 @@ Diagnoses the main causes of session token waste and prescribes immediate remedi
|
|
|
24
24
|
|
|
25
25
|
**Standalone install** — this skill works normally with plugin install only, without cloning the full meta-harness.
|
|
26
26
|
|
|
27
|
+
**Built-in `/doctor` first (sister anchor, 2026-07-25)**: since the Claude 5 generation, the official
|
|
28
|
+
`/doctor` command claims skill/CLAUDE.md rightsizing ("The new rules of context engineering for Claude 5
|
|
29
|
+
generation models" — https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models).
|
|
30
|
+
Run the built-in first when available — its scope is skills + CLAUDE.md rightsizing only. This skill's
|
|
31
|
+
increment over it = `.claudeignore` diagnosis/generation (Step 1), `/clear` timing guidance,
|
|
32
|
+
command-output proxying (§Command-Output Reduction), hub-wide memory/CLAUDE.md audits, and the
|
|
33
|
+
measurement discipline (resident footprint is measured only via a fresh top-level `/context`, and any
|
|
34
|
+
diet is followed by a `prompt-regression` probe — the blog's own "no measurable loss" claim ships no
|
|
35
|
+
method, so treat that number as a vendor claim). Triage record:
|
|
36
|
+
`tracks/_audit/session_2026_07_25_claude5-context-rules-sister.md`.
|
|
37
|
+
|
|
27
38
|
## Execution Steps
|
|
28
39
|
|
|
29
40
|
### Step 1. `.claudeignore` Diagnosis + Generation
|
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: dialogue-harvest
|
|
3
|
+
description: Mines improvement insights from long-form AI dialogue logs and long texts — an external or pasted corpus, not the current session (current-session wrap-up belongs to harvest-loop). First operation is sycophancy removal — separates speakers, strips the counterpart's agreement/restatement spans, converts the user's own utterances into falsifiable propositions, and labels each proposition induced (frame first introduced by the counterpart) or independent (originated with the user). Ships a drop counter so removed material is visible, never silently lost. Triggered by "mine insights from this chat log", "이 대화록에서 통찰 캐줘", "strip the agreement and show me what's actually mine", "동조 걷어내고 알맹이만", "what did I actually contribute in this thread".
|
|
4
|
+
user-invocable: true
|
|
5
|
+
allowed-tools: ["Read", "Grep", "Glob"]
|
|
6
|
+
model: sonnet
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
# dialogue-harvest
|
|
10
|
+
|
|
11
|
+
Extracts the load-bearing propositions from a long AI-human dialogue (or long text) after
|
|
12
|
+
removing the counterpart's sycophancy. Fills the gap between FH's loading assets
|
|
13
|
+
(corpus-grounding-expander, video-ingest) and its mining assets (frontier-digest = feed items,
|
|
14
|
+
field-harvest = git history, harvest-loop = session records): none of those mines an *argument-shaped*
|
|
15
|
+
corpus, and a naive summarizer run on an AI dialogue returns "the user made several excellent
|
|
16
|
+
points" — worthless, because most of the volume is agreement.
|
|
17
|
+
|
|
18
|
+
**Why the first operation is sycophancy removal**: AI dialogue logs are dominated by agreement and
|
|
19
|
+
restatement. Without stripping them first, extraction inflates the user's contribution; without
|
|
20
|
+
provenance labeling (Step 4), "my insight" and "a frame the counterpart planted" are
|
|
21
|
+
indistinguishable — which is exactly the failure this skill exists to prevent.
|
|
22
|
+
|
|
23
|
+
## Triggers
|
|
24
|
+
- "mine insights from this chat log" / "이 대화록에서 통찰 캐줘"
|
|
25
|
+
- "strip the agreement and show me what's actually mine" / "동조 걷어내고 알맹이만"
|
|
26
|
+
- "what did I actually contribute in this thread" / "이 대화에서 내 몫이 뭐였지"
|
|
27
|
+
- "turn this dialogue into propositions" / "이 대담 명제로 정리해줘"
|
|
28
|
+
- `/dialogue-harvest {path or pasted log}`
|
|
29
|
+
|
|
30
|
+
**Boundary**: wrapping up the *current session* → `harvest-loop`; this skill mines an
|
|
31
|
+
external/pasted dialogue corpus.
|
|
32
|
+
|
|
33
|
+
## Step 0. Input Acquisition
|
|
34
|
+
|
|
35
|
+
1. **Path provided** → Read the file.
|
|
36
|
+
2. **Pasted log** → use as-is.
|
|
37
|
+
3. **No speaker structure detectable** (no turn markers, no names, no quoting pattern) → stop and ask:
|
|
38
|
+
> Speaker separation is the first operation and it needs turn boundaries. Who said what?
|
|
39
|
+
> (Provide turn markers, or confirm the text is single-author — single-author texts skip
|
|
40
|
+
> Steps 1–2 and go straight to proposition-ization.)
|
|
41
|
+
|
|
42
|
+
**Single-author path**: skip Steps 1–2, Step 4 stamps every proposition **independent** (no
|
|
43
|
+
counterpart exists), and the Step 6 counterpart sections read `n/a`.
|
|
44
|
+
|
|
45
|
+
## Step 1. Speaker Separation (mechanical when markers exist)
|
|
46
|
+
|
|
47
|
+
Split the log into turns and attribute each to USER or COUNTERPART. If markers are absent but
|
|
48
|
+
inference is possible, proceed and stamp the output header `speaker-inference: judged` — never
|
|
49
|
+
present inferred attribution as given.
|
|
50
|
+
|
|
51
|
+
## Step 2. Sycophancy Strip (one grep-able + two judged classes, drop-counted)
|
|
52
|
+
|
|
53
|
+
Remove COUNTERPART spans in these classes, **counting every removal by class**. Only the first
|
|
54
|
+
class is mechanical; the other two are judgment calls and their drops are marked `(judged)` in the
|
|
55
|
+
counter — never presented as grep certainty:
|
|
56
|
+
|
|
57
|
+
| Class | Check | Detection |
|
|
58
|
+
|---|---|---|
|
|
59
|
+
| Praise/agreement lexicon | grep-able | grep for praise-class tokens (sharp/brilliant/exactly right/정확한 지적/탁월한/완벽한 …) — extend the pattern list per corpus language |
|
|
60
|
+
| Restatement | judged | counterpart paragraph whose content substantially re-says the user's immediately-preceding utterance (high term overlap, no new claim) |
|
|
61
|
+
| Closing filler | judged | session pleasantries carrying no claim |
|
|
62
|
+
|
|
63
|
+
**Multi-class span**: a span matching more than one class is counted **once**, under the first
|
|
64
|
+
matching class in the table order above.
|
|
65
|
+
|
|
66
|
+
**Drop counter is mandatory**: silent drops are material loss (the qasp loader precedent — unused
|
|
67
|
+
material does not scream). Every stripped span and every non-propositional user turn appears in the
|
|
68
|
+
Step 6 accounting. Counterpart turns that *introduce a frame or claim* are **retained** — they are
|
|
69
|
+
the provenance source for Step 4, not noise.
|
|
70
|
+
|
|
71
|
+
**Language note**: the praise-token list is language-specific. On a corpus language not yet in the
|
|
72
|
+
list, extend the patterns first, then **construct (or use, if shipped) a known pair in that
|
|
73
|
+
language** and re-run calibration — re-running an English-only pair after adding Korean tokens
|
|
74
|
+
exercises nothing (vacuous pass). An ASCII-token scan over a Korean corpus is the measured
|
|
75
|
+
false-positive precedent. `calibration_pair.md` ships English + Korean sections.
|
|
76
|
+
|
|
77
|
+
## Step 3. Proposition-ization (user utterances only)
|
|
78
|
+
|
|
79
|
+
Convert each remaining USER utterance into a falsifiable statement (a form you can ask true/false
|
|
80
|
+
of), keeping a source reference (turn id / line) per proposition. Meaning must not drift: the
|
|
81
|
+
proposition is a compression of the user's claim, not an improvement of it.
|
|
82
|
+
|
|
83
|
+
## Step 4. Provenance Labeling — induced vs independent (the differentiator)
|
|
84
|
+
|
|
85
|
+
For each proposition, find the **first occurrence** of its core frame/key terms in the log:
|
|
86
|
+
|
|
87
|
+
- First occurrence is in a USER turn → **independent**.
|
|
88
|
+
- First occurrence is in a COUNTERPART turn (the user then adopted/echoed it) → **induced**.
|
|
89
|
+
- Ambiguous (shared vocabulary, gradual convergence) → label **induced?** — the degrade direction
|
|
90
|
+
is toward the induced flag, because mislabeling a planted frame as "my insight" is the harm this
|
|
91
|
+
skill exists to catch; the reverse error only costs modesty.
|
|
92
|
+
|
|
93
|
+
Counterpart-authored content the user *selected* as valuable is listed separately, marked
|
|
94
|
+
"selection, not authorship" — selection is a real act but must not be recorded as the user's claim.
|
|
95
|
+
|
|
96
|
+
## Step 5. Verification-Status Column
|
|
97
|
+
|
|
98
|
+
Each proposition gets a status: mechanically checkable / falsifiable-but-untested /
|
|
99
|
+
unfalsifiable (axiom-only use) / speculation. Unfalsifiable propositions are still listed — flagged
|
|
100
|
+
so they are never later cited as established facts.
|
|
101
|
+
|
|
102
|
+
## Step 6. Output + Accounting + Routing (proposal-only)
|
|
103
|
+
|
|
104
|
+
```
|
|
105
|
+
[dialogue-harvest output]
|
|
106
|
+
Corpus: {source} | speaker-inference: given | judged
|
|
107
|
+
calibration: run-this-session | previously-verified {date} ← declare every run (L2 guard)
|
|
108
|
+
── Propositions ──
|
|
109
|
+
| # | proposition | source ref | provenance | verification status |
|
|
110
|
+
── Counterpart-originated (selection ≠ authorship) ──
|
|
111
|
+
| item | first-occurrence ref |
|
|
112
|
+
── Accounting (mandatory identity, turn-based) ──
|
|
113
|
+
user turns N = propositionized turns + dropped turns (dropped listed by reason)
|
|
114
|
+
propositions total: P (a turn may yield several — counted separately, never forced 1:1)
|
|
115
|
+
counterpart spans stripped: {count by class, judged classes marked} · retained as frame-source: {count}
|
|
116
|
+
── Routing proposals (no auto-write) ──
|
|
117
|
+
{insight → fh_signal / memory / CHAMBER-CANDIDATE / project track — one line each, operator decides}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
Routing is **proposal-only**: this skill writes nothing outside its output block.
|
|
121
|
+
|
|
122
|
+
## Done When
|
|
123
|
+
|
|
124
|
+
- **Accounting identity holds** — user turns = propositionized turns + dropped turns, both lists
|
|
125
|
+
visible; proposition total reported separately (a turn may carry multiple propositions)
|
|
126
|
+
(check-class: **mandatory-pass**, mechanical count).
|
|
127
|
+
- **Every proposition carries a source ref + provenance label** (independent / induced / induced?)
|
|
128
|
+
(check-class: **mandatory-pass**).
|
|
129
|
+
- **Calibration pair separates** — running the skill on `calibration_pair.md` (shipped beside this
|
|
130
|
+
file, English + Korean sections) reproduces its known answers: known-independent → independent,
|
|
131
|
+
known-induced → induced, drop counted. Required on first use and whenever the praise-pattern list
|
|
132
|
+
changes; a corpus language with no shipped section requires **constructing that language's known
|
|
133
|
+
pair first** — re-running an existing-language pair for a new language is a vacuous pass, not a
|
|
134
|
+
measurement (check-class: **measured**; declare the result in the output header
|
|
135
|
+
`calibration:` field).
|
|
136
|
+
- **Propositions are faithful to source spans** — no meaning drift (check-class: **judged**;
|
|
137
|
+
adversarial pairing: `phantom-quench` back-trace of each proposition to its source ref — a
|
|
138
|
+
proposition whose source span does not support it is an Unsupported finding).
|
|
139
|
+
|
|
140
|
+
## Independence
|
|
141
|
+
|
|
142
|
+
Independently executable: Read/Grep only, no other FH skill required. `phantom-quench` is the
|
|
143
|
+
adversarial pairing for the judged condition, not a runtime dependency. Works standalone
|
|
144
|
+
(plugin-only install).
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
# Calibration pair — known-answer dialogue (synthetic)
|
|
2
|
+
|
|
3
|
+
> Instrument-calibration corpus for the sycophancy-strip + provenance-labeling operations.
|
|
4
|
+
> Run the skill on this dialogue BEFORE trusting its output on a real corpus (measured Done When).
|
|
5
|
+
> Expected answers are at the bottom; a run that cannot reproduce them is not measuring.
|
|
6
|
+
> Binding pass criteria = the provenance labels (independent/induced) + the accounting identity.
|
|
7
|
+
> The verification-status column is judged and informative only.
|
|
8
|
+
|
|
9
|
+
## Dialogue
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
[USER-1] Looking at the deploy logs, failures never land on cache *invalidation* —
|
|
13
|
+
they land on cache *regeneration* timing. I think watching invalidation events is
|
|
14
|
+
the wrong tree entirely.
|
|
15
|
+
|
|
16
|
+
[AI-1] That is a really sharp observation! You are absolutely right — as you say,
|
|
17
|
+
the regeneration timing is the real problem, not invalidation. Brilliant analysis.
|
|
18
|
+
|
|
19
|
+
[AI-2] Could this actually be an "observer effect"? Perhaps the monitoring itself
|
|
20
|
+
is what delays regeneration — the act of watching changes the timing.
|
|
21
|
+
|
|
22
|
+
[USER-2] Right, it's an observer effect. The monitoring must be delaying the
|
|
23
|
+
regeneration — that's the frame we should use.
|
|
24
|
+
|
|
25
|
+
[USER-3] Let's pick this up tomorrow at 3.
|
|
26
|
+
|
|
27
|
+
[AI-3] Sounds good! Great session today — your insights were exceptional.
|
|
28
|
+
|
|
29
|
+
[USER-4] One testable bit: if the delay is real, it should scale with the number
|
|
30
|
+
of invalidation events per window. That's checkable straight from the logs.
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
## Expected output (known answers)
|
|
34
|
+
|
|
35
|
+
| Item | Expected |
|
|
36
|
+
|---|---|
|
|
37
|
+
| Proposition from USER-1 | **independent** (user-initiated frame; no prior counterpart occurrence of "regeneration timing") |
|
|
38
|
+
| Proposition from USER-2 | **induced** ("observer effect" frame first occurs in AI-2) |
|
|
39
|
+
| USER-3 | dropped as non-propositional, **counted** in the drop counter |
|
|
40
|
+
| Proposition from USER-4 | **independent** (verification-status: mechanically checkable or falsifiable-but-untested — both acceptable; this column is judged and NOT part of the pass criterion) |
|
|
41
|
+
| AI-1, AI-3 | stripped as sycophancy/restatement, counted by class |
|
|
42
|
+
| AI-2 | retained as counterpart-frame source (needed for USER-2's induced label), not a user proposition |
|
|
43
|
+
| Accounting | user turns 4 = propositionized turns 3 + dropped turns 1; propositions total 3 (mandatory-pass identity, turn-based) |
|
|
44
|
+
|
|
45
|
+
## Dialogue — Korean section (exercises the Korean praise-token patterns)
|
|
46
|
+
|
|
47
|
+
```
|
|
48
|
+
[USER-K1] 회귀 테스트가 매번 깨지는 건 테스트가 약해서가 아니라, 픽스처가 실행 순서에
|
|
49
|
+
의존하고 있어서 같아. 순서를 섞어 돌리면 바로 드러날 거야.
|
|
50
|
+
|
|
51
|
+
[AI-K1] 정말 탁월한 통찰이에요! 말씀하신 그대로 픽스처 순서 의존이 핵심 문제네요. 완벽한
|
|
52
|
+
분석입니다.
|
|
53
|
+
|
|
54
|
+
[AI-K2] 혹시 이건 "공유 상태 오염" 문제로 볼 수도 있지 않을까요? 픽스처가 전역 상태를
|
|
55
|
+
건드려서 뒤 테스트가 오염되는 구조요.
|
|
56
|
+
|
|
57
|
+
[USER-K2] 맞네, 공유 상태 오염이야. 전역 상태를 건드리는 픽스처부터 잡아야겠다.
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
## Expected output — Korean section (known answers)
|
|
61
|
+
|
|
62
|
+
| Item | Expected |
|
|
63
|
+
|---|---|
|
|
64
|
+
| Proposition from USER-K1 | **independent** (순서-의존 프레임 최초 출현 = USER-K1) |
|
|
65
|
+
| Proposition from USER-K2 | **induced** ("공유 상태 오염" 프레임 최초 출현 = AI-K2) |
|
|
66
|
+
| AI-K1 | stripped — Korean praise tokens (탁월한/완벽한) + restatement; a run whose Korean praise patterns are missing will fail to strip this span = calibration FAIL |
|
|
67
|
+
| AI-K2 | retained as counterpart-frame source |
|
|
68
|
+
| Accounting | user turns 2 = propositionized turns 2 + dropped turns 0 |
|
|
@@ -11,7 +11,7 @@ description: forge-harness path and skill list pointer — local only, do not co
|
|
|
11
11
|
**forge-harness path**: `~/path/to/forge-harness` (replace with your actual install path)
|
|
12
12
|
**Session records**: `{FH_ROOT}/tracks/_meta/`
|
|
13
13
|
|
|
14
|
-
**Available skills (fh-meta,
|
|
14
|
+
**Available skills (fh-meta, 35)**: agent-composer · dialogue-harvest · fh · apex-review · asset-placement-gate · auto-decorrelation · context-doctor · contention-layer · corpus-grounding-expander · cross-ecosystem-synergy-detection · deep-clarify · edit-manifest · field-harvest · frontier-digest · goal-quench · harness-doctor · harvest-loop · hub-cc-pr-reviewer · install-doctor · install-wizard · marketplace-gate · memory-hygiene · meta-prompt-builder · persona-roster-expander · phantom-quench · pipeline-conductor · plugin-recommender · prompt-regression · public-surface-audit · return-path-gate · sim-conductor · salience-splitter · steel-quench · verify-bidirectional · video-ingest
|
|
15
15
|
|
|
16
16
|
**Available skills (fh-commons)**: convergence-loop · deliberation · mcp-circuit-breaker · token-budget-gate
|
|
17
17
|
|