@chrono-meta/fh-gate 1.4.95 → 1.4.96
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +18 -0
- package/CHEATSHEET.md +1 -1
- package/knowledge/shared/harness-core/fh_detail_protocols.md +12 -0
- package/knowledge/shared/harness-core/ship_readiness_gate.md +7 -4
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +43 -1
- package/package.json +6 -1
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-commons/agents/quench-challenger.md +49 -23
- package/plugins/fh-commons/skills/convergence-loop/SKILL.md +14 -0
- package/plugins/fh-commons/skills/deliberation/SKILL.md +14 -0
- package/plugins/fh-commons/skills/mcp-circuit-breaker/SKILL.md +10 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +36 -0
- package/plugins/fh-meta/agents/beginner.md +4 -1
- package/plugins/fh-meta/agents/challenger.md +7 -1
- package/plugins/fh-meta/agents/expert.md +1 -1
- package/plugins/fh-meta/agents/fact-checker.md +7 -1
- package/plugins/fh-meta/agents/hub-persona-auditor.md +2 -1
- package/plugins/fh-meta/agents/main-player.md +4 -1
- package/plugins/fh-meta/agents/persona-innovator.md +10 -2
- package/plugins/fh-meta/skills/agent-composer/SKILL.md +2 -2
- package/plugins/fh-meta/skills/apex-review/SKILL.md +5 -0
- package/plugins/fh-meta/skills/asset-placement-gate/SKILL.md +38 -8
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +16 -2
- package/plugins/fh-meta/skills/context-doctor/SKILL_detail.md +45 -10
- package/plugins/fh-meta/skills/corpus-grounding-expander/SKILL.md +14 -5
- package/plugins/fh-meta/skills/cross-ecosystem-synergy-detection/SKILL.md +93 -30
- package/plugins/fh-meta/skills/deep-clarify/SKILL.md +28 -9
- package/plugins/fh-meta/skills/fh/SKILL.md +4 -0
- package/plugins/fh-meta/skills/frontier-digest/SKILL.md +64 -8
- package/plugins/fh-meta/skills/frontier-digest/SKILL_detail.md +20 -7
- package/plugins/fh-meta/skills/goal-quench/SKILL.md +48 -15
- package/plugins/fh-meta/skills/goal-quench/SKILL_detail.md +58 -11
- package/plugins/fh-meta/skills/harness-doctor/SKILL_detail.md +109 -33
- package/plugins/fh-meta/skills/harvest-loop/SKILL.md +6 -1
- package/plugins/fh-meta/skills/hub-cc-pr-reviewer/SKILL.md +126 -17
- package/plugins/fh-meta/skills/install-doctor/SKILL.md +50 -14
- package/plugins/fh-meta/skills/install-wizard/SKILL.md +26 -7
- package/plugins/fh-meta/skills/install-wizard/SKILL_detail.md +68 -21
- package/plugins/fh-meta/skills/memory-hygiene/SKILL.md +64 -17
- package/plugins/fh-meta/skills/meta-prompt-builder/SKILL.md +38 -4
- package/plugins/fh-meta/skills/persona-roster-expander/SKILL.md +15 -7
- package/plugins/fh-meta/skills/plugin-recommender/SKILL.md +39 -11
- package/plugins/fh-meta/skills/plugin-recommender/SKILL_detail.md +24 -7
- package/plugins/fh-meta/skills/prompt-regression/SKILL.md +54 -11
- package/plugins/fh-meta/skills/salience-splitter/SKILL.md +120 -7
- package/plugins/fh-meta/skills/salience-splitter/SKILL_detail.md +46 -13
- package/plugins/fh-meta/skills/sim-conductor/SKILL_detail.md +28 -3
- package/plugins/fh-meta/skills/steel-quench/SKILL.md +3 -1
- package/plugins/fh-meta/skills/verify-bidirectional/SKILL.md +72 -14
- package/scripts/count_check.sh +47 -1
- package/scripts/degrade_direction_scan.sh +276 -6
- package/scripts/degrade_probe_capability.sh +105 -0
- package/scripts/package_coverage_check.sh +8 -0
- package/scripts/psa_probe_capability.sh +78 -0
- package/scripts/public_surface_scan_files.sh +8 -0
- package/scripts/selfcheck.sh +15 -0
- package/scripts/test_capability_entrypoint_shipping.sh +132 -0
- package/scripts/test_count_check_readme_format_lanes.sh +75 -0
- package/scripts/test_degrade_scan_shell_probes.sh +415 -0
- package/scripts/validate_yaml.sh +146 -0
- package/templates/degrade_direction_scan.sh +276 -6
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: persona-roster-expander
|
|
3
3
|
description: Expands a named persona seed into a tiered, judgment-mapped cast — tiering each persona by a domain safety rule, mapping each to a decision-lens in the user's vocabulary, then proposing additional voices with sourced anchors.
|
|
4
4
|
user-invocable: true
|
|
5
|
-
allowed-tools: ["Read", "Grep", "WebSearch", "WebFetch"]
|
|
5
|
+
allowed-tools: ["Read", "Grep", "WebSearch", "WebFetch", "Write", "Agent"]
|
|
6
6
|
model: sonnet
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -15,7 +15,12 @@ caller-supplied safety rule so the expansion stays faithful to the domain's cons
|
|
|
15
15
|
|
|
16
16
|
> Origin: harvested from the-bible (2026-06-20) — an operator persona seed (priest/nun/angel/devil/
|
|
17
17
|
> God/Jesus/Holy-Spirit/apostles) tiered relay-vs-lens by a relay-safety rule and mapped to
|
|
18
|
-
> engineering-judgment lenses, +4 sourced proposals.
|
|
18
|
+
> engineering-judgment lenses, +4 sourced proposals. **The grounds, inline, because the harvest
|
|
19
|
+
> record does not ship**: an ungated persona cast drifts into inventing its own authority — a voice
|
|
20
|
+
> given a lens will speak past what the domain lets it assert — so tiering by a *caller-supplied*
|
|
21
|
+
> rule is what keeps the expansion faithful, and the lens mapping is what makes a named voice
|
|
22
|
+
> usable as a decision instrument instead of flavor. Full harvest record — **hub-local, not
|
|
23
|
+
> distributed in the npm package**:
|
|
19
24
|
> `tracks/_contrib/field_harvest_2026-06-20_gate-locality-and-grounding-capabilities.md`.
|
|
20
25
|
|
|
21
26
|
## Triggers
|
|
@@ -41,14 +46,17 @@ caller-supplied safety rule so the expansion stays faithful to the domain's cons
|
|
|
41
46
|
4. **Propose 2–4 additions** filling lenses the seed doesn't cover (delegate net-new *name*
|
|
42
47
|
generation to the `persona-innovator` agent — this skill's distinct value is the tiering +
|
|
43
48
|
lens-mapping, not naming). For the strongest 2–3, find a real anchor (a sourced
|
|
44
|
-
example/reference)
|
|
45
|
-
|
|
46
|
-
|
|
49
|
+
example/reference); **any remaining proposal ships with a literal `stub:` prefix** — unanchored
|
|
50
|
+
and unlabeled is not an allowed output (see Done When).
|
|
51
|
+
5. **Emit the tiered, lens-mapped cast** (named + proposed) by **writing it to a structured file**
|
|
52
|
+
the system can load (e.g. `personas.json`), then confirm the written file parses. Leaving the
|
|
53
|
+
cast in the response text only does not complete this step.
|
|
47
54
|
|
|
48
55
|
## Done When
|
|
49
56
|
- **Every persona has a tier + lens label + invoke-condition.** *Check class: mandatory-pass (binary — all three fields present per persona).*
|
|
50
|
-
- **
|
|
51
|
-
- **The
|
|
57
|
+
- **Every proposal is either anchored or explicitly labeled a stub.** An anchored proposal names a source/reference that resolves; a proposal without one is emitted with a literal `stub:` prefix and is **not counted as a candidate**. *Check class: mandatory-pass (binary — each proposal carries either a resolving anchor or a `stub:` label; an unanchored, unlabeled proposal is FAIL). Anchor resolution is checked mechanically (fetch/look up the cited reference), not by judgment.* (This is the Done-When floor; Step 4's "strongest 2–3" is the effort target that sits above it — a 4th proposal may ship as `stub:` without violating either.)
|
|
58
|
+
- **The cast is materialized as a loadable artifact** (e.g. `personas.json`) that a consumer can parse — the file exists on disk and a parse of it succeeds against the declared schema (per persona: tier / lens / invoke-condition; per proposal additionally `anchor` or `stub:`). *Check class: mandatory-pass (binary — file exists and parses; a cast that exists only in the response text is UNMET).*
|
|
59
|
+
- **The tiering respects the caller's safety rule** (no persona exceeds its tier's allowed emission). *Check class: judged, pair: dispatch `fh-meta:challenger` at the tier-escalation angle — "find a persona whose lens mapping lets it emit beyond its tier". The pairing is a **different-agent** read; the author's own re-read does not satisfy it. If that agent is unreachable, record `pair: unavailable (<reason>)` — the condition stays UNMET and is **reported as a named residual**, never self-scored closed. It does **not** block delivery: a roster is a reversible artifact, so the degrade direction here is declare-and-ship, not fail-closed (that direction is reserved for irreversible surfaces — publish, delete, history-rewrite). Shipping with an UNMET pairing is honest; silently marking it met is the defect.*
|
|
52
60
|
|
|
53
61
|
## Guards
|
|
54
62
|
- **Caller-supplied safety rule is mandatory** — the skill tiers by the domain's rule, it does not
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: plugin-recommender
|
|
3
3
|
description: Given a task description, searches internal and external open-source ecosystems (including Codex marketplace and Claude Code marketplace) to find and recommend suitable plugins with installation guidance. Recommendation is quality-validation based (marketplace-listed + performance-validated), not source-origin based. Activates on "recommend a plugin", "what tool should I use?", "is there a plugin for this?", "recommend a tool". Also checks for duplicate installations.
|
|
4
4
|
user-invocable: true
|
|
5
|
-
allowed-tools: ["
|
|
5
|
+
allowed-tools: ["Read", "Grep", "Bash", "WebSearch", "WebFetch"]
|
|
6
6
|
model: sonnet
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -59,7 +59,9 @@ Tier is **independent of platform origin** (Anthropic / OpenAI / community). A w
|
|
|
59
59
|
|
|
60
60
|
2. **[Priority 1.5] Internal GHE Sister Assets (partially completed work — Tier 2)**: Check if user's task domain already exists in internal sister asset clusters. Prioritize direct use or adoption of sister assets if the user's task falls within these cluster domains.
|
|
61
61
|
|
|
62
|
-
3. **[Priority 2
|
|
62
|
+
3. **[Priority 2] Organization's Internal GHE (Tier 1–4)**: Search internal GHE with keywords like `claude-plugin`, `gemini-plugin` + user keywords / API search. Replace with your organization's internal GHE orgs.
|
|
63
|
+
|
|
64
|
+
4. **[Priority 2.5] Project Reference/Contribution Path**:
|
|
63
65
|
|
|
64
66
|
For cases where referencing or contributing to the project itself is more appropriate than installing a plugin. Provide guidance when there's intent for medium-term contribution rather than immediate use.
|
|
65
67
|
|
|
@@ -71,9 +73,7 @@ Tier is **independent of platform origin** (Anthropic / OpenAI / community). A w
|
|
|
71
73
|
- `plugin-recommender` [Priority 2.5]: When no immediately usable plugin is available → guide to project-level reference/contribution path (alternative to installing)
|
|
72
74
|
- `cross-ecosystem-synergy-detection`: Discover hidden synergies among already-installed skills (post-install utilization optimization)
|
|
73
75
|
|
|
74
|
-
Activation condition: Automatically entered when Step 2 [Priority 1]~[Priority 2] search yields no suitable plugin.
|
|
75
|
-
|
|
76
|
-
4. **[Priority 2] Organization's Internal GHE (Tier 1–4)**: Search internal GHE with keywords like `claude-plugin`, `gemini-plugin` + user keywords / API search. Replace with your organization's internal GHE orgs.
|
|
76
|
+
Activation condition: Automatically entered when Step 2 [Priority 1]~[Priority 2] search yields no suitable plugin. **This block therefore runs only after [Priority 2] has actually been executed** — the list above is in execution order (1 → 1.5 → 2 → 2.5 → 3), and the priority labels are names, not the run sequence.
|
|
77
77
|
|
|
78
78
|
5. **[Priority 3] External Open-Source Ecosystem**: WebSearch / WebFetch — "best github actions for X", "claude plugin for Y", etc. Simplification guard: defer external install if internal assets suffice.
|
|
79
79
|
|
|
@@ -91,16 +91,18 @@ When queried for a specific capability (e.g., "adversarial reviewer for bash cod
|
|
|
91
91
|
0. **Platform built-ins (Tier 0)** — does a built-in skill/command already cover the capability? Check the live session's available-skills list before any plugin search. A built-in that covers ~80% beats installing a plugin for the rest
|
|
92
92
|
1. **Installed locally** — `.claude/agents/`, `plugins/` in current cwd
|
|
93
93
|
2. **FH native skills** — always-loaded knowledge in `plugins/fh-meta/` and `plugins/fh-commons/`
|
|
94
|
-
3. **Claude Code marketplace** — `claude
|
|
95
|
-
4. **Codex marketplace** — `npx @openai/codex list
|
|
94
|
+
3. **Claude Code marketplace** — `claude plugin list --available --json` (marketplace plugins; `--available` **requires** `--json`) + `claude plugin marketplace list` to see which marketplaces are even configured. There is **no** CLI keyword-search subcommand — filter the JSON yourself, and fall back to the known CC registry (see verified targets above) for anything not in a configured marketplace
|
|
95
|
+
4. **Codex marketplace** — `npx --yes @openai/codex plugin list --available --json` (optionally `-m <marketplace>`) + `npx --yes @openai/codex plugin marketplace list`. Same limitation: no keyword search, filter the JSON
|
|
96
96
|
5. **npm ecosystem** — `@chrono-meta/`, `@anthropic/`, and other known-quality scoped packages
|
|
97
97
|
|
|
98
|
+
⚠️ **A failed or empty discovery lane is NOT "no candidates".** Both CLIs above list only *configured* marketplaces, and a non-zero exit / empty array means the lane did not answer — not that nothing exists. Report each lane's state explicitly (`EXECUTED` / `EMPTY` / `FAILED: <stderr>`) and never render a `FAILED` lane as a zero result; a lane that could not run must be re-run or replaced by the web-search fallback (Priority 3) before you tell the user nothing was found.
|
|
99
|
+
|
|
98
100
|
**Discovery priority**: built-in (Tier 0) > installed > FH native > Tier 1 (any platform) > Tier 2 > Tier 3 > Tier 4
|
|
99
101
|
**Tier 0 guard**: FH native wins over a built-in only when the FH skill adds governance the built-in lacks (e.g. `/goal` → `goal-quench` adds budget+quality gates; code diff review stays with built-in `/code-review`, FH-asset coherence with `hub-cc-pr-reviewer`)
|
|
100
102
|
|
|
101
103
|
**When sim-conductor chains here for persona discovery**: apply the same platform-aware search scoped to persona/simulation/review capability tags. Return discovered agents with their Tier rating so sim-conductor can decide whether to install or use a built-in brief.
|
|
102
104
|
|
|
103
|
-
For discovery bash commands (`claude
|
|
105
|
+
For discovery bash commands (`claude plugin list --available --json`, `npx --yes @openai/codex plugin list --available --json`, npm scoped search), see `SKILL_detail.md §Discovery-Bash`.
|
|
104
106
|
|
|
105
107
|
### Step 2.6: Quality Validation Signals
|
|
106
108
|
|
|
@@ -149,7 +151,8 @@ When user selects desired plugin from recommendation list, help with installatio
|
|
|
149
151
|
## Constraints
|
|
150
152
|
|
|
151
153
|
- **Recommendations, not guarantees**: Does not guarantee plugin performance, stability, or security.
|
|
152
|
-
- **Inbound supply-chain risk
|
|
154
|
+
- **Inbound supply-chain risk — EVERY tier, not just 3/4**: public skill registries have shipped malicious skills at scale (pointers: [HN 47370624](https://news.ycombinator.com/item?id=47370624) — 824 malicious skills reported on ClawHub; [HN 48678603](https://news.ycombinator.com/item?id=48678603) — Snyk ToxicSkills study. Secondary sources — verify before quoting figures). ⚠️ **The channel in both incidents was the marketplace listing itself — i.e. Tier 1/2.** An earlier version of this bullet scoped the warning to Tier 3/4, which inverted its own evidence: it warned loudest exactly where vetting exists and stayed silent where the cited attacks landed. Flag the risk before recommending **any** candidate; a high Tier means *listed and maintained*, never *audited for intent*.
|
|
155
|
+
- **Install string must be resolved, not relayed**: the `<name>` in an install command frequently comes from a web-search result. Before surfacing it, confirm the name resolves to the intended repository (owner/repo matches the source you are citing) — a plausible near-name is the cheapest form of this attack, and installation is irreversible on the machine that runs it.
|
|
153
156
|
- **User consent required**: Does not auto-install without explicit consent.
|
|
154
157
|
- **Search scope limitations**: Only searches within configured search space.
|
|
155
158
|
|
|
@@ -213,9 +216,34 @@ sim-conductor needs persona X (no installed/built-in match)
|
|
|
213
216
|
|
|
214
217
|
```
|
|
215
218
|
All Steps 0~5 completed
|
|
216
|
-
|
|
217
|
-
|
|
219
|
+
— mandatory-pass: each step produced its stated output, or is
|
|
220
|
+
explicitly marked N/A with the reason
|
|
221
|
+
|
|
222
|
+
+ Recommendation list table output (top 2~3 items, Tier + Platform + synergy
|
|
223
|
+
grade included)
|
|
224
|
+
— mandatory-pass: the table exists with all three columns populated
|
|
225
|
+
|
|
226
|
+
+ Every discovery lane reported with an explicit state
|
|
227
|
+
(EXECUTED / EMPTY / FAILED: <stderr>)
|
|
228
|
+
— measured: count lanes attempted vs lanes reporting a state; the two
|
|
229
|
+
numbers must match. A FAILED lane rendered as "no candidates" is a FAIL
|
|
230
|
+
of this condition, not a pass — the CLI lanes have no keyword search and
|
|
231
|
+
see only configured marketplaces, so an empty result is routinely a
|
|
232
|
+
non-answer rather than an absence (`not found` != `0`)
|
|
233
|
+
|
|
234
|
+
+ Install completed after user selection (or install skipped / 5-B migration
|
|
235
|
+
path guided)
|
|
236
|
+
— mandatory-pass: explicit user consent recorded before any install ran
|
|
237
|
+
|
|
218
238
|
+ Duplicate detection results reported
|
|
239
|
+
— mandatory-pass: `claude plugin list` output consulted in this run
|
|
240
|
+
|
|
241
|
+
+ Every install string surfaced to the user resolves to the repository it is
|
|
242
|
+
cited from (§Constraints — resolved, not relayed)
|
|
243
|
+
— judged; adversarial pairing: in the same run, resolve one candidate name
|
|
244
|
+
you already know is correct AND check one near-name variant. If the
|
|
245
|
+
procedure cannot separate that pair, the resolution check is
|
|
246
|
+
UNCALIBRATED and no install string may be presented as verified
|
|
219
247
|
```
|
|
220
248
|
|
|
221
249
|
## Failure Response
|
|
@@ -57,16 +57,24 @@ gh auth status # default host (github.com) only
|
|
|
57
57
|
# Unauthenticated → guide github.com PAT generation above
|
|
58
58
|
```
|
|
59
59
|
|
|
60
|
-
**Claude Code marketplace search
|
|
60
|
+
**Claude Code marketplace discovery** (there is **no** `search` subcommand — `claude mcp search` does not
|
|
61
|
+
exist and exits 1 with `unknown command 'search'`; list, then filter yourself):
|
|
61
62
|
```bash
|
|
62
|
-
claude
|
|
63
|
+
claude plugin marketplace list --json # which marketplaces are configured at all
|
|
64
|
+
claude plugin list --available --json # installed + available marketplace plugins (--available REQUIRES --json)
|
|
63
65
|
```
|
|
64
66
|
|
|
65
|
-
**Codex marketplace
|
|
67
|
+
**Codex marketplace discovery** (`list-agents` does not exist either — the real noun is `plugin`):
|
|
66
68
|
```bash
|
|
67
|
-
npx @openai/codex list
|
|
69
|
+
npx --yes @openai/codex plugin marketplace list
|
|
70
|
+
npx --yes @openai/codex plugin list --available --json # add -m <marketplace> to scope
|
|
68
71
|
```
|
|
69
72
|
|
|
73
|
+
> **Lane-state reporting (mandatory).** Neither CLI supports a keyword query, and both see only
|
|
74
|
+
> *configured* marketplaces. Record each lane as `EXECUTED` / `EMPTY` / `FAILED: <stderr>` and carry
|
|
75
|
+
> that state into the recommendation. **A `FAILED` lane must never be rendered as "no candidates"** —
|
|
76
|
+
> re-run it or fall back to Priority 3 web search before reporting an empty result.
|
|
77
|
+
|
|
70
78
|
**npm ecosystem search (scoped packages):**
|
|
71
79
|
```bash
|
|
72
80
|
npm search @chrono-meta [keyword]
|
|
@@ -113,11 +121,20 @@ If either duplicate condition met → skip install → report "Already active"
|
|
|
113
121
|
#### 5-1 through 5-3. Install Steps
|
|
114
122
|
|
|
115
123
|
1. Confirm intent: "Would you like to install the `[plugin-name]` plugin?"
|
|
116
|
-
2. On agreement
|
|
124
|
+
2. On agreement — **the two commands take different arguments**: `marketplace add` takes a *source*
|
|
125
|
+
(`<owner/repo>`, a URL, or a local path), never a plugin name; only `install` takes the plugin name.
|
|
117
126
|
```bash
|
|
118
|
-
claude plugin marketplace add [
|
|
119
|
-
|
|
127
|
+
# Usage: claude plugin marketplace add [options] SOURCE
|
|
128
|
+
# SOURCE = owner/repo, a URL, or a local path — NEVER a plugin name
|
|
129
|
+
SOURCE="owner/repo"
|
|
130
|
+
PLUGIN="plugin-name" # optionally "plugin-name@marketplace" to disambiguate
|
|
131
|
+
|
|
132
|
+
claude plugin marketplace add "$SOURCE"
|
|
133
|
+
# Usage: claude plugin install|i [options] PLUGIN
|
|
134
|
+
claude plugin install "$PLUGIN"
|
|
120
135
|
```
|
|
136
|
+
Skip the `marketplace add` line when the plugin already resolves from a configured marketplace
|
|
137
|
+
(check `claude plugin marketplace list`).
|
|
121
138
|
3. Post-install initial configuration guidance:
|
|
122
139
|
- **API token input**: Guide token generation path for external service APIs (Jira/Confluence/Slack — specify each service's token page URL + env var or plugin config storage location)
|
|
123
140
|
- **MCP connection**: If plugin uses MCP server, guide auto-update of `.mcp.json` or `claude mcp add` command
|
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: prompt-regression
|
|
3
|
-
description:
|
|
3
|
+
description: Statically checks harness assets against a known-answer probe set after rule/skill changes — inspects the changed source for the trigger phrases, chain links and gate conditions each probe expects, and reports PASS/FAIL/SKIP per probe. Source inspection only; it does not run live sessions, so it catches assets that no longer SAY the right thing, not models that stop DOING it. Triggers on "prompt regression", "did my changes break anything", "regression check", "test harness changes".
|
|
4
4
|
user-invocable: true
|
|
5
|
-
allowed-tools: ["Read", "Bash", "Glob", "Grep"]
|
|
5
|
+
allowed-tools: ["Read", "Write", "Bash", "Glob", "Grep"]
|
|
6
6
|
model: sonnet
|
|
7
7
|
complexity_routing:
|
|
8
8
|
base: sonnet
|
|
@@ -58,8 +58,15 @@ ls .claude/regression/probes.md 2>/dev/null || echo "NO_CUSTOM_PROBES"
|
|
|
58
58
|
```
|
|
59
59
|
|
|
60
60
|
**If custom probes exist**: load and use them. The hub repo ships its golden probe set
|
|
61
|
-
(known-answer offline eval,
|
|
62
|
-
|
|
61
|
+
(known-answer offline eval, **33 probes** with check classes — 27 mandatory-pass · 1 measured ·
|
|
62
|
+
5 judged, of which 2 rows are inert deletion anchors, so live coverage is 31) at exactly this path
|
|
63
|
+
— when present it is canonical and supersedes the default matrix below.
|
|
64
|
+
|
|
65
|
+
> Count from the file, not from this line: `grep -cE '^\| *`[A-Z][A-Z0-9-]*-[0-9]+` *\|'
|
|
66
|
+
> .claude/regression/probes.md`. Both this number and the tally inside `probes.md` read `32` until
|
|
67
|
+
> 2026-08-12, when the actual count was 33 — a probe had been added to section A without touching
|
|
68
|
+
> either summary. A count duplicated in two files rots independently; if the two disagree, the file
|
|
69
|
+
> wins and both get corrected.
|
|
63
70
|
|
|
64
71
|
**If no custom probes** (e.g. Mode C install without the hub repo): use the default
|
|
65
72
|
probe matrix below.
|
|
@@ -93,7 +100,23 @@ If change scope is `CLAUDE.md` core (10+ lines changed): run **full suite** (all
|
|
|
93
100
|
|
|
94
101
|
---
|
|
95
102
|
|
|
96
|
-
### Step 4.
|
|
103
|
+
### Step 4. Evaluate Affected Probes (static source inspection)
|
|
104
|
+
|
|
105
|
+
**What this step is, stated plainly.** Every check below reads the changed *source* and asks whether
|
|
106
|
+
the text a probe expects is still there. No live session is started, no model is prompted, no output
|
|
107
|
+
is compared against a recorded transcript. The description said "running standard prompt probes …
|
|
108
|
+
comparing outputs against saved baselines", which reads as live execution; the honest label is a
|
|
109
|
+
**known-answer static check**.
|
|
110
|
+
|
|
111
|
+
**What it therefore cannot catch** — say this in the report, do not leave it implied:
|
|
112
|
+
- a rule that is still present but has stopped *firing* (salience loss, ordering, competition
|
|
113
|
+
from another rule)
|
|
114
|
+
- a trigger phrase present in the file but shadowed by a higher-priority route
|
|
115
|
+
- any behavior change that leaves the source text identical
|
|
116
|
+
|
|
117
|
+
Those need a dispatched blind sim at the target tier (`sim-conductor`, or the target-tier sim gate in
|
|
118
|
+
`.claude/rules/fh_4axis_gate.md`). A green report here means *the assets still say the right thing*.
|
|
119
|
+
Treating it as behavioral evidence is the misread this section exists to prevent.
|
|
97
120
|
|
|
98
121
|
For each affected probe, evaluate:
|
|
99
122
|
|
|
@@ -150,24 +173,44 @@ If all probes pass:
|
|
|
150
173
|
|
|
151
174
|
After a deliberate behavior change (not a regression — an intentional improvement), update the baseline:
|
|
152
175
|
|
|
176
|
+
Prompt user first: *"Probe `G-GATE-03` now expects the new gate format. Update baseline? (y/n)"*
|
|
177
|
+
Only update on explicit `y` — never auto-update.
|
|
178
|
+
|
|
179
|
+
On `y`, actually perform the write. The previous version of this step was three lines of which two
|
|
180
|
+
were comments — it created the directory and stopped, so "baseline updated" was reported by a step
|
|
181
|
+
that had written nothing.
|
|
182
|
+
|
|
153
183
|
```bash
|
|
154
|
-
# Baseline stored as markdown in .claude/regression/
|
|
155
184
|
mkdir -p .claude/regression
|
|
156
|
-
# Write updated probe expectations
|
|
157
185
|
```
|
|
158
186
|
|
|
159
|
-
|
|
187
|
+
Then **use the `Write` tool** on `.claude/regression/probes.md` (create it from the SKILL.md default
|
|
188
|
+
matrix if absent) to apply the approved edit. Each approved change edits the probe's row in place —
|
|
189
|
+
Expected Behavior, Scope, and Class — and:
|
|
160
190
|
|
|
161
|
-
|
|
191
|
+
- update the `**Count**:` tally line **in the same edit** if a row was added or removed, including
|
|
192
|
+
the per-section breakdown and the class counts (a stale tally is exactly how the `32`/`33`
|
|
193
|
+
mismatch was introduced)
|
|
194
|
+
- append a one-line note under `**Baseline**:` recording the date and the reason for the change
|
|
195
|
+
- never rewrite rows the user did not approve
|
|
196
|
+
|
|
197
|
+
Confirm afterwards by re-reading the file and reporting the row's new content plus the recount —
|
|
198
|
+
`grep -cE '^\| *`[A-Z][A-Z0-9-]*-[0-9]+` *\|' .claude/regression/probes.md`. Do not report a
|
|
199
|
+
baseline update whose result you have not read back.
|
|
162
200
|
|
|
163
201
|
---
|
|
164
202
|
|
|
165
203
|
## Done When
|
|
166
204
|
|
|
167
|
-
- All affected probes are evaluated (PASS / FAIL / SKIP)
|
|
205
|
+
- All affected probes are evaluated (PASS / FAIL / SKIP) and the three counts sum to the number of
|
|
206
|
+
probes selected in Step 3 — class: measured (`PASS + FAIL + SKIP == selected`)
|
|
168
207
|
- Regression report is output with clear PASS/FAIL verdict — class: mandatory-pass
|
|
208
|
+
- The report states its own scope: **static source inspection, no live session run** — a green
|
|
209
|
+
verdict is never presented as behavioral evidence — class: mandatory-pass
|
|
169
210
|
- If FAIL: specific file + line fix is recommended — class: judged, paired with verify-bidirectional (the fix recommendation is re-checked, not trusted as-is)
|
|
170
|
-
- Baseline updated only on explicit user approval
|
|
211
|
+
- Baseline updated only on explicit user approval, and the written file is **read back** and its
|
|
212
|
+
probe count reported — class: mandatory-pass (HITL). Approval alone does not close this: a step
|
|
213
|
+
that prompts, gets `y`, and writes nothing satisfied the old wording.
|
|
171
214
|
|
|
172
215
|
---
|
|
173
216
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: salience-splitter
|
|
3
3
|
description: 'Splits an over-loaded always-loaded context asset — a SKILL.md, CLAUDE.md, or memory index — into a lean always-loaded layer + an on-demand layer, using a governance-semantic criterion (not length, but when the content is needed), connected by imperative pointers. Based on paper §9.5 Protocol-Priority Split pattern. Diagnoses, classifies, splits, and verifies in one pass. Renamed from skill-splitter (old name still routes here). Triggers: "SKILL.md too large", "split this skill", "skill is bloated", "skill file too long", "CLAUDE.md 너무 커".'
|
|
4
4
|
user-invocable: true
|
|
5
|
-
allowed-tools: ["Read", "Write", "Edit", "Bash", "Grep", "Glob"]
|
|
5
|
+
allowed-tools: ["Read", "Write", "Edit", "Bash", "Grep", "Glob", "Agent"]
|
|
6
6
|
model: sonnet
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -46,6 +46,43 @@ model: sonnet
|
|
|
46
46
|
|
|
47
47
|
**One test**: *"If a consumer agent had only SKILL.md, could they recognize the trigger, understand the full step sequence, and make the key decisions?"* → Yes = correct split. No = something behavioral is missing from SKILL.md.
|
|
48
48
|
|
|
49
|
+
### Two floors this skill must not cross (behavioral — resident on purpose)
|
|
50
|
+
|
|
51
|
+
**① A cut is measured, never eyeballed** — and *which* measurement depends on what layer you are
|
|
52
|
+
cutting from. Both branches forbid the same thing ("I read it and it looks redundant" is not a
|
|
53
|
+
measurement); they differ in cost because the surfaces differ:
|
|
54
|
+
|
|
55
|
+
| Cutting from | Measurement | Why this one |
|
|
56
|
+
|---|---|---|
|
|
57
|
+
| **Resident layer** — CLAUDE.md, a memory index (loaded every session, for every task) | **Ablation harness**: pre-registered question set · **isolated arm B** · `reps>=3` · runner precondition `bash scripts/ablation_calibrate.sh` exits 0 (canon = `scripts/probe_scope_check.sh` header) · verdict line → `.claude/regression/ablation_verdicts.md` | A wrong cut here degrades *every* session silently, and the cost is unattributable after the fact |
|
|
58
|
+
| **SKILL.md → SKILL_detail.md** (loaded only when the skill is invoked) | **Cold-start sim**: `sim-conductor` Area D-skill on SKILL.md alone must reach grade F (see Done When) | The failure is scoped to one skill's invocation and the sim reproduces exactly that condition |
|
|
59
|
+
|
|
60
|
+
**Arm B answering confidently wrong is a KEEP, not a pass** — fluency is not recall, and that is the
|
|
61
|
+
whole reason the arm is isolated. The same reading applies to the sim: a consumer that improvises
|
|
62
|
+
plausibly without the moved rule is grade P, not F.
|
|
63
|
+
|
|
64
|
+
This skill is the one that executes resident removal, so the procedure is named *here* rather than
|
|
65
|
+
assumed — an earlier version left the judgment "mentally", which is exactly the eyeball path both
|
|
66
|
+
branches exist to replace. **Do not read the resident row onto the SKILL.md row**: requiring a full
|
|
67
|
+
ablation for every ambiguous SKILL.md section would price routine splitting out of existence, which
|
|
68
|
+
is how a floor turns into a bypass trainer (cross-family review 2026-08-11 flagged exactly that
|
|
69
|
+
over-block in the first draft of this section).
|
|
70
|
+
|
|
71
|
+
**② Every split must re-ask whether the destination is inside the gate.** Moving content to a new
|
|
72
|
+
path can move it **out of the 4-axis gate's pathspec** — the gate-locality class has recurred four
|
|
73
|
+
times, most sharply when `SKILL_detail.md` fell outside a literal `SKILL\.md` term and 27.7% of the
|
|
74
|
+
skill-spec surface went ungated (it leaked twice for real). `.claude/rules/fh_4axis_gate.md` names
|
|
75
|
+
*this skill* as the producer that **widens that hole every time the diet succeeds**. So coverage is
|
|
76
|
+
re-verified mechanically at split time, not assumed:
|
|
77
|
+
|
|
78
|
+
```bash
|
|
79
|
+
bash scripts/gate_pathspec_check.sh # exit 0 = known-pair coverage holds, 1 = a pair broke
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
A destination path the gate does not match is not a "later" item: **ship the pathspec update in the
|
|
83
|
+
same commit as the split** — the split and its coverage are one change, and separating them is how
|
|
84
|
+
the surface silently shrinks.
|
|
85
|
+
|
|
49
86
|
---
|
|
50
87
|
|
|
51
88
|
## Step Overview
|
|
@@ -64,13 +101,57 @@ Step 3 — Draft SKILL_detail.md
|
|
|
64
101
|
One ## §SectionName header per pointer in SKILL.md
|
|
65
102
|
Move removed content under its section
|
|
66
103
|
Front-matter: name, description, load: on-demand
|
|
104
|
+
→ Gate check (Floor ②): destination path inside the 4-axis pathspec?
|
|
105
|
+
bash scripts/gate_pathspec_check.sh — not-matched = update the pathspec in THIS commit
|
|
67
106
|
|
|
68
107
|
Step 4 — Verify
|
|
69
108
|
phantom-quench: every §pointer in SKILL.md resolves to ## §SectionName in SKILL_detail.md
|
|
70
109
|
sim-conductor Area D-skill: consumer agent with SKILL.md only → must reach grade F
|
|
110
|
+
(grade scale — sim-conductor SKILL.md §Area-D: F = Functional/PASS · P = Partial · B = Broken)
|
|
71
111
|
→ Any pointer mismatch or grade P/B → fix before commit
|
|
72
112
|
```
|
|
73
113
|
|
|
114
|
+
**Step 4 pointer↔header comparison — run it, don't eyeball it.** Both sides need normalizing or the
|
|
115
|
+
check produces ~100% false positives: real headers are written `## §Name — description`, so a naive
|
|
116
|
+
compare never matches. And extraction must be anchored to the **declared pointer form**
|
|
117
|
+
(see "Imperative Pointer Format" below — a blockquote line whose bolded lead is followed by
|
|
118
|
+
`See`), not to "a mention of the
|
|
119
|
+
detail file": any prose that *discusses* a pointer, quoted or not, is otherwise collected as one.
|
|
120
|
+
That is not hypothetical — a paragraph on this page names an example section, and a
|
|
121
|
+
backtick-only filter counted it as a live pointer:
|
|
122
|
+
|
|
123
|
+
```bash
|
|
124
|
+
S="plugins/{plugin}/skills/{name}" # the skill dir being split
|
|
125
|
+
# pointer side: any bolded-lead "See" line (the marker word varies — **Detail**, **Template**, …),
|
|
126
|
+
# and ALL §names on it (a documented variant puts two pointers on one line)
|
|
127
|
+
diff <(grep -E '^> \*\*[A-Za-z]+\*\*: See ' "$S/SKILL.md" \
|
|
128
|
+
| grep -oE '§[A-Za-z0-9_-]+' | sed 's/^§//' | sort -u) \
|
|
129
|
+
<(grep '^## §' "$S/SKILL_detail.md" | sed 's/^## §//; s/ *—.*//' | sort -u)
|
|
130
|
+
# rc=0 → every pointer resolves and every §header has a pointer. rc=1 → the diff names both directions
|
|
131
|
+
# ("<" = pointer with no section · ">" = section with no pointer, i.e. an orphan).
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
**Calibrate before trusting it** (this check was wrong twice while being written — each narrowing
|
|
135
|
+
looked reasonable and produced false orphans): run it across sibling skills and confirm it
|
|
136
|
+
*discriminates*. Measured 2026-08-11 over 6 FH skills: 5 exit 0, `steel-quench` exits 1 on a real
|
|
137
|
+
orphan (`§Phase0` — no pointer anywhere in its SKILL.md), and this page's own prose example is not
|
|
138
|
+
collected. A version of this check that flags everything, or nothing, is not measuring.
|
|
139
|
+
|
|
140
|
+
**Orphan check covers `^## `, not `^## §` — and the two checks have different jobs.** The diff above
|
|
141
|
+
compares §-prefixed headers against pointers; a header written *without* the `§` prefix is invisible
|
|
142
|
+
to it. Measured across this repo's `SKILL_detail.md` files: **139 `## ` headers vs 97 `## §`**, so 42
|
|
143
|
+
(30%) sit outside the comparison entirely and would pass forever. So run both, and read them as one
|
|
144
|
+
rule:
|
|
145
|
+
|
|
146
|
+
```
|
|
147
|
+
diff (§ side) → pointer↔header agreement, both directions
|
|
148
|
+
grep '^## ' → coverage net: any header the diff could not see
|
|
149
|
+
resolution → give it the § prefix AND a pointer (then the diff covers it), or merge it away
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
A non-§ header listed by the grep and left alone is **not** a pass — it is the orphan the §-only
|
|
153
|
+
scan used to hide.
|
|
154
|
+
|
|
74
155
|
> **Detail**: See `SKILL_detail.md §Verification-Checklist` — pre-commit checklist table (8 checks) — read when running Step 4 verification.
|
|
75
156
|
|
|
76
157
|
> **Detail**: See `SKILL_detail.md §Split-Execution` — step-by-step trimming procedure, SKILL_detail.md front-matter format, orphan-section check — read when executing Steps 2–3.
|
|
@@ -120,17 +201,49 @@ Run on a SKILL.md when **any one** of:
|
|
|
120
201
|
|
|
121
202
|
## Done When
|
|
122
203
|
|
|
204
|
+
**Common to all three scopes** (§Scope declares SKILL.md · CLAUDE.md · memory index — the conditions
|
|
205
|
+
below are the ones that hold whatever was split; scope-specific conditions follow):
|
|
206
|
+
|
|
123
207
|
```
|
|
124
208
|
Step 1 classification table produced
|
|
209
|
+
(mandatory-pass — the table exists and every section carries a verdict)
|
|
210
|
+
+ Every AMBIGUOUS section carries the measurement its layer requires (Floor ①)
|
|
211
|
+
(measured — resident layer: ablation verdict line in .claude/regression/ablation_verdicts.md
|
|
212
|
+
(pre-registered set, isolated arm B, reps>=3) · SKILL.md layer: cold-start sim grade F.
|
|
213
|
+
Eyeball judgment = NOT met on either branch)
|
|
214
|
+
+ Gate coverage re-verified for every destination path introduced by the split
|
|
215
|
+
(mandatory-pass — `bash scripts/gate_pathspec_check.sh` exits 0, or the pathspec update
|
|
216
|
+
ships in the same commit)
|
|
217
|
+
+ No behavioral rule lives only in the on-demand layer
|
|
218
|
+
(judged — pair: an isolated consumer read that has ONLY the always-loaded layer and must reach
|
|
219
|
+
a decision the moved rule governs; author re-reading their own split does not satisfy this)
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
**Scope-specific — SKILL.md split:**
|
|
223
|
+
```
|
|
125
224
|
+ SKILL.md trimmed: triggers · principles · step overview · decision tables · Done When retained
|
|
126
|
-
+
|
|
127
|
-
|
|
225
|
+
+ Imperative pointer for every removed section; SKILL_detail.md carries one ## §header per pointer
|
|
226
|
+
(mandatory-pass — Step 4 normalized diff exits 0, both directions)
|
|
128
227
|
+ phantom-quench: 0 phantoms (all §pointers resolve)
|
|
129
228
|
→ Fallback (skill unavailable): run §Verification-Checklist manually from SKILL_detail.md
|
|
130
229
|
+ sim-conductor Area D-skill: grade F (consumer completes core task from SKILL.md alone)
|
|
131
|
-
|
|
230
|
+
(measured — grade scale, sim-conductor SKILL.md §Area-D: F = Functional/PASS · P = Partial · B = Broken)
|
|
231
|
+
→ Fallback (skill unavailable): manually confirm "trigger → step overview → key decision → Done When" present
|
|
132
232
|
```
|
|
133
233
|
|
|
134
|
-
**
|
|
135
|
-
**
|
|
136
|
-
|
|
234
|
+
**Scope-specific — CLAUDE.md split:** the resident file keeps the rule *and* an imperative
|
|
235
|
+
`> **Detail**: See …` pointer; the detail destination is inside the gate pathspec (Floor ②); a cold
|
|
236
|
+
top-level session can still act on the rule without opening the detail file
|
|
237
|
+
(judged — pair: a blind sim at the tier the rule must survive on, **not** the author's re-read).
|
|
238
|
+
⚠️ Resident-footprint claims are measured only by a **fresh top-level `/context`** — file char counts
|
|
239
|
+
are not resident measurements, and an agent-view window cannot measure it at all.
|
|
240
|
+
|
|
241
|
+
**Scope-specific — memory index split:** every demoted entry is reachable from the archive by the
|
|
242
|
+
recall path the index declares (mandatory-pass — grep the archive for the demoted entry's own nouns
|
|
243
|
+
and hit it); the hot index keeps one line per surviving entry.
|
|
244
|
+
|
|
245
|
+
**Not done**: an on-demand section with no pointer from the always-loaded layer (orphan) — scan
|
|
246
|
+
`^## `, not `^## §` (30% of real headers carry no §).
|
|
247
|
+
**Not done**: consumer grade P (Partial) or B (Broken) after split — a behavioral rule was moved out
|
|
248
|
+
when it should have stayed.
|
|
249
|
+
**Not done**: a section CUT on "it looked redundant" with no ablation verdict line.
|
|
@@ -55,13 +55,28 @@ For each section in target SKILL.md:
|
|
|
55
55
|
|
|
56
56
|
### Ambiguous content test
|
|
57
57
|
|
|
58
|
-
|
|
58
|
+
The D-skill cold-start question frames the decision:
|
|
59
59
|
|
|
60
60
|
> *"If a consumer agent had only SKILL.md and typed the trigger phrase, would they need this content in the first 2 steps?"*
|
|
61
61
|
> - YES → Always-loaded
|
|
62
|
-
> - NO →
|
|
62
|
+
> - NO → **candidate** for on-demand — not yet a verdict
|
|
63
63
|
|
|
64
|
-
|
|
64
|
+
**Do not answer it in your head.** AMBIGUOUS is precisely the case where the eyeball answer is
|
|
65
|
+
unreliable, so the verdict is measured, per SKILL.md §Floor ① and CLAUDE.md's ablation procedure:
|
|
66
|
+
|
|
67
|
+
```bash
|
|
68
|
+
bash scripts/ablation_calibrate.sh # runner precondition — must exit 0 before any arm runs
|
|
69
|
+
# then run the pre-registered arms per scripts/probe_scope_check.sh (header = canon:
|
|
70
|
+
# arms · isolation · reps>=3 · pre-registration · the two leak channels)
|
|
71
|
+
# verdict line → .claude/regression/ablation_verdicts.md
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Read the result correctly: a section is CUT only when the **isolated arm B answers the pre-registered
|
|
75
|
+
questions correctly**. Arm B answering **confidently wrong is a KEEP** — fluency is not recall, and
|
|
76
|
+
that is the failure the isolation exists to expose.
|
|
77
|
+
|
|
78
|
+
Behavioral rules skip the probe and stay Always-loaded (even if rarely triggered, the consumer needs
|
|
79
|
+
to know the constraint exists) — that is a KEEP by classification, not an unmeasured cut.
|
|
65
80
|
|
|
66
81
|
---
|
|
67
82
|
|
|
@@ -148,19 +163,35 @@ Opening line:
|
|
|
148
163
|
|
|
149
164
|
For each pointer in SKILL.md: create matching `## §SectionName` header, paste removed content under it.
|
|
150
165
|
|
|
151
|
-
Orphan check: every
|
|
166
|
+
Orphan check: every section header in SKILL_detail.md must have a corresponding pointer in SKILL.md.
|
|
167
|
+
Scan `^## ` — **not** `^## §`: measured across this repo, 139 `## ` headers vs 97 `## §`, so a §-only
|
|
168
|
+
scan is structurally blind to 42 (30%) of real sections and passes them forever. If a header has no
|
|
169
|
+
pointer → add the pointer or merge with an adjacent section.
|
|
152
170
|
|
|
153
171
|
### Step 4: Verification sequence
|
|
154
172
|
|
|
155
173
|
```bash
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
#
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
#
|
|
163
|
-
#
|
|
174
|
+
S=plugins/{plugin}/skills/{name}
|
|
175
|
+
|
|
176
|
+
# 1+2+3 in one comparison. Three normalizations, each one learned from a false positive:
|
|
177
|
+
# · headers are written "## §Name — description" → strip the em-dash tail
|
|
178
|
+
# · the marker word varies (**Detail**, **Template**) → match any bolded lead + "See"
|
|
179
|
+
# · one line may carry TWO pointers (documented variant, §Pointers)
|
|
180
|
+
# → collect ALL §names on the line
|
|
181
|
+
# Anchoring on the pointer LINE (not "any mention of the detail file") is what keeps prose that
|
|
182
|
+
# merely DISCUSSES a pointer out — measured: a backtick-only filter counted an explanatory
|
|
183
|
+
# example as a live pointer, and a **Detail**-only filter called two real pointers orphans.
|
|
184
|
+
diff <(grep -E '^> \*\*[A-Za-z]+\*\*: See ' "$S/SKILL.md" \
|
|
185
|
+
| grep -oE '§[A-Za-z0-9_-]+' | sed 's/^§//' | sort -u) \
|
|
186
|
+
<(grep '^## §' "$S/SKILL_detail.md" | sed 's/^## §//; s/ *—.*//' | sort -u)
|
|
187
|
+
# rc=0 → every pointer resolves AND every §header has a pointer (both directions in one run)
|
|
188
|
+
# rc=1 → the diff output names which side is missing what. Fix before commit.
|
|
189
|
+
|
|
190
|
+
# 4. Orphan scan over the WIDER header form (see above — §-only misses 30%)
|
|
191
|
+
grep -n '^## ' "$S/SKILL_detail.md"
|
|
192
|
+
|
|
193
|
+
# 5. Gate coverage for the destination path (SKILL.md §Floor ②)
|
|
194
|
+
bash scripts/gate_pathspec_check.sh
|
|
164
195
|
```
|
|
165
196
|
|
|
166
197
|
Then run:
|
|
@@ -180,6 +211,8 @@ Use before committing a completed split:
|
|
|
180
211
|
| All §pointers resolve | phantom-quench: 0 phantoms |
|
|
181
212
|
| Cold-start grade F | sim-conductor Area D-skill: consumer reaches core completion |
|
|
182
213
|
| No behavioral rule in SKILL_detail only | Any rule governing "what counts as X" present in SKILL.md |
|
|
183
|
-
| No orphan
|
|
214
|
+
| No orphan sections | Every `## ` header in SKILL_detail.md has a pointer from SKILL.md (scan `^## `, not `^## §` — §-only is blind to 30% of real headers) |
|
|
215
|
+
| Gate coverage holds | `bash scripts/gate_pathspec_check.sh` exits 0, or the pathspec update ships in the same commit |
|
|
216
|
+
| Ablation verdict for every AMBIGUOUS cut | Verdict line in `.claude/regression/ablation_verdicts.md`; arm B correct = CUT, arm B confidently wrong = KEEP |
|
|
184
217
|
| SKILL.md ≤ 50% of original | Line count check |
|
|
185
218
|
| SKILL.md flows without gaps | Reading SKILL.md alone gives complete step sequence understanding |
|
|
@@ -335,14 +335,39 @@ cross_model_coverage: [external | cross-session | NONE] # NONE on risk≥mediu
|
|
|
335
335
|
```bash
|
|
336
336
|
cd "$HARNESS_ROOT"
|
|
337
337
|
BRANCH="fix/sim-$(date +%Y%m%d)-m-tier"
|
|
338
|
-
git
|
|
339
|
-
# [process M-tier items]
|
|
340
|
-
|
|
338
|
+
git switch -c "$BRANCH"
|
|
339
|
+
# [process M-tier items] — collect the files this run actually edited
|
|
340
|
+
# Explicit list; never a glob, never -A, never `.`. Use an ARRAY, not a space-separated string:
|
|
341
|
+
# `git add -- $FILES` relies on bash word-splitting an unquoted parameter expansion, and zsh does
|
|
342
|
+
# NOT split one (SH_WORD_SPLIT off by default). Under zsh the whole string arrives as ONE pathspec,
|
|
343
|
+
# git errors "did not match any files", and the commit that follows stages nothing — the explicit-
|
|
344
|
+
# path discipline this line exists to enforce silently stops applying in the operator's own shell.
|
|
345
|
+
FILES=(path/to/edited1.md path/to/edited2.md)
|
|
346
|
+
git add -- "${FILES[@]}"
|
|
347
|
+
git diff --cached --name-only # confirm the staged set IS the intended set before committing
|
|
341
348
|
git commit -m "fix(sim-conductor): resolve M-tier findings from simulation YYYY-MM-DD"
|
|
342
349
|
git push -u origin "$BRANCH"
|
|
343
350
|
# PR creation requires explicit user request per CLAUDE.md PR principle
|
|
344
351
|
```
|
|
345
352
|
|
|
353
|
+
> **Never `git add -p` here — this skill runs to completion unattended.** `-p` is interactive: with
|
|
354
|
+
> no TTY it either blocks forever, or (with stdin closed) exits **0 having staged nothing**. Measured
|
|
355
|
+
> 2026-08-12 on the documented chain: `add -p` → rc 0, nothing staged → `git commit` → **rc 1**,
|
|
356
|
+
> while `git switch -c` had **already created the branch** — so the run leaves a branch behind, the
|
|
357
|
+
> fixes uncommitted, and a success-shaped exit on the staging step. A silent no-op in the middle of
|
|
358
|
+
> an auto-commit chain is worse than a hang, because only the hang is visible.
|
|
359
|
+
>
|
|
360
|
+
> **Why an explicit list rather than `-A` or `.`**: `sim-conductor` frequently runs in a checkout
|
|
361
|
+
> shared with other work. `git add -A` sweeps up whatever else is dirty — measured above: a
|
|
362
|
+
> co-resident `b.txt` was left untouched by `git add -- a.txt` and would have been swallowed by
|
|
363
|
+
> `-A`. Note the boundary: `git add -- <file>` stages **the whole file**, including lines this run
|
|
364
|
+
> did not write, so if another session has appended to a file you also edited, commit that append
|
|
365
|
+
> first rather than absorbing it.
|
|
366
|
+
>
|
|
367
|
+
> `git diff --cached --name-only` before the commit is the mechanical check that the staged set
|
|
368
|
+
> equals the intended set. Do not skip it — it is what turns "I passed the right paths" from a
|
|
369
|
+
> belief into an observation.
|
|
370
|
+
|
|
346
371
|
---
|
|
347
372
|
|
|
348
373
|
## §PathB-Detail — External Environment Fallback
|