@chrono-meta/fh-gate 3.2.0 → 3.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/registry/agent_cards.json +1 -1
- package/.claude/rules/fh_4axis_gate.md +25 -0
- package/.claude-plugin/marketplace.json +3 -3
- package/AGENTS.md +2 -2
- package/CATALOG.md +4 -4
- package/CHEATSHEET.md +1 -1
- package/CLAUDE.md +4 -4
- package/docs/OUTPUT_EVIDENCE.md +1 -1
- package/docs/STANDARDS_ALIGNMENT.md +1 -1
- package/docs/codex-compat.md +1 -1
- package/knowledge/shared/harness-core/agents_md_runtime_details.md +2 -2
- package/knowledge/shared/harness-core/field_verdict_crossfamily_gate.md +30 -3
- package/knowledge/shared/harness-core/iso_ai_standards_crosswalk.md +1 -1
- package/knowledge/shared/harness-core/skill_quality_rubric.md +1 -1
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +31 -0
- package/knowledge/shared/rules/modes_and_value.md +2 -2
- package/package.json +2 -1
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-commons/README.md +38 -0
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/CHANGELOG.md +95 -0
- package/plugins/fh-meta/skills/agent-composer/SKILL.md +2 -2
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +40 -0
- package/plugins/fh-meta/skills/frontier-digest/SKILL.md +1 -1
- package/plugins/fh-meta/skills/frontier-digest/SKILL_detail.md +43 -5
- package/plugins/fh-meta/skills/{hub-cc-pr-reviewer → harness-pr-reviewer}/SKILL.md +75 -3
- package/plugins/fh-meta/skills/{hub-cc-pr-reviewer → harness-pr-reviewer}/SKILL_detail.md +2 -2
- package/plugins/fh-meta/skills/harvest-loop/SKILL_detail.md +2 -2
- package/plugins/fh-meta/skills/install-doctor/SKILL.md +1 -1
- package/plugins/fh-meta/skills/install-wizard/SKILL.md +1 -1
- package/plugins/fh-meta/skills/meta-prompt-builder/SKILL.md +1 -1
- package/plugins/fh-meta/skills/pipeline-conductor/SKILL.md +1 -1
- package/plugins/fh-meta/skills/plugin-recommender/SKILL.md +1 -1
- package/plugins/fh-meta/skills/sim-conductor/SKILL.md +1 -1
- package/plugins/fh-qp/.claude-plugin/plugin.json +1 -1
- package/scripts/finding_fleet.sh +391 -13
- package/scripts/finding_pipeline.sh +370 -11
- package/scripts/finding_verifier.sh +31 -2
- package/scripts/finding_verify.py +198 -13
- package/scripts/frontier_digest_autopilot.sh +3 -3
- package/scripts/gate_shape_scan.sh +16 -2
- package/scripts/test_finding_pipeline_lanes.sh +1556 -4
- package/scripts/test_gate_shape_scan_lanes.sh +11 -0
- package/scripts/test_marker_crossfamily_lanes.sh +75 -3
- package/scripts/test_marker_standpoint_lanes.sh +31 -6
- package/templates/.git-hooks/pre-commit +152 -6
- package/templates/local_fh_context.md +1 -1
- package/templates/regression_guard.sh +1 -1
|
@@ -172,7 +172,7 @@ Default composition table by task type.
|
|
|
172
172
|
| Plugin recommendation | plugin-recommender (S) | — |
|
|
173
173
|
| Install conflict diagnosis | install-doctor (S) | — |
|
|
174
174
|
| Onboarding install | install-wizard (S) — ⚠️ interactive; `--dry-run` for bg parallel | — |
|
|
175
|
-
| Hub PR review |
|
|
175
|
+
| Hub PR review | harness-pr-reviewer (S) — requires PR number first | — |
|
|
176
176
|
| **Decision-maker approval review** | apex-review (S) — CTO/tech lead/QA lead personas + HTML deck | — |
|
|
177
177
|
| **Project local skills** | LOCAL_SKILL_REGISTRY lookup → relevant project skill (A/S) | Per project |
|
|
178
178
|
|
|
@@ -360,7 +360,7 @@ After fan-in report, evaluate conditions and auto-suggest the next Wave:
|
|
|
360
360
|
|
|
361
361
|
| Condition | Wave suggestion |
|
|
362
362
|
|---|---|
|
|
363
|
-
| ① M-tier > 0 | **Wave next-M**: fact-checker (A) →
|
|
363
|
+
| ① M-tier > 0 | **Wave next-M**: fact-checker (A) → harness-pr-reviewer (S) |
|
|
364
364
|
| ② persona-innovator naming candidates > 0 | **Wave next-I**: delegate to user + asset-placement-gate (S) |
|
|
365
365
|
| ③ External absorption signal High > 0 | **Wave next-E**: persona-innovator Mode E (A) + meta-prompt-builder (S) |
|
|
366
366
|
| ⑤ Design conflict / 2+ conflicting suggestions | **Wave next-D**: deliberation (S) — verdict folds back into Step 4-b |
|
|
@@ -229,6 +229,46 @@ format-checked `residency=(CLEAN|TAINTED|NOT_SCANNED)(...)` token on `DEGRADED_*
|
|
|
229
229
|
`declined`. Fixtures: `scripts/test_marker_crossfamily_lanes.sh` (`r1`–`r13`). This is the SAME
|
|
230
230
|
typed field Step 6 already emits into — one channel, not a second marker line.
|
|
231
231
|
|
|
232
|
+
## Step 4.6 — `evidence=` : say WHAT the panel received (mandatory with `panel(...)`, 2026-09-12)
|
|
233
|
+
|
|
234
|
+
`residency=` says the payload was safe to send. `evidence=` says **what was sent** — and that is the
|
|
235
|
+
axis the outside world just measured as the dominant one. `arXiv:2609.10969` (48 task templates →
|
|
236
|
+
2,880 scenarios, **fixed call budget** so the comparison is about the axis and not about spending
|
|
237
|
+
more): a cross-model vote over **shared** evidence approved **62.9 %** of unsafe proposals against
|
|
238
|
+
**22.9 %** with an **independent source** — source effect **40.9 pp** vs **11.3 pp** for model
|
|
239
|
+
diversity, i.e. **3.6×**.
|
|
240
|
+
|
|
241
|
+
Decide it from **how this run actually dispatched**, not from intent:
|
|
242
|
+
|
|
243
|
+
| What happened | Token |
|
|
244
|
+
|---|---|
|
|
245
|
+
| Every member got the **same** payload (one prompt file / one diff fanned out) | `evidence=SHARED(<name the artifact>)` |
|
|
246
|
+
| Each member produced its **own** evidence — own checkout, own run, own extraction | `evidence=INDEPENDENT(<what each got>)` |
|
|
247
|
+
| Some of each | `evidence=MIXED(<which member got which>)` |
|
|
248
|
+
|
|
249
|
+
🟥 **This skill's default flow is `SHARED`, by construction.** Step 4 builds **one** payload and fans
|
|
250
|
+
it out, so a panel dispatched the ordinary way through this skill is the **62.9 % arm** — the axis it
|
|
251
|
+
raises is model family, which the paper prices at 11.3 pp. That is not a reason to stop using it: a
|
|
252
|
+
finder is worth its recall, and this skill's value was never precision
|
|
253
|
+
(`[[feedback_decorrelation_is_a_cheap_finder]]`). It **is** a reason to stop recording that run as if
|
|
254
|
+
it carried the stronger axis. Write `SHARED` and mean it.
|
|
255
|
+
|
|
256
|
+
🟢 **`SHARED` is a legal, expected answer.** The hook does not block it; it blocks a `panel(...)` that
|
|
257
|
+
does not say which it was, a value outside the closed three, a malformed or duplicated token, and a
|
|
258
|
+
**vacuous body** on `SHARED`/`MIXED` (name the artifact — *"same staged diff"*, *"one findings.jsonl"*,
|
|
259
|
+
*"identical prompt file"*). `INDEPENDENT` is self-describing and is not body-checked, deliberately:
|
|
260
|
+
over-blocking the honest strong answer would train the override.
|
|
261
|
+
|
|
262
|
+
**Where it lands**: the same grounds line as `residency=`, in the same `crossfamily:` field —
|
|
263
|
+
`crossfamily: panel(codex) — residency=CLEAN(files=3) · evidence=SHARED(same staged diff to both) · R1, 3 findings`.
|
|
264
|
+
Hook: `validate_crossfamily_leg` (grace `EVIDENCE_TOKEN_GRACE_DATE`, no retroactivity).
|
|
265
|
+
Fixtures: `scripts/test_marker_crossfamily_lanes.sh` `e1`–`e10`.
|
|
266
|
+
|
|
267
|
+
⚠️ **Named gap, not built today**: this skill has **no independent-source mode**. Giving it one
|
|
268
|
+
(each member reconstructs the evidence itself rather than reading the author's payload) is the change
|
|
269
|
+
that would move it onto the 40.9 pp axis; it is a design question, not a wiring one, and it is
|
|
270
|
+
recorded as a signal rather than improvised here.
|
|
271
|
+
|
|
232
272
|
🟥 **A stripped file is invisible to the reviewer, and a reviewer's default read of invisible is
|
|
233
273
|
"absent," not "redacted"** (governor dogfood, 2026-09-05, same-day live use of this Step): two
|
|
234
274
|
TAINTED files were stripped from a real payload and the CLEAN remainder sent cross-family — the
|
|
@@ -215,7 +215,7 @@ Present Step 4 menu options [1]–[5]. Do not skip to [5] silently — surface t
|
|
|
215
215
|
## Simplification Guards
|
|
216
216
|
|
|
217
217
|
- Video Tier-3 probe fails (any of `yt-dlp` / `curl_cffi` / `ffmpeg` missing, or timedtext returns 429) → fall through to operator summary; never assume `yt-dlp` works
|
|
218
|
-
-
|
|
218
|
+
- 🟢 **arxiv transport, inverted 2026-09-12 (operator decision, condition met: 429 on 09-09·10·11·12)**: the **date-sorted `arxiv.org/list/cs.SE/recent` listing is PRIMARY**; the export API is the **fallback**; HN-only only if both fail. Same sort axis (submission date) is why the listing is promotable and WebSearch is not. 🟥 Cost of the promotion, not an oversight: the listing's ID↔title pairing is the weaker instrument, so **per-item `abs/{id}` re-resolution is now a PRIMARY-path requirement** and the progress line must name which transport produced the items. Never substitute WebSearch (recall channel, 5/5 REPEATs measured 2026-09-04); report `arxiv FAILED (429)`, never `0 items`. 🟥 **On the listing path, pair ID↔title INSIDE one list entry, never by zipping two ordered lists — measured 2026-09-12: 51 IDs vs 50 titles, so pairs drift partway down and mispair while looking well-formed (3 items caught, 0 shipped). Until that is what happened, per-item `abs/{id}` re-resolution is MANDATORY and a title mismatch DROPS the item.** Detail: `SKILL_detail.md §Execution form`
|
|
219
219
|
- On curl timeout, skip that item and continue with the rest
|
|
220
220
|
- If synthesis result exceeds 400 characters, retain top 3 items and truncate the rest
|
|
221
221
|
- Without `--save`, do not create files (conversation output only)
|
|
@@ -28,14 +28,52 @@ load: on-demand
|
|
|
28
28
|
> fell back to **WebSearch**, and the 09-04 run measured what that does: 5/5 papers it returned were
|
|
29
29
|
> REPEATs. Search sorts by *canonicity*, the export API by *submission date* — a search fallback is
|
|
30
30
|
> a **recall channel, not a discovery channel**, and swapping it in silently converts «unrun» into
|
|
31
|
-
> «nothing new» (`not found ≠ 0`).
|
|
32
|
-
>
|
|
33
|
-
>
|
|
31
|
+
> «nothing new» (`not found ≠ 0`).
|
|
32
|
+
>
|
|
33
|
+
> 🟢 **TRANSPORT ORDER INVERTED 2026-09-12 (operator decision) — the listing is now PRIMARY.**
|
|
34
|
+
> Condition stated by the operator was *«if the errors keep coming»*, and they did, measurably:
|
|
35
|
+
> `429` appears in the digest record on **09-09 · 09-10 · 09-11 · 09-12 — four consecutive days**,
|
|
36
|
+
> past the `N≥3` bar twice over. Paying two failed calls plus a backoff every single day to reach
|
|
37
|
+
> items that the listing reaches anyway is not a fallback, it is a toll. New order:
|
|
38
|
+
> 1. **PRIMARY — WebFetch the date-sorted HTML listing** `https://arxiv.org/list/cs.SE/recent`
|
|
34
39
|
> (and `https://arxiv.org/list/cs.AI/recent` if the first is empty of agent/harness items), keep
|
|
35
|
-
> items whose title/abstract match the three query phrases, cap 6
|
|
36
|
-
>
|
|
40
|
+
> items whose title/abstract match the three query phrases, cap 6. **Same sort axis as the API**
|
|
41
|
+
> (submission date) — that property is *why* this one is promotable and WebSearch is not.
|
|
42
|
+
> 2. **FALLBACK — the export API query** (back off ≥30s between attempts). Kept on the ladder on
|
|
43
|
+
> purpose: if arXiv redesigns the listing markup, the channel must not disappear with it.
|
|
44
|
+
> 3. only if both fail → HN-only, and the progress line says `arxiv FAILED (<which transports, why>)`,
|
|
45
|
+
> never `0 items`.
|
|
46
|
+
>
|
|
47
|
+
> 🟥 **What this costs, stated rather than discovered later.** The listing path has the **weaker
|
|
48
|
+
> instrument** — its ID↔title pairing is the defect documented immediately below, and promoting it
|
|
49
|
+
> moves that defect from the fallback onto the **main** path. Therefore the two rules below are no
|
|
50
|
+
> longer conditional-on-being-in-fallback: **per-item `abs/{id}` re-resolution is a PRIMARY-path
|
|
51
|
+
> requirement**, one fetch per shipped ID, and a title mismatch drops the item. That is a real
|
|
52
|
+
> per-run cost increase (≈N extra fetches for N shipped items) and it is the price of the promotion,
|
|
53
|
+
> not an oversight. The old order paid 2 failed calls + a backoff instead; this pays N verified reads.
|
|
54
|
+
> 🟥 The progress line must now name **which transport produced the items** (`arxiv via listing` /
|
|
55
|
+
> `arxiv via export-api`) — otherwise a reader cannot tell which instrument's limits apply to the run.
|
|
37
56
|
> 🟥 WebSearch is **not** on this ladder. If a run used it anyway, the arxiv leg is reported as
|
|
38
57
|
> `RECALL-ONLY (WebSearch)` and its items are excluded from the NEW/REPEAT count.
|
|
58
|
+
>
|
|
59
|
+
> 🟥 **On the listing path, do NOT pair ID and title by position — pair them inside one block
|
|
60
|
+
> (measured 2026-09-12, this runner, and it is the phantom class this skill already has on record).**
|
|
61
|
+
> The listing yielded **51 IDs and 50 titles**, so two parallel extractions drift partway down and
|
|
62
|
+
> every pair after the drift point is wrong *while looking well-formed*. Concretely measured that run:
|
|
63
|
+
> `2609.10550` was attached to *"When Passing Tests Hides Vulnerabilities"* by the positional parse,
|
|
64
|
+
> but its own `abs` page is *"Optimizing AI Inference Across the Deployment Stack"* (2026-07-01);
|
|
65
|
+
> `2609.11060` and `2609.11008` mispaired the same way. 3 items dropped, 7 verified, **0 shipped
|
|
66
|
+
> unverified** — it was caught only because every shipped ID was re-resolved.
|
|
67
|
+
> **Two rules, and the second is the floor:**
|
|
68
|
+
> ① read each entry as **one unit** (the ID and the title that sit in the *same* list item), never as
|
|
69
|
+
> two ordered lists zipped together. A count mismatch between the lists is not a warning to
|
|
70
|
+
> reconcile — it is proof the zip is already wrong.
|
|
71
|
+
> ② until ① is demonstrably what happened, **per-item `abs/{id}` re-resolution is MANDATORY on this
|
|
72
|
+
> path**, not optional — one fetch per shipped ID, confirming the title. Any ID whose `abs` title
|
|
73
|
+
> disagrees is **dropped**, never shipped with a note.
|
|
74
|
+
> ⚠️ There is no hook here and none is possible: the parse happens inside the session that read the
|
|
75
|
+
> page, so this is salience over a manual floor. The floor that actually held on 2026-09-12 was ② —
|
|
76
|
+
> keep it even after ① looks solved (`[[feedback_anchor_can_be_decorative]]`).
|
|
39
77
|
|
|
40
78
|
### HackerNews (Algolia API)
|
|
41
79
|
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
---
|
|
2
|
-
name:
|
|
2
|
+
name: harness-pr-reviewer
|
|
3
3
|
description: Checks a submitted PR against the environment's baseline assets (CLAUDE.md, memory, naming, asset classification) and attaches a review comment with a merge recommendation. 5 steps — diff read, 8-area consistency check, self-catch, comment, merge recommendation.
|
|
4
4
|
user-invocable: true
|
|
5
5
|
allowed-tools: ["Bash", "Read", "Grep", "Glob"]
|
|
@@ -15,13 +15,21 @@ complexity_routing:
|
|
|
15
15
|
|
|
16
16
|
> **Note:** The original developer is the forge-harness original developer (development source + meta-monitoring home). In external user install environments, the install environment user themselves is the baseline integrity gate operator (following path B generalization baseline / `SKILL_detail.md §External User Environment Adaptation Path` §).
|
|
17
17
|
|
|
18
|
-
#
|
|
18
|
+
# harness-pr-reviewer — Hub Gate Operation Rule Automation
|
|
19
19
|
|
|
20
20
|
When a PR is submitted, checks consistency against the user environment's baseline assets (CLAUDE.md · memory · naming · asset classification) and attaches a review comment. 5-step: diff read → 8-matrix check → self-catch → comment attachment → merge recommendation.
|
|
21
21
|
|
|
22
22
|
## Activation Triggers
|
|
23
23
|
|
|
24
|
-
|
|
24
|
+
> **Renamed from `hub-cc-pr-reviewer` (2026-09-12, operator decision).** The old name said where it
|
|
25
|
+
> was born (the hub's own cc reviewing its own PRs). What it actually is — measured on qasp/pmh PRs in
|
|
26
|
+
> 2026-09 — is the **standalone doorway through which a field harness verifies its own submitted PR
|
|
27
|
+
> with FH's review capability**. Old-name utterances ("hub-cc-pr-reviewer", "hub review", "hub cc
|
|
28
|
+
> review") still route here; downstream forks that carry the old directory name (PMH) keep working
|
|
29
|
+
> until they sync. No redirect stub directory is shipped (same policy as `phantom-quench`).
|
|
30
|
+
|
|
31
|
+
|
|
32
|
+
1. **PR #N input**: *"Review PR #N"* / *"Check PR #N"* / *"hub review"* (old-name alias) / *"harness PR review"* / *"baseline consistency check"*
|
|
25
33
|
2. **Action leader cc → hub sync point**: Large decision area PR catch (following Option C Hybrid policy — memory creation / CLAUDE.md change / CATALOG round / skill v0.x evolution / policy change / asset synergy branch judgment)
|
|
26
34
|
3. **Hub cc session entry**: Layer A auto-read recent external commit catch (auto-discover new PRs)
|
|
27
35
|
|
|
@@ -75,6 +83,54 @@ Self-precision catch areas after first cc review (following previous PR self-cat
|
|
|
75
83
|
|
|
76
84
|
Self-catch areas 0 items = skip this entire catch matrix — do not pad with token-filling to make the section look populated.
|
|
77
85
|
|
|
86
|
+
🟥 **Self-catch is a CUE, not a verdict.** Its output is the *input* to Step 3.5 below. A self-catch
|
|
87
|
+
that finds nothing does not clear Axis 2 or Axis 3 — measured 2026-09-08/11: a reviewer session with
|
|
88
|
+
no external cue ran the load-bearing gate **0/51**, and a PMH session reviewing qasp PR #13 dispatched
|
|
89
|
+
the cross-family panel only after the operator pushed (pmh-dev #77). The self-check *is* the reviewer;
|
|
90
|
+
it cannot be its own decorrelation.
|
|
91
|
+
|
|
92
|
+
### Step 3.5. Axis 2 · Axis 3 — mandatory dispatch lane (mechanical trigger, typed degrade)
|
|
93
|
+
|
|
94
|
+
Axis 1 above is wired at mandatory strength (`--pr` mode, typed verdict). Until 2026-09-12 Axis 2
|
|
95
|
+
(adversarial panel) and Axis 3 (phantom-quench) had **no dispatch block in this skill at all** — only
|
|
96
|
+
the self-catch matrix — so they fired only when someone said so. This step closes that with a trigger
|
|
97
|
+
that is decided from the PR, never from the reviewer's feel.
|
|
98
|
+
|
|
99
|
+
**Trigger — mandatory if ANY of these holds** (compute all three; record which fired):
|
|
100
|
+
|
|
101
|
+
```
|
|
102
|
+
(i) Step 2 matrix has ≥1 ❌ (Inconsistent)
|
|
103
|
+
(ii) the PR touches a load-bearing FH asset: SKILL.md · SKILL_detail.md · .claude/rules/*.md ·
|
|
104
|
+
knowledge/shared/rules/*.md · templates/** · scripts/**/*.sh · scripts/**/*.py ·
|
|
105
|
+
plugins/*/agents/** (same list as CLAUDE.md §FH Improvement 4-Axis Auto-Gate)
|
|
106
|
+
→ gh pr diff "$PR" --name-only | grep -E '(^|/)SKILL(_detail)?\.md$|^\.claude/rules/|^knowledge/shared/rules/|^templates/|^scripts/.*\.(sh|py)$|^plugins/[^/]+/agents/'
|
|
107
|
+
(iii) a merge recommendation (Step 5) will be issued for this PR
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
None of the three → record `axis2: not-triggered(reason)` · `axis3: not-triggered(reason)` in the
|
|
111
|
+
comment capsule and continue. This is the *only* non-run path, and it is stated, never silent.
|
|
112
|
+
|
|
113
|
+
**Axis 2 — cross-family adversarial panel.** Dispatch **`auto-decorrelation`** on the PR diff (reuse,
|
|
114
|
+
not a new engine — its Step 1 load-bearing predicate, Step 4.5 residency screen and Step 6 degrade
|
|
115
|
+
ladder all apply unchanged). What this step owns is only the *decision to dispatch*. The result lands
|
|
116
|
+
as the typed `crossfamily:` value: `panel(<families>)` with the source-grounded findings, or one of
|
|
117
|
+
`declined` · `DEGRADED_SINGLE_FAMILY` · `DEGRADED_PANEL_UNUSED` · `UNKNOWN` **with grounds** — e.g.
|
|
118
|
+
consent OFF, no different-family sidecar reachable, quota exhausted. A degrade value is allowed; an
|
|
119
|
+
*absent* value is not. 🟥 A judgment-type question put to the panel (by-design vs fail-open, is this
|
|
120
|
+
a regression) needs **reps ≥ 3**; a 1-of-3 split is recorded as *unresolved*, never as CONCUR
|
|
121
|
+
(`field_verdict_crossfamily_gate.md` §4-2).
|
|
122
|
+
|
|
123
|
+
**Axis 3 — phantom-quench.** Run `/phantom-quench` over the PR's changed documentation surfaces
|
|
124
|
+
(any `*.md` in the diff, plus every path/citation the PR body asserts). Mechanical N/A only when
|
|
125
|
+
`gh pr diff --name-only` has zero `*.md` files **and** the PR body cites no path — record
|
|
126
|
+
`axis3: N/A(no doc surface)`. A phantom finding is a ❌ in the comment capsule, and it flips the Step 5
|
|
127
|
+
recommendation the same way an Axis 1 M-tier block does.
|
|
128
|
+
|
|
129
|
+
**Degrade direction.** This is a review surface (reversible): tooling-down does not block the
|
|
130
|
+
*comment*, but the merge recommendation must carry the degrade value verbatim and read
|
|
131
|
+
**NOT-CONVERGED** while Axis 2 is a `DEGRADED_*`/`UNKNOWN` value on a load-bearing PR. Silent
|
|
132
|
+
same-family pass is the one outcome this step forbids.
|
|
133
|
+
|
|
78
134
|
### Step 4. Review Comment Attachment
|
|
79
135
|
|
|
80
136
|
**Mandatory before any `gh pr comment`: run `/public-surface-audit` over the composed comment text.**
|
|
@@ -176,6 +232,22 @@ All 5 Steps completed
|
|
|
176
232
|
passes; they mean Axis 1 did not examine this PR and the recommendation
|
|
177
233
|
may not cite it as green
|
|
178
234
|
|
|
235
|
+
+ Step 3.5 trigger computed from the PR (matrix ❌ count · load-bearing path
|
|
236
|
+
grep · merge-recommendation requested) and recorded with which clause fired
|
|
237
|
+
— measured: the three booleans appear in the comment capsule. Absent = the
|
|
238
|
+
step did not run, which is the pre-2026-09-12 defect (pmh-dev #77)
|
|
239
|
+
|
|
240
|
+
+ Axis 2 dispatched via auto-decorrelation when triggered, landing a typed
|
|
241
|
+
`crossfamily:` value (panel(...) or a DEGRADED_*/UNKNOWN value WITH grounds)
|
|
242
|
+
— mandatory-pass: the value exists and is one of the closed enum. A
|
|
243
|
+
load-bearing PR whose value is DEGRADED_*/UNKNOWN leaves the merge
|
|
244
|
+
recommendation NOT-CONVERGED; judgment-type panel questions carry reps ≥ 3
|
|
245
|
+
or are recorded unresolved
|
|
246
|
+
|
|
247
|
+
+ Axis 3 phantom-quench run over the PR's doc surfaces when triggered, or
|
|
248
|
+
N/A(no doc surface) decided by `--name-only` grep, never by feel
|
|
249
|
+
— mandatory-pass: a phantom finding is a ❌ that flips the recommendation
|
|
250
|
+
|
|
179
251
|
+ /public-surface-audit run over the composed comment text BEFORE any
|
|
180
252
|
gh pr comment
|
|
181
253
|
— mandatory-pass, fail-closed: CLEAN attaches; REVIEW/LEAK redact-and-rescan;
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
|
-
name:
|
|
3
|
-
description: On-demand detail for
|
|
2
|
+
name: harness-pr-reviewer-detail
|
|
3
|
+
description: On-demand detail for harness-pr-reviewer — step bash commands, comment template, sister-asset utilization, external-environment adaptation, disable path, and persona synergy handling. Read when executing a step or operating in an external/own-PRS/deep-insight environment.
|
|
4
4
|
load: on-demand
|
|
5
5
|
---
|
|
6
6
|
|
|
@@ -35,7 +35,7 @@ Meaning of isolation: The Critic reads the synthesizer conclusion but does not i
|
|
|
35
35
|
- FAIL after re-synthesis → auto-persist as `fh_signal` on hold (no additional retries)
|
|
36
36
|
- Maximum retries: **1**
|
|
37
37
|
|
|
38
|
-
**Post-Core-Skill Critic Verdict Connection**: Following core skills can have Critic called inline after completion: harness-doctor · verify-bidirectional ·
|
|
38
|
+
**Post-Core-Skill Critic Verdict Connection**: Following core skills can have Critic called inline after completion: harness-doctor · verify-bidirectional · harness-pr-reviewer · context-doctor · sim-conductor. Trigger: immediately after completion announcement + "steel-quench" / "re-validate" / "run Critic" utterance.
|
|
39
39
|
|
|
40
40
|
---
|
|
41
41
|
|
|
@@ -150,7 +150,7 @@ grep -rl "complexity_routing" plugins/*/skills/*/SKILL.md
|
|
|
150
150
|
|
|
151
151
|
# 2. Aggregate escalation records from fh_signal files
|
|
152
152
|
grep -rh "" tracks/_meta/fh_signal_*.md 2>/dev/null | \
|
|
153
|
-
grep -oE "(harness-doctor|verify-bidirectional|
|
|
153
|
+
grep -oE "(harness-doctor|verify-bidirectional|harness-pr-reviewer|context-doctor|sim-conductor|agent-composer|harvest-loop|steel-quench)" | \
|
|
154
154
|
sort | uniq -c | sort -rn
|
|
155
155
|
```
|
|
156
156
|
|
|
@@ -169,7 +169,7 @@ grep -i -n "pr\|pull request\|audit\|review\|weekly" CLAUDE.md 2>/dev/null | hea
|
|
|
169
169
|
```
|
|
170
170
|
|
|
171
171
|
Judgment:
|
|
172
|
-
- Existing PR convention present → possible priority conflict with `
|
|
172
|
+
- Existing PR convention present → possible priority conflict with `harness-pr-reviewer` ⚠️
|
|
173
173
|
- Existing weekly audit present → possible format conflict with `harvest-loop` ⚠️
|
|
174
174
|
|
|
175
175
|
### 2-2. Skill Trigger Conflicts
|
|
@@ -322,7 +322,7 @@ Only activate the cluster matching the utterance — loading all skills degrades
|
|
|
322
322
|
| **A — Harvest/Evolution** | harvest · session wrap-up · pattern · reverse absorption · fh evolution | field-harvest · harvest-loop · contention-layer · verify-bidirectional |
|
|
323
323
|
| **B — Diagnosis (doctor)** | diagnose · health · check · inspect · token waste · install conflict · doctor | harness-doctor · context-doctor · sim-conductor · install-doctor · install-wizard |
|
|
324
324
|
| **C — Compose (composer)** | agent composition · parallel · compose prompt · context card | agent-composer · meta-prompt-builder |
|
|
325
|
-
| **D — Audit/Review** | review · audit · PR review · steel quench · adversarial · lint · placement |
|
|
325
|
+
| **D — Audit/Review** | review · audit · PR review · steel quench · adversarial · lint · placement | harness-pr-reviewer · steel-quench · apex-review · harness-doctor (--lint) · marketplace-gate · asset-placement-gate |
|
|
326
326
|
| **E — Explore/Frontier** | trends · plugin recommendation · synergy · frontier | frontier-digest · plugin-recommender · cross-ecosystem-synergy-detection |
|
|
327
327
|
| **F — Common (always-on)** | — | convergence-loop · deliberation (fh-commons) |
|
|
328
328
|
|
|
@@ -53,7 +53,7 @@ When Goal is provided in natural language (e.g., "I need to report to the team l
|
|
|
53
53
|
| "analyze", "diagnose", "something seems off" | harness-doctor → context-doctor | Structural + contextual diagnosis simultaneously |
|
|
54
54
|
| "simulate", "validate", "meta" | sim-conductor (D-code or Area B) | Multi-perspective validation |
|
|
55
55
|
| "install", "setup", "onboarding" | plugin-recommender → install-wizard | Recommend then install |
|
|
56
|
-
| "review", "code check" | sim-conductor D-code →
|
|
56
|
+
| "review", "code check" | sim-conductor D-code → harness-pr-reviewer | Code review chain |
|
|
57
57
|
|
|
58
58
|
Output format:
|
|
59
59
|
```
|
|
@@ -246,7 +246,7 @@ Resume is scope-bound — the scope from Step 0 is preserved across resume calls
|
|
|
246
246
|
After a `CLEAN (--full)` or `PENDING` sweep, the following are natural follow-ons:
|
|
247
247
|
|
|
248
248
|
- `field-harvest` — harvest patterns surfaced during the sweep
|
|
249
|
-
- `
|
|
249
|
+
- `harness-pr-reviewer` — if sweep was run pre-PR, feed results into PR review
|
|
250
250
|
- `agent-composer` — if multiple fix tasks are needed across the Pending list, compose agents to resolve them in parallel
|
|
251
251
|
- `return-path-gate --all` — if Step 0.5 surfaced OPEN chains, run a full chain closure audit
|
|
252
252
|
|
|
@@ -98,7 +98,7 @@ When queried for a specific capability (e.g., "adversarial reviewer for bash cod
|
|
|
98
98
|
⚠️ **A failed or empty discovery lane is NOT "no candidates".** Both CLIs above list only *configured* marketplaces, and a non-zero exit / empty array means the lane did not answer — not that nothing exists. Report each lane's state explicitly (`EXECUTED` / `EMPTY` / `FAILED: <stderr>`) and never render a `FAILED` lane as a zero result; a lane that could not run must be re-run or replaced by the web-search fallback (Priority 3) before you tell the user nothing was found.
|
|
99
99
|
|
|
100
100
|
**Discovery priority**: built-in (Tier 0) > installed > FH native > Tier 1 (any platform) > Tier 2 > Tier 3 > Tier 4
|
|
101
|
-
**Tier 0 guard**: FH native wins over a built-in only when the FH skill adds governance the built-in lacks (e.g. `/goal` → `goal-quench` adds budget+quality gates; code diff review stays with built-in `/code-review`, FH-asset coherence with `
|
|
101
|
+
**Tier 0 guard**: FH native wins over a built-in only when the FH skill adds governance the built-in lacks (e.g. `/goal` → `goal-quench` adds budget+quality gates; code diff review stays with built-in `/code-review`, FH-asset coherence with `harness-pr-reviewer`)
|
|
102
102
|
|
|
103
103
|
**When sim-conductor chains here for persona discovery**: apply the same platform-aware search scoped to persona/simulation/review capability tags. Return discovered agents with their Tier rating so sim-conductor can decide whether to install or use a built-in brief.
|
|
104
104
|
|
|
@@ -542,7 +542,7 @@ Verdicts: PASS · CONDITIONAL_PASS (S/R only, or Area B cadence skip) · FAIL (M
|
|
|
542
542
|
|
|
543
543
|
## Operations Notes
|
|
544
544
|
|
|
545
|
-
**AI-AI Loop Bias Defense**: sim-conductor →
|
|
545
|
+
**AI-AI Loop Bias Defense**: sim-conductor → harness-pr-reviewer → sim-conductor loop shares LLM cognitive blind spots. Internal convergence is "provisional convergence" — elevated only by external review or human gate. Area A/B convergence without human gate = incomplete.
|
|
546
546
|
|
|
547
547
|
---
|
|
548
548
|
|