@chrono-meta/fh-gate 1.4.86 → 1.4.88
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/rules/fh_4axis_gate.md +1 -1
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +8 -0
- package/CLAUDE.md +30 -6
- package/knowledge/shared/harness-core/capability_composition_contract.md +416 -0
- package/knowledge/shared/harness-core/fh_detail_protocols.md +4 -0
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +128 -0
- package/knowledge/shared/rules/auto_project_mapping.md +64 -1
- package/knowledge/shared/rules/operations.md +26 -0
- package/package.json +6 -2
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/skills/prompt-regression/SKILL.md +1 -1
- package/scripts/degrade_direction_scan.sh +34 -2
- package/scripts/gate_pathspec_check.sh +91 -1
- package/scripts/package_coverage_check.sh +27 -0
- package/scripts/selfcheck.sh +112 -1
- package/scripts/session_close_check.sh +18 -4
- package/scripts/test_degrade_scan_shell_probes.sh +11 -2
- package/scripts/test_dispatch_log_lanes.sh +41 -1
- package/scripts/test_node_infra_delta_lanes.sh +211 -0
- package/scripts/test_session_close_chain_lanes.sh +349 -0
- package/scripts/test_wizard_snippet_merge_lanes.sh +202 -0
- package/templates/.claude/rules/session.md +28 -5
- package/templates/.git-hooks/pre-commit +9 -2
- package/templates/degrade_direction_scan.sh +34 -2
- package/templates/regression_guard.sh +25 -0
|
@@ -868,3 +868,131 @@
|
|
|
868
868
|
just declared success. Adversarial review does not substitute for first use; today is the mirror:
|
|
869
869
|
self-repair does not substitute for adversarial review. Axis 2 shipped NOT CONVERGED with the open
|
|
870
870
|
leak channel declared in the header rather than closed by assertion.'
|
|
871
|
+
|
|
872
|
+
- date: 2026-08-03
|
|
873
|
+
agent: codex gpt-5.5 (cross-family, stdin form) + fh-meta:challenger (isolated Agent) + one headless
|
|
874
|
+
Sonnet blind sim (claude -p, tools stripped, repo-external cwd)
|
|
875
|
+
task: Axis 2 + target-tier sim for the ablation known-pair calibration — scripts/ablation_calibrate.sh
|
|
876
|
+
(new), scripts/probe_scope_check.sh header, .claude/regression/probes.md rows G-GATE-08/09
|
|
877
|
+
mode: two families dispatched on the same diff; consolidated here as one entry with measured counts
|
|
878
|
+
outcome: accepted
|
|
879
|
+
evidence: 'Codex 4 critical / 4 high / 3 medium; challenger 2S / 7A / 2B. The two families found
|
|
880
|
+
DIFFERENT classes and only partly overlapped. Codex: stderr merged into the scored stream, so a
|
|
881
|
+
single input echo passes P, N and T at once (every control token lives in the input); no runner
|
|
882
|
+
exit-status check; unanchored grep -q 47 accepting 470/1147; baseline gated on a shared fail flag.
|
|
883
|
+
Challenger (executed against the shipped script, not reasoned): control T conflated "could not read"
|
|
884
|
+
with "read it and the answer was not there" and certified a fully tool-enabled runner as calibrated;
|
|
885
|
+
the two arms shared one cwd so the negative arm could read the answer off its neighbour; probes.md
|
|
886
|
+
Count line shipped stale INSIDE this diff (30/mp25 vs actual 32); G-GATE-09 typed mandatory-pass
|
|
887
|
+
while being mechanically unobservable; the public script cited gitignored tracks/** records four
|
|
888
|
+
times as its evidence. All reproduced before acting. Sonnet blind sim of the repaired header: PASS —
|
|
889
|
+
recovered calibrate-first, pre-registration, reps>=3, the allowlist form, and answered "no" to
|
|
890
|
+
"one rep, arm B answered, can you record CUT?".'
|
|
891
|
+
note: 'Second time in one day that the miss landed where self-repair had just declared success. New
|
|
892
|
+
this round: the ANCHOR ITSELF was decorative on first try — reverting three fixes left the lane file
|
|
893
|
+
15/15 green. Two different causes: two lanes did not carry an input that exposes the defect (they
|
|
894
|
+
went red for an unrelated reason and read as coverage), and one revert silently did not apply. A
|
|
895
|
+
revert is an instrument too; measure that it took effect before believing the anchor.'
|
|
896
|
+
|
|
897
|
+
- date: 2026-08-03
|
|
898
|
+
agent: fh-meta:challenger (isolated Agent) ×6 across three branches + codex gpt-5.5 (cross-family)
|
|
899
|
+
×1 + headless Sonnet target-tier sims ×6 (claude -p, tools stripped, repo-external cwd)
|
|
900
|
+
task: session 2 — Axis 2 for the CLAUDE.md gate-scope pointer fix (4R), the templates/ domain-scoped
|
|
901
|
+
sections fix (1R), the probe_scope_check carryovers (1R), plus the cross-family leg on the
|
|
902
|
+
load-bearing pointer change and the before/after known-pair sims for each salience change
|
|
903
|
+
mode: consolidated entry with measured counts, per the log's own allowance
|
|
904
|
+
outcome: accepted
|
|
905
|
+
evidence: 'Pointer branch 0S3A → 1S5A → 1S6A → 1S2A; template branch 0S/3A; probe-scope branch
|
|
906
|
+
0S/3A/5B. The cross-family leg found in ONE pass what four same-family rounds had walked past —
|
|
907
|
+
the record asserted "REDUCED, not withdrawn" in one place and "withdrawn" in three others, 40
|
|
908
|
+
lines apart — and named the mechanism: the file performs visible self-correction in a persuasive
|
|
909
|
+
narrative voice, which makes an internal contradiction easy to wave through. Three separate
|
|
910
|
+
rounds independently pushed toward REDUCTION rather than tightening, and all three reductions
|
|
911
|
+
held: a second tool control, a prose/mechanism parity check, and a `[ -r ]` guard were each
|
|
912
|
+
measured inert or finding-generating and removed.'
|
|
913
|
+
note: 'The dominant new-defect class this session was MY OWN ANCHORS BEING DECORATIVE — four
|
|
914
|
+
instances, four different causes (a lane carrying no input that exposes the defect; a revert that
|
|
915
|
+
silently did not apply; a needle matching unrelated prose; a stub short-circuiting the layer the
|
|
916
|
+
guard lives in). None was self-caught. Harvested to [[feedback_anchor_can_be_decorative]].'
|
|
917
|
+
|
|
918
|
+
- date: 2026-08-04
|
|
919
|
+
agent: codex (gpt-5.5) — cross-family adversarial sidecar
|
|
920
|
+
invoked_by: FH self-dev session (card carry-over autopilot)
|
|
921
|
+
task: >-
|
|
922
|
+
Two separate adversarial audits. (1) A pre-commit hook trigger split plus a new known-pair
|
|
923
|
+
section in scripts/gate_pathspec_check.sh. (2) The §ⓐ restriction-union merge spec in
|
|
924
|
+
knowledge/shared/harness-core/capability_composition_contract.md.
|
|
925
|
+
outcome: accepted
|
|
926
|
+
evidence: >-
|
|
927
|
+
Audit (1): 3 defects, same-family inline pass found 0. The load-bearing one was a worktree-vs-
|
|
928
|
+
staged FAIL-OPEN — the checker read the worktree copy of CLAUDE.md, so staging a broken asset
|
|
929
|
+
line and then repairing the worktree WITHOUT staging passed the gate; reproduced before fixing
|
|
930
|
+
(index 0 occurrences / worktree 1 / anchor exit 0) and re-measured closed after (exit 1 both
|
|
931
|
+
directions). The other two: an `agents/` needle that passed on a partial class drop, and a stale
|
|
932
|
+
comment instructing the next maintainer to restore the exact coupling the change removed.
|
|
933
|
+
Codex explicitly CLEARED the LOADBEARING separation and the bash-3.2 quoting — a negative result
|
|
934
|
+
recorded because "found nothing here" is also evidence.
|
|
935
|
+
Audit (2): 5 defects, inline pass found 4 of them independently. The one net-new finding was the
|
|
936
|
+
worst of the set — §ⓐ.3 check 2 was written `merged ⊒ l` while §ⓐ.1 defined `a ⊑ b` as "a is at
|
|
937
|
+
least as strict as b", i.e. the spec's only mechanical direction check asserted the INVERSE of
|
|
938
|
+
its own invariant. Fixed by removing the glyph and writing every strictness claim as a
|
|
939
|
+
`permits(...) ⊆` relation.
|
|
940
|
+
note: >-
|
|
941
|
+
Second consecutive session where the cross-family leg produced net-new findings that same-family
|
|
942
|
+
rounds walked past, and where the net-new finding was the most severe one. Counted 5 dispatches
|
|
943
|
+
in the SubagentStop tally (2 codex audits plus 3 harness-side spawns); consolidated into this one
|
|
944
|
+
entry per the "counts and outcomes, not one stub per dispatch" rule.
|
|
945
|
+
Instrument note: the invocation form matters — `| codex exec -` (stdin) as recorded in
|
|
946
|
+
[[feedback_sidecar_liveness_not_passive_wait]]; the argv form hangs nondeterministically.
|
|
947
|
+
|
|
948
|
+
- date: '2026-08-04'
|
|
949
|
+
agent: codex/gpt-5.5 (cross-family sidecar, auto-decorrelation)
|
|
950
|
+
context: >-
|
|
951
|
+
FH self-dev — two load-bearing gate changes (PR #253 degrade-scan S5 rule, PR #254/#255 Axis-1
|
|
952
|
+
pathspec parity + anchor wiring). Pre-push gate demanded a cross-family leg on both.
|
|
953
|
+
dispatches: '3 (2 completed, 1 killed on timeout and re-run)'
|
|
954
|
+
outcome: accepted
|
|
955
|
+
evidence: >-
|
|
956
|
+
Round 1 (PR #253): REFUTED the author's own fix. The first repair widened the S5 rule with a
|
|
957
|
+
NAMED filter list; codex produced reproduced counter-examples in BOTH directions — still-missed
|
|
958
|
+
real disarms (`grep -Ec`, `grep --count`, `grep -Fcx`, `| cat`) and a NEW false positive the
|
|
959
|
+
repair introduced (`| tail -n +2` yields one line, not a disarm). It also caught that a fixture
|
|
960
|
+
labelled known-NEGATIVE was in fact a positive, because its upstream succeeded and therefore
|
|
961
|
+
never exercised the failure path it claimed to pin. All four re-run locally before acceptance.
|
|
962
|
+
Round 2 (PR #255 wiring): no blocker; one residual raised ("a distribution that SHOULD be
|
|
963
|
+
complete but drops a subject reports SKIP, not FAIL"), which was then TESTED and confirmed
|
|
964
|
+
uncovered by any existing anchor. Recorded as a WEAK signal — the auditor stated it could not
|
|
965
|
+
trace F1-F3 because those implementations were not in the diff it was given, and that residual
|
|
966
|
+
was closed locally by measurement, not by the auditor.
|
|
967
|
+
note: >-
|
|
968
|
+
Third consecutive session where the cross-family leg produced net-new findings the same-family
|
|
969
|
+
round walked past. New this session: the two rounds differed in SIGNAL STRENGTH (refutation vs
|
|
970
|
+
nothing-found), and both markers record which one it was — recording them at equal weight would
|
|
971
|
+
launder a weak pass into a strong one.
|
|
972
|
+
Instrument note: the FIRST attempt derailed into its own sidecar-availability exploration and
|
|
973
|
+
timed out (exit 124, ~7 min). Re-running with an explicit "Do NOT run shell commands, do NOT use
|
|
974
|
+
tools, answer ONLY from the text below" preamble fixed it. Stdin form `codex exec -m gpt-5.5
|
|
975
|
+
--skip-git-repo-check -` as recorded in [[feedback_sidecar_liveness_not_passive_wait]].
|
|
976
|
+
These were CLI sidecars, not Agent-tool spawns, so the SubagentStop hook tallied ZERO — the
|
|
977
|
+
④-e check passed on a true zero while three real dispatches had happened. Named gap: the tally
|
|
978
|
+
is blind to CLI-invoked sidecars.
|
|
979
|
+
|
|
980
|
+
- date: '2026-08-05'
|
|
981
|
+
agent: (unattributed — SubagentStop tally only)
|
|
982
|
+
context: >-
|
|
983
|
+
Close chain of the 2026-08-04 session, which crossed midnight. The tally file recorded ONE
|
|
984
|
+
dispatch dated 2026-08-05; this session made no Agent-tool spawn that it is aware of.
|
|
985
|
+
dispatches: '1 (tallied, origin not attributable)'
|
|
986
|
+
outcome: sustained
|
|
987
|
+
evidence: >-
|
|
988
|
+
Recorded rather than guessed. The substantive sidecar work of this session is the 2026-08-04
|
|
989
|
+
entry above (3 codex CLI dispatches). MEASURED here, after the claim had already been written
|
|
990
|
+
into a commit message and the session card: a `codex exec` call does NOT increment the tally
|
|
991
|
+
(before=1, after=1 on `grep -c '^2026-08-05$'`), and the hook matcher is `SubagentStop`, i.e.
|
|
992
|
+
Agent-tool spawns only. So the CLI-blindness claim holds — but its evidence arrived AFTER
|
|
993
|
+
publication, which is the ordering CLAUDE.md §Instrument-Calibration forbids.
|
|
994
|
+
note: >-
|
|
995
|
+
What produced the 4 tallies dated 2026-08-04 and this 1 dated 2026-08-05 is UNKNOWN. Not
|
|
996
|
+
invented: a fabricated outcome would poison the 60/40 promotion gate worse than a missing one,
|
|
997
|
+
which is the stated reason the hook only tallies and never writes an entry. Open question for
|
|
998
|
+
the next session that touches the tally: which tool events actually fire SubagentStop here.
|
|
@@ -91,12 +91,75 @@ governance depth to condition 4 (mechanical-anchor / 4-axis gate) — the litmus
|
|
|
91
91
|
|
|
92
92
|
| # | Item | Source → target | Effect |
|
|
93
93
|
|---|---|---|---|
|
|
94
|
-
| 1 | Session rules | `templates/.claude/rules/session.md` → `{project}/.claude/rules/session.md` | Session-start auto-read, backup, rule hierarchy |
|
|
94
|
+
| 1 | Session rules | `templates/.claude/rules/session.md` → `{project}/.claude/rules/session.md` **→ then §6-b prune + substitute (the copy is not the install)** | Session-start auto-read, backup, rule hierarchy |
|
|
95
95
|
| 2 | Context filter | `templates/.claudeignore` → `{project}/.claudeignore` | Token footprint control |
|
|
96
96
|
| 3 | Env card | `templates/fh_env_context.jsonc` → `{project}/.claude/rules/fh_env_context.jsonc` | Environment context for sessions |
|
|
97
97
|
| 4 | **MCP tool gating** (conditional — offered only when the project mounts an external MCP server: `.mcp.json`/`mcp.json` present, or the user is adding one) | `templates/.claude/rules/mcp_tool_gating.md` → `{project}/.claude/rules/mcp_tool_gating.md` | Name-keyed ask/allow tiers for external MCP tools — server annotations are unreliable (measured 2026-06-11: a live server shipped all-None hints incl. its irreversible send tool); §3 table filled at mount time |
|
|
98
98
|
| 5 | **Official-plugin scan** (recommend list only — **never auto-install**) | plugin-recommender Tier 0/1 pass on the project's stack | No-reinvention acceleration: matching `*-lsp` for the project's language + workflow plugins (code-review · commit-commands · feature-dev …) from `claude-plugins-official` — see `knowledge/shared/plugin-catalog/recommended_plugins.md` §Category 0.5. Each install user-approved |
|
|
99
99
|
|
|
100
|
+
### 6-b. Prune + substitute — the step that was missing, and the measurement that says so
|
|
101
|
+
|
|
102
|
+
**A copy is not an install.** Item 1 above used to end at the copy, and the field result was measured
|
|
103
|
+
on 2026-08-04 across every repo that inherited this template:
|
|
104
|
+
|
|
105
|
+
| Repo | `{FH_ROOT}` raw | `[CUSTOMIZE]` raw |
|
|
106
|
+
|---|---|---|
|
|
107
|
+
| A (bash/python wiki engine) | 1 | 3 |
|
|
108
|
+
| B | 2 | 6 |
|
|
109
|
+
| C (a *template* repo — identical is correct here) | 2 | 6 |
|
|
110
|
+
| D | **0** | **0** — the one that was customized by hand |
|
|
111
|
+
|
|
112
|
+
Three of four shipped with the hub-path placeholder unresolved. The earlier repair went to the
|
|
113
|
+
*template* (marking domain-scoped sections for deletion, 2026-08-04) and measured a real gain — but
|
|
114
|
+
only *when someone asked*. The field failure is that **there was no moment of asking**. This is that
|
|
115
|
+
moment; it is a step, not a better sentence.
|
|
116
|
+
|
|
117
|
+
Run immediately after each `templates/` copy, in the target repo:
|
|
118
|
+
|
|
119
|
+
```bash
|
|
120
|
+
TGT={project}/.claude/rules/session.md
|
|
121
|
+
# 1. SUBSTITUTE — the hub path this project should point at
|
|
122
|
+
FH_ROOT_ABS=${FH_ROOT_ABS:-$HOME/projects/forge-harness} # adjust if the hub lives elsewhere on this machine
|
|
123
|
+
sed -i.bak "s|{FH_ROOT}|$FH_ROOT_ABS|g" "$TGT" && rm -f "$TGT.bak" # -i.bak = the form BSD requires and GNU accepts
|
|
124
|
+
|
|
125
|
+
# 2. PRUNE — one ask per SECTION-HEADING marker, then delete the whole section (not just the marker)
|
|
126
|
+
grep -n '<!-- \[CUSTOMIZE\] DOMAIN-SCOPED' "$TGT"
|
|
127
|
+
|
|
128
|
+
# 3. VERIFY — Done-When. BOTH must print 0.
|
|
129
|
+
grep -c '{FH_ROOT}' "$TGT"
|
|
130
|
+
grep -c '<!-- \[CUSTOMIZE\] DOMAIN-SCOPED' "$TGT"
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
**Step 3 is the Done-When** *[mandatory-pass]* — and the two counts are deliberately narrow.
|
|
134
|
+
|
|
135
|
+
- **Why `{FH_ROOT}`**: an unresolved substitution token. It is the exact thing three of four inherited
|
|
136
|
+
repos shipped raw.
|
|
137
|
+
- **Why only the `<!-- ... -->` form of `DOMAIN-SCOPED`**: those three markers head a deletable
|
|
138
|
+
section. The template also *mentions* the phrase in prose ("TWO are DOMAIN-SCOPED and belong only
|
|
139
|
+
to…"), which is documentation and must survive — counting every occurrence would make the target
|
|
140
|
+
unreachable.
|
|
141
|
+
- **Why `[CUSTOMIZE]` is NOT counted**: the template's own instruction line reads "Change sections
|
|
142
|
+
marked with [CUSTOMIZE] comments to match your project". A zero target for that string can never be
|
|
143
|
+
met while the file explains itself, and an unreachable Done-When trains exactly one behaviour —
|
|
144
|
+
**deleting the marker instead of doing the work** — which is the defect this step exists to prevent.
|
|
145
|
+
`[CUSTOMIZE]` stays an editing hint, judged by the installer, not a counted gate.
|
|
146
|
+
|
|
147
|
+
> That distinction is not theory. The first draft of this step counted all three strings and demanded
|
|
148
|
+
> zero; running it on a fresh copy of the template returned **2 after a full prune** — both from lines
|
|
149
|
+
> that must stay. The instrument was checked against a known pair before this step shipped, which is
|
|
150
|
+
> the only reason the unreachable target did not become the rule (`§Instrument-Calibration`).
|
|
151
|
+
|
|
152
|
+
Do **not** silence a count by deleting the marker alone: the markers are attached to sections, and a
|
|
153
|
+
marker removed from a section that stayed is the same defect with its evidence erased.
|
|
154
|
+
|
|
155
|
+
**Do not split the template into core+domain modules.** The 2026-08-04 measurement pointed at the
|
|
156
|
+
*install step*, not at the file's shape, and this repo's own history says a new mechanism introduced
|
|
157
|
+
speculatively generates its own defects. The markers plus this step are the fix.
|
|
158
|
+
|
|
159
|
+
**Residency**: run this inside the target repo only. One of the counted repos is company-adjacent —
|
|
160
|
+
it was counted and left untouched, and its contents are not reproduced here or anywhere outside its
|
|
161
|
+
own environment.
|
|
162
|
+
|
|
100
163
|
**Field-asset scaffold (on-demand — NOT a `templates/` install)**: harness-ification surfaces field-specific skills/agents *as needed* — the field's domain work produces them. FH accelerates their *creation*: not the domain content (the field team authors that), but the **gate-compliant structure**. When a skill-worthy recurring pattern appears (3+ reps → `#skill-candidate` tag · `field-harvest` signal · a repeated manual workflow), offer to scaffold a skeleton:
|
|
101
164
|
|
|
102
165
|
- **Skill skeleton** — the frontmatter field set (`name`/`description`/`user-invocable`/`allowed-tools`) that is universal across active FH skills (regen: `grep -L user-invocable plugins/*/skills/*/SKILL.md` returns none; `model:` is deliberately NOT universal since 2026-07-10 — hard pins retired per `sonnet_floor_doctrine.md`, skills session-inherit and express depth as dispatch) + **four stubs, all mandatory in the skeleton**: ① Trigger stub (≥3 phrases — initiate) ② Done-When stub (with check-class — complete) ③ halt stub (budget/convergence guard) ④ persist stub (what state reaches the next run) + Step skeleton — the 5-question loop discipline by construction (`loop_engineering.md`; validate=#3 is the check-class declaration in ②). **Parameterizes** `contention-layer` Step 4's skeleton template (per its need-driven note — `origin: field-scaffold`, `contention-parents`→source-pattern pointer).
|
|
@@ -41,6 +41,32 @@ The hub audits and improves itself weekly.
|
|
|
41
41
|
|
|
42
42
|
---
|
|
43
43
|
|
|
44
|
+
## Flaky-lane triage — re-running erases the evidence you need
|
|
45
|
+
|
|
46
|
+
**Before you re-run a red CI job, pull the failing lane name out of its log.** A re-run overwrites the
|
|
47
|
+
run's `conclusion`, and GitHub's summary then shows the whole history as green — including the run
|
|
48
|
+
that failed. The failure is not hidden behind a filter; it no longer exists as a queryable state.
|
|
49
|
+
|
|
50
|
+
Measured 2026-08-03: a `sync_from_be_lanes.sh` flake was recorded and, when someone went looking the
|
|
51
|
+
next day, **all 14 runs that day reported `success`**. The re-runs had rewritten the conclusions and
|
|
52
|
+
the failing job's lane name had never been captured anywhere. There was nothing left to diagnose, so
|
|
53
|
+
nothing was diagnosed — a guessed diagnosis was correctly refused rather than written down.
|
|
54
|
+
|
|
55
|
+
Order, on the next occurrence:
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
gh run list --workflow <wf> --limit 30 --json databaseId,conclusion,headSha # find the red run FIRST
|
|
59
|
+
gh run view <id> --log-failed | grep -E '❌|FAIL|not ok' | head -20 # capture the lane NAME
|
|
60
|
+
# only now consider re-running
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
Record the lane name in the session log even if the re-run goes green — **a flake that cannot be named
|
|
64
|
+
cannot be fixed**, and the second occurrence is the one that gets fixed only if the first was named.
|
|
65
|
+
|
|
66
|
+
This is prose on purpose. N=1: one flake, one lost diagnosis. The mechanization threshold in this repo
|
|
67
|
+
is N≥3 or the same class recurring on another surface; a scheduled log-capture job built on a single
|
|
68
|
+
occurrence would be the speculative build the evidence-threshold rule forbids.
|
|
69
|
+
|
|
44
70
|
## 3-Phase Maturity Roadmap
|
|
45
71
|
|
|
46
72
|
The hub's long-term evolution path is managed as a 3-stage model: **Phase I (entering maturity) → Phase II (frontier following) → Phase III (frontier leading)**. Phase transition gates have the simplification principle ("A good harness gets simpler over time") as a common condition.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@chrono-meta/fh-gate",
|
|
3
|
-
"version": "1.4.
|
|
3
|
+
"version": "1.4.88",
|
|
4
4
|
"description": "FH runtime adapters — run FH governance, skills, and agents via Claude or Codex with machine-parseable gates.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"keywords": [
|
|
@@ -51,6 +51,7 @@
|
|
|
51
51
|
"CHEATSHEET.md",
|
|
52
52
|
"CLAUDE.md",
|
|
53
53
|
".claude/registry/agent_cards.json",
|
|
54
|
+
".claude/registry/README.md",
|
|
54
55
|
"docs/CONTRIBUTING.md",
|
|
55
56
|
"bin/fh-codex-doctor.js",
|
|
56
57
|
"bin/fh-gate.js",
|
|
@@ -148,6 +149,9 @@
|
|
|
148
149
|
"scripts/test_halffix_lanes.sh",
|
|
149
150
|
"scripts/test_ollama_panel_lanes.sh",
|
|
150
151
|
"scripts/test_tag_version_lanes.sh",
|
|
151
|
-
"templates/subagent-tally-hook.json"
|
|
152
|
+
"templates/subagent-tally-hook.json",
|
|
153
|
+
"scripts/test_session_close_chain_lanes.sh",
|
|
154
|
+
"scripts/test_node_infra_delta_lanes.sh",
|
|
155
|
+
"scripts/test_wizard_snippet_merge_lanes.sh"
|
|
152
156
|
]
|
|
153
157
|
}
|
|
@@ -58,7 +58,7 @@ ls .claude/regression/probes.md 2>/dev/null || echo "NO_CUSTOM_PROBES"
|
|
|
58
58
|
```
|
|
59
59
|
|
|
60
60
|
**If custom probes exist**: load and use them. The hub repo ships its golden probe set
|
|
61
|
-
(known-answer offline eval,
|
|
61
|
+
(known-answer offline eval, 32 probes with check classes) at exactly this path — when
|
|
62
62
|
present it is canonical and supersedes the default matrix below.
|
|
63
63
|
|
|
64
64
|
**If no custom probes** (e.g. Mode C install without the hub repo): use the default
|
|
@@ -164,9 +164,41 @@ for f in "${FILES[@]}"; do
|
|
|
164
164
|
# yields "9\n0" and the `-gt` guard goes silent. Verified as a known pair (both directions) in
|
|
165
165
|
# scripts/test_degrade_scan_shell_probes.sh; narrowing without that anchor would just trade a
|
|
166
166
|
# noisy probe for a blind one.
|
|
167
|
+
#
|
|
168
|
+
# WIDENED 2026-08-04 — the narrowing had gone one step too far, and the proof is that this
|
|
169
|
+
# probe was BLIND to the live instance that bit us the day before. PR #251 fixed
|
|
170
|
+
# `grep -c … | tr -d ' ' || echo 0` in session_close_check.sh ④-e; running this scanner over
|
|
171
|
+
# the pre-#251 file produced ZERO hits (re-measured: fixed rule = 2, old rule = 0).
|
|
172
|
+
#
|
|
173
|
+
# THE ANCHOR WAS ON THE WRONG PROPERTY. The old regex asked "is there an upstream pipe" and
|
|
174
|
+
# required the counter to be the LAST stage. Neither is the discriminator. The discriminator
|
|
175
|
+
# is: **does the failing side still EMIT?** `grep -c` prints "0" and exits 1 on no-match, so it
|
|
176
|
+
# disarms with no upstream pipe at all, and any trailing stage that passes that "0" through
|
|
177
|
+
# keeps the "0\n0" intact.
|
|
178
|
+
#
|
|
179
|
+
# A first fix widened to a NAMED filter list (tr|head|sed|…). Cross-family review (gpt-5.5,
|
|
180
|
+
# 2026-08-04) refuted it with reproduced counter-examples, and the refutation held on
|
|
181
|
+
# measurement in BOTH directions:
|
|
182
|
+
# · still missed, all verified to yield "0\n0": `grep -Ec …`, `grep --count …`, `grep -Fcx …`
|
|
183
|
+
# (combined/long flag forms the `-c` literal never matched) and `| cat`.
|
|
184
|
+
# · newly over-matched: `| tail -n +2` after the counter yields ONE line — not a disarm.
|
|
185
|
+
# Name-based approximation produced under-match and over-match simultaneously. So the trailing
|
|
186
|
+
# chain is no longer enumerated: the rule keys on the COUNTER (grep with a `c` in any flag
|
|
187
|
+
# cluster or `--count`; or `wc` behind a real pipe) plus the `|| echo 0` fallback, and accepts
|
|
188
|
+
# any chain between them.
|
|
189
|
+
#
|
|
190
|
+
# PRESERVED from the 2026-07-28 narrowing, still pinned as known-negatives: (a) `a || b ||
|
|
191
|
+
# echo 0` has no pipeline, and (b) a stage that emits NOTHING on failure (`| jq -r … ||
|
|
192
|
+
# echo 0`) — there the fallback supplies the only line, as intended.
|
|
193
|
+
# KNOWN RESIDUAL, stated not hidden: a trailing stage that can swallow the counter's output
|
|
194
|
+
# (`| tail -n +2`, `| sed -n '/[1-9]/p'`) is now flagged though it is not a disarm. That is a
|
|
195
|
+
# deliberate recall-over-precision trade on an ADVISORY probe, taken because the measured cost
|
|
196
|
+
# of the other direction was a real defect shipping. Revisit if non-fixture FPs appear.
|
|
197
|
+
# Known pair: scripts/test_degrade_scan_shell_probes.sh (9 positives / negatives silent).
|
|
198
|
+
# Reverting this line turns the positive lane red — checked by applying the revert.
|
|
167
199
|
while IFS= read -r m; do
|
|
168
|
-
emit "$f" "${m%%:*}" "S5:pipefail-fallback(sh)" "\`|| echo 0\` fallback
|
|
169
|
-
done < <(grep -nE '[^|]\|[[:space:]]*([a-z]+[[:space:]]+)*(grep[
|
|
200
|
+
emit "$f" "${m%%:*}" "S5:pipefail-fallback(sh)" "\`|| echo 0\` fallback after a COUNTER (grep with -c / --count, or wc) — the counter PRINTS \"0\" and exits non-zero on no-match, so the fallback appends a SECOND line, the value becomes \"0\\n0\", the integer test dies with 'integer expression expected' (stderr only) and the guard goes silent; split it and sanitize: n=\$(...); n=\${n:-0}; case \$n in *[!0-9]*) n=0;; esac. VERIFY the trailing stage can actually emit — if it swallows the count (| tail -n +2), this is a known false positive"
|
|
201
|
+
done < <(grep -nE '(([^|]\|[[:space:]]*([a-z]+[[:space:]]+)*wc[^|]*)|(grep([[:space:]]+-[A-Za-z]*c[A-Za-z]*|[[:space:]]+--count)[^|]*))([^|]*\|[^|]*)*\|\|[[:space:]]*echo[[:space:]]+[\"'"'"']?0' "$f" 2>/dev/null \
|
|
170
202
|
| grep -vE '^[0-9]+:[[:space:]]*#' \
|
|
171
203
|
| grep -vE '#[[:space:]]*noqa[:[:space:]]*degrade')
|
|
172
204
|
fi
|
|
@@ -107,7 +107,12 @@ spec_matches() { # $1 = path
|
|
|
107
107
|
for pair in \
|
|
108
108
|
"plugins/fh-meta/skills/frontier-digest/SKILL_detail.md|tracks/_meta/x.md|PATHSPEC covers SKILL_detail.md" \
|
|
109
109
|
"plugins/fh-meta/skills/frontier-digest/SKILL.md|README.md|PATHSPEC still covers SKILL.md" \
|
|
110
|
-
"CLAUDE.md|CLAUDE.local.md|PATHSPEC covers CLAUDE.md but not the local override"
|
|
110
|
+
"CLAUDE.md|CLAUDE.local.md|PATHSPEC covers CLAUDE.md but not the local override" \
|
|
111
|
+
"scripts/selfcheck.sh|README.md|PATHSPEC covers scripts/*.sh (seam #1 — the HEAVY term already gated it, this array did not)" \
|
|
112
|
+
"scripts/sub/nested.sh|package.json|PATHSPEC covers scripts/ recursively (git/bash globs cross '/')" \
|
|
113
|
+
"templates/regression_guard.sh|package-lock.json|PATHSPEC covers templates/*.sh — i.e. the guard can guard ITSELF" \
|
|
114
|
+
"templates/.git-hooks/pre-commit|.gitignore|PATHSPEC covers the git-hook floor (the hook that hard-blocks commits)" \
|
|
115
|
+
"plugins/fh-meta/agents/challenger.md|tracks/_meta/y.md|PATHSPEC covers agent definitions (seam #3)"
|
|
111
116
|
do
|
|
112
117
|
IFS='|' read -r pos neg label <<< "$pair"
|
|
113
118
|
ok=1
|
|
@@ -155,6 +160,91 @@ else
|
|
|
155
160
|
fail=$((fail + 1))
|
|
156
161
|
fi
|
|
157
162
|
|
|
163
|
+
# ── 5. CLAUDE.md asset-list parity — the FIXED per-class needle table ─────────
|
|
164
|
+
# WHY A TABLE, AND WHY FIXED. A parity check was attempted for this on 2026-08-03 and CUT the same
|
|
165
|
+
# day: it guessed its needles per run, produced four findings in one adversarial round, and — the
|
|
166
|
+
# decisive part — never fired on `CLAUDE.md` at all, because `CLAUDE\.md` was not in the hook's
|
|
167
|
+
# GATE_IMPL trigger. Both defects are addressed here and neither is optional: the needles below are
|
|
168
|
+
# FIXED (a literal table, reviewed as data), and the hook gained a SEPARATE trigger
|
|
169
|
+
# (`ASSETLIST_IMPL`) so this anchor actually runs when the list it guards is edited. The trigger is
|
|
170
|
+
# deliberately NOT `GATE_IMPL`: that variable also feeds `$LOADBEARING`, and enrolling every
|
|
171
|
+
# CLAUDE.md prose edit into cross-family review is a different job. Retrying without both the
|
|
172
|
+
# fixed table AND a trigger repeats the cut.
|
|
173
|
+
#
|
|
174
|
+
# WHAT IT GUARDS. `CLAUDE.md` §FH Improvement 4-Axis Auto-Gate carries a canonical asset list, and
|
|
175
|
+
# `.claude/rules/fh_4axis_gate.md` carries another. When a class is added to the hook but not to a
|
|
176
|
+
# prose list (or dropped from one list only), the divergence is silent — the gate still blocks, so
|
|
177
|
+
# nothing goes red, while the resident text tells readers a class is not covered. That is the exact
|
|
178
|
+
# shape measured on 2026-08-03, when two classes were added to both lists with no anchor behind them.
|
|
179
|
+
#
|
|
180
|
+
# SCOPE, deliberately narrow: this asserts the asset-list SENTENCE names each class. It does not
|
|
181
|
+
# assert the hook and the sentence are equal sets — that stronger claim needs a parse of the regex
|
|
182
|
+
# into classes, which is the guessing this table replaces. Narrow and mechanical beats broad and
|
|
183
|
+
# self-generating; the enumeration sweep in §4 is where breadth lives.
|
|
184
|
+
CLAUDE_MD="$REPO_ROOT/CLAUDE.md"
|
|
185
|
+
# Read the STAGED blob as well as the worktree, and require BOTH to declare the class.
|
|
186
|
+
# Cross-family review (2026-08-04, codex/gpt-5.5) reproduced the fail-open: reading only the worktree
|
|
187
|
+
# lets a commit stage a broken asset line, repair the worktree WITHOUT staging, and pass — the hook
|
|
188
|
+
# sees CLAUDE.md in $STAGED, runs this anchor, and validates content that is not what is being
|
|
189
|
+
# committed. Measured before the fix: index 0 occurrences / worktree 1 / anchor exit 0.
|
|
190
|
+
# Checking only the index would invert the same hole (a broken UNSTAGED edit would pass a manual
|
|
191
|
+
# run), so both are required. `git show :CLAUDE.md` fails outside a repo or with nothing in the
|
|
192
|
+
# index; that is a skip of the staged leg, never a pass of it.
|
|
193
|
+
_asset_line_of() { # $1 = file-or-'-' content source label; reads stdin
|
|
194
|
+
grep -m1 '^\*\*FH 자산을 수정하면\*\*' || true
|
|
195
|
+
}
|
|
196
|
+
ASSET_LINE=$(_asset_line_of < "$CLAUDE_MD" 2>/dev/null || true)
|
|
197
|
+
# Distinguish "not in the index" (a real skip) from "git failed" (an instrument error). The advisory
|
|
198
|
+
# degrade lint flagged the first draft here, and correctly: a bare `|| true` let the staged leg
|
|
199
|
+
# silently abstain on ANY git failure, which half-reopens the fail-open this section was just fixed
|
|
200
|
+
# for. Tracked files always have an index entry, so for this repo the abstain branch is unreachable
|
|
201
|
+
# in normal operation — it exists for a clone where CLAUDE.md is untracked.
|
|
202
|
+
if git -C "$REPO_ROOT" ls-files --error-unmatch CLAUDE.md >/dev/null 2>&1; then
|
|
203
|
+
ASSET_LINE_STAGED=$(git -C "$REPO_ROOT" show :CLAUDE.md 2>/dev/null | _asset_line_of)
|
|
204
|
+
if [ -z "$ASSET_LINE_STAGED" ]; then
|
|
205
|
+
echo " ❌ CLAUDE.md is tracked but its staged blob yielded no declaration line —"
|
|
206
|
+
echo " instrument error (or the staged content dropped the line entirely), NOT a pass."
|
|
207
|
+
fail=$((fail + 1))
|
|
208
|
+
fi
|
|
209
|
+
else
|
|
210
|
+
ASSET_LINE_STAGED="" # untracked clone — worktree leg is the only measurable one
|
|
211
|
+
fi
|
|
212
|
+
if [ -z "$ASSET_LINE" ]; then
|
|
213
|
+
echo " ❌ CLAUDE.md asset-list declaration line not found — instrument error, NOT a pass"
|
|
214
|
+
echo " (looked for a line starting '**FH 자산을 수정하면**' in $CLAUDE_MD)"
|
|
215
|
+
fail=$((fail + 1))
|
|
216
|
+
else
|
|
217
|
+
# label|needle (ERE, matched against the declaration line only)
|
|
218
|
+
for entry in \
|
|
219
|
+
'SKILL.md|SKILL\.md' \
|
|
220
|
+
'SKILL_detail.md|SKILL_detail\.md' \
|
|
221
|
+
'.claude/rules|\.claude/rules/' \
|
|
222
|
+
'knowledge/shared/rules|knowledge/shared/rules/' \
|
|
223
|
+
'templates/|templates/' \
|
|
224
|
+
'CLAUDE.md|CLAUDE\.md' \
|
|
225
|
+
'AGENTS.md|AGENTS\.md' \
|
|
226
|
+
'scripts/**/*.sh|scripts/\*\*/\*\.sh' \
|
|
227
|
+
'agent definitions (plugins/*/agents)|plugins/\*/agents/' \
|
|
228
|
+
'agent definitions (.claude/agents)|\.claude/agents/'
|
|
229
|
+
do
|
|
230
|
+
lbl="${entry%%|*}"; needle="${entry#*|}"
|
|
231
|
+
_hit=1
|
|
232
|
+
printf '%s' "$ASSET_LINE" | grep -qE "$needle" || _hit=0
|
|
233
|
+
# The staged leg only votes when it exists — an empty index blob is "not measured", not "absent".
|
|
234
|
+
if [ -n "$ASSET_LINE_STAGED" ]; then
|
|
235
|
+
printf '%s' "$ASSET_LINE_STAGED" | grep -qE "$needle" || _hit=0
|
|
236
|
+
fi
|
|
237
|
+
if [ "$_hit" -eq 1 ]; then
|
|
238
|
+
echo " ✅ CLAUDE.md asset list declares: $lbl"; pass=$((pass + 1))
|
|
239
|
+
else
|
|
240
|
+
echo " ❌ CLAUDE.md asset list no longer declares: $lbl"
|
|
241
|
+
echo " The hook still gates it, so nothing goes red — the resident text now tells readers"
|
|
242
|
+
echo " a gated class is not covered. Re-add it, or retire the class from the hook too."
|
|
243
|
+
fail=$((fail + 1))
|
|
244
|
+
fi
|
|
245
|
+
done
|
|
246
|
+
fi
|
|
247
|
+
|
|
158
248
|
echo
|
|
159
249
|
if [ "$fail" -eq 0 ]; then
|
|
160
250
|
echo "gate_pathspec_check: PASS ($pass pairs)"
|
|
@@ -65,6 +65,12 @@ ACCEPTED_ABSENT=(
|
|
|
65
65
|
# user's own cross-context wiring — the template it is copied FROM is what ships.
|
|
66
66
|
".claude/rules/local_fh_context.md"
|
|
67
67
|
".claude/regression/probes.md"
|
|
68
|
+
# Its sibling, and absent for the same reason one layer up: this file records which sections of
|
|
69
|
+
# THIS repo's resident CLAUDE.md were measured load-bearing. A consumer's CLAUDE.md is their
|
|
70
|
+
# own text, so our verdicts are not merely useless to them — cited from a shipped doc they would
|
|
71
|
+
# read as claims about THEIR file. The shipped docs name it as the place verdicts live in the
|
|
72
|
+
# harness repo, which is a pointer for contributors, not a promise of a shipped artifact.
|
|
73
|
+
".claude/regression/ablation_verdicts.md"
|
|
68
74
|
"scripts/sync-to-be.sh"
|
|
69
75
|
"scripts/sync_guard_check.sh"
|
|
70
76
|
# Return path (companion store → hub) and its anchor. Same reason as the forward path above: the
|
|
@@ -77,6 +83,27 @@ ACCEPTED_ABSENT=(
|
|
|
77
83
|
# regression run must not compare against this harness's probe set). Shipping the reader without
|
|
78
84
|
# its corpus would put a script in the package that can only ever report "instrument error".
|
|
79
85
|
"scripts/probe_scope_check.sh"
|
|
86
|
+
# The known-pair precondition for `probe_scope_check.sh`'s ablation procedure. Absent for the same
|
|
87
|
+
# reason as its subject — the procedure ablates THIS repo's resident CLAUDE.md, so a consumer has
|
|
88
|
+
# nothing to point it at — plus one of its own: every run spends API calls against the consumer's
|
|
89
|
+
# account. Shipping a script whose only effect on a consumer's machine is a bill is worse than
|
|
90
|
+
# omitting it.
|
|
91
|
+
"scripts/ablation_calibrate.sh"
|
|
92
|
+
# Its anchor. Absent for exactly one reason — its SUBJECT is absent — and that pairing is the whole
|
|
93
|
+
# rule: shipping an anchor whose subject is not in the package is the defect that put a red
|
|
94
|
+
# selfcheck in front of every 1.4.85/1.4.86 consumer. selfcheck's block SKIPs when the subject is
|
|
95
|
+
# missing, so the package stays green without pretending the lanes ran.
|
|
96
|
+
"scripts/test_ablation_calibrate_lanes.sh"
|
|
97
|
+
# Anchor for probe_scope_check.sh, which is ACCEPTED_ABSENT above for want of its corpus. Same
|
|
98
|
+
# pairing rule: an anchor whose subject does not ship must not ship either, or the consumer's
|
|
99
|
+
# selfcheck goes red on a subject they do not have.
|
|
100
|
+
"scripts/test_probe_scope_lanes.sh"
|
|
101
|
+
# Measures what the LIVE `claude` CLI does with several SessionStart hooks on one matcher —
|
|
102
|
+
# so every run needs the CLI, auth, and spends tokens on the consumer's account. Same reason
|
|
103
|
+
# `ablation_calibrate.sh` is absent: shipping a script whose only effect on a consumer's
|
|
104
|
+
# machine is a bill is worse than omitting it. selfcheck reports it NOT EXERCISED (exit 2)
|
|
105
|
+
# where the CLI is missing, so the package stays green without pretending the lanes ran.
|
|
106
|
+
"scripts/test_sessionstart_multihook_lanes.sh"
|
|
80
107
|
)
|
|
81
108
|
|
|
82
109
|
out=$(python3 - "${ACCEPTED_ABSENT[@]}" <<'PY'
|
package/scripts/selfcheck.sh
CHANGED
|
@@ -202,6 +202,71 @@ else
|
|
|
202
202
|
fail=1
|
|
203
203
|
fi
|
|
204
204
|
|
|
205
|
+
# And the ablation calibrator, for the third instance of the same reason. Its whole job is telling
|
|
206
|
+
# apart states that look identical from outside — "the arm could not answer" vs "the runner is dead"
|
|
207
|
+
# vs "the runner read the answer off disk" — and round 2 of its own adversarial review found that
|
|
208
|
+
# three of its round-1 fixes had no discriminating lane at all. Unwired lanes are green on one
|
|
209
|
+
# machine and nowhere else, which is the case this file exists to prevent. Stub runners, no API
|
|
210
|
+
# spend, so running them here costs nothing. Not hermetic w.r.t. the filesystem: two isolation lanes
|
|
211
|
+
# create and remove an empty, per-PID, git-invisible directory in the worktree.
|
|
212
|
+
# probe-scope lanes — the subject had 15 sibling checkers with an anchor and none of its own, so
|
|
213
|
+
# three repairs shipped on 2026-08-03 that could each be reverted with nothing turning red. Same
|
|
214
|
+
# SKIP/FAIL shape as the blocks above: a missing SUBJECT is a legitimate skip, a present subject with
|
|
215
|
+
# a missing anchor is a real failure.
|
|
216
|
+
if [ ! -f scripts/probe_scope_check.sh ]; then
|
|
217
|
+
echo "SKIP test_probe_scope_lanes.sh (subject scripts/probe_scope_check.sh absent)"
|
|
218
|
+
elif [ -f scripts/test_probe_scope_lanes.sh ]; then
|
|
219
|
+
if ! bash scripts/test_probe_scope_lanes.sh; then
|
|
220
|
+
fail=1
|
|
221
|
+
fi
|
|
222
|
+
else
|
|
223
|
+
echo "FAIL test_probe_scope_lanes.sh: probe_scope_check.sh present but its anchor is missing"
|
|
224
|
+
fail=1
|
|
225
|
+
fi
|
|
226
|
+
|
|
227
|
+
# gate-pathspec anchor — wired here 2026-08-04. It was reachable ONLY from templates/.git-hooks/
|
|
228
|
+
# pre-commit, i.e. only in a clone where the operator had run `git config core.hooksPath`. Every
|
|
229
|
+
# other clone, every CI run, and the npm package carried the anchor file and never executed it —
|
|
230
|
+
# the built-but-not-wired shape, one layer up: the anchor for the gate had no anchor of its own.
|
|
231
|
+
# That mattered the same day: PR #254 added five known-pairs to it, all of which would have been
|
|
232
|
+
# unexecuted outside the author's machine.
|
|
233
|
+
# Subject = the two implementations it reads (the hook's HEAVY term and the guard's GUARD_PATHSPEC).
|
|
234
|
+
# Absent subject → package/partial surface → legitimate SKIP; present subject with the anchor gone
|
|
235
|
+
# → FAIL, same shape as every block above.
|
|
236
|
+
# NAMED RESIDUAL (cross-family, gpt-5.5, 2026-08-04): if a distribution that SHOULD be complete
|
|
237
|
+
# accidentally drops one subject, this reports SKIP, not FAIL — silent non-coverage. Measured the
|
|
238
|
+
# same day: removing `templates/.git-hooks` from package.json `files[]` and running
|
|
239
|
+
# scripts/package_coverage_check.sh still PASSED, so no existing anchor catches that omission
|
|
240
|
+
# either. Deliberately NOT patched with a stricter branch: the only discriminator available
|
|
241
|
+
# ("templates/ exists but the hook does not") would be built on an UNMEASURED assumption about how
|
|
242
|
+
# a narrower package is shaped, and this repo's rule is not to build before the constraint is
|
|
243
|
+
# measured. What is cheap and honest is naming WHICH subject is missing, so a SKIP is diagnosable
|
|
244
|
+
# instead of opaque. Revisit when a real partial distribution is observed.
|
|
245
|
+
_gps_missing=""
|
|
246
|
+
[ -f templates/.git-hooks/pre-commit ] || _gps_missing="templates/.git-hooks/pre-commit"
|
|
247
|
+
[ -f templates/regression_guard.sh ] || _gps_missing="${_gps_missing:+$_gps_missing, }templates/regression_guard.sh"
|
|
248
|
+
if [ -n "$_gps_missing" ]; then
|
|
249
|
+
echo "SKIP gate_pathspec_check.sh (subject absent: $_gps_missing) — not-checked, NOT a pass"
|
|
250
|
+
elif [ -f scripts/gate_pathspec_check.sh ]; then
|
|
251
|
+
if ! bash scripts/gate_pathspec_check.sh; then
|
|
252
|
+
fail=1
|
|
253
|
+
fi
|
|
254
|
+
else
|
|
255
|
+
echo "FAIL gate_pathspec_check.sh: the gate implementations are present but their coverage anchor is missing"
|
|
256
|
+
fail=1
|
|
257
|
+
fi
|
|
258
|
+
|
|
259
|
+
if [ ! -f scripts/ablation_calibrate.sh ]; then
|
|
260
|
+
echo "SKIP test_ablation_calibrate_lanes.sh (subject scripts/ablation_calibrate.sh absent)"
|
|
261
|
+
elif [ -f scripts/test_ablation_calibrate_lanes.sh ]; then
|
|
262
|
+
if ! bash scripts/test_ablation_calibrate_lanes.sh; then
|
|
263
|
+
fail=1
|
|
264
|
+
fi
|
|
265
|
+
else
|
|
266
|
+
echo "FAIL test_ablation_calibrate_lanes.sh: ablation_calibrate.sh present but its anchor is missing"
|
|
267
|
+
fail=1
|
|
268
|
+
fi
|
|
269
|
+
|
|
205
270
|
if [ ! -f scripts/fh_node_check.sh ]; then
|
|
206
271
|
echo "SKIP test_node_check_lanes.sh (subject scripts/fh_node_check.sh absent)"
|
|
207
272
|
elif [ -f scripts/test_node_check_lanes.sh ]; then
|
|
@@ -213,6 +278,52 @@ else
|
|
|
213
278
|
fail=1
|
|
214
279
|
fi
|
|
215
280
|
|
|
281
|
+
# The infra-delta half of the same subject. Separate suite, same pairing rule: it exists only because
|
|
282
|
+
# fh_node_check.sh does, so its absence beside a present subject is a FAIL, not a skip.
|
|
283
|
+
if [ ! -f scripts/fh_node_check.sh ]; then
|
|
284
|
+
echo "SKIP test_node_infra_delta_lanes.sh (subject scripts/fh_node_check.sh absent)"
|
|
285
|
+
elif [ -f scripts/test_node_infra_delta_lanes.sh ]; then
|
|
286
|
+
if ! bash scripts/test_node_infra_delta_lanes.sh; then
|
|
287
|
+
fail=1
|
|
288
|
+
fi
|
|
289
|
+
else
|
|
290
|
+
echo "FAIL test_node_infra_delta_lanes.sh: fh_node_check.sh present but its anchor is missing"
|
|
291
|
+
fail=1
|
|
292
|
+
fi
|
|
293
|
+
|
|
294
|
+
# SessionStart multi-hook + install-wizard snippet merge. Subject for both = the shipped settings
|
|
295
|
+
# snippets; a clone without them is a legitimate SKIP, a clone with them and no anchor is not.
|
|
296
|
+
for _pair in \
|
|
297
|
+
"templates/settings.SessionStart.snippet.json|scripts/test_sessionstart_multihook_lanes.sh" \
|
|
298
|
+
"templates/settings.SessionStart.snippet.json|scripts/test_wizard_snippet_merge_lanes.sh"
|
|
299
|
+
do
|
|
300
|
+
_subj="${_pair%%|*}"; _anc="${_pair#*|}"
|
|
301
|
+
if [ ! -f "$_subj" ]; then
|
|
302
|
+
echo "SKIP ${_anc##*/} (subject $_subj absent)"
|
|
303
|
+
elif [ -f "$_anc" ]; then
|
|
304
|
+
# THREE-valued, like the session-close anchors above — and for a third reason they do not have.
|
|
305
|
+
# test_sessionstart_multihook_lanes.sh measures what the LIVE `claude` CLI does with several
|
|
306
|
+
# SessionStart hooks on one matcher. It declares exit 2 = NOT EXERCISED (no CLI / no auth /
|
|
307
|
+
# opt-out). A CI runner structurally cannot have that CLI, so collapsing 2 into fail=1 makes
|
|
308
|
+
# every Linux run red forever — over-blocking, which is how a red CI stops being read at all
|
|
309
|
+
# (the same reasoning that keeps the session-close check advisory on ordinary pushes).
|
|
310
|
+
# 2 does NOT set fail, and it prints a line that cannot be misread as a pass. On a machine that
|
|
311
|
+
# DOES have the CLI the suite runs in full and a real failure still exits 1.
|
|
312
|
+
# Measured 2026-08-04: the first draft of this wiring flattened 2 into fail=1 and turned CI red
|
|
313
|
+
# while the suite had correctly reported "NOT EXERCISED — the `claude` CLI is not on PATH" —
|
|
314
|
+
# i.e. it rebuilt, ten lines below the comment warning against it, the exact flattening defect.
|
|
315
|
+
bash "$_anc"; _rc=$?
|
|
316
|
+
if [ "$_rc" -eq 2 ]; then
|
|
317
|
+
echo "NOT EXERCISED ${_anc##*/}: this environment cannot run the measurement (exit 2 — never a pass)"
|
|
318
|
+
elif [ "$_rc" -ne 0 ]; then
|
|
319
|
+
fail=1
|
|
320
|
+
fi
|
|
321
|
+
else
|
|
322
|
+
echo "FAIL ${_anc##*/}: $_subj present but its anchor is missing"
|
|
323
|
+
fail=1
|
|
324
|
+
fi
|
|
325
|
+
done
|
|
326
|
+
|
|
216
327
|
# Two guards that read the AUTHOR's own actions rather than the repo's files. Both were added
|
|
217
328
|
# 2026-07-31; the pipe-verdict lane shipped in PR #209 WITHOUT this wiring, which is itself the
|
|
218
329
|
# half-fix class the second guard exists to catch — found by running that guard on this repo.
|
|
@@ -249,7 +360,7 @@ else
|
|
|
249
360
|
fail=1
|
|
250
361
|
fi
|
|
251
362
|
|
|
252
|
-
for _anchor in scripts/test_session_close_lanes.sh scripts/test_card_drift_probe.sh; do
|
|
363
|
+
for _anchor in scripts/test_session_close_lanes.sh scripts/test_card_drift_probe.sh scripts/test_session_close_chain_lanes.sh; do
|
|
253
364
|
if [ ! -f scripts/session_close_check.sh ]; then
|
|
254
365
|
echo "SKIP ${_anchor##*/} (subject scripts/session_close_check.sh absent)"
|
|
255
366
|
elif [ -f "$_anchor" ]; then
|