@ccoalm/ccl-skills 0.10.0 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +4 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +178 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +127 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +11 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dispatch-owner-skills.md +9 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/problem-resolution-and-learning.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/tag-and-prod-pipeline-gate.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +9 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +7 -23
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +29 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +35 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +817 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +81 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +521 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +41 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +116 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +123 -32
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -8
- package/dist/assets/release.json +47 -37
- package/package.json +1 -1
|
@@ -4,7 +4,7 @@ Use this reference for the owner-dispatch discipline behind the technical design
|
|
|
4
4
|
|
|
5
5
|
## Why the complete owner set loads at design time
|
|
6
6
|
|
|
7
|
-
When the technical design gate is triggered and the deliverable's substance spans more than one owning skill — backend/platform/architecture (`go-microservice-architecture` / `python-service-architecture`
|
|
7
|
+
When the technical design gate is triggered and the deliverable's substance spans more than one owning skill — backend/platform/architecture (`go-microservice-architecture` / `python-service-architecture` for those stacks; for a stack with no `*-architecture` owner, this workflow's architecture gate plus the relevant `platform-*` skills — see the stack owner map below), LLM/inference/RAG/eval (`llm-inference-integration`), the storage/data-contract owner, the client stack owner (`web-react-dev` / `app-cross-platform-dev` / `miniapp-product-dev`), `testing-strategy`, `product-ui-ux-design` — those owning skills own BOTH (a) producing the design substance and (b) the review. Invoking the router does not discharge them.
|
|
8
8
|
|
|
9
9
|
- Before producing or reviewing, enumerate the concerns the deliverable touches and load the COMPLETE owner set for those concerns during design, not only at review. Do not load owners for untouched concerns, and do not discover required owners one at a time mid-review — if a new touched concern appears, add it to the inventory and load its owner before continuing.
|
|
10
10
|
- A cold dispatched worker is the worst case (see `multi-agent-delegation`).
|
|
@@ -33,3 +33,11 @@ This gate is enforceable in code, not only prose:
|
|
|
33
33
|
- ccl-skills itself ships no config (its own edits are shared-skill/process edits, exempt from owner-load).
|
|
34
34
|
|
|
35
35
|
The closeout-acquire rule (`status` must read `opted-in: yes`, or install, or record why exempt) stays in the entrypoint because it is a closeout item, not a map mechanic.
|
|
36
|
+
|
|
37
|
+
## Stack dev owners and the CLI implementation carve-out
|
|
38
|
+
|
|
39
|
+
The per-stack implementation owners are `go-microservice-dev` (Go), `python-service-dev` (Python), and `nodejs-service-dev` (Node.js). A newly added stack dev owner joins this map in the same landing that adds the owner, and the entrypoint's Development routing line carries the same list — the two must not drift.
|
|
40
|
+
|
|
41
|
+
- **CLI/tooling implementation without terminal-UI concerns goes to that language's dev owner** and must not fall back to `terminal-cli-dev` merely because the deliverable is a command. `terminal-cli-dev` keeps the interface contract — the rendered surface plus the command/subcommand/flag/help/exit contract it owns even when nothing is rendered — and each stack dev owner carries the reciprocal skip leg. A CLI in a language with no dev owner stays with `terminal-cli-dev` as the default CLI owner.
|
|
42
|
+
- **Node.js has no `*-architecture` sibling by decision, not by oversight.** Its architecture, service-boundary, data-ownership, and reliability decisions run through this workflow's architecture gate plus the relevant `platform-*` owners, with `nodejs-service-dev` owning the Node-side implementation mechanics. Do not borrow `go-microservice-architecture` or `python-service-architecture` for a Node.js service: their stack-specific RPC, storage, and codegen contracts do not transfer, and treating a missing sibling as "route to the nearest architecture skill" is the failure this rule prevents.
|
|
43
|
+
- A stack whose dev owner exists but whose architecture decisions have no owner is a **routing vacuum, not a valid `not-applicable`**. Name the owner that absorbs them, as Node.js does here; recording the absence itself as the reason restates the symptom and leaves the stack unreachable from this gate.
|
|
@@ -54,6 +54,8 @@ Update the smallest correct durable place:
|
|
|
54
54
|
- Go backend implementation, testing, codegen, DB, Redis, MQ, or protobuf issue: update `go-microservice-dev` references.
|
|
55
55
|
- Python backend, AI-service host, worker, SDK/package, or batch-job architecture issue: update `python-service-architecture`.
|
|
56
56
|
- Python implementation, pytest, packaging, schema, ORM/migration, Redis, queue, async, or service-wiring issue: update `python-service-dev` references.
|
|
57
|
+
- Node.js implementation, runtime/module/type path, package-manager or lockfile, async lifecycle, stream, worker, outbound-client, or `node:test`/runner issue: update `nodejs-service-dev` references.
|
|
58
|
+
- Node.js architecture or service-boundary issue: there is no Node architecture sibling skill by decision — update this workflow's architecture gate text or the relevant `platform-*` owner, and record which one absorbed it rather than filing it under a sibling stack's architecture skill.
|
|
57
59
|
- UI/product interaction issue: update the relevant design or frontend skill.
|
|
58
60
|
- One-off business/domain issue: do not turn it into a generic skill rule.
|
|
59
61
|
|
|
@@ -30,7 +30,7 @@ Use this for implementation of Python backend products, services, microservices,
|
|
|
30
30
|
- Convert domain-specific source patterns into reusable mechanics: route shape, schema validation, repository/unit-of-work boundary, transaction scope, cache key strategy, idempotency, job lease, config object, trace context, fake client, or test style.
|
|
31
31
|
- If an observed pattern only works for one product domain, discard it instead of turning it into a rule.
|
|
32
32
|
- Resolve conflicts by choosing the safer generic default: explicit schemas over dicts, typed settings over ad hoc environment reads, reviewed migrations over blind autogeneration, bounded async concurrency over unbounded gather, dependency injection over import-time clients, focused pytest tests over live-infra tests by default, and fail-closed for auth/permission/data-integrity paths.
|
|
33
|
-
- When adding or revising durable Python implementation guidance, check whether the lesson is generic backend service practice that should also update `go-microservice-dev`, or belongs in a shared workflow skill instead. If the rule depends on Python tooling, FastAPI/Flask/Django, Pydantic, asyncio, pytest, or Python package layout, keep it here and do not force a Go mirror.
|
|
33
|
+
- When adding or revising durable Python implementation guidance, check whether the lesson is generic backend service practice that should also update the sibling stack owners `go-microservice-dev` and `nodejs-service-dev`, or belongs in a shared workflow skill instead. Record each sibling as `update`, `unchanged`, or `route-to-shared` rather than leaving it unexamined. If the rule depends on Python tooling, FastAPI/Flask/Django, Pydantic, asyncio, pytest, or Python package layout, keep it here and do not force a Go or Node mirror.
|
|
34
34
|
|
|
35
35
|
## Development Workflow
|
|
36
36
|
|
|
@@ -18,3 +18,12 @@ After pushing a tag, read back:
|
|
|
18
18
|
- Produced image/digest/version evidence when available.
|
|
19
19
|
|
|
20
20
|
If the tag target is wrong after push, do not force-move a published production tag. Stop and escalate to the release owner for the corrective version/tag path.
|
|
21
|
+
|
|
22
|
+
## The version pointer is under the same immutability, one step earlier
|
|
23
|
+
|
|
24
|
+
A published version cannot be changed or reused, so the source tree's version pointer may never sit **below** the highest already-released version. Treat that as a checked invariant, not a convention:
|
|
25
|
+
|
|
26
|
+
- **Check it at merge time, not only at tag time.** A tag-time check blocks the bad release but leaves the wrong pointer on the integration branch until a person happens to notice, and the next bump then lands on a corrupted base.
|
|
27
|
+
- **The commit that lowers it is usually not a release commit.** The version line is a both-sides-changed hunk, so the observed shape is a conflict resolved the wrong way inside a change about something else entirely — the commit subject gives no warning, and reviewers reading it for its stated purpose skip the hunk.
|
|
28
|
+
- **Repair every site.** The version is stated in the manifest and again in the lockfile (twice, in current npm lockfile versions); a partial repair leaves the sites disagreeing, which is its own release defect.
|
|
29
|
+
- **Compare against the released record, not against a base branch.** "Did this branch lower it" is a different, weaker question than "does the tree point under something already published"; derive the record from release tags or the registry.
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md
CHANGED
|
@@ -49,10 +49,10 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
49
49
|
- Task-retrospective extraction must inspect the whole delivery chain, not only the final fix. For incidents, regressions, contract drift, weak UI, missed tests, bad reviews, or repeated user corrections, trace the failure through definition, implementation, verification, review/MR or release readiness, and retrospective quality. If the root cause includes this extraction workflow allowing a shallow summary, update this skill or its references before claiming the lesson is landed.
|
|
50
50
|
- **A retrospective over a LARGE multi-batch / multi-phase session has a second axis beyond the per-delivery chain: distinct lesson-TYPE axes that must each be covered or explicitly marked `no-new-lesson` — (a) per-artifact CONTENT lessons (the specific bug / contract value / domain rule; for research/writing/design programs this axis is the METHOD/CRAFT — how the work was done well), (b) PROGRAM/PROCESS lessons (how the multi-batch effort was structured and driven), (c) WORKFLOW/META lessons (did the retro or extraction itself recur shallow, under-trigger, or stop at the most salient content lesson), and (d) SUSTAIN lessons (what went RIGHT and how the next run reuses it — counts only with mechanism + non-luck evidence + owner routing; **axes (a)-craft and (d) read from the produced-artifact class — an enumeration driven by correction turns cannot reach them and will come back falsely empty**). Landing only the loudest content lesson and declaring the session "fully summarized / 复盘完成" is incomplete.** "LARGE" is not a vibe — it fires when the session already carries a coverage/program structure: a source register or named batch-progress standard was applied, OR the work spanned multiple explicit phases/batches/verticals. A user re-ask after a "done" claim = same-scope correction signal — classify first; never manufacture a lesson. Per-axis detail, re-ask classification, DO-CONFIRM card, `covered-through` watermark: `references/source-to-skill-extraction.md` (Task Retrospective Extraction).
|
|
51
51
|
- **A long operational delivery session also needs a separate non-lesson delivery-state axis (in addition to the content/program/meta lesson axes above).** When the session changed operational delivery state across multiple repositories, branches, MRs, pipelines, releases, or deployable artifacts, the source register must carry that axis — changed artifact set, branch/worktree state, remote/MR state, CI or local verification state, cancelled/retried pipeline state, unresolved risks, and the next concrete action — before "whole-session retro complete" is claimed; closeout records either the axis rows' locator (sanitized labels in the shared landing, real per-repo evidence in scratch/private archive) or `artifact/status axis: not-applicable` with a reason. If required rows are absent, the retro can be reported only as `interim`, even when the extracted lesson text is correct. The row-family fields and closeout-evidence forms: `references/source-to-skill-extraction.md` (Task Retrospective Extraction).
|
|
52
|
-
- A blocked verification item is not closed by naming the blockage. Before marking a test, device, browser, service, credential, or environment layer unavailable, attempt the normal remediation path for that layer
|
|
53
|
-
- **A landed CONCLUSION that a tool / capability / lane is unavailable, impossible, or must permanently fail-closed
|
|
54
|
-
- A blocked source read is not closed by naming the blockage. If Figma, code, document, API, or repository reads time out, return partial output, or fail transport, switch to a smaller or different read strategy before extracting rules — and when the source is **missing rather than unreadable**, change WHERE you enumerate instead
|
|
55
|
-
- **Large reads can lose the middle with no reliable signal — the trigger is read-OUTPUT size, so chunk proactively.** A read whose OUTPUT exceeds ~256 lines / ~10 KiB can be silently head+tail truncated (no marker guaranteed), so a single `cat`/whole-file read does not count as coverage even when it returns no error. Whenever you need a **complete** view — whole-file coverage, a no-findings/absence claim, or a load-bearing section read — chunk it under **both ~200 lines AND ~8 KiB** and confirm a mid-file section was ingested.
|
|
52
|
+
- A blocked verification item is not closed by naming the blockage. Before marking a test, device, browser, service, credential, or environment layer unavailable, attempt the normal remediation path for that layer — restart the client daemon, run the documented fallback (full rung list and sandbox-denial triage: the next ladder). Only record `unavailable` after remediation fails, with command evidence, residual risk, and the next concrete unblock action.
|
|
53
|
+
- **A landed CONCLUSION is a hypothesis until an operation that could have falsified it has been run** — both a claim that a tool / capability / lane is unavailable, impossible, or must permanently fail-closed ("fail-closed is the safe default" does not waive the in-env attempt) and a DIAGNOSIS of why an observed failure happened — a search hit proves the text EXISTS, not that it RAN on the path that failed, so run the falsifying operation first — exercise the suspected mechanism on the failing path for an observation only IT predicts, or build a paired control differing in exactly ONE variable — **only within existing sandbox/permission, non-destructive, synthetic-target, and credential-safety boundaries** (never unsafe mutation, prod/live credentials, secret-bearing state, or a permission-boundary bypass — trading this rule for the security/authority/data-loss axis). Where no safe attempt is available after remediation the record is `pending` with remediation and residual risk — never `unavailable`, `fail-closed`, or a stated cause — and an unfalsified cause is `hypothesis`, kept off shared surfaces, because withdrawing a landed cause costs more than testing it. **When REVIEWING a change that asserts impossibility/unavailability or rests on a diagnosis, independently run the same falsification attempt before accepting it** — an inherited "it can't be done" or "this is why it broke" is hypothesis-grade (see the named-convention primary-source re-verify rule). Both forms and failure shapes: `references/validation-and-landing.md` (Behavioral Validation).
|
|
54
|
+
- A blocked source read is not closed by naming the blockage. If Figma, code, document, API, or repository reads time out, return partial output, or fail transport, switch to a smaller or different read strategy before extracting rules — and when the source is **missing rather than unreadable**, change WHERE you enumerate instead. Both ladders: `references/source-to-skill-extraction.md#blocked-verification-and-source-read-remediation`. Failed or timed-out reads do not count as coverage.
|
|
55
|
+
- **Large reads can lose the middle with no reliable signal — the trigger is read-OUTPUT size, so chunk proactively.** A read whose OUTPUT exceeds ~256 lines / ~10 KiB can be silently head+tail truncated (no marker guaranteed), so a single `cat`/whole-file read does not count as coverage even when it returns no error. Whenever you need a **complete** view — whole-file coverage, a no-findings/absence claim, or a load-bearing section read — chunk it under **both ~200 lines AND ~8 KiB** and confirm a mid-file section was ingested. Detail: `references/source-to-skill-extraction.md#read-in-chunks-large-reads-lose-the-middle`.
|
|
56
56
|
- Think across the full delivery lifecycle before editing: product intent, design/UX, implementation, debugging, test strategy, launch acceptance, iteration feedback, team onboarding, and normal users without source access. A rule that improves only one slice while leaving another slice ambiguous is incomplete or belongs in a narrower skill.
|
|
57
57
|
- Evidence must come before new rules. Do not add a new conceptual layer, workflow gate, or strong claim first and then backfill supporting sources. If a useful rule appears before source review, keep it as a working hypothesis and do not land it until evidence confirms it, narrows it, or routes it elsewhere. For subjective design, UX, frontend/client, product, architecture, or review rules, unverified external expertise is not enough to land executable guidance.
|
|
58
58
|
- **Product-agnostic / industry-practice skills require an external authoritative source class in the evidence plan, not internal corpus alone.** An extraction sourced only from one internal corpus (an SOP, one repo, one project doc) shows what *this org* does, not whether the skill matches the public state of the art.
|
|
@@ -150,10 +150,10 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
150
150
|
- **One self-detectable firing point does exist and must be used: the moment YOUR OWN output names 沉淀 / 提炼 / 复盘 / "distil this into a skill", OR **enumerates what an external source has that we lack** (a gap list vs another pack; see `references/firing-point-placement.md`), that naming is a trigger to RECOGNISE the owner and load it** — not a licence to widen scope: shared-skill edits still need the authority you already have, so when the user's request covered only a status review or a narrow fix, record the extraction as `pending` with the owner named and ask rather than self-authorising a shared-skill change off your own suggestion.
|
|
151
151
|
- The mechanical backstops are (a) the closeout gate — a committed skill-change with neither a visible in-session `skill-extraction-workflow` invocation nor the round's durable charter/target-output record is `interim` (per the closeout gate's evidence forms) — and (b) **user-signal escalation**: you generally cannot self-count misses you did not notice, so a user-pointed-out under-trigger is a recurrence check (was there a similar miss earlier this session, even on another task?) and, if so, escalates to tightening the always-on discipline rather than landing another narrow per-case trigger.
|
|
152
152
|
- **Firing-point-placement corollary:** when the SAME meta-class (a precise gate walked past at the routing → pre-code/design transition) recurs at a *new* lifecycle sub-point despite prior bootstrap-salience + the closeout gate, the durable lever is **moving the owning gate's firing point ONTO the transition itself** (pre-substance-draft AND pre-first-impl-edit) and sharpening *name→invoke* — naming/knowing an owner is NOT invoking/loading it, and a named-but-unloaded owner's mechanical rules never fire — at the SAME transition, NOT another bootstrap/per-case bullet or more prose.
|
|
153
|
-
- **Record-field corollary (the forgery surface):**
|
|
153
|
+
- **Record-field corollary (the forgery surface):** any field that NAMES an owner — a checklist row, a CLI flag, a "decision:" slot — is fillable without invoking that owner, and filling it is what *feels* like discharging the gate, so it carries an explicit invoke bar on its triggered values.
|
|
154
154
|
- The self-detect firing point's authority boundary and observed shape, the record-field corollary expansion (the invoke-bar coverage set-diff mechanics, the delegation-dispatch worked case), the worked recurrence-chain, and the landed owner-dispatch implementation: `references/firing-point-placement.md`.
|
|
155
155
|
- **Run your own adversary to convergence BEFORE any "done / fixed / passing / covered / converged / complete" claim — your own such claim is the least-trustworthy thing you emit.** For any non-trivial completion/coverage/convergence claim, you must have already run — **yourself, not deferred to the user** — the verification or adversarial pass that would catch its failure, to a **clean fresh result** (a first clean pass on the current candidate, never a "confirm my fix" pass), OR **downgrade the claim to `interim` and name what you ran vs. didn't**. "Covered / converged / already handled" is a claim, not a status — back it with firing-path or clean-pass evidence or do not emit it; this self-adversary duty never narrows the mandatory dual-track challenge (it is the always-on generalization of self-audit-to-convergence, not a replacement for the gate).
|
|
156
|
-
That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified, and re-reading your own prose can only ever confirm that the prose is self-consistent with itself. **A mutation you did not APPLY is a hypothesis, not evidence** — bound its blast radius (never
|
|
156
|
+
That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified, and re-reading your own prose can only ever confirm that the prose is self-consistent with itself. **A mutation you did not APPLY is a hypothesis, not evidence** — bound its blast radius (never buy a RED by disabling a guard against a shared or live dependency; where no isolated path exists record the property `unverified`). **Prove the oracle can fail before trusting its clean verdict** — point it at something you know is broken and watch it report that; a check that can only ever say clean is no evidence, and whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration. **A failing anchor is first a question about the ANCHOR, not a verdict on the implementation** (§Self-audit). **A validated oracle is still clean only over the DIMENSIONS it crossed** — proving it can fail says nothing about the axis you never varied, so a clean run is reported with the dimensions it covers, and the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering — `testing-strategy` owns that list) before values. If no contradicting observation exists, the property is `unverified` and must be labelled that way rather than counted as audited.
|
|
157
157
|
A scoped "X verified; Y not run" is an interim checkpoint, **not** `done`/`complete`/`landed`: `Y not run` blocks a done/complete/landed claim unless a **risk owner — the user/maintainer, never the agent self-accepting — explicitly accepts the gap AND it is tracked to that owner** (agent self-labeling "risk accepted" or "deferred" does not qualify; scoping is a downgrade, never a license to call the narrowed slice done). **Recurrence signal:** a user prompting you to keep digging / verify / disputing a "covered/converged/done" is a premature-completion signal — on the **2nd** such correction in a session (even across different tasks) escalate to tightening this discipline, not just fixing the one case (per the repeated-correction escalation above).
|
|
158
158
|
The full self-adversary method — the mutation enumeration, the applied-mutation discipline, the independent-oracle validation, the re-owe-after-fixes rule, the graded-verdict calibration, and the recognition-dependent honesty caveat: `references/dual-track-review-gate.md` §Self-audit.
|
|
159
159
|
- Automatically trigger durable learning when extraction work exposes a reusable failure — **and when ordinary delivery work does, capture it here too, but without extraction taking over the delivery**: let the active owner (`product-rd-workflow` / `defect-diagnosis` / `testing-strategy` / …) handle the immediate work first, then route the durable lesson here. **For a premature-stop correction after affirmative continuation**, immediate recovery means first rerun the active owner's current continuation/blocking gate in full (for product R&D, Pre-Final Continuation Gate steps 1–6) against current state, then follow its observable outcome — proceeding only when a literal binding exists (the original proposed-next action/scope plus literal assent, preserved in the visible conversation or quoted exactly in trusted host-owned session/compaction state — never reconstructed, broadened, or substituted — or the user's correction literally naming the paused action and scope) — a semantic compaction paraphrase or a bare "why did you stop" complaint is not path-(b) authority, and a `blocked:` recovery without the step-1 evidence and a specific missing authority/ambiguity is invalid — asking again when neither binds, the user intervened, or scope/gates changed, and never copying real conversation text into a shared repository record. Do not let correction RCA or extraction extend a still-authorized delivery, and do not let stale assent bypass a newly pending or inconclusive gate. After delivery recovery, correction RCA plus the durable prevention landing and verification are still due before the turn can be reported complete; otherwise report `interim`. The full binding rules, the `continuing:`-line form, and the invalid-`blocked:`-recovery rule: `references/resume-paused-delivery.md`.
|
|
@@ -172,7 +172,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
172
172
|
- For any correction where the missed step was covered by any rule in this workflow that a reasonable reader would apply to the scenario, add or tighten a closeout gate that would have blocked the exact premature final answer.
|
|
173
173
|
- **"The owner skill already states the rule" / "avoid monotonic growth" does NOT license a memory-only or no-op landing when the gate demonstrably failed to fire.** If a rule exists yet the failure still happened (and would recur for another agent or project), adequate *content* is not adequate *enforcement*: land the firing mechanism — the trigger, closeout step, validator, or merged clause that makes the existing rule actually catch this case, in the owning shared skill — or prove it now fires. Retreating to a personal memory or "no change, content is fine" while the gate stays un-fired is the dodge this prevents (memory is supplement only, per the memory-only-insufficient rule).
|
|
174
174
|
- For changed upstream owners, `check-ccl-skills.sh` (via `scripts/impact-chain-gate.rb`) is the mechanical closeout over every added source-register row; the declaration format (behavioral-evidence / observed-failure / firing-path fragments), anchor rules, wording-only classification, the `RED-baseline` floor, and the author-declaration trust model: `references/external-practice-controls.md#behavioral-evidence-and-attestation`.
|
|
175
|
-
- For shared-skill changes, classify the diff before finalization: `shared-skill change`
|
|
175
|
+
- For shared-skill changes, classify the diff before finalization: `shared-skill change` is defined in `references/dual-track-review-gate.md` (it also covers the plugin-shipped command/behavior surfaces outside `skills/`); `wording-only` means punctuation, grammar, typo, formatting, or synonym substitution with no change to trigger, scope, routing, validation, condition, example, owner, or acceptance meaning, and touching no frontmatter (any `description`/frontmatter edit, even a pure typo fix, is a routing-surface change, never wording-only); all other changes are non-wording.
|
|
176
176
|
- Every shared-skill change, including wording-only edits, must include a recorded independent review row before commit.
|
|
177
177
|
- Non-wording shared-skill changes must also include a recorded challenge row AND a recorded behavioral-evidence row before commit (actual behavior/routing deltas require a true `RED-baseline`; unchanged controls use paired `semantic-control`; status rules in `references/dual-track-review-gate.md`). A non-wording owner package must include at least one `RED-baseline` row, so stable-control labels cannot self-clear the package. **For a DESTRUCTIVE/irreversible change, a `RED-baseline` row must show the protected predicates were mutated, not merely that negative cases ran.** Executing the must-NOT-touch cases is necessary but not sufficient: a negative probe that would still pass with its protecting predicate removed is evidence of nothing (recurring shape: a degenerate fixture short-circuits every probe on an unrelated conservative branch, so the safety predicate is never reached and the green suite certifies the hole).
|
|
178
178
|
- So the row records, per protected predicate, the removal that was **applied** and observed to turn the suite RED **for the right reason** — a bare non-zero exit does not qualify (a mutant that breaks syntax or fixture setup also exits non-zero and would bank a broken build as proof of sensitivity); the failure must be attributable to the named protected assertion, and attribution is **differential** (the owning assertion passes in the unmutated control and fails under the mutant, with no non-owning assertion failing) rather than a substring match on aggregate output. An unapplied "this mutation would fail it" is a hypothesis. `testing-strategy` owns the encoded form of that walk (route, don't copy) — for a destructive artifact the walk belongs inside the suite so a later fixture change cannot silently re-blind it.
|
|
@@ -256,7 +256,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
256
256
|
- Before extracting rules from broad or mixed sources, summarize coverage, contradictions, thin areas, and source-quality gaps.
|
|
257
257
|
- For tool failures during source reading, keep a retry ledger: failed method, observed error, fallback method, source rows recovered, source rows still missing, and whether the fallback evidence is strong enough for the target rule. Do not report a pressure test, re-read, or extraction as complete when the only successful step was rephrasing existing notes.
|
|
258
258
|
- When changing a source-derived skill's conceptual layer, refresh the relevant source classes before editing. If the source was not refreshed, label the change as wording/routing cleanup, not source extraction.
|
|
259
|
-
- Record coverage using
|
|
259
|
+
- Record coverage using one of the depth labels defined under `references/source-to-skill-extraction.md#extraction-charter`; invent none outside that full set.
|
|
260
260
|
- If evidence is thin, constrain the output: fewer rules, explicit low-confidence notes, wider "do not use when" boundaries, and a clear list of evidence that would improve the skill.
|
|
261
261
|
|
|
262
262
|
4. Extract candidate rules.
|
|
@@ -301,7 +301,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
301
301
|
- Compare lifecycle impact against the target-output map. If an affected stage has no owning target or no-output reason, validation fails even if all listed targets have correct diffs.
|
|
302
302
|
- For UI/UX, Figma, frontend, app, miniapp, or client extraction, validation must include the judgment-delta matrix per the Core Rules `What to extract` section (each layer marked new / confirmed / narrowed / routed / no new evidence with source observations; visual direction/tokens rows carry token provenance and the design/dev/test owner decision). If the matrix is absent or lacks the required visual-token fields, the work can be reported only as source inventory or execution hardening, not design-judgment extraction.
|
|
303
303
|
- UI/UX validation must include the minimum pressure set from `references/uiux-judgment-extraction.md`: long content, empty/no-data, error/retry, slow/weak network, permission/disabled, narrow/responsive, accessibility text scaling, orientation or viewport change where relevant, interruption/return recovery, and repeated-use/cache-hit behavior; mark irrelevant cases explicitly instead of silently skipping them.
|
|
304
|
-
- Source-map and final-response claims match actual evidence depth. Broad coverage language is a blocker unless the evidence ledger supports it.
|
|
304
|
+
- Source-map and final-response claims match actual evidence depth. Broad coverage language is a blocker unless the evidence ledger supports it. The same bar governs CAUSAL claims: a stated cause names the falsifying operation run against it or is `hypothesis`.
|
|
305
305
|
- Pressure-test authenticity check: if the final claim says "tested", "pressure-tested", "re-read", "re-extracted", or equivalent, validation must show the source artifact reopened in that validation pass. If no source artifact was reopened, downgrade the claim to "static review" or rerun the test correctly.
|
|
306
306
|
- Repeated-correction lessons have a visible prevention point in the owning skill, reference, checklist, validator, or final-response constraint recorded in a skill/reference. A chat-only apology, summary, or ad-hoc wording in the current reply does not count as durable landing.
|
|
307
307
|
- Every landed source-derived rule has an `observed` ledger entry with concrete observations. Any rule that started as a hypothesis is either converted to observed evidence, narrowed, routed, discarded, or explicitly left out of the landed skill.
|
|
@@ -524,29 +524,12 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
|
|
|
524
524
|
|
|
525
525
|
The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
|
|
526
526
|
|
|
527
|
-
Five is the generic `code-review` transport ceiling, not this extraction lane's
|
|
528
|
-
spend. Non-wording Agent-autonomous extraction calls go through
|
|
529
|
-
`scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1`: the
|
|
530
|
-
initial review plus at most one challenge. At round 2 the autonomous lane ends.
|
|
531
|
-
An authenticated human may request later review, but that is separately
|
|
532
|
-
attributed human-requested evidence outside this chain/budget, never an
|
|
533
|
-
additional Agent round. Unused generic capacity never authorizes automatic continuation.
|
|
534
|
-
The v3 closeout validator rejects referenced receipts whose recorded budget is
|
|
535
|
-
not the wrapper-fixed value and checks budget and ordering consistency within the caller-supplied
|
|
536
|
-
set. It cannot authenticate that the wrapper produced those receipts or that
|
|
537
|
-
the caller retained every earlier chain or receipt. The wrapper does not mint or
|
|
538
|
-
persist
|
|
539
|
-
`review_chain_id` or `autonomous_review_index`: the caller still supplies both,
|
|
540
|
-
and could start a fresh-looking chain after the final round. The validator detects bad
|
|
541
|
-
order inside the referenced set but cannot detect a prior chain the caller
|
|
542
|
-
omitted, so complete caller-owned ledger retention—and treating an Agent reset
|
|
543
|
-
as a contract violation—remains part of the boundary rather than a property the
|
|
544
|
-
local scripts prove.
|
|
527
|
+
Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **The lane spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There the autonomous lane ends. An authenticated human may request later review, but that is separately attributed human-requested evidence outside this chain/budget, never an additional Agent round. Unused generic capacity never authorizes automatic continuation. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
|
|
545
528
|
|
|
546
529
|
**Self-hosted chains break on every fix; the budget is summed across chains, never per chain.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break spends the SAME Agent-autonomous budget. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
|
|
547
530
|
|
|
548
|
-
1. **Batch dispositions; never re-chain per finding — and hold
|
|
549
|
-
2. **Sum spent rounds across all chains before opening one more.** Count every prior external round in the caller-retained ledger — every chain, finished or broken — against
|
|
531
|
+
1. **Batch dispositions; never re-chain per finding — and hold every fix until the round-2 challenge has run.** Triage the whole batch through the disposition bar and deep-self-review once, then hold, never deciding by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, and round 2 binds the round-1 candidate, so a fix applied between the two forfeits the double-receipt terminal and costs a fresh human-authorized chain to recover — an observed failure, not a hypothetical: a round that landed its review fixes before the challenge had to be closed by a user-granted continuation chain. So the rule through round 2 is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, then apply the held batch, MR/PR-listed, and let round 3's succession challenge — owed exactly when the batch moved the candidate — be what inspects it.
|
|
532
|
+
2. **Sum spent rounds across all chains before opening one more; the lane's cap is three rounds across two chains.** Count every prior external round in the caller-retained ledger — every chain, finished or broken — against that cap. The only restart the lane funds is the single succession challenge, opened with `--predecessor-chain-result-file` so the ended chain is carried rather than laundered into a fresh-looking loop; any OTHER restart opens with a fresh review by contract, trading a challenge round for a review round, and must not be opened autonomously at the cap. Effective exhaustion is reached when the remaining rounds cannot fund the closeout floor for any continuation; treat it exactly like the cap.
|
|
550
533
|
3. **Front-load packet quality in chain 1.** The first chain's packet must already be the full-context diff (`--unified` wide enough to carry whole files, e.g. `-U200`) with the plan frozen alongside the candidate; narrow packets breed packet-boundary pseudo-findings whose fixes break chains and burn rounds on artifacts of the packet itself.
|
|
551
534
|
4. **At the cap — or at effective exhaustion — the designed terminal is disposition, never another chain.** Apply or disposition the final batch, name every post-review fix in the MR/PR description, record the honest terminal state (`continuation_authorization_required` when the lane's final round ran and itself returned findings; otherwise — including a chain broken before its challenge could run — an interim record naming the last externally reviewed candidate and every later delta), and hand continuation or merge to the human. The post-review batch sits only on the pending MR/PR branch beside that record — the human's authenticated continuation, waiver, or merge decision is what certifies it, and it is never reported as reviewed. This is the bounded outcome working as designed, so do not report it as convergence and do not launder it through a fresh-looking chain.
|
|
552
535
|
|
|
@@ -607,13 +590,14 @@ convergence.
|
|
|
607
590
|
|
|
608
591
|
For a focused single-skill change:
|
|
609
592
|
- **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
|
|
610
|
-
- **Round 2 — challenge**: after implementer triage — fixes stay HELD:
|
|
593
|
+
- **Round 2 — challenge**: after implementer triage — fixes stay HELD: applying any fix before this round breaks the chain, so the challenge runs on the frozen round-1 candidate (self-hosted-chain rule; enumeration item 1 above) — attack the highest-risk unresolved surface with an unprimed prompt.
|
|
594
|
+
- **Round 3 — succession challenge, owed only when the fix batch moved the candidate**: apply the held batch, commit it, then ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate now hashes to. Unchanged (every finding accepted, pre-existing, or source-refuted) ⇒ the lane ends at round 2 and owes nothing. Changed ⇒ open ONE succeeding chain with `--predecessor-chain-result-file <round-2 receipt>` and challenge the landing candidate on a focus distinct from round 2's. This is the final Agent-initiated external round; its findings feed the post-budget checkpoint rather than an automatic further round, and the batch lands MR/PR-listed.
|
|
611
595
|
|
|
612
|
-
Broad extractions use the same
|
|
596
|
+
Broad extractions use the same lane budget, round 3 included on the same condition. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
|
|
613
597
|
|
|
614
598
|
### Anti-patterns
|
|
615
599
|
|
|
616
|
-
- **
|
|
600
|
+
- **Landing a fix batch no round ever saw**. The round-1 fix-up itself may introduce bugs, so a batch that moved the candidate owes the succession challenge of round 3 above — the earlier absolute ("always re-challenge after a non-trivial fix-up") was unreachable while the budget was two rounds, and an unreachable obligation reads as satisfied. A candidate the batch did not move owes nothing: the condition is the candidate hash, not the author's sense of how big the fix was.
|
|
617
601
|
- **Iterating external review until zero findings**. Stop Agent reviewer calls at the configured budget. Stabilized or repeated findings are recorded, triaged, and may cause a method/design change or a parked dependent slice; implementation and independent work continue.
|
|
618
602
|
- **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
|
|
619
603
|
- **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one. The observed extreme is chain multiplication: a chain broken by your own fix and restarted at index 1 is the same budget, not a new loop — sum rounds across chains per the self-hosted-chain rule above.
|
|
@@ -139,6 +139,22 @@ Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds
|
|
|
139
139
|
- 给列表内的外部 CLI 传长 flag 之前,**必须先在本机读过该 CLI 的 help 输出**;未读过就用,属于把记忆当一手源。
|
|
140
140
|
- 一条命令因无法识别的 flag 失败时,**不得**先改用法之外的东西——先怀疑该 flag 在本机这个版本不受支持。语料里多次出现「以为是用法错、其实是这版没这个 flag」。
|
|
141
141
|
|
|
142
|
+
### 记录半边:把外部契约钉进控制时,同时钉住它的时效
|
|
143
|
+
|
|
144
|
+
上两节讲**读法**(读一手源、把谓词挂在自己拥有的东西上);这一节讲**记录**。一旦本仓的某个控制**编码了**对外部契约的一次读取——闸强制它、脚本依赖它、recipe 假设它——控制自身必须带三样东西,且写在控制旁边,不是写在 commit message 或某轮私档里:
|
|
145
|
+
|
|
146
|
+
- **出处到条款级**:读的是哪一份一手源的哪一节,不是「按官方文档」。
|
|
147
|
+
- **核验时点与被核对象的版本**:哪天核的、对着工具/宿主的哪个版本或哪个 pin 核的。
|
|
148
|
+
- **失效条件**:必须写明什么变化会使这条结论不再成立(pin 的大版本被抬、宿主换了解析链、上游把该字段改成会读的)。
|
|
149
|
+
|
|
150
|
+
缺这三样,下一个读到该控制的 agent 分不清它是**经过判断的决定**还是**没人敢动的化石**,唯一的复核方式是把整轮外部调研重做一遍——这正是「继承来的读法」能一路存活的原因:不是没人想核,是核一次的代价被隐藏了。三样都在时,复核退化成一次针对性的对照。
|
|
151
|
+
|
|
152
|
+
- A control whose annotation lacks any of the three must be treated as carrying an unverified reading: the next change to it re-reads the primary source instead of inheriting the claim, and a reviewer must not accept the inherited reading as established merely because the control has shipped for a long time.
|
|
153
|
+
|
|
154
|
+
判据窄,别泛用:只对**控制依赖的外部契约**成立(宿主/工具/格式/协议的行为)。纯内部不变量不适用;「谁提出的」那一档归 `attribution-verification.md` 的源质量与 pending 语义,与本条不同 facet。
|
|
155
|
+
|
|
156
|
+
本轮两个活体实例(都是读代码时先撞上「这为什么在这」才发现的):`verify-packed.mjs` 用机械闸禁止 `plugin.json` 带 `version`——这是一条关于宿主版本解析链的主张,落地时正确、但当时没留出处与时点;以及 `check-release-version.py` 依赖 `actions/checkout` 的 `fetch-depth: 0` 会取到 tag。两处已按上述三样补齐。
|
|
157
|
+
|
|
142
158
|
## 只据内部语料的「深度提炼」:失败形态
|
|
143
159
|
|
|
144
160
|
一次仅以内部 launch-SOP 为源的「深度提炼」会落出**看起来完整**的规则集,却从未核过这套技能声称代表的**公开实践**;下一个用户于是问「参考网上优秀实践了么 / did you check industry practice」。
|
|
@@ -119,7 +119,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
119
119
|
- File: `~/.<host>/skills/.extraction-work/<project>-completion.md`
|
|
120
120
|
- Final state: which batches done, which deferred, which sources unavailable.
|
|
121
121
|
- Lessons: what surprised; what would change in next extraction; what to add to skill-extraction-workflow.
|
|
122
|
-
- For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row
|
|
122
|
+
- For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. Ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate hashes to first: if the held fix batch moved it, the ledger owes the succession challenge bound to that hash, and the same script is the merge-side gate that refuses a landing whose evidence binds a different candidate. When the whole candidate exceeds one packet, split the review rather than the pull request: `--print-manifest --partition <paths> ...` renders a landing partition manifest, and one validated ledger per partition plus the committed manifest is what the gate binds. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row and does not fabricate a multi-round ledger — but note the cost the merge-side gate imposes on it: that gate accepts only a validator-checked ledger, because it cannot authenticate a hand-writable receipt, so a wording-only change that touches the bound paths still owes the two-round chain before it can land.
|
|
123
123
|
|
|
124
124
|
### 5. Provenance migration
|
|
125
125
|
|
|
@@ -28,6 +28,8 @@ The at-add-time check above decides *where* a rule lands (merge vs new bullet);
|
|
|
28
28
|
| Agent complies but the artifact's SHAPE is wrong (bloated review packet, buried verdict, register row restating the source) | Positive recipe/contract: state what the artifact IS — its parts, in order | A "don't"-list about the shape — the measured backfire above |
|
|
29
29
|
| Agent omits a required element from an artifact it already produces (missing status/evidence cell, absent map row) | A REQUIRED slot in the template/validator it must fill (our closeout rows and register gate are this form) | Prose reminders near the template |
|
|
30
30
|
| Behavior should differ by situation | Conditional keyed to an observable predicate ("fan-out → name the tier") | Unconditional rule + exemption clauses |
|
|
31
|
+
| Every step is right and the aggregate is still wrong — a batch, destructive, or otherwise high-stakes operation names a target that does not exist, misses a required item, or applies two conflicting changes, and the damage is done before anything checks | **plan-validate-execute**: the step emits a structured plan artifact, a script validates the plan, and only a valid plan is executed. Errors name the offending item *and* the set it was checked against ("field `signature_date` not found. Available: …"), because a verdict the agent cannot act on sends it back to guessing. The plan is also what makes the operation reversible while it is still cheap — iterating on the plan touches nothing (Anthropic, *Skill authoring best practices*, "Create verifiable intermediate outputs"; verified 2026-09-01) | Prose care ("verify each target before applying"), or a validator whose output is a bare pass/fail |
|
|
32
|
+
| Instruction latitude does not match the operation's fragility — a fixed-sequence fragile operation written as a heuristic (the agent improvises a variant), or a genuinely multi-solution judgement written as one exact script (the agent follows it off a cliff when the context differs) | Pick the tier from the operation, not from the author's confidence: **high** freedom (prose heuristics) when several approaches are valid and context decides; **medium** (parameterised template or pseudocode) when a preferred pattern exists and configuration varies; **low** (one exact invocation, few or no parameters, stated as not-to-be-modified) when the operation is fragile, consistency is critical, or a specific sequence must hold (same source, "Set appropriate degrees of freedom"; verified 2026-09-01) | One uniform specificity across a whole skill |
|
|
31
33
|
|
|
32
34
|
The rows are not mutually exclusive — a real failure often carries several axes (e.g. a required field omitted only under pressure). Give each axis its own form **on its own target**: the discipline axis keeps a prohibition aimed at the *act of skipping/violating*; the shape axis gets the recipe/slot describing *what the output is*. Do NOT re-express the shape requirement as a prohibition merely because a discipline axis is also present — a "don't restate X"-style prohibition riding along with a recipe is exactly the measured backfire. (This combination guidance is an inference beyond the source's separately-tested arms; like any load-bearing form choice it is subject to the behavioral-evidence row below.)
|
|
33
35
|
|
|
@@ -493,3 +493,32 @@ Round 073-receipt-bundling rows (new table so the entry renders as a table row a
|
|
|
493
493
|
| Core Rules own each invariant and a Workflow step only points at it: same-facet text living in both surfaces is converged toward the canonical surface rather than restated, and the sweep enumerates candidates instead of fixing whichever one a diff happens to touch | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/SKILL.md#do not make normal users route through a source name | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. The Step 6 cross-section facet-ownership check already owned this rule and fires on any edit touching Core Rules or a step; this round ran it as a five-candidate enumeration rather than a spot fix (the user challenged an earlier framing that treated the duplications as a word-budget offset ledger). Verdicts: content placement (Core Rules content-placement bullet vs two Step 5 bullets) converged to a pointer; the sibling-generalization mini-map field list (Core Rules owner-generalization group vs Step 4) converged to a pointer keeping only the step-order clause; capability naming (Core Rules naming bullet vs Step 5 vs two Step 6 checklist lines) converged to one Step 5 pointer plus one merged Step 6 residual-search check; representative sampling (Core Rules full-ask prohibition vs Step 3 labeling duty) judged complementary and left unchanged; the source-register row schema (Step 3 vs references/source-register.md) left unchanged as out of this contract's Core-Rules-versus-step scope and useful where an author builds rows. Zero-loss obligation map for the four rewritten passages: entrypoint-owns-trigger/routing/non-negotiables and references-own-detail both survive in the Core Rules content-placement bullet; the mini-map field list, the `update`/`unchanged`/`route-to-shared` vocabulary, and the smallest-common-owner routing survive verbatim in the Core Rules owner-generalization bullet, with `then add stack-specific implementation notes only where needed` kept in the step; the capability-name examples, the never-name-after-source clause, and the do-not-route-users-through-a-source-name clause all survive in the step, and provenance-labelling survives in the Core Rules naming and provenance bullets; the Step 6 residual-search list gains the page-name and scenario-label terms the two merged lines carried separately, and keeps `absent from executable guidance or clearly marked as provenance`. Net effect on the frozen entrypoint: base_body_words=16759 head_body_words=16750 (-9), so the round funded its own additions and left the entrypoint smaller than it found it. |
|
|
494
494
|
| A base-relative gate's design-time premise is measured against EVERY base its landing faces resolve, never the current round's base alone: the set is read off CI's own base-resolution expression rather than guessed - one face per pull-request target branch plus the pushed branch's previous tip on a push build - and each resolution is a separate run of the author-dogfood leg, a difference that accumulated before the gate existed surfaces as one violation on the face nobody measured, and the repair is to shrink the frozen surface rather than add a cross-base exemption | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#enumerate the landing faces, never assume one | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key; the changed file is `references/dual-track-review-gate.md`. **Observed failure (RED, recorded incident + re-computable).** The reference line ratchet landed in the prior round measured only against the round's own base. Reproduction on the pre-fix candidate: `CCL_SKILL_BASE_REF=origin/main bash skills/skill-extraction-workflow/scripts/check-size-budget.sh .` printed `reference_line_block: ... dual-track-review-gate.md: over-limit reference grew base_lines=662 head_lines=668` and `reference_line_budget_blocking_failed`, while the same command with `CCL_SKILL_BASE_REF=origin/dev` printed `reference_line_budget_blocking_ok`. The gate behaved exactly as documented; the defect was that leg (a) said `the SAME base resolution CI uses` (singular), so the author measured one face and the six lines that older rounds had added to a frozen surface only became visible on the promotion face. **Compliance (with-change).** Same two commands on the head candidate: base=origin/main `base_lines=662 head_lines=661 allowed_lines=662` and base=origin/dev `base_lines=668 head_lines=661 allowed_lines=668`, both ending `reference_line_budget_blocking_ok`. Running BOTH faces is this round's own dogfood of the rule it lands. **Zero-loss obligation map for the five consolidations that funded the addition** (all within the same file per the frozen-reference funding rule; 668 to 660 lines, and 125978 to 126158 bytes - the line unit the ratchet measures fell, the byte count rose by the amount the new obligation text exceeds the recovered duplication, which is stated here rather than hidden). (1) The four-line raw-CLI preamble collapses to one line carrying all three of its propositions - diagnostics only, never review or challenge evidence, never a replacement for the owner wrapper on a non-wording lane. (2) The standalone do-not-iterate-to-zero-findings paragraph merges into the convergence-bar paragraph, keeping the design-tradeoff clause, the pre-existing clause, the over-correct-or-scope-creep clause, and both contrasts - the bar differs from zero findings AND from no-new-P0-P1. (3) The R0-evidence value menu was stated verbatim three times; two occurrences become pointers to the review-pass row that keeps the menu, matching the pointer form the adjacent Item-9 field already used. (4) The Rules bullet restating the behavioral-evidence table's own two rows is dropped; its one clause absent from the table, that an author cannot self-assert semantic-control, moves into the semantic-control cell itself. (5) The trailing what-a-fallback-is-worth paragraph folds into the reviewer-ladder item that already owns that question, keeping the ad-hoc-run bar, the remediation-versus-evidence distinction, and the interim rule. No obligation was dropped and no paragraph was re-wrapped to buy lines. **A sixth consolidation was attempted and reverted, which is the reusable finding.** The premise-verification bullet in the self-audit section reads as a verbatim restatement of the two paragraphs above it, and deleting it looked free; `register_firing_path_unresolved` then failed because an earlier round's register row anchors its firing path on that exact line. A rule line can be another row's evidence, so an append-only ledger makes some prose non-deletable: check the anchor set before treating any rule line as redundant, and restore rather than EXEMPT when the deletion was to fund your own budget. **Other owners.** `references/attention-budget-ratchet.md` is `unchanged: already-covered` with a real firing path - its five-invariant preamble already says the budget gates are `the budget-gate instantiation of the design-time operability check in dual-track-review-gate.md - run that check's four legs too`, so the tightened leg (a) reaches ratchet authors through that pointer and restating it there would be the same-facet drift the Step 6 check forbids. `scripts/check-size-budget.sh` is `not-applicable`: the gate is correct as shipped and judges whatever base it is handed, so multi-base topology belongs to the caller, not the script. **Residual risk, stated rather than hidden.** The firing path is a walked design-time enumeration in prose, not a mechanical multi-base run; `Makefile` still defaults `CCL_SKILL_DEFAULT_BASE_REF` to the integration branch, so a local check still measures one face unless the author enumerates. A mechanical all-faces target is deferred with an owner - it would be a new gate surface owing its own four legs, oracle self-proof and suite registration, which is a round of its own rather than a rider on this one. |
|
|
495
495
|
| A design-time obligation over a SET is discharged by a manifest a reviewer can diff against the set's authoritative source, never by an asserted walk: the manifest names one entry per member with the value that member resolved to and the check's verdict on it, and a set with no finite manifest is a blocking residual routed to the existing non-blocking or risk-owner-deferral exits rather than reported as coverage | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#recorded manifest, never by a walk you assert | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key; this is the post-review fix batch of the same round, kept as its own commit so the externally reviewed candidate stays identifiable on the branch. **Observed failure (RED).** The first formulation of the preceding row's rule said to run the design-time leg `once per resolution CI can produce`. Both lanes of the dual-track chain independently reached the same defect on the frozen candidate: the review lane found the required set is not statically enumerable because an unrestricted pull-request trigger can resolve any target branch, so an author must either guess a subset and falsely claim completion or cannot satisfy the rule; the challenge lane found the same accumulated cross-base violation can therefore still ship despite apparent compliance, since the unchanged per-base gate cannot detect the omitted face. Two independent lanes converging on self-certifiability is the self-adjudication shape this file already names - the classification verb had no named output that produces the classification. **Fix.** The obligation now takes a manifest derived from the workflow itself, one entry per landing branch carrying the resolved base ref and the gate verdict, which a reviewer can diff against the workflow; and an unenumerable set is routed to the exits the premise-check leg already defines instead of being claimed as covered. No new mechanism is introduced - the fix converts an author-adjudicated claim into a reviewer-checkable artifact using exits that already exist. **Chain state.** Wrapper-fixed budget of one review plus one challenge, both bound to the same frozen candidate digest `70fd34678f99c444bf5b7c808a283380e99201eada6d71d076fd08e2f2ec1789`; fixes were held un-applied across both rounds. The final round itself returned findings, so the honest terminal state is `continuation_authorization_required` and this batch is post-review, not reviewed. A third review finding asked for the workflow file and captured command outputs inside the packet; it is dispositioned per this file's reviewer-verification-scope rule - deterministic-gate claims are verified by CI re-running the gates on the branch, never accepted from the implementer's prose - with the packet-composition miss recorded as a process defect for the next round rather than re-litigated here. **Supersedes two cells of the preceding row.** (i) Its compliance figures read `head_lines=661`, drafted before a later edit in the same round took the file to 660; the correct values are base=origin/main `base_lines=662 head_lines=660 allowed_lines=662` and base=origin/dev `base_lines=668 head_lines=660 allowed_lines=668`, both `reference_line_budget_blocking_ok`. (ii) Its rule cell describes the face set as read off CI's base-resolution expression; the manifest form in this row is the one that governs. The correction is recorded here rather than by editing that row: the ledger is append-only and the gate enforces it mechanically - a row edited after its own round committed no longer survives at HEAD, its round loses its only row, and the gate fails closed. That is the append-only contract catching an in-place fix, which is what supersede-by-pointer exists for. |
|
|
496
|
+
| A published version is immutable, so the source tree's version pointer must never sit below the highest already-released version; the invariant is checked at merge time against two independent records the repository owns - the release-tag set and the version the merge target already declares - because the tag record is deletable and no depth of fetch recovers a tag removed from the remote, and because the commit that lowers the pointer is typically an unrelated change whose rebase resolved a both-sides-changed hunk backwards | `release-coordination` | Enforced by `scripts/check-release-version.py`, wired into the `test-repo-gates` leaf CI runs. This row carries no impact-chain declaration: the rule's owner is `release-coordination`, which is outside the curated upstream-owner set, and the mechanism has no surface inside this register's owner package | updated | Observed miss: an extraction commit about review budgets rewrote all three version sites from 0.9.0 to 0.8.0 while 0.9.0 was already published and immutable; a human restored it later. Paired RED on the replayed tree: the four pre-existing PR-time repo gates each exit 0 with clean tokens, the new gate exits 1 naming the declared version, the released version, its tag, and all three sites to repair. Six applied mutations of the gate are each killed differentially by their own leg (numeric ordering, tagless-is-unevaluated, the below-floor comparison, the nested lockfile site, the base floor, the unreadable-file catch); the control tree is clean at 18 legs. The base floor was added post-chain after both reviewer lanes independently attacked the tag record's deletability; it makes the gate partly base-relative, so its design-time landing faces are recorded rather than asserted: pull_request into dev (base floor 0.10.0), pull_request into main (0.10.0), push on dev previous tip (0.10.0), push on main previous tip (0.9.0), and an unresolvable base (floor absent, tag floor alone) - all five exit 0 on the candidate, and the verdict is stable across them because the tag floor dominates. Residual recorded rather than closed: a version published without a tag, or the loss of both records, lowers the floor with them; closing that would put a registry query inside a merge-time gate. `skills/release-coordination/references/tag-and-prod-pipeline-gate.md` carries the rule beside the existing do-not-force-move-a-published-tag paragraph |
|
|
497
|
+
| A repository control that ENCODES a reading of an external contract carries that reading's expiry beside itself — the primary-source clause, the date and the tool/host version it was verified against, and the condition that invalidates it — because without them the next reader cannot separate a deliberate decision from a fossil, and re-verification costs the whole external read again | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; result-class: failure; firing-path: file:skills/skill-extraction-workflow/references/external-practice-controls.md#must be treated as carrying an unverified reading | updated | RED-baseline of the recorded-incident kind, observed in this round against the pre-rule state: the packed-artifact verifier carried a machine-enforced ban on a `version` key in plugin.json with no source, date or expiry, so reading the control could not distinguish a decision from a fossil - and applying the new rule to it is what surfaced that the ban read only one of the two plugin manifests, leaving the load-bearing one unguarded across every release shipped so far. The paired observation is that the same control had passed the full gate suite and both reviewer lanes in earlier rounds without that gap being visible. Merged into `external-practice-controls.md`'s inherited-reading section as its record half; the predicate half (own the predicate, not the upstream's vocabulary) already lived there. Scoped narrowly to external-contract claims: attribution's who-said-it tier stays with `attribution-verification.md`. Two live instances found while reading, both annotated: the packed-artifact verifier's ban on a `version` key in plugin.json, and the new release-version gate's dependence on `actions/checkout` fetch-depth. Both reviewer lanes then caught the first annotation failing the rule it illustrates - it carried a date but no host version - which is the rule's own dogfood arriving as a finding; both annotations now pin the verified host or action major and state which upstream release re-opens them. Annotating the first surfaced that the ban read only the Codex manifest, leaving the Claude one — the manifest whose version actually pins host updates — unguarded; the ban now covers both. The owning entrypoint `skills/skill-extraction-workflow/SKILL.md` is deliberately `unchanged`: both rules merge into references its Core Rules already route to - the inherited-reading section for the record half, and the form-by-failure table its rule-consolidation pointer already owns - so restating either at the entrypoint would be the same-facet drift the Step 6 cross-section check forbids. Also landed: two rows in `rule-consolidation.md`'s form-by-failure table (plan-validate-execute for batch/destructive steps; freedom tier matched to operation fragility), both verified against the Anthropic skill-authoring best-practices page on 2026-09-01 |
|
|
498
|
+
| A gate whose verdict is computed per OWNER but caused by one ROW must name the offending row, not only the owner: an owner-scoped diagnostic sends the author to inspect the row they just wrote, which is usually the correct one, while an unrelated row that the round merely EDITED is the actual offender — editing an already-committed ledger row makes it an added row of this round, so its own original anchor stops resolving | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; result-class: failure; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | updated | Observed in the preceding round: a correct new row failed the owner because an unbounded string replace had also modified a pre-existing ledger row; three iterations rewrote the good anchor before debug output patched into the gate found the real offender, and the round then recorded a non-existent encoding defect as its root cause. The diagnostic now lists each failing row's first cell. SUPERSEDES that round's claim that a non-ASCII firing-path anchor can never validate: `locator_valid` resolves blobs through `blob_at`, which reads without forcing an encoding, while only the separate `raw_blob_at` forces binary. Paired fixtures holding every other variable fixed — uniqueness, the 16-character bar, list-line shape, whitelisted verb — show a Chinese anchor validating normally, and that positive control is now a suite leg (`case-ref-non-ascii-anchor-accepted`). Four applied mutations kill differentially: reverting the diagnostic to owner-only kills the offending-row leg; forcing `blob_at` to binary — the implementation the previous round wrongly believed existed — kills the Han pair; dropping the control-character scrub kills the control-byte leg; and special-casing the one Han literal kills the NFC pair. That last mutation is why the acceptance proof is a differential TABLE rather than one passing literal: the challenge lane observed that a single-literal leg is satisfied by a special-cased implementation, so each non-ASCII anchor is paired with an ASCII control matched on everything the gate predicates on — uniqueness, the 16-character bar, list-line shape, a whitelisted normative verb — and the two verdicts must agree across Han, precomposed NFC, and decomposed NFD. The diagnostic scrubs C0/C1 controls, bidi overrides and zero-width characters before truncating, because it renders contributor-controlled ledger text onto a terminal: both reviewer lanes independently found that raw ANSI or carriage-return bytes could erase or forge the surrounding diagnostic. Suite at 88 cases; its own case-count guard caught every addition rather than absorbing them silently. `skills/skill-extraction-workflow/SKILL.md` unchanged because the anchor contract it states was already correct |
|
|
499
|
+
| A landed conclusion — a claim that a capability is unavailable or impossible, or a DIAGNOSIS of why an observed failure happened — is a hypothesis until an operation that could have falsified it has been run; the two admissible operations are the suspected mechanism exercised on the path that actually ran with the reaching call site named, and a paired control differing in exactly ONE variable with every other precondition of the tested predicate enumerated and confirmed equal, and an unfalsified cause is labelled `hypothesis` and kept off the shared surfaces | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#until an operation that could have falsified it; revalidate-when: a long-session probe puts the tempting hit mid-task rather than in a one-shot answer, so the with-change half of this row can be recorded instead of left inconclusive | updated | `skill-extraction-workflow/SKILL.md` is the owner key; supporting changes in `skill-extraction-workflow/references/validation-and-landing.md` (method detail) and this register. Observed failure (RED) in the preceding round: a gate rejected a ledger anchor, a `force_encoding(BINARY)` call found by grep in the gate script was recorded as the cause, and that cause landed in a merged pull-request body, a durable note and the round charter; the suspected line was never on the checked path and the real cause was an unbounded string replace in the round's own edit. The pre-rule tree obligated a falsifying attempt only for unavailability/impossibility conclusions and for done/covered/converged claims, so a causal diagnosis fell between the two and rode through a full dual-track gate. Applied differential control in THIS round, scoped honestly: the new `defect-diagnosis` anchor line was first written with "does not enter", `impact-chain-gate.rb` rejected that owner because "does not" is excluded from its normative vocabulary, and changing that single token to "must not" with the rest of the line, the row and the diff held constant turned the same command green. That is the executed-path proof of this round's own reading of that predicate and evidence that the anchor gate is sensitive - it is NOT evidence that the rule moves agent behavior, and the independent review was right to say so. The RED half of this row is therefore the recorded incident alone; the with-change half is NOT established, and the row is honest about that rather than resting on the mutation. BEHAVIORAL ATTRIBUTION NOT ESTABLISHED - read this before treating the row as efficacy evidence: both halves of the baseline are executed and recorded, but the paired control complied too, so no delta is demonstrated. Detail: with the changed text as the operating layer an agent kept the same tempting cause hypothesis-grade, named the competing causes and asked for the discriminating observation - but the paired control carrying the base text of the same passages did the same, so no delta was demonstrated. Probe detail: two headless arms differing only in whether the changed or the base text of these passages was the operating layer, scored against a rubric frozen before either run, three independently designed scenarios were run, and the control passed all three: a neutral question, one pressing for a confident cause paragraph to paste into a pull request, and - the shape that targets this rule most directly - a synthetic repository with a known-by-construction ground truth, where a grep-visible force_encoding sits on a dead path while the real cause is a uniqueness predicate, run once as an investigation task and once as a closeout task whose stated cause was inherited, false, and already followed by a green pipeline. In the closeout arm the control verified the inherited cause and corrected it unprompted. The first probe pair additionally ran with host auto-memory and repository instructions enabled and was contaminated - the control cited this repository own earlier round - so only the reruns with both disabled count; the synthetic-repository fixture and its frozen rubric are reproducible by a third party. The probe is therefore recorded inconclusive rather than as a demonstrated behavioral delta: a bounded task does not reproduce the context load the observed failure happened under, and the honest reading of three passing controls is that the base text already yields the wanted behavior on tasks this size - the rule earns its place on the recorded incident and on two independent review lanes endorsing its wording, not on a measured delta. The standing RED is the recorded incident above; what would settle the with-change half is a long-session probe where the tempting hit appears mid-task. Landed by consolidation rather than addition: the entrypoint's body-word count falls from 16750 to 16749 with the merged rule in place, the collapsed passages being ones a cited reference already carries verbatim |
|
|
500
|
+
| A hypothesised cause is stated together with the observation that would FALSIFY it, and that observation is collected first: a search hit, a log line or a plausible implementation detail proves the text exists, not that it ran on the failing path, so the diagnosis must name the reaching call site or build a one-variable paired control before the cause may become a fix, a commit or merge-request body, a durable note, or a report | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/SKILL.md#the observation that would falsify it | updated | `defect-diagnosis/SKILL.md` is the owner key. Downstream executable owner of the preceding row's upstream rule. Its hypothesise step previously asked only for the observation expected IF THE HYPOTHESIS IS TRUE, which the observed failure satisfied exactly - the grep hit was that expected observation - so the step licensed the confirmationist probe rather than blocking it. The upstream rule governs what may enter a landing; this row governs the diagnostic step itself, which is where the executed-path form is actually applied. Same RED as the row above, plus the paired-control half: three probes built in the preceding round to test the wrong cause each differed in more than one precondition of the predicate under test - anchor length against the threshold, uniqueness within the file, vocabulary membership - so no verdict among them was attributable |
|
|
501
|
+
| A stack dev owner that is absent from the lifecycle coordinator's dispatch enumeration is unreachable from every multi-stage delivery in that stack, and the coordinator's own no-owner fallback clause then silently reverses the ownership a previous round established on the executor side | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#merely because it is a command; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-multistage-to-product-rd | `updated` | Owner key `product-rd-workflow/SKILL.md`. RED (measured on origin/dev): `grep -rc nodejs-service-dev skills/product-rd-workflow/` returned 0 across the whole package while the Development-routing line named the Go and Python dev owners and then fell back with "a CLI in a language with no such dev owner stays with `terminal-cli-dev`" — so the Node CLI ownership landed by rows 371/372 was reversed by this clause. Candidate adds the Node owner to the routing enumeration and to the CLI carve-out list, adds a stack-owner map to the product-rd dispatch-owner-skills reference, fills the Node architecture vacuum in the delivery-lifecycle reference and the problem-resolution-and-learning reference per the maintainer's decision not to create a Node architecture sibling, and funds every addition inside the severe-debt entrypoint by in-file consolidation (head_bytes= 76200 vs base 76200). Correction to an earlier draft of this round: that draft moved the two CLI carve-out sentences into the reference, but the routing pointer-integrity suite pins both branches on the entrypoint, so both were restored verbatim and the budget repaid elsewhere in the same line. Independent review then found the round's routing-bank case does NOT guard the coordinator omission — a multi-stage Node request selects the coordinator from its own wording and stays green with the Node owner removed from the dispatch map — so that fixture is advisory only. The actual probe is three checks pinned over the stack enumeration, the CLI carve-out list, and the design-gate architecture destination; each was verified by an APPLIED mutation that turned exactly its own check RED with no non-owning assertion failing, against a green unmutated control. |
|
|
502
|
+
| A stack implementation owner with no `*-architecture` sibling must name the owner that absorbs its architecture decisions, or the stack is reachable for implementation and unreachable for design | `nodejs-service-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/nodejs-service-dev/SKILL.md#Do not borrow the Go or Python architecture skill | `updated` | Owner key `nodejs-service-dev/SKILL.md`. RED (measured on origin/dev): the package named neither sibling stack owner and carried no generalization section (0 hits for `go-microservice-dev`/`python-service-dev`/`Generalization`), and `AsyncLocalStorage` resolved only inside the maintainer source map — a file whose own preamble says it is "not required reading for ordinary Node.js implementation work", so the extracted constraint had no executable landing. Candidate adds the architecture-owner statement, a Bun/Deno runtime boundary, a Generalization Discipline section carrying the inward and outward sibling duty, the request-context rule (primary source: Node async_context docs, which recommend the built-in store over custom `async_hooks`), and an outbound-HTTP-client section grounded in undici's own docs (global dispatcher backs built-in `fetch`; per-origin pools default to unlimited connections; connect/headers/body timeouts are layered with no overall deadline; an unconsumed response body holds its pooled connection). Version-sensitive numbers are deliberately not hardcoded. Independent review round 2 raised a P1 on the first draft of the context rule: prescribing a declared default for an absent store is unsafe when the stored value is security-bearing, because a tenant, subject, or permission scope that defaults executes the operation under the wrong identity. The rule now splits by what the value authorizes — diagnostic values may default, security-bearing ones fail closed — and states that the store is a propagation mechanism, never the authorization decision. This was the draft-time security axis the pre-cover sweep should have caught before handoff, not the reviewer. |
|
|
503
|
+
| A defect-routing table that enumerates stack owners must carry every stack that has an owner, and must say where a stack's architecture defects go when that stack has no architecture sibling | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/SKILL.md#must not be filed under a sibling stack's architecture skill | `updated` | Owner key `defect-diagnosis/SKILL.md`. RED (measured on origin/dev): 0 hits for `nodejs-service-dev` in the package while the prevention-routing table listed Go and Python implementation and architecture rows, so a Node defect had no durable landing place. Same-class member of the sweep this round performed after the set-difference predicate surfaced four owners with the identical shape. |
|
|
504
|
+
| A test-layer owner that hands implementation mechanics to stack skills must name every stack owner, or the layer choice lands with no executor for that stack | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#packaged-artifact verification must use | `updated` | Owner key `testing-strategy/SKILL.md`. RED (measured on origin/dev): 0 hits for `nodejs-service-dev` while the post-layer-choice pointers named Go, Python, app, miniapp, web, terminal, and LLM owners. The addition is funded inside this severe-debt entrypoint by consolidating the phrase "after the test layer is chosen", which was repeated on all 8 stack-pointer lines, into one statement on the scope line: net -2 bytes vs base with the repeats removed. |
|
|
505
|
+
| A skill that hands a host stack back to its owner must cover every stack that has one, including the disposition of that stack's architecture decisions | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/SKILL.md#never to a Node architecture sibling | `updated` | Owner key `llm-inference-integration/SKILL.md`. RED (measured on origin/dev): 0 hits for `nodejs-service-dev` while the hand-back clause named the Go and Python architecture and dev owners on identical terms. Same-class member of this round's sweep. |
|
|
506
|
+
| A test-code implementation pointer list is a routing obligation, not a convenience enumeration: a stack missing from it has approved test cases and no owner to implement them | `test-artifact-management` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/test-artifact-management/SKILL.md#must go to the owning stack skill | `updated` | Owner key `test-artifact-management/SKILL.md`. RED (measured on origin/dev): 0 hits for `nodejs-service-dev` in the six-owner implementation pointer. Same-class member of this round's sweep; the line is also restated as a normative routing obligation rather than a suggestion. |
|
|
507
|
+
| Member-join obligation inheritance enumerated from a remembered list of obligation types reproduces the miss: the list that shipped named four routing-surface forms and omitted the coordinator's dispatch enumeration and the shared body section, which are exactly what went missing on the next join — so the predicate must be a mechanical set difference over where an established sibling is named | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#do not enumerate the inherited obligations from a remembered list | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failure: row 371 ran the join-side sweep for `nodejs-service-dev` and dispositioned the remaining class obligations as unchanged/not-applicable, recording `not-applicable: nodejs has no *-architecture sibling` — which restates a routing vacuum as a reason. One round later the coordinator enumeration, the sibling generalization loop, and four peer routing tables were all still missing the member. RED baseline (measured, both directions): on origin/dev the new predicate `comm -23 <(grep -rl python-service-dev …) <(grep -rl nodejs-service-dev …)` emitted every one of the five files this round had to change (the lifecycle router entrypoint plus its delivery-lifecycle and problem-resolution references, and both sibling stack dev entrypoints); on the candidate all five are absent from the difference, and the 30 remaining paths are candidates requiring disposition, four of which were confirmed real and swept this round. The predicate replaces the type list rather than extending it, per the same-class-recurrence rule that a proxy predicate be re-expressed over an invariant the control owns. The reference stays level at 691 lines by folding the two new clauses into the existing bullet. The same owner's routing pointer-integrity suite gains four coordinator-dispatch assertions and a line-scoped assertion helper; its self-pin count moves 21 -> 24, so silently deleting one is itself caught. (Two earlier drafts of this row recorded 25 and then 23; each described a candidate that a later round superseded, and the number is corrected here rather than left to describe a tree that no longer exists.) The adversarial challenge rejected the first form of these checks: whole-file substring greps pass when the Node tokens are moved into unrelated prose while being deleted from the line that carries the obligation, and the architecture check passed for a sentence that declared the vacuum without naming any destination. The assertions are now scoped to the carrying line and require the complete relationship (both owner enumerations on the routing line; both the platform owners and the Node implementation owner on the design-gate line). Verified by applied mutation: moving the token out of the routing line into stray prose leaves the OLD whole-file form green (the hole, reproduced) and turns exactly the new line-scoped assertion RED; stripping the destination from the architecture sentence turns exactly its two assertions RED; the unmutated control and the restored tree are green, with no non-owning assertion failing in either mutant. A later review round on this same candidate reported that the diff DELETED the sibling owner's causal-falsification rule. The report's direction was right and its attribution was not: the rule had landed on the target branch after this round's worktree was cut, so `git diff <target>` rendered "the branch is behind" as "the branch deleted it" — the stale-branch shape, on a real collision set (both rounds changed the same two files). Resolution followed the update-before-merge recipe rather than a text edit: pin the target, compute the pre-update merge-base, list the collision set, rebase (the branch was never pushed, so no shared history was rewritten), and resolve the ledger conflict by keeping BOTH sides' rows in append-only order rather than taking one side. Content-layer verification then confirmed the sibling round's rule and both of its ledger rows survive on the rebased candidate, this round's seven rows and its routing rule are intact, and all eighteen deleted lines are this round's own intended rewrites. Recorded because `--stat` cannot see an in-file revert, and because a finding whose ATTRIBUTION is wrong can still name a real defect: the fix belonged to the baseline, not to the prose the finding pointed at. The landing round then produced two findings of ONE shape — an assertion weaker than the contract it states — and they were swept as a class rather than patched individually: the per-requirement form asserted each token against the whole set of selector-matching lines, so two requirements could be satisfied on two different lines while the contract said one line must carry both, and neither architecture assertion pinned the positive ownership clause, so deleting it while keeping both owner tokens left the architecture decision ownerless. The four assertions collapse into two that filter the matched lines through every requirement in turn, with the ownership clause pinned alongside the two owners. Applied-mutation evidence, both legs measured: deleting the ownership clause while keeping both owner tokens leaves the pre-fix form green and turns the new assertion RED; splitting the requirements across two selector-matching lines does the same. The first control leg for the ownership mutation was invalid — it included the newly added requirement in the supposedly OLD form, so both arms went red and proved nothing — and was rerun with the actual pre-fix requirement set before the result was recorded. The landing round also closed a conflict the review chain surfaced and an earlier draft of this round had deferred: `testing-strategy` claimed `terminal-cli-dev` for "command-line ... implementation", which collides with the CLI carve-out that the coordinator and every stack dev owner carry, so a Node, Go, or Python CLI test could route to a skill that does not own the implementation. Measured as pre-existing on the target branch — the claim and both sibling stack pointers are older than this round — it was resolved by DELETING the over-broad claim rather than adding a fourth copy of the hand-off rule: the line now scopes to the terminal interface contract, PTY/ANSI/keyboard rendering, and full-screen TUI testing, which costs 13 bytes less inside a severe-debt entrypoint and leaves the hand-off stated once, where it is already pinned. A same-line assertion pins the narrowed scope, verified by an applied mutation that restores the old wording and turns exactly that assertion RED. Separately, `frozen_at_sha: root` was raised in three rounds and stays refuted on measurement rather than on argument: 143 of the 158 existing routing cases carry that exact value, so the new cases follow the bank's dominant convention, not a deviation from it. |
|
|
508
|
+
| Supersedes by pointer the two budget rows above where they state the extraction lane is two rounds and that a restart is never fundable: the lane is one review plus one challenge on the frozen candidate PLUS one succession challenge owed exactly when the held fix batch moved the candidate, and the trigger is the candidate hash rather than any disposition label the author writes | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in references/dual-track-review-gate.md, references/extraction-quickstart.md, scripts/validate_extraction_review_state.py, scripts/review_ledger_binding.py, and both fast-lane suites). Observed failure, found by reading the gate against itself: the anti-pattern clause demanded "always do at least one re-challenge after a non-trivial fix-up" while the cadence plus the hold-fixes rule made that round unreachable under the two-round budget — an obligation no reader could satisfy, which reads as satisfied. Underneath it, the closeout validator required every counted round to bind the ledger candidate, so once the held batch was applied the ledger could not close on what lands at all: the batch left the gate's jurisdiction and the only remaining lane was human MR review, which this same document disqualifies as dual-track. Predicate choice recorded because the obvious one is wrong: "a P0/P1 was dispositioned fixed" is a self-adjudication clause of exactly the shape this reference already forbids, so the trigger is the candidate hash the control owns — changing a label cannot fake it, changing bytes is the thing being detected, and a round whose findings were all accepted or pre-existing pays nothing. Wrapper unchanged (22 lines, no allowlist, so the new flag transits; it is stateless and cannot count chains). Per-chain `challenge_budget` stays 1, so all 14 historical receipts in specs/*/evidence/review remain valid byte-for-byte; the lane bound is the new constant pair. RED-baseline (applied, differential): with the validator reverted to base, the succession-ready fixture reds at "controller receipt 1 does not bind the ledger candidate" — the exact invariant this round re-phases — and the restored validator greens; the pre-existing over-budget case stayed red for the right reason after the bound moved, which is what caught a non-succession third receipt slipping through, fixed by requiring any post-chain round to be a succession. |
|
|
509
|
+
| A caller-built ledger that nothing reads at merge time cannot enforce anything: the landing candidate is bound to the evidence that inspected it by a merge-side gate that recomputes the packet with the controller's own freeze, and evidence is kept outside the reviewed paths so committing the ledger cannot move the hash the ledger records | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md` (contract sentence updated net-negative under the entrypoint word ratchet; mechanics land in references/staged-review-contract.md and scripts/review_gate.py; the merge-side half and its CI wiring belong to the extraction lane's row above). Observed failure, measured not assumed: grep for extraction_review_state / closeout / receipt / review_chain across .github/workflows/ and check-ccl-skills.sh returned zero references, so the whole dual-track ledger was agent-run and agent-supplied with no merge-time reader — adding a third round alone would have changed nothing mechanically. Contract change: a fix touching a selected owner's SKILL.md or references/**.md moves selected_skills_sha256 and ends the chain by construction (verified in review_gate.py, not inferred: a changed candidate alone only adds a material_candidate_change self-review trigger and does NOT end the chain, and the owner digest excludes scripts/, so script-only fixes continue in-chain), therefore one succeeding chain may open at index 1 in challenge mode against a moved candidate. The owner digest is the single binding allowed to move — controller digest, owner-selection source, owner names, scope digest, and stage/depth/risk-tags/budget must all still match, an unmoved candidate is rejected as a repeat round, and prior focuses carry forward. RED-baseline (applied, differential): six new controller cases red against the unchanged controller and green after, with 257 pre-existing cases unaffected; the rejection cases were re-anchored after they were caught passing for the wrong reason — the old controller emits the same review_chain_invalid code for these arguments, so reason_code alone could not tell the mechanism from its absence and each now pins the succession diagnostic. The merge-side consumer of this contract carries its own suite in the extraction lane. |
|
|
510
|
+
| Supersedes by pointer the merge-side row above on two counts its own review found: the bound paths include the workflow directory, because with only the skill tree bound the CI step that runs the gate could be deleted without moving the candidate; and the gate has no evidence branch keyed on a self-declared field, because this gate cannot authenticate that a controller minted any file it reads | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, scripts/check-ccl-skills.sh, scripts/test_review_ledger_binding.sh, references/dual-track-review-gate.md). Observed failure: this change's own review lane, rounds 1 and 2 against the frozen candidate, returned seven findings; four are dispositioned here. (a) A run with no base printed an unevaluated notice and exited 0, so a base-wiring mistake would have read as a passing required check, and the diagnostic named an environment variable the code never read — now fail-closed by default, with the variable actually read and an explicit --allow-unevaluated for events that genuinely have no base. (b) The wording-only branch accepted any file carrying the candidate hash and a non-empty proof field: hand-writable, so deleted rather than patched, and the cost is recorded — a wording-only change now owes the same evidence. (c) Nothing invoked the gate against a real checkout; a smoke now runs from the repository gate suite, deliberately dropping an inherited CCL_SKILL_BASE_REF because that checker runs against synthetic clones inside other suites where the leaked base resolves against the wrong repository — a hazard this repository has recorded before and which reproduced here on the first wiring attempt. (d) A stale sentence still told broad extractions to stop at two rounds. RED-baseline (applied, differential): each case reds the suite on its own assertion before the fix — the no-base case at the exit status, the receipt-shaped file at acceptance, the real-checkout smoke at packet freeze — and the fourteen-case suite is green after. |
|
|
511
|
+
| Supersedes by pointer the succession row above: succession is one-shot. A succession receipt may neither be continued inside its own chain nor become the next succession's predecessor, so the relaxation cannot be daisy-chained into unbounded autonomous rounds | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md` (entrypoint unchanged this round; lands in scripts/review_gate.py and its suite). Observed failure: the adversarial round of this change's own review lane found that a succession chain was minted at index 1 with a positive budget and nothing made it terminal — its receipt could enter the ordinary prior-result path for index 2, and that chain's terminal receipt could then be another succession's predecessor, composing without bound. The controller-side cap matters independently of the extraction lane's ledger bound, because other lanes consume the same controller. Both directions are now refused with their own diagnostics. RED-baseline (applied, differential): two new cases red against the pre-fix controller and green after, with the full suite at 259 passing and no pre-existing case disturbed. Bootstrap limit recorded rather than worked around: a round that edits the controller itself cannot use succession on its own predecessor, because the predecessor cannot preserve a controller digest the round just moved — observed live on this very candidate, which is the rule working, not a defect. |
|
|
512
|
+
| Supersedes by pointer the merge-side rows above with the residual those rows did not state: a gate that lives inside the candidate cannot authenticate itself, so the pinned CI step and the pinned fail-closed branch raise the bar without closing it, and the terminal authority is the platform's required-check configuration plus human review of the gate's own diff | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/extraction-quickstart.md#still owes the two-round chain before it can land | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/contract-anchors.tsv, scripts/review_ledger_binding.py, references/extraction-quickstart.md). Observed failure: the review and challenge rounds of this candidate's own second chain both found that a pull request may delete the workflow step or replace its command while keeping the required job name, and that the gate runs the candidate's own validator, so handcrafted receipts plus a validator that exits zero would pass. Neither is closable from inside the candidate. What lands is the honest pair: a contract anchor on the fail-closed branch so hollowing it also edits a registry another required check verifies — the workflow step is deliberately NOT pinned, because that registry addresses skill files and a cross-tree row reds the anchor checker against its own synthetic fixtures, which is how the attempt was caught — and a docstring that states what the gate proves — that what merges is the candidate an external round inspected — rather than implying it proves the round was honest. The wording-only path was also reconciled: the quickstart promised a single-review row with no ledger while the gate accepts only a validator-checked ledger, and the doc now records the cost rather than leaving the contradiction. RED-baseline (applied): the anchor gate reds on deletion of the pinned literal and is green at 10 anchors on the final candidate. |
|
|
513
|
+
| Supersedes by pointer the succession row above: inherited challenge focuses live in their own receipt field, because `prior_challenge_focuses` carries in-chain arithmetic that consumers derive from the round index, and a succession round arriving at index 1 with inherited focuses is a count the index cannot explain | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md` (entrypoint unchanged this round; lands in scripts/review_gate.py and its suite). Observed failure: the review round of this candidate's own second chain found that no test proved the completion checkpoint accepts the chain-index-1 challenge a succession produces. Adding that case turned it red immediately — the checkpoint rejected every succession round, because it computes the expected focus count as index minus two and the succession round carried its predecessor's focuses at index one. The ledger's terminal step was therefore unreachable for exactly the chain shape the third round exists to produce, and the defect would have surfaced only while closing a real ledger. Inherited focuses now populate `predecessor_challenge_focuses`; the distinctness rule still tests the union, so a succession cannot reuse a focus the ended chain already spent. RED-baseline (applied, differential): the new completion case red before the field split and green after, with the full suite at 260 passing and no pre-existing case disturbed. |
|
|
514
|
+
| The merge gate resolves its base to a commit id before that value reaches any git command, because an option-shaped base is read by git as an option and `git diff` with no revision compares the index to the working tree — which in a clean checkout reports no paths and passes the gate having compared nothing | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and its suite). Observed failure: the adversarial round of this candidate's third review chain constructed the input — `CCL_SKILL_BASE_REF=--quiet` — and traced it to a clean pass. The environment path is the one that matters because no argument parser stands in front of it, which is exactly how a base arrives in CI. RED-baseline (applied, differential): three unresolvable bases (`--quiet`, `--name-only`, a ref that does not exist) each red the suite before the fix and are refused with a named diagnostic after, while a resolvable revision expression still reaches the normal verdict. The same round's second case is also pinned: the suite now compares `git status --porcelain --ignored` across a gate run, so a regression in the bytecode or packet-cleanup discipline reds instead of passing a hash back. |
|
|
515
|
+
| Succession terminality is the predecessor receipt's own arithmetic — challenge index, remaining rounds, and the allowed flag — not its round index alone, because a forged receipt can carry a terminal index while every other field still says the chain has rounds left | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md` (entrypoint unchanged this round; lands in scripts/review_gate.py and its suite). Observed failure: the adversarial round of this candidate's third review chain showed that a tracked challenge receipt with budget one and index two, but `challenge_index=0` and a positive remaining count, satisfied the terminality test and could be succeeded — extending the lane past the rounds its chain had actually spent. RED-baseline (applied, differential): the forged-terminal predecessor case reds against the pre-fix controller and is refused after, with the suite at 261 passing and no pre-existing case disturbed. |
|
|
516
|
+
| The merge gate reads committed evidence only — it enumerates from HEAD and refuses a dirty evidence tree — because what merges is the committed tree, so a ledger read from the working tree can be evidence about something nobody can find after the merge | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, its suite, and .github/workflows/ci.yml). Observed failure: the fourth review chain of this candidate found that the gate globbed the working tree and passed whatever it found to the validator without proving any of it was committed, and separately that scoping the CI step to pull requests leaves a skipped step in a merge-queue run — where a skipped step does not fail its job. Both land: enumeration comes from `git ls-tree HEAD`, an uncommitted evidence tree is refused outright, and the step now also runs for `merge_group`. The residual is stated rather than closed: direct pushes to a protected branch are refused by the platform, not by this gate. Non-blocking item deferred with its reason: the controller returns the predecessor chain id without comparing it to the new chain id, so a self-succession is refused one layer up by the closeout validator rather than at mint time; fixing it would move the controller digest and forfeit this round's ability to close its own ledger with the succession round it introduces, so it is recorded here for the next round that touches the controller. RED-baseline (applied, differential): the uncommitted-evidence case reds before the fix and is refused with its own diagnostic after; the option-shaped-base and tree-perturbation cases from the previous batch stay green. |
|
|
517
|
+
| A gate that reports no-change must say so rather than surfacing an empty packet as a freeze error: the repository-gate smoke diffs against the parent commit, and the commit that lands the evidence touches no reviewed path, so the honest answer is `no-change`, not a broken gate | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/check-ccl-skills.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and scripts/check-ccl-skills.sh). Observed failure: self-inflicted and caught by this repository's own suites rather than by review — committing the closeout evidence made the parent-commit diff empty over the reviewed paths, the smoke asked for a candidate, the packet freeze refused an empty packet, and the checker exited before its later verdicts, reddening two unrelated suites (route drift and sync pointers) that only read those verdicts. The shape is worth keeping: a diagnostic that cannot distinguish "nothing to check" from "the check is broken" turns one benign state into a cascade. RED-baseline (applied): the smoke reds against the pre-fix gate on the evidence commit and is green after, with route-drift and sync-pointer suites recovering in the same run. |
|
|
518
|
+
| An assertion about a subject with two legitimate outputs must accept both, or it is an assertion about repository state rather than about the subject: the real-checkout case accepted only a candidate hash, so the commit that landed the evidence — which touches no reviewed path — turned the benign no-change answer into a red suite | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/test_review_ledger_binding.sh). Observed failure: the succession round of this candidate's own lane raised it as P2 and it fired within the same sitting — the ledger commit made the parent-commit diff empty over the reviewed paths, and the assertion reddened on an output the gate is designed to produce. The shape generalizes past this case: an oracle that admits one of its subject's several valid outputs measures the fixture, not the behaviour, and goes red on a change that is correct. RED-baseline (applied): the assertion reds on the ledger commit before the fix and is green after, with the gate's own behaviour unchanged in both runs. |
|
|
519
|
+
| A base-relative gate compares against the fork point, not the base branch's tip: measured against the tip, every unrelated merge on the target branch restates what this branch is and voids evidence that is still correct | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and its suite). Observed failure, live and not hypothetical: while this round's own pull request was open, the target branch gained an unrelated merge, and the gate immediately reported that no evidence bound the landing candidate — thirty files appeared changed, all of them somebody else's. Measured against the fork point the candidate was byte-identical to the one the ledger already bound, which is what proved the ledger right and the comparison point wrong. The operational consequence had this shipped unfixed is worth naming: a busy target branch would demand a fresh review lane every time anyone else merged, which is a treadmill rather than a gate. RED-baseline (applied, differential): the suite advances its synthetic base branch with an unrelated commit and asserts the candidate is unchanged; that case reds against the tip-relative gate and is green against the fork-point one, with the rest of the suite unaffected. |
|
|
520
|
+
| A probe that cannot tell "this checkout cannot be measured" from "the thing being measured is broken" is deleted, not patched again: the repository checker runs against synthetic fixtures inside other suites, and a smoke that reds there fails suites that have nothing to do with it | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/check-ccl-skills.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/check-ccl-skills.sh). Supersedes by pointer the row above that added this smoke. Observed failure, three times in one round and each time in a suite that does not own the gate: the catalog suite, then route-drift and sync-pointers together, then the source-register lifecycle suite — every one of them runs the checker against a synthetic repository where the smoke legitimately cannot operate, and the third failure additionally exposed that its capture was not `set -e` safe, so the checker died silently mid-run before printing the verdicts those suites read. Two patches had already narrowed the predicate (skip without a controller, skip without a parent commit) and a third would have narrowed it again, which is the signal this repository already records: when the same class recurs, question whether the capability should exist. It should not. The gate's own suite owns the real-checkout path with a case that runs it against the checkout it ships in, and the CI step is the enforcement point — verified green on this candidate's own pull request. What is lost is stated rather than glossed: nothing else runs the gate during `make test-repo-gates`, so a break in it surfaces at the CI step rather than locally. RED-baseline (applied): the lifecycle suite reds against the smoke-bearing checker and is green after its removal, with the gate's own eighteen-case suite unchanged in both runs. |
|
|
521
|
+
| The merge gate binds every tracked path minus exactly what this round ADDS under a round's evidence directory, and refuses a candidate tree that is not committed: a whitelist binds only the paths some round happened to review, so unreviewed executable content rode along on a valid ledger; a written-down `specs/` exclusion would additionally hide edits to the committed review history itself; and a packet frozen from a dirty working tree produces a hash no clean checkout recomputes | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, its suite, and .github/workflows/ci.yml). Observed failure: the succession round of the previous round's own lane raised the whitelist half and it was deferred with its reason -- every fix moves the candidate and voids the ledger, and that lane's budget was spent -- so it was recorded for the next round that touches this gate. This round's own review and challenge then each found that the first inversion traded one hole for another, and both land here rather than as further deferrals. The review found that the frozen packet includes untracked files, so a scratch file inside the widened set produces a candidate hash only that working copy can reproduce: the author records it in the ledger and the merge-side run then reports that nothing binds the landing candidate, refusing valid work. The challenge found that excluding all of `specs/` excludes the committed review history, so a pull request could delete or rewrite an earlier round's plan and receipts with no evidence required. The exclusion is therefore computed from the round's own diff rather than written down: only added paths under a round's evidence directory stay outside, because a receipt inside the bound set would move the hash it records, while every modification and deletion under `specs/` is bound like any other file. The succession round then broke that shape too -- an arbitrary added file under an evidence directory, a script included, was excluded for the same reason -- which is the third occurrence of one class: the rule kept naming a LOCATION and letting the location stand in for `this is a receipt`. Rather than narrow the path a fourth time, the predicate moved to an invariant this gate owns: a path is excluded only when its committed blob parses as a JSON object carrying a 64-hex `candidate_sha256`, so a script, a fixture, or an unbound JSON file committed there is bound like anything else. Backward compatibility was measured, not assumed: recomputed at the previous round's own fork point, the old and new path sets produce the identical candidate hash its committed ledger records. Supersedes by pointer the merge-queue half of the row above, which recorded that the CI step now also runs for `merge_group`: the workflow's `on:` never subscribed to that event, so in a merge-queue run the workflow would not start at all and the condition read as coverage while providing none. Restoring the trigger was rejected rather than done, because a merge_group HEAD combines several queued pull requests while each committed ledger binds one individual candidate, so no ledger binds the aggregate and every otherwise-valid queued request would be refused -- the trigger would make the sentence true and the system worse. The unreachable branch is removed and the real coverage boundary is stated where the step lives. RED-baseline (applied, differential, two mutants each attributed to its own cases and nothing else): restoring the whole-subtree exclusion reds exactly the three committed-history cases -- rewriting an earlier plan, rewriting an earlier receipt, deleting an earlier receipt; removing the committed-tree refusal reds exactly the three dirty-tree cases; degenerating the receipt predicate to always-true reds exactly the three smuggled-file cases; three unbound executable paths (a root Makefile, a README, a release script) each red against the original whitelist and are refused after; the added-evidence case stays green throughout, which is what proves the self-reference exclusion survived. Control is 34 passing with no case disturbed. A fourth round then broke the content predicate too -- a JSON file carrying any 64-hex `candidate_sha256` is accepted as a receipt -- and that one is NOT fixed, deliberately. Four shapes of this exclusion have now been broken in four rounds, and every one of them was a proxy for `this is a controller-generated receipt` over a file the candidate itself supplies, which is the already-recorded boundary that a gate living inside the candidate cannot authenticate what it reads. This repository's own standard is that a class recurring across rounds is a question about the design rather than a fifth patch, so the residual is recorded for a person: accept it, or replace the exclusion mechanism outright -- binding the tree as of the commit before the evidence lands would need no exclusion predicate at all. What did close is real: history can no longer be rewritten unnoticed, and a script or binary can no longer ride in under an evidence directory. |
|
|
522
|
+
| A succession may not carry the chain id of the chain it succeeds, and the controller refuses it at mint rather than leaving the refusal to the closeout validator: the validator only sees a lane it reads whole, while the controller mints one receipt at a time, so a caller that never closes a ledger never reaches that check | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md` (entrypoint unchanged this round; lands in scripts/review_gate.py and its suite). Observed failure: the previous round recorded this as a non-blocking deferral with its reason -- fixing it would have moved the controller digest and forfeited that round's ability to close its own ledger with the succession round it introduced. The severity recorded then is the one that holds now, and it is narrower than it first reads: this is not an open bypass, because `validate_extraction_review_state.py` already refuses a succession whose chain id equals the wrapper chain's. What lands is the same refusal at the point the receipt is made, which is the only place it applies to a controller run that never reaches a closeout. The equality direction is not inferred: the existing validator refusal uses the same predicate and the same words, so the intended semantics is that the two ids must differ. RED-baseline (applied, differential): a succession minted with its predecessor's own chain id reds against the pre-fix controller and is refused with its own diagnostic after, with the suite moving from 261 to 262 passing and no pre-existing case disturbed. |
|
|
523
|
+
| The merge gate binds a landing candidate larger than one review packet through a committed landing partition manifest: path partitions whose changed files together equal the candidate's exactly once, each recomputed with the controller's own freeze and bound by its own validated ledger, so the reviewer's byte ceiling is no longer the pull request's ceiling | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, its suite, references/dual-track-review-gate.md, references/extraction-quickstart.md, and a .github/workflows/ci.yml comment). Observed failure: the gate defined the landing candidate as one packet hash, and the controller caps a packet at what one reviewer can read whole, so a candidate larger than that could not be frozen and no ledger could ever bind it -- a release whose whole diff was three times the ceiling had to land as eight separate merges, each splitting the pull request where the review side already permitted splitting the packet. The two identities are different sizes: base..HEAD has no natural byte limit, a reviewer's input does. A manifest committed under a round's evidence directory names the partitions and, for each, the hash that `--print-candidate --paths` already answers; the gate refuses any manifest whose parts do not add up to the whole -- a changed file in no partition, a changed file in two, a partition whose recorded hash no longer reproduces, a base other than the fork point, an aggregate hash that does not reproduce its partitions, or a partition path shaped like a pathspec. The manifest carries a top-level 64-hex `candidate_sha256` (the aggregate identity) so it satisfies the existing receipt predicate and committing it moves no partition; the exclusion predicate is unchanged and the accepted caller-controlled-evidence residual is not widened. `--print-manifest --partition ...` renders the manifest with every hash computed by the gate, so the canonical form lives in one place. Merge-queue aggregation of several pull requests into one HEAD is a different aggregate and stays unsolved, as the workflow comment now states. RED-baseline (applied, differential): the partition cases red 18 against the pre-fix gate with the existing 35 undisturbed; five in-place mutants -- coverage equality, disjointness, aggregate recomputation, base equality, partition-hash recomputation -- each red exactly their own cases (2, 2, 1, 1, 1) and nothing else, restored and verified after each. The candidate that triggered the round, measured at 623,458 bytes, renders as six freezable partitions. |
|
|
524
|
+
| The partition manifest refuses a wildcard partition path and requires the partition union to EQUAL the reviewed changed set, not merely contain it: git reads `*`, `?`, `[` and `\\` as glob syntax even in a non-magic pathspec, so a wildcard partition chooses its own coverage, and under a narrowed `--paths` scope a partition can reach changed files outside the reviewed set with no uncovered file and no overlap to refuse | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and its suite). Observed failure: the round's own dual-track lane found both -- the independent review reported that `validate_partition_path` rejected only leading pathspec magic while `lane-*` passed through to git as a glob, and that `partition_coverage` checked only `changed_all - owner`, so with `--paths skills/a` a partition naming `.` covered changed files outside the scope and passed; the adversarial challenge independently hit the wildcard class on the same frozen candidate. Both are the same shape: the manifest was allowed to influence what git enumerated on its behalf. Fix: refuse the metacharacters as a path (never handed to git), and refuse a union larger than the reviewed set with the surplus named. Held until the challenge ran, then applied as one batch that moved the candidate, so the lane owes and runs one succession challenge bound to what lands. RED-baseline (applied, differential): the four wildcard shapes plus the rendering case red against the pre-fix gate and are refused after; the out-of-scope case reds against the pre-fix gate with a freeze error on the oversized `.` partition and is refused before freezing after; disabling the wildcard check reds exactly the five wildcard cases and disabling the equality check reds exactly the out-of-scope case, suite otherwise at 60 passing. |
|
|
@@ -255,9 +255,9 @@ Rules:
|
|
|
255
255
|
|
|
256
256
|
The SKILL.md class-wide COMPLETE-set rule (example of a class-wide change: "every stack `*-dev`/`*-architecture` should advertise a localized-refactor trigger"; example of a member-pair landing: Python+Go) fires on change-side sweeps. This section is its join-side dual.
|
|
257
257
|
|
|
258
|
-
- When a NEW skill joins an already-swept class (e.g. a new stack `*-dev` joining the stack-implementation-owner class), enumerate the
|
|
258
|
+
- When a NEW skill joins an already-swept class (e.g. a new stack `*-dev` joining the stack-implementation-owner class), **do not enumerate the inherited obligations from a remembered list of obligation types — derive them mechanically.** Take the file set where an established sibling is named and subtract the file set where the newcomer is named: `comm -23 <(grep -rl '<sibling>' skills/ eval/ | sort) <(grep -rl '<newcomer>' skills/ eval/ | sort)`. Every path in that difference is a candidate obligation site and gets a disposition — `inherit` or `not-applicable: <reason>` — before the member lands. Run it against each established sibling, not just one. **The set difference is the predicate; any type list is only illustration** — the forms it surfaces include advertised triggers, skip-leg reciprocity on routing counterparties, pinned pointer-integrity anchors, bank fixtures, the lifecycle **COORDINATOR's** dispatch enumerations (a router's development-routing list, a technical-design-gate owner table, a problem-routing table), and **body-level structural sections the siblings share** — notably a cross-sibling generalization clause, which is N-way: the newcomer must name every peer *and* every peer must name the newcomer, or the loop breaks in the direction nobody reads. A remembered type list is how the miss recurs: the list that shipped with this section named four routing-surface forms, and the two it omitted — the coordinator's dispatch enumeration and the shared body section — are the ones that went missing on the next join.
|
|
259
259
|
- A sibling-generalization map that only answers the copied-content question ("does the new skill duplicate sibling text?") must not be counted as obligation-inheritance coverage; the two questions are independent, and the join-side miss survives a clean copied-content map.
|
|
260
|
-
- Validation: the closeout map lists each
|
|
260
|
+
- Validation: the closeout map lists each path from the set difference with a status, exactly as the change-side sweep lists each member. Failure shapes: (a) a new stack implementation owner landed with "siblings unchanged because the new skill routes to their existing contracts" while lacking the CLI carve-out reciprocity and localized-refactor trigger pair its siblings carry; (b) the same owner, one round later, was still absent from the lifecycle coordinator's development-routing enumeration — so multi-stage deliveries in that stack reached the coordinator with no owner to dispatch to, and the coordinator's own fallback clause ("a CLI in a language with no dev owner stays with the default CLI owner") silently reversed the ownership the previous round had just established. Both were found by review or audit, never by the join-side map.
|
|
261
261
|
|
|
262
262
|
## Capability Naming And Provenance
|
|
263
263
|
|
|
@@ -59,6 +59,8 @@ If the validator reports `missing_required_command`, keep the failure visible an
|
|
|
59
59
|
- For new skills or major workflow changes, use `writing-skills` for RED-baseline/test-first methodology — **and the firing point is BEFORE drafting the body, not only before finalizing**. Eval-first authoring for a NEW skill (or a new hard-rule section): (1) write the evaluation scenarios first — **at least three** for a new skill (both the vendor's published authoring guide and the high-star practice pack converge on three-plus scenarios before body text; a single-rule edit may scope down to that rule's own scenario); (2) run them WITHOUT the skill and record the observed failures verbatim — and for a discipline-slip failure (the agent knows the rule and skips it under pressure), capture the agent's rationalizations word-for-word: each verbatim excuse is the raw material for one rationalization-vs-reality row and one red-flag line in the skill text (the discipline-slip form in `rule-consolidation.md`'s form-by-failure table); an invented hypothetical excuse does not qualify — counter only what a run actually said, and don't add rows for excuses no run produced; **a no-skill control that does not exhibit the failure is a stop signal — do not author guidance for a failure you cannot observe** (record the null finding instead; this is the pre-draft face of "Evidence must come before new rules"); (3) draft the **minimal** content that addresses the observed failures, then re-run the same scenarios WITH the skill; (4) when a later run, review round, or live miss surfaces a NEW rationalization for an existing discipline gate, add its explicit counter row to that gate's table and re-run the tempting scenario — counter tables accrete from observed excuses across rounds, never from imagination. The code-level RED-GREEN-REFACTOR method (write the failing case first, watch a fresh agent violate the rule WITHOUT the skill, then add the skill and watch it comply) is owned by `superpowers:writing-skills` + `superpowers:test-driven-development` — **if installed, route there; otherwise apply the RED-baseline rule inline** (manually record the without-change failure and the with-change compliance). This is the skill-authoring face of **eval-driven development** (for a behavior/routing change, run the scenario before you finalize; never special-case the scenario just to make it pass) — borrow the *principle*, not a claim of production-grade eval rigor.
|
|
60
60
|
- **What makes a `RED-baseline` valid is executed-and-recorded vs narrated — not recorded vs live.** Any evidence form (before-after diff, golden trace, or pressure scenario) is valid when it actually records the without-change failure AND the with-change compliance with a locator + expected-vs-actual (see `dual-track-review-gate.md`). Prose that merely *describes* an expected failure without running it is not a baseline; a pressure scenario you actually executed and recorded is.
|
|
61
61
|
- **A claimed impossibility is falsified in-env before it lands (authoring and reviewing alike; relocated from `SKILL.md`).** Deferring work behind a conservative-sounding stub ("not implemented", "needs a future tested wrapper") is the avoidance form of a blocked-verification claim: it *feels* safe but ships an unverified impossibility as durable behavior — distinct from a real `unavailable`/`pending`-with-remediation+residual-risk record, which is what remains after a safe attempt was genuinely impossible (the attempt itself stays within the safety boundaries the `SKILL.md` rule names). Failure shape: a fallback reviewer lane shipped as a permanent fail-closed stub on an unverified "permission probe not implemented" premise that a ~60-second live tool run refuted, re-enabling the lane.
|
|
62
|
+
- **The falsifying operation has exactly two admissible forms, and a search hit is neither (detail for the `SKILL.md` landed-conclusion rule).** (1) **Executed path** — run the suspected mechanism on the path that actually failed and collect an observation that only THIS cause predicts. Naming the call site is the entry bar, not the operation: **reachability alone rules out non-execution and nothing else** — a wrapper can demonstrably run while a downstream stall is what caused the failure, so "I watched the suspected code execute" clears no cause. The probe must produce a discriminating consequence (the mechanism's own signature in the output, a value only it would set, a timing or state it alone explains); if the run cannot separate the suspect from the alternatives still on the table, it has not falsified anything and form (2) is required. (2) **Single-variable control** — a paired probe whose arms differ in exactly ONE variable. Enumerate EVERY precondition of the predicate under test (each threshold, each uniqueness or vocabulary requirement, each surrounding-state assumption) and confirm both arms equal on all but the varied one; an arm differing in two preconditions yields a verdict attributable to neither, and it fails silently because the verdict still reads as decisive. A shared surface here is the source-register row, the commit or merge-request body, a durable note or memory, and the final response's causal account. The control leg a deferred registration owes (`source-to-skill-extraction.md`) is this rule's narrow instance, and the differential attribution a `RED-baseline` row owes (`dual-track-review-gate.md`) is its encoded form — both assume the arms are otherwise equal, which is the part that has to be enumerated rather than assumed.
|
|
63
|
+
Failure shape: a gate rejected a ledger anchor; a `force_encoding(BINARY)` call found by grep in the gate script was recorded as the cause and landed in a merged PR body, a durable memory note, and a round charter. The suspected line was never on the checked path — that check resolves its blobs through a different helper — and the real cause was an unbounded string replace in the round's own edit, which had modified a pre-existing ledger row. Three paired probes built to test the theory each differed in more than one precondition (anchor length against the threshold, uniqueness within the file, vocabulary membership), so no verdict was attributable; the first of them produced the wrong cause. Withdrawing it cost corrections on three surfaces, against about a minute for the executed-path question.
|
|
62
64
|
- **A headless code-writing RED/GREEN needs a *fair* violation-tempting scenario + an *independently-valid* objective measure — a capable agent complies on a clean task.** When the rule governs how an agent WRITES code, a clear well-specified prompt usually makes even the no-rule baseline produce compliant code (zero delta → no RED), so a clean-task baseline proves nothing. Surface the RED with **realistic inherited pressure** — a real legacy/house convention the rule must override, or a genuinely ambiguous spec — **NOT an explicit instruction to emit the anti-pattern**: leading the baseline directly into the violation launders a constructed failure into "RED" and is fabrication (the same defect as the rule above), so record why the prompt is fair and mark any direct-leading scenario synthetic/advisory, not RED. Score both runs with an **objective measure that detects the ANTI-PATTERN** — the rule's own checker counts ONLY if it was independently validated first (held-out positive/negative fixtures + a documented residual boundary, so RED genuinely fails); a same-change checker that merely whitelists the expected GREEN form makes GREEN trivially true. Run the agent **cross-model / fresh-context** so the baseline isn't primed by your session. If even the fair tempting baseline complies, that is an honest finding — the rule's marginal value is in edge/legacy cases, not the common one — record it, don't manufacture a RED.
|
|
63
65
|
- **Optional real-agent RED-baseline (F4 Tier-3).** For a routing-surface or hub-skill change you can execute the baseline through the in-repo F4 Tier-3 harness (`scripts/eval-golden-trace.rb`; contract in `eval-routing.md`) instead of a hand-recorded scenario: run it manually **twice** — first on the checkout WITHOUT your change, then WITH it (the harness does **not** auto-checkout or auto-diff; you run both and preserve both reports as the without→with evidence). Caveats that keep it honest: (a) it is **advisory + non-deterministic + small-N** (a few hub traces, structured *routing* assertions only, statuses PASS/FAIL/INCONCLUSIVE) — a without-change run that returns PASS or INCONCLUSIVE does **not** establish a RED; only an actually-observed miss does; (b) prefer a **pre-existing frozen trace** — a trace you author alongside the change is self-authored evidence (the runner checks `frozen_at_sha` ancestry but does **not** detect a trace added/tuned in the same change), so the reviewer must confirm it is not trivially fail-before/pass-after; (c) it stays **optional and never mandated** — a properly executed-and-recorded pressure scenario remains a valid `RED-baseline` for most changes.
|
|
64
66
|
- Create at least one pressure scenario: a realistic prompt where the agent should use the skill and avoid a known failure mode.
|
|
@@ -1444,6 +1444,11 @@ if [ -n "$root_worktree_toplevel" ]; then
|
|
|
1444
1444
|
fi
|
|
1445
1445
|
ruby "$impact_chain_gate" "$root"
|
|
1446
1446
|
|
|
1447
|
+
# No merge-side ledger-binding probe runs here, deliberately: this checker also
|
|
1448
|
+
# runs against synthetic fixtures inside other suites, where such a probe cannot
|
|
1449
|
+
# operate and reddens suites that do not own it. That gate's own suite covers the
|
|
1450
|
+
# real-checkout path and the CI step enforces it. Do not re-add one here.
|
|
1451
|
+
|
|
1447
1452
|
# Whole-ledger firing-path resolution. The impact-chain gate above is
|
|
1448
1453
|
# diff-scoped by design, so a row's firing path is machine-checked exactly once
|
|
1449
1454
|
# — at the commit that added it. Nothing re-checks it afterwards, while the
|