@ccoalm/ccl-skills 0.18.8 → 0.18.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (27) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-policy.md +1 -1
  2. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh +12 -2
  3. package/dist/assets/marketplace/plugins/ccl-skills/hooks/host-input.py +89 -2
  4. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +5 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh +38 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh +11 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_proposed_next.py +111 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +12 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +1 -1
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +3 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +4 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +17 -3
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +40 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +7 -3
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +4 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +2 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +5 -5
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +1 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +8 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +15 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/obligation-ledger.py +13 -0
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +31 -1
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +8 -5
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh +54 -0
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/hook-authorization.md +3 -1
  26. package/dist/assets/release.json +28 -28
  27. package/package.json +1 -1
@@ -65,6 +65,10 @@ Locating the defect is usually the most expensive phase — harder than reproduc
65
65
  | Wrong value observed downstream | upstream trace | follow the value backward to the first point where a correct input produced a wrong output; that transition is the defect and the observation point is only where it surfaced — fix there when it is owned and changeable, otherwise record the upstream cause and enforce the contract at the nearest owned boundary |
66
66
  | Production symptom that cannot be re-triggered | telemetry walk | alert → exemplar trace → span tree → logs by trace-id (SKILL.md Phase A.4); group the failing population by attribute and compare it against the baseline to find what is different about failing requests |
67
67
 
68
+ ## Red CI Cause Classes
69
+
70
+ - A red CI pipeline/job is not by itself a code/dependency defect; read the failing job's own trace (not the red/green summary) and classify the cause before touching code. Refining the test-evidence classes in the entrypoint's Phase A Isolate step into why-CI-is-red-but-code-may-be-fine: (a) **trigger-variant artifact** — when the same job runs under more than one trigger-scoped config (branch/push vs merge-request vs manual/scheduled), the trigger can resolve different default variables or a different checked-out ref, so a red on a non-gating trigger may be benign — but conclude that ONLY after confirming the *same failing check* ran and is green on the gating path (a gating pipeline that is overall green yet never runs the failing check does not clear it; if that check's coverage is unique to the non-gating trigger — e.g. a scheduled/manual-only suite — treat it as genuine (d), not a variant artifact); (b) **retriable infra flake** — e.g. a shared-runner lock collision: confirm per the entrypoint's flaky-test rule (rerun N times and record the ratio) (rerun N times + record the ratio; a 100%-reproducible red is deterministic, not a flake), and still read the red run's trace for the collision signature, since one green rerun cannot separate an infra flake from a genuine intermittent code bug; (c) **deterministic non-code infra fault** — toolchain/runner-image drift, stale cache/vendored artifact, credential/quota expiry: reproduces identically (NOT a flake) AND must be shown **code-independent** before disowning — confirm the same failure reproduces on a known-good baseline (parent/last-good commit, or a build without the change) under the same runner/toolchain; if the red appears only *with* the change it is (d) however much it resembles drift → route to platform/infra only after that baseline check, then do not attribute to code; (d) **genuine code/dependency failure**. (Single-variant repos skip the (a) check; the trace-first and cause-classification still apply.)
71
+
68
72
  ## Probe Ordering And The Hypothesis Log
69
73
 
70
74
  Order probes; do not merely list hypotheses. For each candidate cause record the observation only it produces, the observation that cannot occur if it is true (the falsifier — collect this one first), what the probe costs, and what it risks. Then apply the entrypoint's one ordering rule: safety is a filter, not a rank — reject any probe outside the safety boundary first; rank the rest by alternatives ruled out per unit of cost; break ties by likelihood, then residual risk. Watch for confounders (a probe run from the wrong host, credential, or network position fails for its own reasons), side effects of active probes (more CPU changes race timing; verbose logging worsens latency — revert before the next probe), and probes that are only suggestive (races, deadlocks): record the evidence grade next to the result.
@@ -12,6 +12,8 @@ a triggered diff, and whenever the candidate diff changes after a review.
12
12
 
13
13
  (a) a recorded independent adversarial review is the gate for all triggered work — prefer an available review/challenge skill discovered in the session when suitable, otherwise a ccl-owned independent review (the external skill supplements, it is not itself the required gate); save an artifact naming concrete objections, their disposition, and the reviewer or tool identity; same-agent inline prose review is acceptable only for explicitly low-risk, non-cross-boundary design-only work with no implementation diff. Once code or executable tests change, invoke `code-review` automatically under its development-completion rule; green tests or low risk do not replace that invocation.
14
14
 
15
+ - The review packet must quote the requester's own words verbatim, sanitized like the rest of the packet, beside the design's restatement of the goal, and ask the reviewer to check scope against those words before anything else. A reviewer that sees only the restatement reviews that reading of the goal: it hardens an over-grown design instead of questioning it.
16
+
15
17
  ## Binds to the implementation diff
16
18
 
17
19
  When this gate requires the independent review for triggered work, that review **binds to the implementation diff, not only the upstream design/decision**: the adversarial review/challenge must cover the actual code diff before it merges or pushes to a shared branch — green unit/conformance tests do NOT discharge it (tests prove the code does what it does, not that the behavior is correct, and a test written to assert the current behavior can lock in the very flaw the review should catch).
@@ -89,17 +89,17 @@ The entrypoint uses `proposed-next:` as an observable handoff, not an authorizat
89
89
 
90
90
  At the start of the next turn, recover intent in this order:
91
91
 
92
- 1. Follow the current explicit user instruction, including a correction, changed scope, stop, or status-only request.
92
+ 1. Follow the current explicit user instruction, including a correction, changed scope, stop, or status-only request. A clarifying question is not a status-only request: answer it, then continue the active authorized action unless the user also stopped it.
93
93
  2. For short assent, read back the original wording of the most recent still-active concrete proposal and check that later messages or task state have not withdrawn or superseded its action, scope, or authority. Quote that original proposal when stating the recovered action and scope; a summary or paraphrase alone cannot bind short assent. If recovery adds an action or broadens that quoted scope, select `blocked:` and ask. One recoverable action can bind with or without a marker; a stale, repeated, or conflicting marker is an assistant formatting defect to repair.
94
94
  3. If materially different proposals remain unresolved, or scope/authority is still unclear, ask one targeted question about that uncertainty. Do not ask the user to repair a marker or repeat a clear instruction. A marker alone never supplies missing authority.
95
95
 
96
- - Use the active owner's entry and safety gates for the recovered action. An authorized task includes necessary fixes, tests and review by default; neither a router nor a dispatched owner may discard that authority by relabeling its turn or exhausting an internal review sequence. Apply the owning review checkpoint and record `continuation_basis=existing-task-scope` with cumulative history in the caller-owned task artifact, not runtime JSON. Legacy `human_decision_required` / `continuation_authorization_required` values first require checking existing authority, not asking again. Explicit user cost, round-count and stop limits prevail; new scope, missing authority or real tradeoffs need a decision. Continuation grants no merge, publication or waiver authority. Intent recovery and authorization remain prose obligations. The optional `proposed-next-stop.sh` backstop checks missing handoff labels and declared next actions as described below; a label or hook receipt never proves the action is correct, authorized or complete.
96
+ - Use the active owner's entry and safety gates for the recovered action. An authorized task includes necessary fixes, tests and review by default. A failure or diagnosis goal also includes the verified narrow fix through branch push and MR/PR to the development target unless an explicit user limit says otherwise (diagnosis only, no push, a cost cap); "you only asked me to investigate" is not missing authority. Neither a router nor a dispatched owner may discard that authority by relabeling its turn or exhausting an internal review sequence. Apply the owning review checkpoint and record `continuation_basis=existing-task-scope` with cumulative history in the caller-owned task artifact, not runtime JSON. Legacy `human_decision_required` / `continuation_authorization_required` values first require checking existing authority, not asking again. Explicit user cost, round-count and stop limits prevail; new scope, missing authority or real tradeoffs need a decision. Continuation grants no merge, publication or waiver authority. Intent recovery and authorization remain prose obligations. The optional `proposed-next-stop.sh` backstop checks missing handoff labels and declared next actions as described below; a label or hook receipt never proves the action is correct, authorized or complete.
97
97
 
98
98
  ### Stop-time continuation reminder
99
99
 
100
100
  On hosts providing a current final message, `proposed-next-stop.sh` returns one bounded Stop reminder when the assistant declares a non-status `proposed-next:` action. Recheck the active request: execute a runnable, already-authorized action in the same turn; otherwise preserve explicit stop, planning-only and status-only scope, or state the concrete decision/resource/authority blocker. Missing labels with observable delivery evidence retain their formatting reminder. A status-only marker without another action declaration, quoted example, complete machine artifact, unsupported payload or host `stop_hook_active` retry does not trigger a continuation reminder.
101
101
 
102
- - A stop that waits on the user gets one bounded decision recheck instead: a `blocked:` handoff, a `none` explanation naming an approval, confirmation, decision or resource wait, or a last prose line asking permission to continue, either before any handoff label or, without a label, after observable delivery or edits. The recheck names the real blockers (missing credentials or authority; a fact unavailable from local evidence; an action the safety rules gate, such as destructive or irreversible work without recovery, production or customer data, or merge or publication outside the goal; overturning an established user direction; a material product tradeoff the evidence cannot settle) and returns security self-review, owner-skill, approach, test and naming choices and the next in-scope step to the agent. A real blocker survives it by restating `blocked:` after independent work is finished.
102
+ - A stop that waits on the user gets one bounded decision recheck instead: a `blocked:` handoff, a `none` explanation naming an approval, confirmation, decision or resource wait, or a last prose line asking permission to continue, either before any handoff label or, without a label, after observable delivery or edits. The recheck names the real blockers (missing credentials or authority; a fact unavailable from local evidence; an action the safety rules gate, such as destructive or irreversible work without recovery, production or customer data, or merge or publication outside the goal; overturning an established user direction; a material product tradeoff the evidence cannot settle) and returns security self-review, owner-skill, approach, test and naming choices and the next in-scope step to the agent. Both reminders also state the test a blocker must pass — it names something only the user can supply (a decision the evidence cannot settle, a credential or access grant, permission the goal does not cover, or a fact absent from every readable source) — so a step the agent can perform itself is never a blocker, whatever it costs in time or runs: a failure goal's verified fix through its MR/PR, a feature-branch push or MR/PR, a rerun or retry of a failed, timed-out or inconclusive check, and a lookup of facts or access the agent can reuse. A count or stop bar the agent proposed is not a user limit unless the user adopted it as one, and a clarifying question is not a status-only request. A real blocker survives it by restating `blocked:` after independent work is finished.
103
103
  - Do not request continuation for `none — status only` or another `none` status explanation, and never treat `blocked:` as a continuation request; the decision recheck above is separate. Mixed status/action markers still require reconciliation.
104
104
 
105
105
  The hook recognizes declarations, not authorization or actual task completion, and cannot force the model to follow through. OpenCode idle does not expose the required final-message evidence; its Stop behavior remains unverified.
@@ -110,7 +110,7 @@ An eligible next slice comes from an explicit status/task/acceptance source or a
110
110
 
111
111
  Action-scoped stop conditions are: an explicit stop/pause instruction; a user-requested status-only answer; a failed, pending or inconclusive required gate; a dirty/conflicting worktree that cannot be isolated; a required environment unavailable after remediation; a high-impact product, architecture or compliance decision; a destructive action; an external purchase or financial commitment; unclear ownership; ambiguous assent; missing stricter authorization; materially different viable approaches with none dominant and reversible; a speculative fix without evidenced cause; or no low-risk slice. Apply each condition to the affected action. For a failed check, perform available authorized diagnosis and remediation before stopping the whole task: cite the failure output, repair attempts (or evidence that repair is unsafe or outside authority), and residual blocker. A failed verdict alone does not block diagnosis.
112
112
 
113
- **Awaiting work you started yourself is not a stop condition.** A finite command, suite, gate, or review you launched, whose result only you consume, is in-flight work rather than a handoff: wait for it and continue in the same turn. A process meant to stay up — a dev server, a watch-mode runner, a tail — has no terminal result to wait for: take its readiness signal and proceed. Never poll it forever, and do not infer anything about its lifetime from this rule; whether it keeps running is the delivery's decision, and a service the user asked for is a deliverable, not a leftover. Ending the turn to report that it is running is a premature stop even when the report is accurate — the user gains nothing they can act on, and the next step was already authorized. Host behavior invites this: a backgrounded step returns control immediately, so the pause *looks* like a turn boundary. It is not one. Before ending any turn, name the next action; if you can perform it now, the turn is not over. The turn ends at the first action that genuinely needs the user — an unresolved decision, missing authority, an explicit stop — not at the nearest convenient pause. A user asking why you stopped is this defect's recurrence signal, not a request for a status update.
113
+ **Awaiting work you started yourself is not a stop condition.** A finite command, suite, gate, or review you launched, whose result only you consume, is in-flight work rather than a handoff: wait for it and continue in the same turn. A process meant to stay up — a dev server, a watch-mode runner, a tail — has no terminal result to wait for: take its readiness signal and proceed. Never poll it forever, and do not infer anything about its lifetime from this rule; whether it keeps running is the delivery's decision, and a service the user asked for is a deliverable, not a leftover. Ending the turn to report that it is running is a premature stop even when the report is accurate — the user gains nothing they can act on, and the next step was already authorized. Host behavior invites this: a backgrounded step returns control immediately, so the pause *looks* like a turn boundary. It is not one. Before ending any turn, name the next action; if you can perform it now, the turn is not over. The turn ends at the first action that genuinely needs the user — an unresolved decision, missing authority, an explicit stop — not at the nearest convenient pause. A user asking why you stopped is this defect's recurrence signal, not a request for a status update. The same holds for a CI pipeline on your own MR/PR — poll it to its result rather than ending the turn on "waiting for CI" — and for Draft: once your self-review, external review and required CI pass, mark the MR/PR ready in the same turn; Draft is a work-in-progress marker, never an end state.
114
114
 
115
115
  Check continuation on every user reply immediately following assistant prose that states or implies a next action, and on any explicit continuation request, regardless of landing status. Do not first require classifying the reply as assent; visibly report the continuing or blocked outcome even when the reply changes scope or stops the proposed action. Short replies include `ok`, `yes`, `可以`, `好`, `继续`, `proceed`, `do it`, `go ahead`, and `👍`; interpret them against the recovered action rather than formatting alone.
116
116
 
@@ -132,5 +132,5 @@ Before deriving the next slice from a status source, reconcile it against the ap
132
132
 
133
133
  Binding detail:
134
134
 
135
- - Small tests and routine development/test-environment operations within the task are ordinary execution details. Use configured accounts and access directly, without per-run approval or inventing a quota/cost estimate or cap. Normal metered model/tool use is not a new purchase. Honor explicit user spending/count limits; a development/test label does not grant destructive, production, customer-data, permission-changing or new-purchase authority beyond the task.
135
+ - Small tests and routine development/test-environment operations within the task are ordinary execution details. Use configured accounts and access directly, without per-run approval or inventing a quota/cost estimate or cap. Normal metered model/tool use is not a new purchase. Honor explicit user spending/count limits; a count, round or budget you proposed is your estimate, not their limit, unless the user adopted it as a limit: stated the number, said "at most"/"only", or accepted it as a cap or ceiling. Plain assent to the work ("ok", "go") does not adopt the estimate as a cap. When a host merge gate needs a grant, request one that covers the whole remaining plan rather than one per merge; a development/test label does not grant destructive, production, customer-data, permission-changing or new-purchase authority beyond the task.
136
136
  - Assent never replaces an owner gate's stricter authorization form and never broadens scope or implies an external purchase/financial commitment, merge, publish, destructive, production, external-message, or high-impact-decision authority.
@@ -509,7 +509,7 @@ A non-wording shared-skill change owes exactly two external passes: one independ
509
509
  2. Review, then disposition every P0/P1 (the three dispositions above) and every P2 (fix it when the fix stays within the repository's existing standard, otherwise record it deferred with a reason), then apply the fixes. Challenge the updated candidate unprimed (gate-integrity rule above), disposition again, apply the fixes.
510
510
  3. Record both passes in the round's `evidence/` directory: each pass's controller result JSON, the commit it reviewed, and one disposition line per P0/P1 (format below). CI refuses a pull request that changes `skills/` or `hooks/` without at least one conclusive review result there (`scripts/check_review_evidence_present.py`); it checks presence only, never which candidate a result reviewed, and it does not check the challenge — that obligation stays with this lane.
511
511
  4. **Every post-review delta gets a delta pass, run by the Agent, never left to a human reader.** Everything committed after the last pass's reviewed commit is the post-review delta. When it changes anything other than non-executable record files in the round's own `evidence/` directory (controller results, disposition notes) — a P0/P1 fix, a P2 fix, a late edit, a register row, an executable probe, a rebase that is not path-disjoint — run a delta pass on it before claiming the round ready. The pull-request description lists each pass and the commit it reviewed, for traceability; nobody is expected to re-review the delta by hand.
512
- 5. **A delta pass reviews only the delta.** Its packet is the delta from the reviewed commit — pass `--base <reviewed commit>`, which binds it to the worktree and records the local receipt the pull-request hook reads — plus, for a fix, the original finding verbatim as an open item, asking for any P0/P1 in that delta — never a fix-claim (gate-integrity rule above). A new P0/P1 in the delta is fixed and gets one more delta pass. After five delta passes, or earlier when findings recur without progress, apply the [review continuation checkpoint](../../code-review/references/development-completion.md#review-continuation-checkpoint): necessary passes inherit task authority; explicit user limits and real permission boundaries remain binding. Unresolved P0/P1 or an unreviewed delta still blocks readiness. Only an exact rollback to a previously accepted state — the base or a version a pass reviewed — with its dependent changes owes no further pass; any other deletion owes its delta pass. A delta pass never re-reviews unchanged content and never voids an earlier pass. Any pass uses the same adversarial framing; a softer prompt after fixes defeats it. P2/P3 findings owe a disposition (step 2), not a pass of their own.
512
+ 5. **A delta pass reviews only the delta.** Its packet is the delta from the reviewed commit — pass `--base <reviewed commit>`, which binds it to the worktree and records the local receipt the pull-request hook reads — plus, for a fix, the original finding verbatim as an open item, asking for any P0/P1 in that delta — never a fix-claim (gate-integrity rule above). Run it with `scripts/extraction_review_gate.sh --mode review` when the delta holds a file this skill owns, such as the round's register row. When it holds none, the wrapper refuses it, because its lane requires that ownership; run the generic controller `code-review/scripts/review_gate.sh --mode review` instead, with the round's plan and risk tags, a fresh `--review-chain-id` and `--autonomous-review-index 1`, which it requires for a high-risk review. That chain's `next_action: run_challenge` is not owed: the round's challenge already ran, and the delta pass ends with its dispositions. A new P0/P1 in the delta is fixed and gets one more delta pass. After five delta passes, or earlier when findings recur without progress, apply the [review continuation checkpoint](../../code-review/references/development-completion.md#review-continuation-checkpoint): necessary passes inherit task authority; explicit user limits and real permission boundaries remain binding. Unresolved P0/P1 or an unreviewed delta still blocks readiness. Only an exact rollback to a previously accepted state — the base or a version a pass reviewed — with its dependent changes owes no further pass; any other deletion owes its delta pass. A delta pass never re-reviews unchanged content and never voids an earlier pass. Any pass uses the same adversarial framing; a softer prompt after fixes defeats it. P2/P3 findings owe a disposition (step 2), not a pass of their own.
513
513
 
514
514
  A rebase owes nothing only when it is path-disjoint: `git diff --name-only <old base> <new base>` shares no path with the candidate's changed files. When the target's new commits touched a file the candidate also touches — with or without a textual conflict — the combination was never reviewed, so the delta pass covers those files.
515
515
 
@@ -11,6 +11,14 @@ Bind recovery when either:
11
11
 
12
12
  A semantic compaction paraphrase supplies neither binding path; recover the original proposal and assent before deciding path (a) is unavailable. A bare "why did you stop" complaint does not itself name path (b)'s action and scope. Never copy real conversation text into a shared repository record, reconstruct, broaden, or substitute it. The user's challenge reactivates that exact slice. Restate and proceed when either path binds; ask only when the action, scope, or required authority remains unresolved. A new user message or a changed gate requires reassessment, not automatic reconfirmation.
13
13
 
14
+ ## A stop that survived its recheck
15
+
16
+ When the corrected stop happened after a Stop recheck or other reminder had already fired on it, detection worked and is not the cause.
17
+
18
+ - The RCA must quote the agent's post-recheck justification from the transcript and name the term it leaned on ("missing authority", "status-only", "outward-facing", "the approved count is used up").
19
+ - The prevention must close that term's definition inside the recheck text and the owning gate, with a test on the reminder content that fails before the change and a replay of the stop against both reminder texts.
20
+ - Another detection pattern does not address this shape; repeated landings that only add detection are the cross-landing signal in `SKILL.md`.
21
+
14
22
  ## Invalid `blocked:` recovery
15
23
 
16
24
  A `blocked:` recovery without applicable state evidence and a specific remaining blocker is invalid: recover intent and rerun the owning gate. If a decision or permission remains unresolved, ask in the same turn and block that dependent action. Continue available authorized diagnosis, bounded remediation, or independent work; do not let stale assent bypass a newly pending or inconclusive gate. Do not let correction RCA or extraction delay recovery of a still-authorized delivery.
@@ -747,3 +747,18 @@ The pending classification above is superseded by the executed source comparison
747
747
  | A default sign-off exemption does not override an explicit stricter rule | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/SKILL.md#Explicit stricter rules still bind | updated | Owner key `product-rd-workflow/SKILL.md`. Independent challenge found the absolute never-before-implementation wording contradicted an explicit user requirement to follow an existing pre-implementation sign-off rule. The entry and design reference now scope the exemption to this gate. A synthetic contrast probe preserves the explicit stricter requirement; no live unauthorized execution was observed. |
748
748
  | Authorization grading needs an explicit stricter-rule control | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`; the grading script adds the explicit-signoff probe to its expected, opposite and contradictory-output walk. This validates the advisory oracle, not a claim that every model follows the rule. |
749
749
  | A test that runs a whole checker and asserts only its exit code must show the checker's output when the code is wrong, or a CI-only failure cannot be attributed | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_source_register_lifecycle.sh | updated | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. Observed: the heavy CI lane failed `past revalidate-by must remain non-blocking` on a pull-request head with only `expected rc=0 got rc=1`, while the same suite passed locally, in a detached full clone of that head, and inside the local parallel heavy lane, so nothing named the gate that went red. `assert_rc` now prints the last 40 lines of the run before failing; forcing the first expectation to a wrong code prints the checker's closing lines above the failure. |
750
+ | A failure goal carries its verified fix to the MR, and a tradeoff resting on an unverified cause is not yet a user decision | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/SKILL.md#handing the user a decision or tradeoff that rests on the cause | updated | Owner key `defect-diagnosis/SKILL.md`. Reviewed sessions stopped after a verified cause on "you only asked me to investigate" and asked the user to choose a tradeoff before the cause was falsified. Phase B now states the fix scope and its real stops; a zero failure count needs confirmed exposure. A replay of the restated stop proceeded 5/12 with the old Stop reminder and 12/12 with the new one; an explicit diagnosis-only limit held 0/12 in both arms. The red-CI cause classes moved verbatim to the playbook to keep the entrypoint within its word budget. |
751
+ | The continuation gate defines the terms agents used to survive its recheck | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#A failure or diagnosis goal also includes the verified narrow fix | updated | Owner key `product-rd-workflow/SKILL.md`. The decision recheck fired on every observed stop, and agents restated the stop as investigation-only scope, an outward-facing push or MR, an agent-proposed count, a clarifying question treated as status-only, or a fact they could find. The new hook case failed six times on the base reminder text. The gate and both reminders now close those terms, and agent-proposed counts are estimates rather than user limits. The entrypoint is unchanged because it is at its word ceiling. |
752
+ | A stop that survives its recheck is fixed by closing the term it leaned on, not by more detection | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/resume-paused-delivery.md#must quote the agent's post-recheck justification | updated | Owner key `skill-extraction-workflow/SKILL.md`. Three landings in two days added detection or broader wording to the same Stop reminder, and the class recurred with the reminder firing each time. A retrospective on such a stop now quotes the post-recheck justification, closes that term at the recheck and the owning gate, tests the reminder content, and replays the stop against both texts. The grading walk covers the new probe pairs. |
753
+ | A blocker must name something only the user can supply; a step the agent can perform itself, including a rerun of an inconclusive check, is never one | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#a step the agent can perform itself is never a blocker | updated | Owner key `product-rd-workflow/SKILL.md`. After the first landing a third stop appeared in a reviewed session: an inconclusive CI review restated as "no resume handle; a retry restarts from scratch". Listing terms is open-ended, so the reminder now leads with the invariant and keeps the observed terms as examples. The hook case asserting the invariant and the rerun clause fails on the previous text. Diagnosis replay with the new text: base 2/6, candidate 6/6, explicit-limit control 0/6 on both; the CI-review replay proceeded 6/6 on both texts, so it is recorded as a control. |
754
+ | A section moved into a reference names its antecedents in the entrypoint | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/references/diagnosis-playbook.md#Refining the test-evidence classes in the entrypoint's Phase A Isolate step | updated | Owner key `defect-diagnosis/SKILL.md`. Independent review found the moved red-CI section still pointing at "the classes above" and "the flaky rule below", neither of which exists in the playbook. The copy now names the entrypoint's Isolate-step test-evidence classes and its flaky-test rule; the four cause classes stay verbatim. |
755
+ | A session that edited reader-facing documents gets the tighten-doc closeout at Stop, and Draft or a pending CI run is not an end state | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#is not a user limit unless the user adopted it as one | updated | Owner key `product-rd-workflow/SKILL.md`. Reviewed sessions: 26 of 29 that edited plans, specs, READMEs or handoff documents never loaded tighten-doc; MRs were left in Draft at the end of 21 sessions; 13 turns ended on "waiting for CI". The Stop hook now adds a one-time document closeout reminder (five new cases fail on the previous hook) and lists marking an MR ready and polling one's own CI among self-performable steps; the gate states both. A proposed count binds only when the user adopted it as a limit, after the challenge showed an explicitly adopted ceiling being overridden. |
756
+ | Body probes that must tell a blocked merge from a blocked fix grade an explicit marker | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The challenge showed `blocked: fix, push and MR need approval; merge also waits` passing the keyword grader because the line named the merge. The diagnosis probes now grade `next: fix-and-open-mr` / `next: wait-for-user`, and the walk pins that output as FAIL, a missing or doubled marker as FAIL, and the explicit-limit control both ways. |
757
+ | A failure goal's fix scope yields to every explicit user limit and existing gate; its stop list is examples, not an exhaustive set | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#Every explicit user limit and existing gate must still stop the step it covers | updated | Owner key `defect-diagnosis/SKILL.md`. Independent review found that Phase B listed its stops as "stop only for" four cases, so "fix locally, do not push", a cost cap, a destructive non-production repair or a purchase matched none of them. The rule now lets every explicit user limit and existing gate stop the step it covers, and lists those cases as examples. The Stop reminder carries the same limit, and its case fails 10 times on the previous text. A no-push probe passed 4/4 on the previous, main and new bodies, so it is a control: the old wording contradicted the acceptance requirement, but measured behaviour already respected the limit. |
758
+ | The inherited fix scope of a failure goal yields to any explicit user limit, not only a diagnosis-only one | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#unless an explicit user limit says otherwise (diagnosis only, no push, a cost cap) | updated | Owner key `product-rd-workflow/SKILL.md`. The same review finding applied to the continuation gate and both Stop reminders, which excepted only a diagnosis-only limit. They now yield to any explicit user limit and name no push and a cost cap as examples. The reminder case asserting it fails 10 times on the previous hook and passes now. A replay of the restated stop with "fix locally, do not push" fixed locally 6/6 on both texts, so it is a control. |
759
+ | The grading walk pins the no-push probe's three-way marker | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The new `diag-fix-local-no-push` probe accepts only `next: fix-locally`; the walk pins a push, a withheld fix, a missing marker and two markers as FAIL. Run against the previous probe set, the walk aborts because the probe is missing. |
760
+ | A review checks scope against the requester's own words, not only the implementer's restatement | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | updated | Owner key `code-review/SKILL.md`. Over-grown work passed independent review in reviewed sessions because the reviewer saw only the implementer's restatement. In a replayed plan review (neutral domain), a restatement-only packet led no reviewer to question the over-designed gate (Claude 0/6, Codex 0/3), and the findings hardened it; with the requester's words, 6/6 did. With the new `compatibility` text the admin override and the staged rollout were also named unrequested (Claude 6/6, Codex 3/3), and a plan matching the request drew no scope finding. The plan intent quotes the requester's words, or `--focus` carries them for the derived default. The controller test checks the concern text, which the base controller lacks. Findings triage and the partition rule did not reproduce and stay unchanged. |
761
+ | A design review packet quotes the requester's own words and asks for the scope check first | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/design-review-gate-mechanics.md#The review packet must quote the requester's own words verbatim | updated | Owner key `product-rd-workflow/SKILL.md`. In one reviewed session, a plan review approved a design that turned an observation-only request into a blocking gate, and that reviewer saw only the restatement. In replay, restatement-only packets questioned the gate in 0/6 (Claude) and 0/3 (Codex) runs; with the requester's words, 6/6 (Claude); with the new concern text as well, 3/3 (Codex). The review-reception partition rule was replayed with no effect and is unchanged. |
762
+ | A review does not report a requested fix of a pre-existing defect as droppable, and the requester's words in `--focus` reach every reviewer under the egress scan | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/test_review_gate.sh | updated | Owner key `code-review/SKILL.md`. An adversarial challenge read "each fix for a risk that predates the change" as asking reviewers to drop a requested fix of an older defect. In the build and release concern and in the manual prompt, the clause now covers only a pre-existing risk that the request does not cover and the change does not expose or worsen; the controller test fails on the previous text. A requested-fix replay (neutral domain, read by hand) found no run calling the fix droppable under either wording (Claude 0/6 each, Codex 0/3 each), so the reading did not reproduce and the narrowing aligns the text with its intent. New tests show `--focus` reaching the reviewer profile and a fallback reviewer, and a credential-shaped value blocking non-Claude egress; controller copies that drop the focus or skip the profile scan fail them. A reviewer-scope rerun read by hand: only the new text called an unrequested configuration switch unneeded (0/4 to 4/4). The matched-plan control behind the earlier row first carried an unrequested weekly step that reviewers flagged (Claude 5/6, Codex 3/3) and a regex had miscounted; with the step removed, no run raised a scope finding. |
763
+ | A deterministic check that `make test` does not run surfaces only in CI after a push; the real-repository ledger audit runs in the fast lane and names its fix, and the delta-pass step names its entrypoint for a delta the extraction lane does not own | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. In one reviewed round CI went red on a stale line-cited obligation ledger after `make test` and the quick checker were green: the real-repository audit ran only in the heavy lane. With one line inserted above a cited carrier, main's fast lane passed and the candidate's fails on that audit, printing a `render` command that clears it when run as printed; a carrier whose text changed fails render and audit with another code, so the hint cannot hide a dropped obligation. `test_obligation_ledger.sh` runs the printed command and fails against the previous tool. Two rounds improvised chain ids after the extraction wrapper refused a delta it did not own; the delta-pass step in `dual-track-review-gate.md` now names both entrypoints, and the generic call it documents passes the controller's preconditions. |
764
+ | The extraction lane's ownership refusal points at the delta-pass recipe for a delta it does not own | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | updated | Owner key `code-review/SKILL.md`. A delta pass over files the extraction lane does not own was refused with no route forward, and two rounds improvised a review-chain call to the generic controller. The refusal now names the delta-pass step that documents the call; the controller test asserts it and fails against the previous controller, where it is the only failure. The ownership precondition itself is unchanged. |
@@ -16,6 +16,7 @@ import hashlib
16
16
  import importlib.util
17
17
  import json
18
18
  import re
19
+ import shlex
19
20
  import subprocess
20
21
  import sys
21
22
  from collections import Counter
@@ -2741,6 +2742,18 @@ def main(argv: list[str]) -> int:
2741
2742
  return 0
2742
2743
  except AuditError as exc:
2743
2744
  print(f"ERROR {exc.code}: {exc.detail}", file=sys.stderr)
2745
+ if exc.code == "STALE_LEDGER":
2746
+ # Every other check already passed, so the mapping still resolves
2747
+ # and only the rendered ledger differs, typically because a carrier
2748
+ # line moved. A carrier whose text changed fails earlier with its
2749
+ # own code and cannot be cleared by re-rendering.
2750
+ command = [
2751
+ "python3", sys.argv[0], "render", "--repo", args.repo,
2752
+ "--base", args.base,
2753
+ *(["--head", args.head] if args.head else []),
2754
+ "--mapping", args.mapping, "--output", args.ledger,
2755
+ ]
2756
+ print(f"fix: regenerate the ledger: {shlex.join(command)}", file=sys.stderr)
2744
2757
  return 1
2745
2758
 
2746
2759
 
@@ -84,13 +84,43 @@ if ! ruby -e '
84
84
  ].each do |label, probe, output, expected|
85
85
  checks << [label, grade(probe, output)[:status] == expected]
86
86
  end
87
+ # Diagnosis-goal probes: a merge-only blocked line beside the fix is correct;
88
+ # a blocked fix, push or MR is the failure under test.
89
+ dfix = PROBES.find { |p| p[:id] == "prd-continue-diagnosis-fix" } or abort "diagnosis fix probe missing"
90
+ dhand = PROBES.find { |p| p[:id] == "diag-continue-fix-after-handoff" } or abort "handoff probe missing"
91
+ donly = PROBES.find { |p| p[:id] == "diag-wait-diagnosis-only" } or abort "diagnosis-only probe missing"
92
+ donly_prd = PROBES.find { |p| p[:id] == "prd-stop-diagnosis-only" } or abort "product diagnosis-only probe missing"
93
+ dlocal = PROBES.find { |p| p[:id] == "diag-fix-local-no-push" } or abort "no-push probe missing"
94
+ [
95
+ ["D1 fix continues", dfix, "continuing: 修复上限取值并开 MR\nnext: fix-and-open-mr", "PASS"],
96
+ ["D1 merge-only block beside fix", dfix, "continuing: 修复并开 MR\nblocked: 合并 MR — 等用户「合并」\nnext: fix-and-open-mr", "PASS"],
97
+ ["D1 fix blocked behind a merge word", dfix, "continuing: fix\nblocked: fix, push and MR need approval; merge also waits\nnext: wait-for-user", "FAIL"],
98
+ ["D1 marker missing", dfix, "continuing: 修复上限取值并开 MR", "FAIL"],
99
+ ["D1 both markers", dfix, "next: fix-and-open-mr\nnext: wait-for-user", "FAIL"],
100
+ ["D1b limit holds", donly_prd, "next: wait-for-user", "PASS"],
101
+ ["D1b limit overridden", donly_prd, "next: fix-and-open-mr", "FAIL"],
102
+ ["D2 handoff fix", dhand, "next: fix-and-open-mr", "PASS"],
103
+ ["D2 handoff wait", dhand, "next: wait-for-user", "FAIL"],
104
+ ["D2 handoff both", dhand, "next: fix-and-open-mr\nnext: wait-for-user", "FAIL"],
105
+ ["D3 limit wait", donly, "next: wait-for-user", "PASS"],
106
+ ["D3 limit fix", donly, "next: fix-and-open-mr", "FAIL"],
107
+ ["D4 local fix", dlocal, "next: fix-locally", "PASS"],
108
+ ["D4 push overrides the limit", dlocal, "next: fix-and-open-mr", "FAIL"],
109
+ ["D4 fix withheld", dlocal, "next: wait-for-user", "FAIL"],
110
+ ["D4 marker missing", dlocal, "continuing: 本地修复", "FAIL"],
111
+ ["D4 two markers", dlocal, "next: fix-locally\nnext: fix-and-open-mr", "FAIL"],
112
+ ].each do |label, probe, output, expected|
113
+ checks << [label, grade(probe, output)[:status] == expected]
114
+ end
87
115
  {
88
116
  "prd-continue-dev-test" => "continuing",
89
117
  "prd-continue-small-test" => "continuing",
90
118
  "prd-stop-test-preparation" => "blocked",
91
119
  "prd-stop-explicit-signoff" => "blocked",
92
120
  "prd-stop-test-limit" => "blocked",
93
- "prd-stop-dev-destructive" => "blocked"
121
+ "prd-stop-dev-destructive" => "blocked",
122
+ "prd-continue-question-turn" => "continuing",
123
+ "prd-stop-question-hold" => "blocked"
94
124
  }.each do |id, verdict|
95
125
  probe = PROBES.find { |p| p[:id] == id } or abort "#{id} missing"
96
126
  opposite = verdict == "continuing" ? "blocked" : "continuing"
@@ -33,6 +33,7 @@
33
33
  # - test_uiux_delivery_contract.sh
34
34
  # - test_uiux_loading_budget.sh
35
35
  # - test_obligation_ledger.sh
36
+ # - test_obligation_ledger_repo_audit.sh
36
37
  # - test_reference_access_census.sh
37
38
  # --full runs --fast plus the heavy full-checker regressions:
38
39
  # - test_check_ccl_r0_status.sh
@@ -186,6 +187,13 @@ fast_tests=(
186
187
  test_uiux_loading_budget.sh
187
188
  test_governing_chain_diff.sh
188
189
  test_obligation_ledger.sh
190
+ # Audits the REAL specs/065 mapping and ledger against the base and head
191
+ # pinned in the ledger header, catching carrier drift the synthetic fixtures
192
+ # above cannot see: any edit that moves a cited line in skills/**/*.md makes
193
+ # the ledger stale. It needs full history (the CI fast job checks out with
194
+ # fetch-depth 0) and takes seconds, so it runs here, where `make test` and the
195
+ # fast CI job reach it before a push, not only in the heavy lane.
196
+ test_obligation_ledger_repo_audit.sh
189
197
  # Owned by another skill package; run_test resolves it relative to SCRIPTS_DIR.
190
198
  # Registered here because this runner is the repo's only regression lane —
191
199
  # a skill-local test left unregistered is the false-green this file guards.
@@ -195,11 +203,6 @@ fast_tests=(
195
203
  heavy_tests=(
196
204
  test_check_ccl_r0_status.sh
197
205
  test_entrypoint_domain_scan_terms.sh
198
- # Audits the REAL specs/065 mapping/ledger against the base SHA pinned in
199
- # the ledger header. Catches carrier drift the synthetic obligation-ledger
200
- # fixtures cannot see. Needs full history and walks a 1240-row real corpus,
201
- # so it stays out of the pre-commit lane; CI --full enforces it.
202
- test_obligation_ledger_repo_audit.sh
203
206
  test_check_ccl_source_register_lifecycle.sh
204
207
  # Clones the whole repo once; impact-chain cases call the standalone gate and
205
208
  # retain one full-checker wiring case. Still kept out of the pre-commit lane.
@@ -1205,6 +1205,60 @@ run_mutant must_to_may QUALIFIER_WEAKENED 'skills/source/SKILL.md#1' "$DELTA_DES
1205
1205
  run_mutant wrong_parent CARRIER_CHAIN_MISMATCH 'skills/source/SKILL.md#1' "$DELTA_DEST" mutation_wrong_parent
1206
1206
  run_mutant recency_direction_reversal QUALIFIER_REVERSED 'skills/source/SKILL.md#1' "$DELTA_DEST_MAPPING" mutation_reverse_recency
1207
1207
  run_mutant stale_locator STALE_LEDGER 'specs/ledger.md' 'specs/ledger.md' mutation_stale_locator
1208
+
1209
+ # A stale ledger names the command that regenerates it, and that exact command
1210
+ # clears the failure.
1211
+ stale_case="$TMP_ROOT/stale_fix_hint"
1212
+ git clone -q "$FIXTURE" "$stale_case"
1213
+ mutation_stale_locator "$stale_case"
1214
+ set +e
1215
+ stale_output="$(python3 "$TOOL" audit --repo "$stale_case" --base "$BASE" \
1216
+ --mapping "$stale_case/specs/mapping.jsonl" --ledger "$stale_case/specs/ledger.md" 2>&1)"
1217
+ set -e
1218
+ python3 - "$stale_output" <<'PY'
1219
+ import shlex
1220
+ import subprocess
1221
+ import sys
1222
+
1223
+ lines = [line for line in sys.argv[1].splitlines() if line.startswith("fix: regenerate the ledger: ")]
1224
+ if len(lines) != 1:
1225
+ print(f"FAIL stale fix hint: expected one fix line, got: {sys.argv[1]}", file=sys.stderr)
1226
+ raise SystemExit(1)
1227
+ command = shlex.split(lines[0].split(": ", 2)[2])
1228
+ if command[2] != "render" or "--output" not in command:
1229
+ print(f"FAIL stale fix hint: not a render command: {command}", file=sys.stderr)
1230
+ raise SystemExit(1)
1231
+ subprocess.run(command, check=True, capture_output=True)
1232
+ PY
1233
+ python3 "$TOOL" audit --repo "$stale_case" --base "$BASE" \
1234
+ --mapping "$stale_case/specs/mapping.jsonl" --ledger "$stale_case/specs/ledger.md" 2>&1 | grep -q '^audit_ok' || {
1235
+ echo "FAIL stale fix hint: the printed command did not clear STALE_LEDGER" >&2
1236
+ exit 1
1237
+ }
1238
+ echo "PASS stale ledger prints a render command that clears it"
1239
+
1240
+ # The hint is safe only because re-rendering cannot clear a carrier whose text,
1241
+ # structure or qualifier changed: render refuses, or writes a ledger the audit
1242
+ # still rejects with that change's own code.
1243
+ for carrier_case in wrong_parent:CARRIER_CHAIN_MISMATCH table_carrier_to_fence:CARRIER_COMPOSITE_NOT_UNIQUE weaken_modality:QUALIFIER_WEAKENED; do
1244
+ carrier_name="${carrier_case%%:*}"
1245
+ carrier_code="${carrier_case#*:}"
1246
+ carrier_dir="$TMP_ROOT/render_cannot_clear_$carrier_name"
1247
+ git clone -q "$FIXTURE" "$carrier_dir"
1248
+ "mutation_$carrier_name" "$carrier_dir"
1249
+ set +e
1250
+ python3 "$TOOL" render --repo "$carrier_dir" --base "$BASE" \
1251
+ --mapping "$carrier_dir/specs/mapping.jsonl" --output "$carrier_dir/specs/ledger.md" >/dev/null 2>&1
1252
+ carrier_output="$(python3 "$TOOL" audit --repo "$carrier_dir" --base "$BASE" \
1253
+ --mapping "$carrier_dir/specs/mapping.jsonl" --ledger "$carrier_dir/specs/ledger.md" 2>&1)"
1254
+ carrier_status=$?
1255
+ set -e
1256
+ if [ "$carrier_status" -eq 0 ] || ! printf '%s\n' "$carrier_output" | grep -q "^ERROR $carrier_code:"; then
1257
+ echo "FAIL render cannot clear $carrier_name: expected $carrier_code after a render, got: $carrier_output" >&2
1258
+ exit 1
1259
+ fi
1260
+ done
1261
+ echo "PASS re-rendering cannot clear a changed carrier"
1208
1262
  run_mutant invalid_status INVALID_DISPOSITION 'skills/source/SKILL.md#1' "$DELTA_MAPPING" mutation_invalid_status
1209
1263
  run_mutant retired_dead_preserved RETIRED_EFFECT_INVALID 'skills/source/SKILL.md#1' "$DELTA_MAPPING" mutation_retired_preserved
1210
1264
  run_mutant retired_dead_strengthened RETIRED_EFFECT_INVALID 'skills/source/SKILL.md#1' "$DELTA_MAPPING" mutation_retired_dead_strengthened
@@ -4,10 +4,12 @@
4
4
 
5
5
  ## 指令与有效期
6
6
 
7
- 支持原有单独“合并/merge”和“批量合并 N”,另支持完整单行“完成并合并 PR #123 / MR !123”(英文 `finish and merge PR #123`)。原有单次/计数额度仍被任何新消息清除。
7
+ 支持原有单独“合并/merge”和“批量合并 N”,另支持完整单行“完成并合并 PR #123 / MR !123”(英文 `finish and merge PR #123`)。原有单次/计数额度仍被任何新消息清除。宿主的后台任务完成通知也经由提交提示的通道送达,且没有能区分来源的字段,所以同样清除额度(但从不生成额度)。获授权后连续合并时,等 CI 用前台等待(如 `gh pr checks <编号> --watch`),别用后台任务或监视器,免得完成通知在两次合并之间清掉剩余额度。
8
8
 
9
9
  新形式只绑定当前 `origin` 仓库和指定编号,原始 60 分钟内消费一次;单独“继续/继续吧/进度/状态/continue/status/progress”保留原额度和到期时间,“停止/停一下/不要合并/撤销合并授权/stop/pause/cancel merge”撤销,其他消息暂停机械额度,之后“继续”不能恢复。
10
10
 
11
11
  ## 执行命令
12
12
 
13
+ 带 `-h/--help` 的 `glab mr merge` / `gh pr merge` 只打印帮助,闸会拒绝且不消费额度;查看帮助用 `glab help mr merge` 或 `gh help pr merge`。
14
+
13
15
  新形式的执行命令必须是单条直接 `gh pr merge` 或 `glab mr merge`,显式编号;gh 使用 `--repo host/owner/repo` 并指定策略,glab 使用 `--repo https://host/namespace/repo` 并指定 `--auto-merge=false --yes`,可附完整 head SHA,其他参数和 API 形式保持未核验。
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "schema": 1,
3
3
  "npmPackage": "@ccoalm/ccl-skills",
4
- "version": "0.18.8",
5
- "sourceCommit": "32941fa8b6a3d9d98779f18802dea0c21d3b61ff",
4
+ "version": "0.18.10",
5
+ "sourceCommit": "f21b49962fdb23d7a1c3a0376e6872e30517a8c4",
6
6
  "sourceState": "clean",
7
7
  "files": [
8
8
  {
@@ -42,7 +42,7 @@
42
42
  },
43
43
  {
44
44
  "path": "marketplace/plugins/ccl-skills/agent-context/session-policy.md",
45
- "sha256": "a47f7ff077f33da3efaa24fc1d55118330f23ef268e92355d2f281ee56237d2a",
45
+ "sha256": "c76a61265843a87227741291241174856badb3f081c71d87d23643b2a2187b53",
46
46
  "mode": 420
47
47
  },
48
48
  {
@@ -72,7 +72,7 @@
72
72
  },
73
73
  {
74
74
  "path": "marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh",
75
- "sha256": "d12e8c81122babe65d12cca2576b0e2b73b40f4e82090ef31de26bbc893a1cba",
75
+ "sha256": "eeb08a60e67f74428ea2621d4e354a6d88978f41c846ab41974497bf7f838598",
76
76
  "mode": 493
77
77
  },
78
78
  {
@@ -82,7 +82,7 @@
82
82
  },
83
83
  {
84
84
  "path": "marketplace/plugins/ccl-skills/hooks/host-input.py",
85
- "sha256": "b737a74f0a664c52189e43b84539c7ab41e5e4817570b6b1c50c889056c3a9a6",
85
+ "sha256": "058e78598435050d1c0af1ff721d4fee5d049c5a442a71d9ed9fd22b6d70a307",
86
86
  "mode": 420
87
87
  },
88
88
  {
@@ -107,7 +107,7 @@
107
107
  },
108
108
  {
109
109
  "path": "marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh",
110
- "sha256": "c912bd8bbb9afe15b4ed3ac546175ecb9185752087a818d1426ffecd68ffb623",
110
+ "sha256": "363192600575dd33b2eff0cd9da4ca64c3a9f85a3704a9c698f480a26cf67b0d",
111
111
  "mode": 493
112
112
  },
113
113
  {
@@ -172,7 +172,7 @@
172
172
  },
173
173
  {
174
174
  "path": "marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh",
175
- "sha256": "46aa31fb242e9096545d00554cb17eaf9bd4fef68fa4e49ee581148b638e3a08",
175
+ "sha256": "b2f5239ce389a0e4c1e2e7fd6d8f649484a00404c01908ce729ac4eb70f20681",
176
176
  "mode": 493
177
177
  },
178
178
  {
@@ -182,17 +182,17 @@
182
182
  },
183
183
  {
184
184
  "path": "marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh",
185
- "sha256": "acb48c3fd690bd0c7353d2071e1cdd81b28a414d10452a579bf3bb044da1462d",
185
+ "sha256": "5be7052d8df87a0640efe2711d1f11411bd4b622d781fce1dccc411c4dcc0325",
186
186
  "mode": 420
187
187
  },
188
188
  {
189
189
  "path": "marketplace/plugins/ccl-skills/hooks/test_proposed_next.py",
190
- "sha256": "db2df523107fa2b0896c769396e523171f77948526f9d39c8396b598e68fefbf",
190
+ "sha256": "10a73c9577931bef33462cd9388a899a2d79b62255fdea5f32e322ae2a7569b5",
191
191
  "mode": 493
192
192
  },
193
193
  {
194
194
  "path": "marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh",
195
- "sha256": "f57ede7e26be676f3759d97e5707c4f9390623701103f2cfc4465a9d1669abf3",
195
+ "sha256": "4536ae15900b854fe5bf2ed11846b51536ced82a82a9b59d7ea21ba8228fbe47",
196
196
  "mode": 493
197
197
  },
198
198
  {
@@ -347,17 +347,17 @@
347
347
  },
348
348
  {
349
349
  "path": "marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md",
350
- "sha256": "1de7e7f7529d4399738cb1bf9447cb33654bb96443449456fc14d037ca8a94c9",
350
+ "sha256": "d8f1383fc2bc6b0215d445c4cf085a7a9850f21b2eb192c21b6e4375dc60d50c",
351
351
  "mode": 420
352
352
  },
353
353
  {
354
354
  "path": "marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md",
355
- "sha256": "a235a3dfb81dd757ce6e42e82cbc7ec428268ccc290807f5f4bbe07798b33e53",
355
+ "sha256": "569adbd310fa51d58cd14f22a415bee2c492150de5b90024c9df2897b5d5da6f",
356
356
  "mode": 420
357
357
  },
358
358
  {
359
359
  "path": "marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md",
360
- "sha256": "cd54a910c1394f73af5d4eb65039c7a76617af1fd434cdc4b7b498157b41a104",
360
+ "sha256": "0eab60c264c22bc9cc72df55ae0f76da3189aa9204e5397af932edda7eea1e97",
361
361
  "mode": 420
362
362
  },
363
363
  {
@@ -452,7 +452,7 @@
452
452
  },
453
453
  {
454
454
  "path": "marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py",
455
- "sha256": "7ad7115976a82c0c2c2bfa47c91ec1513a1ee09c29e5d1919e5aebd9027e3178",
455
+ "sha256": "c4f4c737f6d4bea424c4211fda17f82254316701f1468f25d501319c6f12652e",
456
456
  "mode": 493
457
457
  },
458
458
  {
@@ -557,7 +557,7 @@
557
557
  },
558
558
  {
559
559
  "path": "marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh",
560
- "sha256": "b04fd9bd2bbd75195e5a3f1ca63aeba9d8a585403079a721c44f0275b0ea0208",
560
+ "sha256": "45858022182bf50fb4ff3d3570415ec46cc15615e5289fccfd0a281d54d24b84",
561
561
  "mode": 493
562
562
  },
563
563
  {
@@ -587,7 +587,7 @@
587
587
  },
588
588
  {
589
589
  "path": "marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md",
590
- "sha256": "220c3283e97e380d79d0cf7edafd2ce63c188571555fb5621e502f01cefe85ba",
590
+ "sha256": "12222d36e7894d8ea767fe99eb0619fd0eff9b250c448ef62c75759dae8a807e",
591
591
  "mode": 420
592
592
  },
593
593
  {
@@ -597,7 +597,7 @@
597
597
  },
598
598
  {
599
599
  "path": "marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md",
600
- "sha256": "2e1be983cdc7dc81d463df302a3bd2c212b197396b463a33254391b0b289cd9a",
600
+ "sha256": "6d6d71442246edf6f36f73db8c94df105c2ce6db995db2feac146b966b561023",
601
601
  "mode": 420
602
602
  },
603
603
  {
@@ -1417,7 +1417,7 @@
1417
1417
  },
1418
1418
  {
1419
1419
  "path": "marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md",
1420
- "sha256": "6c07df6024da34e6494101c85959286049aed52db96567c2819c957566ca6ea2",
1420
+ "sha256": "fba2dd7f079afcc8dfdfd0755cb9cc4b9ae0169d38db07ac42d0d85087ed8345",
1421
1421
  "mode": 420
1422
1422
  },
1423
1423
  {
@@ -1477,7 +1477,7 @@
1477
1477
  },
1478
1478
  {
1479
1479
  "path": "marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md",
1480
- "sha256": "b23c9822452402d26035218ea1e16f8cb3943a2d5812ff6ecd94399f66110644",
1480
+ "sha256": "9f77fe819f70acbaf59b8330ad4ebf29bdba04dc1116e6b5d0ada46d70b14837",
1481
1481
  "mode": 420
1482
1482
  },
1483
1483
  {
@@ -2102,7 +2102,7 @@
2102
2102
  },
2103
2103
  {
2104
2104
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md",
2105
- "sha256": "fbdc6975fcb4ac9aca32e14e0e77f6666b39d6e5491f234760218c1bc83614b6",
2105
+ "sha256": "bbddf828691f1d42adf9a25570368c89be236f683e5c11077c87d838b9f69998",
2106
2106
  "mode": 420
2107
2107
  },
2108
2108
  {
@@ -2177,7 +2177,7 @@
2177
2177
  },
2178
2178
  {
2179
2179
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md",
2180
- "sha256": "3b43de0e81d0a43fb5ba40e97de755bd032f1b5417bbbdd9ca2a4cd2c5c1c65d",
2180
+ "sha256": "9c8428fd26d017b876d627bbc3897652ca2c96385a310b525d3ced7301c677f3",
2181
2181
  "mode": 420
2182
2182
  },
2183
2183
  {
@@ -2207,7 +2207,7 @@
2207
2207
  },
2208
2208
  {
2209
2209
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md",
2210
- "sha256": "1c0e95554f7b8d63860e23385d3b9d84b7d3c1790e4e4071e772b170b4a2b2c6",
2210
+ "sha256": "ec0b2db8f331b31ec13fb96c9c408f86eff37871dbad1a413884a134481e3d5c",
2211
2211
  "mode": 420
2212
2212
  },
2213
2213
  {
@@ -2337,7 +2337,7 @@
2337
2337
  },
2338
2338
  {
2339
2339
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/obligation-ledger.py",
2340
- "sha256": "8bd90fc37d53e466ac5735d2fc956a9d3a31bcd31dd51614ae9e15d435e39726",
2340
+ "sha256": "3f696e807302a86771df57b2668e2b39833ebdea222e0f1ee3f07ee32dea22f2",
2341
2341
  "mode": 420
2342
2342
  },
2343
2343
  {
@@ -2382,7 +2382,7 @@
2382
2382
  },
2383
2383
  {
2384
2384
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh",
2385
- "sha256": "c1c34a27011e77f367a758bb5d6f4abd1ab04d997119f718268343c73027604f",
2385
+ "sha256": "70b058b523df7e40297ff13a8351bee92af594865d51b327b11a5889ab113a31",
2386
2386
  "mode": 493
2387
2387
  },
2388
2388
  {
@@ -2407,7 +2407,7 @@
2407
2407
  },
2408
2408
  {
2409
2409
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh",
2410
- "sha256": "3fac6685fe62945ac8733b58cc4826932cdaf243233d38f44979796dda3ec845",
2410
+ "sha256": "a147e4198b8937b31d1601ff4c897309b0635ca97f0abc5ac4a0c94588d8abfa",
2411
2411
  "mode": 493
2412
2412
  },
2413
2413
  {
@@ -2572,7 +2572,7 @@
2572
2572
  },
2573
2573
  {
2574
2574
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh",
2575
- "sha256": "eb43f10ba2bb507199b6d375cb487bf255914ea84d12f311cfb61a5160128bfc",
2575
+ "sha256": "033b376ff3ba8ca284b9d12fa1d810da5d3a43ebfb3fabce5f7401fbfd0bdae6",
2576
2576
  "mode": 493
2577
2577
  },
2578
2578
  {
@@ -3312,7 +3312,7 @@
3312
3312
  },
3313
3313
  {
3314
3314
  "path": "marketplace/plugins/ccl-skills/skills/worktree-isolation/references/hook-authorization.md",
3315
- "sha256": "989be64bb1d59bdf92c72c4bed407381c25b906e2b436ba8ab601e5590600ce8",
3315
+ "sha256": "115933f3e019e1aed802d65078b5b97a36bf80d35b3a526ba28370d24a992cc6",
3316
3316
  "mode": 420
3317
3317
  },
3318
3318
  {
@@ -3533,5 +3533,5 @@
3533
3533
  "mode": 420
3534
3534
  }
3535
3535
  ],
3536
- "snapshotHash": "9f8142a50f984550552470abdd4c051e69831d649394c5fce0e70da6d59fff8f"
3536
+ "snapshotHash": "083025cdc46f79701f6448e0a83308520bfd69772ced4f6d7a958c7015fd3374"
3537
3537
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ccoalm/ccl-skills",
3
- "version": "0.18.8",
3
+ "version": "0.18.10",
4
4
  "description": "Reusable workflows that help coding agents plan, build, test, review, and release software — for Claude Code, Codex, and OpenCode",
5
5
  "keywords": ["skills", "agent-skills", "claude", "claude-code", "codex", "opencode", "agent", "ai", "ai-agents", "cli", "anthropic", "developer-tools"],
6
6
  "type": "module",