@ccoalm/ccl-skills 0.15.1 → 0.15.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (50) hide show
  1. package/README.md +3 -1
  2. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +1 -1
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +2 -2
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +6 -5
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +26 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +37 -13
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +5 -2
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +77 -5
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +98 -4
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +48 -1
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +230 -16
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +165 -11
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +143 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +572 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +3 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -1
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +1 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +1 -1
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +2 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +3 -1
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +2 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +7 -7
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +1 -1
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +20 -11
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +7 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +1 -1
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +2 -2
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +17 -17
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +4 -4
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +3 -3
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +11 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +83 -48
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +80 -2
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +3 -2
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +190 -2
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +106 -4
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +1 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +3 -1
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -1
  42. package/dist/assets/release.json +54 -49
  43. package/dist/codex-host.d.ts +1 -1
  44. package/dist/codex-host.js +39 -15
  45. package/dist/host-probe.d.ts +16 -0
  46. package/dist/host-probe.js +29 -5
  47. package/dist/operations.js +34 -8
  48. package/dist/unified.d.ts +1 -1
  49. package/dist/unified.js +11 -4
  50. package/package.json +1 -1
@@ -5,6 +5,8 @@ description: Use when implementing, modifying, scaffolding, or testing Node.js b
5
5
 
6
6
  # Node.js Service Development
7
7
 
8
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
+
8
10
  ## Skill Routing
9
11
 
10
12
  - Use this skill for Node.js service implementation: handlers, middleware, adapters, workers, jobs, clients, runtime/toolchain mechanics, and focused tests.
@@ -5,7 +5,9 @@ description: Use when designing, reviewing, debugging, or shipping observability
5
5
 
6
6
  # Platform Observability
7
7
 
8
- This skill owns the **evidence layer**: logs, metrics, traces, log/trace correlation, dashboards, alerts, on-call routing, SLI/SLO/error-budget design, and the framework-level wiring that guarantees a new service is observable on day one.
8
+ Owns logs/metrics/traces, correlation, dashboards, alerts, on-call routing, SLI/SLO/error-budget design and framework wiring for new services.
9
+
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
11
 
10
12
  You do not own:
11
13
  - Traffic routing, retries, timeouts, mTLS, or service mesh policy — go to `platform-service-connectivity`.
@@ -5,7 +5,9 @@ description: 发布 / 灰度 / canary / rollback / rollout / 环境泳道 / prom
5
5
 
6
6
  # Platform Release Engineering
7
7
 
8
- This skill owns **how a change reaches users and how it gets recalled**: lane / environment topology, build-to-deploy pipeline, traffic-shifting strategies (canary, blue-green, mirror), promotion gates that consume observability evidence, secret + dynamic-config distribution, and the rollback contract.
8
+ Owns lane/environment topology, build/deploy pipelines, traffic shifting (canary, blue-green, mirror), promotion gates using observability evidence, secret/dynamic-config distribution and rollback contracts.
9
+
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
11
 
10
12
  You do not own:
11
13
  - What signals exist to judge a release — see `platform-observability`.
@@ -5,6 +5,8 @@ description: 服务互通 / service mesh / service discovery / mTLS / retry / ti
5
5
 
6
6
  # Platform Service Connectivity
7
7
 
8
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
+
8
10
  This skill owns **how requests move between services**: the transport layer (mesh), service discovery, multi-environment routing, retry/timeout/circuit-breaker policy, and the framework middleware that propagates app-level context across hops.
9
11
 
10
12
  You do not own:
@@ -9,7 +9,7 @@ Use this skill as the top-level workflow for new product development, feature de
9
9
 
10
10
  **Entry precedence.** For any product idea, feature delivery, release, cross-cutting refactor, or a **restart/redo of an in-flight delivery**, invoke this workflow first to classify and route — naming a stack/execution skill (e.g. `web-react-dev`, `multi-agent-delegation`) does not by itself skip this workflow's lifecycle gates (design / test / release / acceptance); those still apply unless already covered.
11
11
 
12
- **Continuation-proposal output contract (session-wide for product delivery).** Every assistant message in a delivery session routed by or through this workflow carries exactly one literal line until the user explicitly ends/pauses the delivery or changes scope: `proposed-next: <action and scope>` when the message has imperative/future/next-step wording or a proposed action, otherwise `proposed-next: none — status only`. At the start of every subsequent user turn, read that line before interpreting the reply: absence, multiplicity, or a marker/wording conflict enters `blocked:`/`interim` by default, never `not-applicable`. Coverage detail and the host-layer caveat: `references/pre-final-continuation-gate.md` (Continuation-proposal output contract).
12
+ **Continuation-proposal output contract (session-wide for product delivery).** Use `proposed-next: <action and scope>` to make a next action observable, or `proposed-next: none — status only` for a status-only handoff. A missing, repeated, or conflicting marker triggers intent recovery, not a stop. Follow the current explicit user request; bind short assent to one recoverable concrete proposal and its scope. Repair your own formatting without asking the user to repeat an already-clear instruction. Markers never grant authority. Recovery, ambiguity, and outcome rules: `references/pre-final-continuation-gate.md` (Continuation-proposal output contract).
13
13
 
14
14
  - **Restart/redo of an in-flight delivery** (清除代码重新开发 / 完全重新开始 / 推倒重来 / redo-from-scratch): a restart is a fresh delivery entry; mid-delivery coding momentum is NOT a license to skip re-classification and re-enter. Re-entry means re-ESTABLISH the plan and develop against it, not code from memory — "丢脚手架 / 重来" defaults to discarding CODE, not the design/spec artifact; discarding the design/spec itself needs an explicit user opt-out after clarification (an explicit instruction to drop the design always wins). Recovery mechanics and deviation recording live in `references/implementation-entry-reentry-gate.md` §Baseline Selection.
15
15
  - **Implementation entry / re-entry gate (active plan/spec required by default).** For product R&D deliveries that stay in this workflow, "start development" means first establish the current executable artifact set, then code against it; requests routed straight to another owning skill by the *Go straight to the owning skill* bullet below use that owner's entry rules instead. Use existing specs, implementation plans, assessment reports, issue/MR descriptions, or repo-local task docs only after reading them back or citing artifacts just produced in the active session, then checking freshness, scope, owner skills, acceptance checks, tests, stop conditions, and **landing state** (`local status`, `MR-ready`, `landed`, `release-ready`, or `shared-status-ready`). Full mechanics for every case below live in `references/implementation-entry-reentry-gate.md`.
@@ -164,7 +164,7 @@ At each stage boundary, walk the per-stage entry-state enumeration in [Stage-Ent
164
164
  - High-risk workflows cannot be accepted by happy-path tests alone. Require a risk scenario matrix and replayable incident drills for the relevant classes: duplicate money/quota/write side effects, permission uncertainty, tenant/user data isolation, AI provider/model failure, user repeated submission or unclear final state, and traceable incident explanation.
165
165
  - For UI backed by APIs or generated content, require `testing-strategy` to produce evidence that covers rendered states, contract/error handling, and one real visible flow where feasible. Do not let ideal mocked data stand in for runtime integration evidence.
166
166
  - When live infrastructure is required, keep it explicit and separate from default fast tests.
167
- - Treat review status as an explicit artifact. If an independent review is required but times out, returns empty output, or is otherwise inconclusive, record it as pending and do not describe it as passed.
167
+ - After code/test edits, self-check and invoke `code-review` automatically before completion. Persist review status; timeout, empty output or inconclusive results remain pending, never passed.
168
168
  - For independent review runs, prefer bounded diff/file input over broad repo prompts. Load the repository's local rules such as `AGENTS.md` or equivalent review context when available, skip generated/docs noise unless it is the review target, and make timeout/inconclusive results recoverable through a durable pending record.
169
169
  - For product/spec normalization, standards-to-health-gate work, or any change that edits a shared deterministic gate/verifier (workspace verifier, conformance script, contract-coverage gate, status-source validator, CI harness, continuation-state checker, or a cross-repo contract/status/version/release/compatibility coordination surface), this workflow classifies the artifact before implementation — `spec/plan`, `gate design`, `gate implementation`, `status sync`, or `runtime/code` — without delegating that decision (do not delegate the spec-vs-plan-vs-code decision to `feature-risk-router`), then routes every shared-gate change through `feature-risk-router` and applies its `shared-gate` decision before shared branch push or MR merge; a recorded independent adversarial review names concrete objections, their disposition, and the reviewer/tool identity — prefer the session's review/challenge skill, otherwise a ccl-owned independent review (external tools supplement only; same-agent inline prose review only for explicitly low-risk, non-cross-boundary work). Rule/scope/failure/completion semantics changes require a concrete repo-local persistent artifact before editing; a `gate implementation` runs the plan/status verifier(s) before implementation and before claiming the plan active — an explicit status-source validator takes precedence, otherwise run every authoritative non-alias verifier or record why each is not applicable; a verifier gap or unavailable required review/challenge stays `interim` / pending-review. Details live in `references/shared-gate-artifact-classification.md`.
170
170
  - Do not use landing labels without matching evidence. `landed` requires the relevant local commit or persisted artifact; `MR-ready` requires branch, push, review artifact, known CI/pipeline status when applicable, known mergeability when applicable, and review status that matches reality; `release-ready` requires the relevant release checks, rollback/mitigation, and runtime verification evidence; `shared-status-ready` requires the owning status or product document to match the real branch/MR/review/verification state. For local-only or exploratory slices, report the actual uncommitted/unpushed state and use a local status label instead of treating MR evidence as mandatory.
@@ -187,16 +187,16 @@ At each stage boundary, walk the per-stage entry-state enumeration in [Stage-Ent
187
187
 
188
188
  ### Pre-Final Continuation Gate
189
189
 
190
- Run this gate before finalizing a product R&D turn after any delivery slice lands. The session-wide continuation-proposal output contract above creates a second, independent trigger at the start of every subsequent user turn in that delivery session when either (a) the immediately preceding assistant message carries an action-form `proposed-next:` and the user replies, (b) its marker is absent/multiple or `none` conflicts with imperative/future/next-step wording, or (c) a user reply reads as affirmative/permissive toward an explicit assistant-proposed next action. Paths (a) and (b) are unconditional literal/fail-closed checks. Before any further action or final response, visibly emit exactly one of `continuing: <action and scope>` or `blocked: <proposed action and scope> — <specific stop, missing authority, or ambiguity>`; emitting neither or both is invalid. Path-(c) examples and the full outcome contract: `references/pre-final-continuation-gate.md` (Gate triggers and outcome contract).
190
+ Run this gate before finalizing a product R&D turn after any delivery slice lands, on every user reply immediately following an assistant message that states or implies a next action, and on any explicit continuation request. The reply need not first be classified as assent. Recover the intended action from the current request and relevant conversation; a malformed or absent marker alone never selects `blocked:`. Report `continuing: <action and scope>` for the work you will execute, or `blocked: <action and scope> — <specific blocker>` when no authorized work can proceed. A blocked dependent action may remain pending while a different authorized action continues. Apply landing checks only to actual landing claims; an already-authorized local investigation does not require inventing a prior landing. Details: `references/pre-final-continuation-gate.md` (Gate triggers and outcome contract).
191
191
 
192
192
  1. Confirm the landing state from real evidence (local branch, remote sync, MR/review artifact, CI/pipeline when applicable, review/challenge status when required, status-doc sync, dirty worktree), proving the landing before reading any document (for the remote-backed default, fetch/update the target ref from its remote immediately before classifying the slice landed) per `references/pre-final-continuation-gate.md` (Landing-state proof); content/tree/patch equivalence never by itself proves a slice landed.
193
193
  2. Inspect the current product/status source of truth, issue list, repo-local next-step artifact, unresolved acceptance item, or direct user continuation instruction for the next implied slice. **Reconcile it against the current branch/MR/merge/CI/tag state from step 1 before deriving: contradicting reality means stale — stop, repair the status source first, and do NOT derive from the stale source or a git-log/grep scan** (`references/pre-final-continuation-gate.md` §Status-source reconciliation).
194
194
  - **Deferred-evidence continuation check (`DFE-CONT`).** When real/runtime evidence is due (named by an acceptance item, status source, landing-evidence row, required gate, user correction, or because it is the behavior's only meaningful proof) yet deferred, blocked after remediation, skipped at finalization, or replaced by local/mock verification. Report deferred real evidence as `interim`/outstanding; do NOT report the turn complete while it is outstanding. A local/mock substitution is terminal only when a cited **non-agent** anchor — **agent-authored or agent-co-edited status/router/gate/handoff text never satisfies this** — names the same evidence, declares the deferral terminal, and carries the outstanding command/source forward for the active slice/ref. Never add verifier/config/test hardening motivated only by missing deferred evidence; never auto-continue past the pending gate. **Load `references/pre-final-continuation-gate.md` before treating any deferral as terminal** — it owns the valid/invalid-anchor list and hardening boundary.
195
- - **Affirmative-assent binding rule** lives in `references/pre-final-continuation-gate.md` §Assent binding — **load it before selecting `continuing:` on any assent**, and any concrete next-slice proposal you issue must itself carry the `proposed-next:` marker or a later assent cannot bind — an unmarked referent is ambiguous, never self-cleared; the rule fires only when the immediately preceding assistant message itself states one concrete next action and its scope, and `continuing:` binds to that proposal, never to adjacent status or response-format prose; ambiguous assent, referent, or authority ⇒ `blocked:` with step-4 precedence — restate the proposed action/scope plus the specific ambiguity/authority, cite the step-1 evidence and ask one concise question in the same turn; self-classifying the reply or marker away is never an exit, and the `continuing:` default applies only when assent is unambiguous and no step-4 condition holds (the full rule and its fallback are stated there); the visible `continuing:`/`blocked:` outcome obligation is unchanged.
196
- 3. Continue automatically only when no step-4 stop condition fires and: the next slice comes from an explicit status/task/acceptance source or active user continuation, is low-risk, local-only/already-authenticated, in accepted scope, clearly owned, verifiable with existing commands, and needs no destructive action, external purchase/financial commitment, production access, legal/compliance/product-strategy decision, or high-impact architecture choice. Existing configured internal developer-self-use metered model/tool accounts aren't an external purchase here.
197
- 4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, a failed/pending/inconclusive required/blocking gate, dirty/conflicting worktree that can't be isolated, required environment unavailable after remediation, high-impact product/architecture/compliance decision, destructive action, external purchase/financial commitment, unclear owner, ambiguous assent, missing stricter authorization, materially differing viable approaches (none dominant-and-reversible), a fix lacking evidenced cause, or no low-risk slice. Exactly one dominant reversible approach and no other stop condition firing: do not stop at a recommendation: deliver a tested reviewable draft.
195
+ - **Affirmative-assent binding rule** lives in `references/pre-final-continuation-gate.md` §Assent binding — load it when recovering a short reply. Bind to the current explicit request or one recoverable concrete proposal, including an unmarked proposal; preserve its scope and existing authority. Ask only if action, scope, or required authority remains unresolved after recovery. A status remark or output marker cannot substitute for a proposal or permission; self-classifying the reply or marker away is never an exit from carrying out an already-clear request.
196
+ 3. Continue automatically with a clearly owned, verifiable, low-risk next slice from an explicit task/status/acceptance source or active user continuation, within accepted scope and existing authority. Apply the eligibility and stop conditions in `references/pre-final-continuation-gate.md`. Necessary fixes, tests and review inherit task authorization; a reviewer-budget flag triggers a method checkpoint and cumulative-history record, not renewed permission. Explicit user limits still govern. Existing configured internal developer-self-use metered model/tool accounts aren't an external purchase here.
197
+ 4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, or a concrete blocker for the affected action. Block materially differing viable approaches (none dominant-and-reversible) and a fix lacking evidenced cause; load `references/pre-final-continuation-gate.md` for the full stop conditions. **Scope each blocker to its dependent action or claim.** An unproven cause blocks the speculative patch, not available diagnosis; a pending gate blocks dependent landing/completion, not authorized remediation or independent work. Before ending, perform available in-scope diagnosis, owner discovery, remediation, live-handle monitoring or independent work. Quality-gate failures require diagnosis and available related behavior-preserving cleanup before escalation; preserve readability and compatibility, never game counters (`references/refactoring-discipline.md`). Never bypass the blocked gate, invent a pass, widen scope, or substitute unrelated hardening. Stop the task only at the user's stop/status-only instruction or when no safe authorized work remains. With one dominant reversible approach and no applicable stop condition, do not stop at a recommendation: deliver a tested reviewable draft.
198
198
  5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form and classifies the turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
199
- 6. **Assent-outcome closeout check.** Before every final response in a product-delivery session, walk the literal immediately preceding marker and visible outcome; the agent cannot exclude a status, question, review, or dispatched-owner turn by reclassifying it outside the session. An action marker or a plausibly affirmative user reply requires exactly one already-visible `continuing:`/`blocked:` outcome; until `continuing:` has executed the accepted slice or `blocked:` has named the blocker, the current turn may not use `proposed-next: none — status only` or `not-applicable`. A valid status-only marker permits `not-applicable` only when the preceding prose has no imperative/future/next-step wording and no affirmative reply pending. A missing/multiple/conflicting marker forces a visible `blocked:`/`interim` outcome with one clarifying question; it never produces `not-applicable`. The current assistant message itself must end with exactly one action-form or status-only marker for the next turn. Omitting the marker cannot justify a silent stop.
199
+ 6. **Assent-outcome closeout check.** Every user reply immediately following an assistant message that states or implies a next action requires a visible `continuing:` or `blocked:` outcome before finalizing, even if the reply is not classified as assent; every explicit continuation request does too. Missing markers never waive it. Reconcile the current request, original proposal, scope/authority changes, tool/output evidence, and remaining blockers. Respect a current explicit stop or status-only request; name that reason in the blocked outcome without executing the prior proposal. Otherwise `continuing:` must be followed by execution in the same turn; a promised next step is not execution. If part remains blocked, report its pending state and independent work performed. A status-only handoff cannot discharge an unexecuted accepted action. Repair marker formatting; for short assent, if the original proposal cannot be recovered verbatim, select `blocked:` and ask. Formatting never requires clarification. Do not silently drop an accepted action or claim a pending gate passed.
200
200
 
201
201
  If a user later challenges "why did you stop" or "was the rule too weak", treat it as a product workflow defect: route through `skill-extraction-workflow`, strengthen the smallest owning skill or validation checklist, validate the diff, and only then claim the process issue is solved.
202
202
 
@@ -10,7 +10,7 @@ a triggered diff, and whenever the candidate diff changes after a review.
10
10
 
11
11
  ## Recorded review artifact
12
12
 
13
- (a) a recorded independent adversarial review is the gate for all triggered work — prefer an available review/challenge skill discovered in the session when suitable, otherwise a ccl-owned independent review (the external skill supplements, it is not itself the required gate); save an artifact naming concrete objections, their disposition, and the reviewer or tool identity; same-agent inline prose review is acceptable only for explicitly low-risk, non-cross-boundary work.
13
+ (a) a recorded independent adversarial review is the gate for all triggered work — prefer an available review/challenge skill discovered in the session when suitable, otherwise a ccl-owned independent review (the external skill supplements, it is not itself the required gate); save an artifact naming concrete objections, their disposition, and the reviewer or tool identity; same-agent inline prose review is acceptable only for explicitly low-risk, non-cross-boundary design-only work with no implementation diff. Once code or executable tests change, invoke `code-review` automatically under its development-completion rule; green tests or low risk do not replace that invocation.
14
14
 
15
15
  ## Binds to the implementation diff
16
16
 
@@ -85,21 +85,30 @@ If over-polishing around deferred evidence recurs, or the user flags it as a reu
85
85
 
86
86
  ## Continuation-proposal output contract (session-wide coverage)
87
87
 
88
- The entrypoint defines the contract itself (exactly one literal `proposed-next:` line per assistant message, in one of its two forms). The coverage detail:
88
+ The entrypoint uses `proposed-next:` as an observable handoff, not an authorization token. The contract follows the delivery across status answers, questions, reviews, and dispatched owners, before or after a slice lands. Progress messages need no ritual marker to preserve a user's request.
89
89
 
90
- - It binds every assistant message in a delivery session routed by or through the workflow — including "go straight to the owning skill" paths: status answers, questions, review reports, and work dispatched directly to implementation, diagnosis, or another owning skill. Neither the router nor the executor may classify its own turn out of the contract; if it is unclear whether the workflow was invoked, default to carrying it.
91
- - It applies before any slice lands and is not conditional on entering the Pre-Final Continuation Gate.
92
- - At the start of every subsequent user turn, read that literal line before interpreting the reply: an action marker enters the gate's continuing/blocked classification; absence/multiplicity enters `blocked:`/`interim` by default; a `none` marker that conflicts with proposal wording does the same.
93
- - This is an observable prose contract, not a host-enforced hook; a task that never loads this workflow cannot be mechanically controlled by this text, so do not claim it prevents that host-level omission.
90
+ At the start of the next turn, recover intent in this order:
91
+
92
+ 1. Follow the current explicit user instruction, including a correction, changed scope, stop, or status-only request.
93
+ 2. For short assent, read back the original wording of the most recent still-active concrete proposal and check that later messages or task state have not withdrawn or superseded its action, scope, or authority. Quote that original proposal when stating the recovered action and scope; a summary or paraphrase alone cannot bind short assent. If recovery adds an action or broadens that quoted scope, select `blocked:` and ask. One recoverable action can bind with or without a marker; a stale, repeated, or conflicting marker is an assistant formatting defect to repair.
94
+ 3. If materially different proposals remain unresolved, or scope/authority is still unclear, ask one targeted question about that uncertainty. Do not ask the user to repair a marker or repeat a clear instruction. A marker alone never supplies missing authority.
95
+
96
+ Use the active owner's entry and safety gates for the recovered action. An authorized task includes necessary fixes, tests and review by default; neither a router nor a dispatched owner may discard that authority by relabeling its turn or exhausting an internal review sequence. Apply the owning review checkpoint and record `continuation_basis=existing-task-scope` with cumulative history in the caller-owned task artifact, not runtime JSON. Legacy `human_decision_required` / `continuation_authorization_required` values first require checking existing authority, not asking again. Explicit user cost, round-count and stop limits prevail; new scope, missing authority or real tradeoffs need a decision. Continuation grants no merge, publication or waiver authority. This is a prose contract, not a host-enforced hook; it cannot prove compliance by a task that never loads it.
94
97
 
95
98
  ## Gate triggers and outcome contract
96
99
 
97
- The session-wide contract creates a second, independent trigger at the start of every subsequent user turn in that delivery session when either (a) the immediately preceding assistant message carries an action-form `proposed-next:` and the user replies, (b) its marker is absent/multiple or `none` conflicts with imperative/future/next-step wording, or (c) a user reply reads as affirmative/permissive toward an explicit assistant-proposed next action — regardless of whether a slice landed or the turn is being finalized. Paths (a) and (b) are unconditional literal/fail-closed checks; marker absence enters path (b) without first asking the agent to recognize why it was omitted.
100
+ An eligible next slice comes from an explicit status/task/acceptance source or active user continuation, is low-risk, local-only/already-authenticated, in accepted scope, clearly owned and verifiable with existing commands. It needs no destructive action, external purchase/financial commitment, production access, legal/compliance/product-strategy decision or high-impact architecture choice. Existing configured internal developer-self-use metered model/tool accounts are not an external purchase. Apply the following conditions to each action.
101
+
102
+ Action-scoped stop conditions are: an explicit stop/pause instruction; a user-requested status-only answer; a failed, pending or inconclusive required gate; a dirty/conflicting worktree that cannot be isolated; a required environment unavailable after remediation; a high-impact product, architecture or compliance decision; a destructive action; an external purchase or financial commitment; unclear ownership; ambiguous assent; missing stricter authorization; materially different viable approaches with none dominant and reversible; a speculative fix without evidenced cause; or no low-risk slice. Apply each condition to the affected action, then check for available authorized diagnosis or remediation before stopping the whole task.
103
+
104
+ Check continuation on every user reply immediately following assistant prose that states or implies a next action, and on any explicit continuation request, regardless of landing status. Do not first require classifying the reply as assent; visibly report the continuing or blocked outcome even when the reply changes scope or stops the proposed action. Short replies include `ok`, `yes`, `可以`, `好`, `继续`, `proceed`, `do it`, `go ahead`, and `👍`; interpret them against the recovered action rather than formatting alone.
98
105
 
99
- - Path (c) examples include, but are not limited to, `ok`, `okay`, `yes`, `sure`, `可以`, `好`, `行`, `按这个来`, `继续`, `proceed`, `do it`, `go ahead`, or `👍`; any plausibly affirmative reply enters unless it explicitly says stop/pause.
100
- - Select `continuing:` only when the reply is unambiguously affirmative and the marker or immediately preceding message contains exactly one concrete action and scope.
101
- - An explicit stop/pause selects `blocked:` with that reason; a mixed/ambiguous reply, invalid/conflicting marker, or bare emoji/interjection after status-mixed or multi-proposal prose also enters but must select `blocked:`.
102
- - Before any further action or final response, visibly emit exactly one of `continuing: <action and scope>` or `blocked: <proposed action and scope> — <specific stop, missing authority, or ambiguity>`; emitting neither or both is invalid.
106
+ - Select `continuing: <action and scope>` when that action is clear and authorized, then execute it in the same turn. A tool call and its result or a produced artifact establish execution; the label alone does not.
107
+ - A blocked patch, review, or landing does not block every action. Keep that dependent action/claim pending while continuing available diagnosis, bounded remediation, monitoring of the existing live handle, or independent accepted work. These paths retain their own scope and permission checks; they cannot bypass the blocked gate or substitute unrelated hardening for missing evidence.
108
+ - A failed quality gate calls for a repair that preserves its purpose. Before asking the user to choose a workaround, inspect and perform a safe structural cleanup related to the current change when available, then rerun the gate and affected tests. Follow [refactoring discipline](refactoring-discipline.md#responding-to-quality-gates): preserve behavior, compatibility and readability; do not shrink identifiers or necessary comments, weaken a baseline or rewrite history solely to make the counter pass. If no safe in-scope repair remains, report the evidence and the actual decision needed.
109
+ - Independent work must neither depend on the pending verdict nor modify the candidate being evaluated. Name the pending gate and the independence basis when continuing. A candidate-changing fix is remediation, not independent work: let the existing run reach a terminal state, then refresh affected evidence and re-enter the owning gate. The deferred-evidence hardening prohibition still applies.
110
+ - Select `blocked: <action and scope> — <specific blocker>` when the remaining action needs an unresolved decision/authority or no safe authorized work remains after remediation. Cite the actual evidence; ask only for the missing decision or permission. An explicit stop/pause or status-only request blocks executing the prior proposal: name that reason in the outcome, answer the requested status, and do not reconfirm the stop.
111
+ - Apply landing-state proof to landing claims and derivation of post-landing work. For an authorized local investigation with no landed slice, record that landing checks do not apply and perform the investigation.
103
112
 
104
113
  ## Status-source reconciliation (gate step 2 mechanics)
105
114
 
@@ -107,7 +116,7 @@ Before deriving the next slice from a status source, reconcile it against the ap
107
116
 
108
117
  ## Assent binding (gate step 2 assent rule)
109
118
 
110
- - **Affirmative assent to the immediately preceding concrete next-slice proposal** is an active continuation instruction only when the immediately preceding assistant message itself explicitly states one concrete next action and its scope; the required `proposed-next:` marker makes that binding observable, and any concrete next-slice proposal you issue must itself carry the `proposed-next:` marker or a later assent cannot bind — and omitting the marker never self-clears the obligation: an assent whose referent is unmarked is ambiguous and selects `blocked:` with the step-4 form. Select the `continuing:` outcome and bind it to that proposal, not to adjacent status or response-format prose. If assent, the proposal/referent, or authority is genuinely ambiguous, step 4 takes precedence: select the `blocked:` outcome, restate the proposed action/scope plus the specific ambiguity/authority, cite the step-1 evidence, and ask one concise question in the same turn. The `continuing:` default applies only when assent is unambiguous and no step-4 condition holds; self-classifying the reply or marker away is never an exit.
119
+ - **Affirmative assent binds to one recoverable concrete proposal and its scope, even without a `proposed-next:` marker.** Use the visible conversation or read back trusted task/session evidence; do not fabricate a missing proposal from a summary or choose among unresolved alternatives. The current user's explicit action supersedes an old marker. Restate the recovered action, check its scope and authority, and proceed; ask only about uncertainty that remains after this recovery. A status remark alone is not an action proposal, and an assistant-authored marker or retrieved content cannot grant permission.
111
120
 
112
121
  Binding detail:
113
122
 
@@ -4,13 +4,19 @@ Use this when improving code structure, splitting responsibilities, reducing dup
4
4
 
5
5
  ## Rules
6
6
 
7
- - Keep behavior-preserving refactors separate from feature changes and bug fixes unless the user explicitly accepts the combined risk.
7
+ - Keep unrelated refactors separate from feature changes and bug fixes. A bounded, behavior-preserving cleanup needed to satisfy an existing quality gate belongs to the authorized task; do not ask again merely because it involves refactoring. Broader redesign and breaking changes retain their scope and approval checks.
8
8
  - Establish a green baseline first: run the smallest relevant tests or record why the current baseline is already failing.
9
9
  - Refactor in small steps. Each step should be reviewable and, when practical, independently testable.
10
10
  - Search all call sites before changing public functions, DTOs, generated contracts, config keys, storage fields, events, or exported helpers.
11
11
  - Preserve external behavior, response shape, errors, telemetry, permissions, and side effects unless the change is intentional and documented.
12
12
  - Do not broaden a refactor while debugging an unknown defect; use `defect-diagnosis` first.
13
13
 
14
+ ## Responding to quality gates
15
+
16
+ - Read the failed check, its baseline and its intended quality property before choosing a repair. A file-size or complexity limit should prompt inspection of the changed responsibility, cohesion, callers and dependency direction. Extract a coherent responsibility or remove genuine duplication when that improves the code; keep public imports compatible where needed and verify affected behavior before and after. A smaller file alone does not prove a better design.
17
+ - Do not abbreviate meaningful names, remove necessary explanations, pack statements, fragment responsibilities arbitrarily, or change the threshold/history just to satisfy a counter. A gate with an evidenced defect can be diagnosed and corrected under its owning contract; that is distinct from evading a valid failure.
18
+ - Perform available, in-scope remediation and rerun the failed check before handing the problem back. Ask only for a remaining material tradeoff or missing authority after this work. Force-pushing, waiving the gate and accepting lower readability are not substitutes for inspecting a safe structural repair; a failed gate grants none of those permissions.
19
+
14
20
  ## Impact Analysis
15
21
 
16
22
  Before changing a shared shape or function, identify:
@@ -5,7 +5,7 @@ description: 用 Python 写接口 / FastAPI / Django model / Celery 任务 / pyt
5
5
 
6
6
  # Python Service Dev
7
7
 
8
- Use this for implementation of Python backend products, services, microservices, AI-service hosts, workers, packages, and batch tools. For new backend products, implement the smallest deployable or package shape justified by ownership, data boundary, runtime isolation, scaling, release cadence, and rollback needs. It should adapt to the repository in front of you, but the workflow is independent of any prior codebase.
8
+ Use this for implementation of Python backend products, services, microservices, AI-service hosts, workers, packages, and batch tools. For new backend products, implement the smallest deployable or package shape justified by ownership, data boundary, runtime isolation, scaling, release cadence, and rollback needs. It should adapt to the repository in front of you, but the workflow is independent of any prior codebase. After code/test edits, self-check and invoke `code-review` automatically before completion.
9
9
 
10
10
  ## Skill Routing
11
11
 
@@ -5,7 +5,7 @@ description: 复盘 / 沉淀 / 总结经验 / 补进技能 / 技能缺陷 / 流
5
5
 
6
6
  # Skill Extraction Workflow
7
7
 
8
- Use this skill to turn observed experience into durable agent skills without copying business-specific codebase details. It complements public skill-authoring guidance such as `writing-skills` and `skill-creator`: those define skill format and authoring discipline; this skill defines how to mine, filter, generalize, validate, and land reusable CCL skills.
8
+ Turn observed experience into reusable skills without business-specific details. `writing-skills` and `skill-creator` define format and authoring discipline; this skill covers mining, filtering, generalizing, validating and landing. When shared-skill changes are ready, invoke `code-review` automatically before completion. Apply `code-review/references/development-completion.md` and this owner's required review/challenge gate.
9
9
 
10
10
  ## Start here (30 秒定位)
11
11
 
@@ -137,7 +137,7 @@ Use this skill to turn observed experience into durable agent skills without cop
137
137
  - The mechanical backstops are (a) the closeout gate — a committed skill-change with neither a visible in-session `skill-extraction-workflow` invocation nor the round's durable charter/target-output record is `interim` — and (b) **user-signal escalation**: you generally cannot self-count misses you did not notice, so a user-pointed-out under-trigger is a recurrence check across the whole session (even other tasks) and, on recurrence, escalates to tightening the always-on discipline rather than landing another narrow per-case trigger.
138
138
  - **Firing-point-placement corollary:** when the SAME meta-class (a precise gate walked past at the routing → pre-code/design transition) recurs at a *new* lifecycle sub-point despite prior bootstrap-salience + the closeout gate, the durable lever is **moving the owning gate's firing point ONTO the transition itself** (pre-substance-draft AND pre-first-impl-edit) and sharpening *name→invoke* at the SAME transition — naming/knowing an owner is NOT invoking/loading it, and a named-but-unloaded owner's mechanical rules never fire — NOT another bootstrap/per-case bullet or more prose. **Record-field corollary (the forgery surface):** any field that NAMES an owner is fillable without invoking that owner, and filling it is what *feels* like discharging the gate, so it carries an explicit invoke bar on its triggered values. The self-detect firing point's authority boundary and observed shape, the output-shape and option-set siblings, the invoke-bar set-diff mechanics, the worked recurrence-chain, and the landed owner-dispatch implementation: `references/firing-point-placement.md`.
139
139
  - **Run your own adversary to convergence BEFORE any "done / fixed / passing / covered / converged / complete" claim — your own such claim is the least-trustworthy thing you emit.** For any non-trivial completion/coverage/convergence claim, you must have already run — **yourself, not deferred to the user** — the verification or adversarial pass that would catch its failure, to a **clean fresh result** (a first clean pass on the current candidate, never a "confirm my fix" pass), OR **downgrade the claim to `interim` and name what you ran vs. didn't**. "Covered / converged / already handled" is a claim, not a status — back it with firing-path or clean-pass evidence or do not emit it; this self-adversary duty never narrows the mandatory dual-track challenge. That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified; **a mutation you did not APPLY is a hypothesis, not evidence** (bound its blast radius; where no isolated path exists record the property `unverified`); **prove the oracle can fail before trusting its clean verdict**; **a failing anchor is first a question about the ANCHOR, not a verdict on the implementation**; a clean run is reported with the dimensions it crossed, walked before values; a property with no contradicting observation is `unverified`, never counted as audited. A scoped "X verified; Y not run" is an interim checkpoint, **not** `done`/`complete`/`landed`, unless a **risk owner — the user/maintainer, never the agent self-accepting — explicitly accepts the gap AND it is tracked to that owner** (agent self-labeling "risk accepted" or "deferred" does not qualify; scoping is a downgrade, never a license to call the narrowed slice done). **Recurrence signal:** a user prompting you to keep digging / disputing a "covered/converged/done" is a premature-completion signal — on the **2nd** such correction in a session (even across different tasks) escalate to tightening this discipline, not just fixing the one case (per the repeated-correction escalation above). The full method — mutation enumeration, the applied-mutation discipline and its blast-radius bound, independent-oracle validation, the dimension walk (`testing-strategy` owns the axis list), re-owe-after-fixes, graded-verdict calibration, the failing-anchor section, and the recognition-dependent honesty caveat: `references/dual-track-review-gate.md` §Self-audit.
140
- - Automatically trigger durable learning when extraction work exposes a reusable failure — **and when ordinary delivery work does, capture it here too, but without extraction taking over the delivery**: let the active owner (`product-rd-workflow` / `defect-diagnosis` / `testing-strategy` / …) handle the immediate work first, then route the durable lesson here. **For a premature-stop correction after affirmative continuation**, immediate recovery means first rerun the active owner's current continuation/blocking gate in full (for product R&D, Pre-Final Continuation Gate steps 1–6) against current state, then follow its observable outcome — proceeding only when a literal binding exists (the original proposed-next action/scope plus literal assent preserved verbatim, or the user's correction literally naming the paused action and scope — never reconstructed, broadened, or substituted, and never copied into a shared repository record); a `blocked:` recovery without the step-1 evidence and a specific missing authority/ambiguity is invalid; asking again is required when neither path binds, the user intervened, or scope/gates changed; and neither correction RCA nor stale assent may extend a still-authorized delivery or slip past a gate that is newly pending or inconclusive. After delivery recovery, correction RCA plus the durable prevention landing and verification are still due before the turn can be reported complete; otherwise report `interim`. What counts as a binding (the two paths, and why a compaction paraphrase or a bare "why did you stop" is not one), the `continuing:`-line form, and the invalid-`blocked:`-recovery rule: `references/resume-paused-delivery.md`.
140
+ - Automatically trigger durable learning when extraction work exposes a reusable failure — **and when ordinary delivery work does, capture it here too, but without extraction taking over the delivery**: let the active owner (`product-rd-workflow` / `defect-diagnosis` / `testing-strategy` / …) handle the immediate work first, then route the durable lesson here. **For a premature-stop correction after affirmative continuation**, rerun the active owner's current continuation/blocking gate against current state (for product R&D, Pre-Final Continuation Gate steps 1–6, with landing checks only where applicable). Proceed when a literal binding exists: one recoverable concrete proposal and scope plus the user's assent, even without a marker, or the current user explicitly naming the paused action and scope. Never fabricate or broaden that binding, or copy real conversation text into shared records. Ask only if action, scope, or required authority remains unresolved after recovery; a new user message or changed gate requires reassessment, not automatic reconfirmation. Keep newly blocked dependent actions pending while continuing available authorized diagnosis, remediation, or independent work. Correction RCA and extraction must not delay that recovery or bypass a gate. After delivery recovery, correction RCA plus durable prevention and verification are still due before claiming this process defect complete; otherwise report `interim`. Binding paths, trusted evidence, and invalid-`blocked:` recovery: `references/resume-paused-delivery.md`.
141
141
  - The trigger is a correction about *reusable skill/process behavior*, NOT every bug/QA/review nit handled inside its own owner skill. Do not wait for the user to say "沉淀": if the user points out a missed source, missed sibling skill, shallow rule, overclaim, domain leakage, missing trigger, missing verification, repeated correction, **or that this workflow should have been invoked at all (an under-trigger / "should you have used 提炼/复盘" correction, including outside an active extraction)**, run correction RCA, update the smallest owning skill/reference/validator, and verify the prevention point before finalizing the turn.
142
142
  - When the user asks whether a lesson was durably landed after a failed extraction, verify the actual skill diff or file content first. Do not answer from memory or intent. If the prevention rule is not present in the owning skill, add it or state that it has not been durably landed.
143
143
  - **Consolidate and retire rules; a skill's rule set must not grow monotonically.** Every correction adds a guard, but an N-bullet wall on one theme is itself the over-prescription/unreadability failure, and "just append another bullet" is how it regrows.
@@ -522,16 +522,16 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
522
522
 
523
523
  **A convergence or closure declaration must be written falsifiably.** Name the exact candidate identity it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
524
524
 
525
- The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
525
+ The initial independent review plus Agent-initiated challenges share a **bounded external-review sequence of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. An authorized task includes its necessary fixes, tests and review by default; a sequence limit triggers the checkpoint below, not a new permission request. Candidate edits, commits, rebases, amended plans or renamed slices never erase cumulative spending or broaden that task authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
526
526
 
527
- Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **The lane spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There the autonomous lane ends. An authenticated human may request later review, but that is separately attributed human-requested evidence outside this chain/budget, never an additional Agent round. Unused generic capacity never authorizes automatic continuation. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). An integration branch that accumulated several reviewed rounds is promoted as one pull request without a new ledger: when neither a single ledger nor a manifest binds the promotion, the gate walks HEAD's first-parent chain down to the first commit already on the target and rebinds each round merge in a detached checkout of its second parent against its first parent, judged with the landing tree's own controller and validator rather than the round's (a round could carry a hollowed validator that a later round restores); a merge whose second parent is already on the target is a sync merge and owes nothing; every step must be exactly the automatic merge of its parents (a hand resolution or an extra file in the merge commit is refused as unreviewed), a non-merge commit on the chain is refused, and the chain is consulted only for the default path set. Rounds that appended to the same register therefore no longer force the promotion to be split by round. It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
527
+ Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **Each extraction receipt sequence spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There that receipt sequence ends. Necessary further review follows the checkpoint and original-task authority rule below in another bounded sequence, with complete cumulative history retained. Record a later round as human-requested only when a human actually requested that round. Unused generic capacity alone never justifies another call. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). An integration branch that accumulated several reviewed rounds is promoted as one pull request without a new ledger: when neither a single ledger nor a manifest binds the promotion, the gate walks HEAD's first-parent chain down to the first commit already on the target and rebinds each round merge in a detached checkout of its second parent against its first parent, judged with the landing tree's own controller and validator rather than the round's (a round could carry a hollowed validator that a later round restores); a merge whose second parent is already on the target is a sync merge and owes nothing; every step must be exactly the automatic merge of its parents (a hand resolution or an extra file in the merge commit is refused as unreviewed), a non-merge commit on the chain is refused, and the chain is consulted only for the default path set. Rounds that appended to the same register therefore no longer force the promotion to be split by round. It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
528
528
 
529
- **Self-hosted chains break on every fix; the budget is summed across chains, never per chain.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break spends the SAME Agent-autonomous budget. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
529
+ **Self-hosted chains break on owner edits; sum rounds across chains within each sequence and retain cumulative spending across sequences.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break still consumes its sequence's budget; a later sequence requires the checkpoint below and retains every earlier round. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
530
530
 
531
- 1. **Batch dispositions; never re-chain per finding — and hold every fix until the round-2 challenge has run.** Triage the whole batch through the disposition bar and deep-self-review once, then hold, never deciding by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, and round 2 binds the round-1 candidate, so a fix applied between the two forfeits the double-receipt terminal and costs a fresh human-authorized chain to recover — an observed failure, not a hypothetical: a round that landed its review fixes before the challenge had to be closed by a user-granted continuation chain. So the rule through round 2 is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, then apply the held batch, MR/PR-listed, and let round 3's succession challenge — owed exactly when the batch moved the candidate — be what inspects it.
532
- 2. **Sum spent rounds across all chains before opening one more; the lane's cap is three rounds across two chains.** Count every prior external round in the caller-retained ledger — every chain, finished or broken — against that cap. The only restart the lane funds is the single succession challenge, opened with `--predecessor-chain-result-file` so the ended chain is carried rather than laundered into a fresh-looking loop; any OTHER restart opens with a fresh review by contract, trading a challenge round for a review round, and must not be opened autonomously at the cap. Effective exhaustion is reached when the remaining rounds cannot fund the closeout floor for any continuation; treat it exactly like the cap.
531
+ 1. **Batch dispositions; never re-chain per finding — and hold every fix until the round-2 challenge has run.** Triage the whole batch through the disposition bar and deep-self-review once, then hold, never deciding by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, and round 2 binds the round-1 candidate, so a fix applied between the two forfeits the double-receipt terminal and requires a recorded recovery checkpoint before a fresh bounded sequence. So the rule through round 2 is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, then apply the held batch, MR/PR-listed, and let round 3's succession challenge — owed exactly when the batch moved the candidate — be what inspects it.
532
+ 2. **Sum spent rounds across all chains before opening one more; each sequence's cap is three rounds across two chains.** Record every prior external round — every sequence and chain, finished or broken — and the cumulative count in the caller-owned task artifact. Within one sequence, the only restart is the single succession challenge opened with `--predecessor-chain-result-file`; never append an over-budget receipt or clear earlier spending. At the cap, or when remaining rounds cannot fund that sequence's closeout floor, run the checkpoint before starting another bounded sequence under existing task authority.
533
533
  3. **Front-load packet quality in chain 1.** The first chain's packet must already be the full-context diff (`--unified` wide enough to carry whole files, e.g. `-U200`) with the plan frozen alongside the candidate; narrow packets breed packet-boundary pseudo-findings whose fixes break chains and burn rounds on artifacts of the packet itself.
534
- 4. **At the cap — or at effective exhaustion — the designed terminal is disposition, never another chain.** Apply or disposition the final batch, name every post-review fix in the MR/PR description, record the honest terminal state (`continuation_authorization_required` when the lane's final round ran and itself returned findings; otherwise — including a chain broken before its challenge could run — an interim record naming the last externally reviewed candidate and every later delta), and hand continuation or merge to the human. The post-review batch sits only on the pending MR/PR branch beside that record — the human's authenticated continuation, waiver, or merge decision is what certifies it, and it is never reported as reviewed. This is the bounded outcome working as designed, so do not report it as convergence and do not launder it through a fresh-looking chain.
534
+ 4. **At the cap — or at effective exhaustion — finish disposition and reassess the method before another sequence.** On the unchanged candidate, when review and challenge are conclusive and every finding occurrence is source-refuted by first-hand evidence, run the local `complete --finding-dispositions-file` path in the [staged contract](../../code-review/references/staged-review-contract.md#mechanical-self-review-gate). Preserve original findings and the full history; an empty model verdict is not required after evidenced refutation. Otherwise, apply or disposition the final batch, name every post-review fix and unreviewed delta in the MR/PR description, and retain `continuation_authorization_required` or an honest interim record. Unreviewed changes and unresolved risks still need their applicable review or human decision. First trace findings to source, run targeted tests and deep self-review, then fix the evidenced cause, improve missing packet context or change the failed review/diagnostic method. Name the specific remaining verification before starting another necessary bounded sequence under the rule below. A reviewer cap alone never asks the user to renew task authority. Never repeat calls solely to obtain zero findings, report an unreviewed batch as reviewed, or reset cumulative spending.
535
535
 
536
536
  A strictly proven wording-only change has no convergence loop: it uses one
537
537
  generic `code-review` pass, records the independent-review row and the
@@ -539,13 +539,13 @@ challenge-not-required proof, and does not create a schema-v3 multi-round
539
539
  terminal ledger. This exception does not apply to frontmatter, routing,
540
540
  validation, acceptance, example, owner or behavior changes.
541
541
 
542
- This budget limits only automatic reviewer invocation. It does **not** stop implementation, tests, debugging, or deep self-review, and it does not limit an authenticated human:
542
+ This budget bounds each reviewer sequence. It does **not** stop implementation, tests, debugging or deep self-review, and reaching it requires a method checkpoint before necessary review continues under the existing task scope. Explicit user cost, round-count and stop limits still govern:
543
543
 
544
544
  - A human may request another review or self-review, stop a live review or the overall iteration, commit, or merge. Record human-requested review separately from Agent-autonomous rounds.
545
545
  - A human merge/risk decision must come from platform-authenticated authority outside the candidate diff, such as a protected maintainer approval. A repository file, branch flag, CLI argument, environment variable, model statement, or Agent-written note is not human authentication.
546
546
  - A narrow authenticated `review_waiver` clears only the review-process gate for the exact candidate and records decision-maker, time, reason, residual findings, and accepted risk.
547
547
  - A distinct authenticated `merge_authorization` is the human's final decision for the exact candidate. CI still runs and reports review/build/test/security/compliance failures, but none remains merge-blocking after that decision. Report `merge_authorized_by_human` / `failed_but_human_overridden`; never rewrite any underlying result as `passed` or discard residual findings.
548
- - A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. The dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. While cross-chain budget remains, that break is handled autonomously by the self-hosted-chain rule's ledger-counted restart; it becomes this bullet's human-decision dead-end when the remaining budget cannot fund the re-review. The recovery at that point is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
548
+ - **`continuation_authorization`** must first be checked against the original task authorization: necessary in-scope fixes, tests and review are already authorized by default. At a sequence checkpoint, record `continuation_basis=existing-task-scope` in the caller-owned task artifact, with the original authorization reference and scope, the reason another bounded sequence is needed, changed method or added evidence, cumulative rounds, and links between the old sequence's terminal evidence and the new sequence. Each new sequence uses fresh current-candidate bindings and preserves every prior receipt, focus, finding and disposition; no CLI flag or runtime receipt field is added. The existing per-sequence format, timeout and validation bounds remain unchanged. This is inherited task authority, not a new human request for each round: never relabel these calls as newly human-requested or erase earlier spending. Ask only for scope or authority the original task lacks, an explicit user limit that prevents the next action, or a genuine unresolved product/design/risk decision; continue independent authorized work. Continuation waives no review, test or evidence obligation and grants no merge, publication or risk-acceptance authority. Never infer a lane waiver from silence or from authorization to continue.
549
549
 
550
550
  When a round returns findings, hand them to the implementer before another autonomous review. The implementer verifies each failure path, classifies it as a local fix, false positive, deferred risk, or human decision, and records targeted self-review plus tests. Do not blindly apply every suggestion and do not use the reviewer as the primary defect finder.
551
551
 
@@ -553,12 +553,12 @@ The mechanical reminder is `self_review_gate`, not prose alone. It records outst
553
553
 
554
554
  In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause, two no-progress attempts, or recurring findings trigger a method change, narrower reproduction, redesign, validation switch, or parked decision item; they never auto-stop unrelated runnable work.
555
555
 
556
- At the final Agent-autonomous round, do not start another automatically. If findings remain:
556
+ At the final round of a bounded sequence, perform the checkpoint before another necessary sequence. If findings remain:
557
557
 
558
- - keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`;
558
+ - keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`; these legacy fields describe the spent sequence, so check existing task authority before asking for permission;
559
559
  - record the last externally reviewed candidate and every later candidate delta; stale review evidence never certifies changed content;
560
560
  - mark findings that need product/design/risk authority as `needs_human_decision`, freeze only dependent work, and continue independent runnable slices;
561
- - enter `awaiting_human` only when no independent runnable work remains. This is a scheduling state, not task failure and not a human merge prohibition.
561
+ - enter `awaiting_human` only for an actual missing decision or authority after available authorized work and checkpoint recovery are exhausted. A sequence cap alone is not that blocker.
562
562
 
563
563
  The terminal checkpoint is an extraction closeout record, not a state emitted by
564
564
  `review_gate.py`, and its schema-v3 state is derived from evidence rather than
@@ -574,8 +574,8 @@ stop as race immediately after round 1 rather than spending an illegal challenge
574
574
  after the terminal predicate already fired.
575
575
  It ends in exactly one state:
576
576
 
577
- - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance.
578
- - `continuation_authorization_required`: the final round itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
577
+ - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance. For `completion_basis=source_refuted_findings`, add the same-directory `finding_dispositions: {file, sha256}` reference. Its digest must match the completion receipt; every original occurrence must also retain its separately bound `source_refuted` class evidence with exactly the same evidence array. This proves consistency and coverage, not the truth of the reasoning or permission to accept risk.
578
+ - `continuation_authorization_required`: the final round itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation. Preserve this legacy schema value and first check the original task scope; it does not unconditionally require another user grant.
579
579
  - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
580
580
 
581
581
  Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
@@ -591,16 +591,16 @@ convergence.
591
591
  For a focused single-skill change:
592
592
  - **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
593
593
  - **Round 2 — challenge**: after implementer triage — fixes stay HELD: applying any fix before this round breaks the chain, so the challenge runs on the frozen round-1 candidate (self-hosted-chain rule; enumeration item 1 above) — attack the highest-risk unresolved surface with an unprimed prompt.
594
- - **Round 3 — succession challenge, owed only when the fix batch moved the candidate**: apply the held batch, commit it, then ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate now hashes to. Unchanged (every finding accepted, pre-existing, or source-refuted) ⇒ the lane ends at round 2 and owes nothing. Changed ⇒ open ONE succeeding chain with `--predecessor-chain-result-file <round-2 receipt>` and challenge the landing candidate on a focus distinct from round 2's. This is the final Agent-initiated external round; its findings feed the post-budget checkpoint rather than an automatic further round, and the batch lands MR/PR-listed.
594
+ - **Round 3 — succession challenge, owed only when the fix batch moved the candidate**: apply the held batch, commit it, then ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate now hashes to. Unchanged (every finding accepted, pre-existing, or source-refuted) ⇒ the lane ends at round 2 and owes nothing. Changed ⇒ open ONE succeeding chain with `--predecessor-chain-result-file <round-2 receipt>` and challenge the landing candidate on a focus distinct from round 2's. This is the sequence's final external round; its findings feed the method/authority checkpoint before any necessary next bounded sequence, and the batch lands MR/PR-listed.
595
595
 
596
- Broad extractions use the same lane budget, round 3 included on the same condition. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
596
+ Broad extractions use the same per-sequence bounds, round 3 included on the same condition. Continue necessary implementation and checkpoint-qualified review within the original task scope; retain cumulative history rather than resetting the task.
597
597
 
598
598
  ### Anti-patterns
599
599
 
600
600
  - **Landing a fix batch no round ever saw**. The round-1 fix-up itself may introduce bugs, so a batch that moved the candidate owes the succession challenge of round 3 above — the earlier absolute ("always re-challenge after a non-trivial fix-up") was unreachable while the budget was two rounds, and an unreachable obligation reads as satisfied. A candidate the batch did not move owes nothing: the condition is the candidate hash, not the author's sense of how big the fix was.
601
- - **Iterating external review until zero findings**. Stop Agent reviewer calls at the configured budget. Stabilized or repeated findings are recorded, triaged, and may cause a method/design change or a parked dependent slice; implementation and independent work continue.
601
+ - **Iterating external review until zero findings**. Stop blind repetition at the configured sequence limit and run the checkpoint. Stabilized or repeated findings require source disposition and a method/design or evidence change before necessary review continues under existing task authority; a genuine decision blocks only its dependent slice.
602
602
  - **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
603
- - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one. The observed extreme is chain multiplication: a chain broken by your own fix and restarted at index 1 is the same budget, not a new loop — sum rounds across chains per the self-hosted-chain rule above.
603
+ - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another review round only when recorded verification needs and current risk call for it, or when a human explicitly requests one. A fresh bounded sequence never resets task history or cumulative spending — retain both and apply the self-hosted-chain checkpoint above.
604
604
  - **Re-running with a softer prompt after fixes**. Use the same adversarial framing every round; weakening the prompt to make later rounds "pass" defeats the purpose.
605
605
 
606
606
  ### Recording the loop
@@ -209,15 +209,15 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
209
209
  | **副作用边界触达** | 任何 destructive op(rm -rf / force-push / drop table / cross-team shared-doc overwrite)| stop + 等用户确认 |
210
210
  | **不可解决依赖** | 等外部服务 / 等人审批 / 等数据到 | stop + 报当前状态 + 等依赖解除 |
211
211
 
212
- **Warning(命中即报告 + 用户决策继续 / 切策略 / 停,不自动停)**:
212
+ **Warning(命中即报告并自查 / 调整方法;原任务授权覆盖必要续行,显式用户限制仍优先)**:
213
213
 
214
214
  | Trigger | 启发阈值 | 决策 |
215
215
  |---|---|---|
216
- | **同一失败重复 N 次** | N = 3(同 error 第 4 次出现)| 报告 + 等人;**严格指"identical retry"**,不是 challenge 多轮发现新问题 |
217
- | **预算 warning** | tool call > 100 / token 紧张 / wallclock > 30 分钟 | 报中间状态 + 用户决定继续 / 切策略 / 停;不自动停(大 codebase + flaky 外部依赖 + 长跑但有进展的工作 都可能合理超阈值)|
216
+ | **同一失败重复 N 次** | N = 3(同 error 第 4 次出现)| 报告 + 停止相同重试,核对失败证据后调整方法或补上下文;**严格指"identical retry"**,不是 challenge 多轮发现新问题 |
217
+ | **预算 warning** | tool call > 100 / token 紧张 / wallclock > 30 分钟 | 报中间状态、累计用量和下一步依据,按原授权继续必要工作;仅缺权限、超出范围、真实取舍或用户显式限制阻断该行动时等人,不因启发阈值重新请批 |
218
218
 
219
219
  **与 convergence standard 的区别**(重要):本节阈值针对 **same-error retry**(重复尝试同一失败方案),不替代 [convergence standard](../SKILL.md)(针对 challenge round — 只要每轮还有 P1 required 就继续,不按 round 计数停)。区分:
220
- - **Same-error retry**:每次尝试本质相同方案 → N=3 后升级
220
+ - **Same-error retry**:每次尝试本质相同方案 → 命中上表阈值后停止相同重试并调整方法,不能只换名称继续重试
221
221
  - **Challenge convergence round**:每轮发现不同新层 / 新 P1 → 继续直到 0 P1 required(recursive self-validation 类 extraction 可能走 5-6 轮,每轮抓不同层的新 P1,不该按 round count 停)
222
222
 
223
223
  **Escalation 必含 5 字段**(避免"escalate"沦为"stop"):
@@ -6,11 +6,11 @@ Companion to the **auto-trigger durable learning** Core Rule in `SKILL.md` (unde
6
6
 
7
7
  Bind recovery when either:
8
8
 
9
- - **(a)** the literal original `proposed-next:` action/scope (or literal explicit proposal for a pre-marker incident) and literal assent are preserved in the visible conversation or quoted exactly in trusted host-owned session/compaction state; or
9
+ - **(a)** one concrete original proposal, its scope, and the assent are preserved verbatim in the visible conversation or read back verbatim from trusted host-owned session state; a `proposed-next:` marker is not required; or
10
10
  - **(b)** the current user's premature-stop correction itself literally names the paused action and scope — on path (b), quote those exact user words in the visible `continuing:` line before proceeding.
11
11
 
12
- A semantic compaction paraphrase or a bare "why did you stop" complaint is not path-(b) authority. Never copy real conversation text into a shared repository record, reconstruct, broaden, or substitute it. The user's challenge is a current reactivation signal for that exact slice, not a demand for redundant reconfirmation; restate and proceed only when path (a) or (b) binds, and ask again when neither path binds the slice, the user intervened, or current scope/gates changed.
12
+ A semantic compaction paraphrase supplies neither binding path; recover the original proposal and assent before deciding path (a) is unavailable. A bare "why did you stop" complaint does not itself name path (b)'s action and scope. Never copy real conversation text into a shared repository record, reconstruct, broaden, or substitute it. The user's challenge reactivates that exact slice. Restate and proceed when either path binds; ask only when the action, scope, or required authority remains unresolved. A new user message or a changed gate requires reassessment, not automatic reconfirmation.
13
13
 
14
14
  ## Invalid `blocked:` recovery
15
15
 
16
- A `blocked:` recovery without the step-1 evidence and a specific missing authority/ambiguity is invalid: rerun the gate and emit a well-formed outcome before any action; if ambiguity remains, ask the blocker question in the same turn and stay blocked. Do not let correction RCA or extraction extend a still-authorized delivery, and do not let stale assent bypass a newly pending or inconclusive gate.
16
+ A `blocked:` recovery without applicable state evidence and a specific remaining blocker is invalid: recover intent and rerun the owning gate. If a decision or permission remains unresolved, ask in the same turn and block that dependent action. Continue available authorized diagnosis, bounded remediation, or independent work; do not let stale assent bypass a newly pending or inconclusive gate. Do not let correction RCA or extraction delay recovery of a still-authorized delivery.
@@ -644,3 +644,14 @@ The pending classification above is superseded by the executed source comparison
644
644
  | Changed definition tables can provide a firing path without inventing a normative list rule | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The accepted-definition fixture failed before the gate change with expected rc=0 and actual rc=1. Afterward the complete suite passed: one full-checker wiring case and 112 standalone-gate cases. Fifteen table cases cover acceptance plus stale or unchanged anchors, foreign owners, duplicate anchors, comments, fences, raw HTML, indented code, missing headers, header anchors, and missing RED evidence. Existing owner, round, normative-list, and executable checks remain. |
645
645
  | An impact-chain evidence row cannot use the source register itself as its firing path | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. A synthetic owner bookkeeping edit and a self-citing table row passed the prior gate with actual rc=0 where refusal rc=1 was required. The anchor occurred only once, inside its own locator, so uniqueness alone did not prevent self-certification. Rejecting the source register as a file firing path made that case pass; the full focused suite passed one checker wiring case and 113 standalone-gate cases. Ordinary changed definition tables still pass. |
646
646
  | Table firing anchors exclude raw HTML processing instructions, declarations, CDATA and multiline raw-tag openers | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. An executed paired check of the current table predicate accepted plain data and also accepted data inside processing-instruction, declaration and CDATA blocks. The full gate fixture additionally returned rc=0 for a raw script opener where refusal rc=1 was required. Explicit block terminators now exclude these non-table surfaces; a completed CDATA block followed by a real table remains accepted. The full focused suite passed one checker wiring case and 118 standalone-gate cases. |
647
+ | Diagnosed code and test repairs require implementer self-checks and automatic independent review before completion | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/SKILL.md#Code/test changes require self-checks | updated | Owner key `defect-diagnosis/SKILL.md`. `test_ai_coding_implementation_gates.sh` checks this owner's automatic-review trigger. An applied removal of that trigger failed its owning assertion; unchanged and restored controls passed. The current focused suite also passes with the trigger expressed as a normative list rule and the original introduction retained. This proves source-contract coverage, not automatic invocation in every agent runtime. |
648
+ | LLM and inference code or test changes require implementer self-checks and automatic independent review before completion | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/SKILL.md#Code/test changes require self-checks | updated | Owner key `llm-inference-integration/SKILL.md`. `test_ai_coding_implementation_gates.sh` checks this owner's automatic-review trigger. An applied removal of that trigger failed its owning assertion; unchanged and restored controls passed. The current focused suite also passes with the trigger expressed as a normative list rule and the original introduction retained. This proves source-contract coverage, not automatic invocation in every agent runtime. |
649
+ | Node.js service code and test changes require implementer self-checks and automatic independent review before completion | `nodejs-service-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/nodejs-service-dev/SKILL.md#Code/test changes require self-checks | updated | Owner key `nodejs-service-dev/SKILL.md`. `test_ai_coding_implementation_gates.sh` checks this owner's automatic-review trigger. An applied removal of that trigger failed its owning assertion; unchanged and restored controls passed. The current focused suite also passes with the trigger expressed as a normative list rule and the original introduction retained. This proves source-contract coverage, not automatic invocation in every agent runtime. |
650
+ | Observability code and test changes require implementer self-checks and automatic independent review before completion | `platform-observability` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/SKILL.md#Code/test changes require self-checks | updated | Owner key `platform-observability/SKILL.md`. `test_ai_coding_implementation_gates.sh` checks this owner's automatic-review trigger. An applied removal of that trigger failed its owning assertion; unchanged and restored controls passed. The current focused suite also passes with the trigger expressed as a normative list rule and the original introduction retained. This proves source-contract coverage, not automatic invocation in every agent runtime. |
651
+ | Release-engineering code and test changes require implementer self-checks and automatic independent review before completion | `platform-release-engineering` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-release-engineering/SKILL.md#Code/test changes require self-checks | updated | Owner key `platform-release-engineering/SKILL.md`. `test_ai_coding_implementation_gates.sh` checks this owner's automatic-review trigger. An applied removal of that trigger failed its owning assertion; unchanged and restored controls passed. The current focused suite also passes with the trigger expressed as a normative list rule and the original introduction retained. This proves source-contract coverage, not automatic invocation in every agent runtime. |
652
+ | Service-connectivity code and test changes require implementer self-checks and automatic independent review before completion | `platform-service-connectivity` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-service-connectivity/SKILL.md#Code/test changes require self-checks | updated | Owner key `platform-service-connectivity/SKILL.md`. `test_ai_coding_implementation_gates.sh` checks this owner's automatic-review trigger. An applied removal of that trigger failed its owning assertion; unchanged and restored controls passed. The current focused suite also passes with the trigger expressed as a normative list rule and the original introduction retained. This proves source-contract coverage, not automatic invocation in every agent runtime. |
653
+ | Executable test changes require implementer self-checks and automatic independent review before completion | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#Code/test changes require self-checks | updated | Owner key `testing-strategy/SKILL.md`. `test_ai_coding_implementation_gates.sh` checks this owner's automatic-review trigger. An applied removal of that trigger failed its owning assertion; unchanged and restored controls passed. The current focused suite also passes with the trigger expressed as a normative list rule and the original introduction retained. This proves source-contract coverage, not automatic invocation in every agent runtime. |
654
+ | Repeated identical verification failures require a reported method checkpoint; existing task authority continues to cover necessary work while genuine missing decisions remain blockers | `multi-agent-delegation` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/multi-agent-delegation/SKILL.md#Repeated identical verification failures | updated | Owner key `multi-agent-delegation/SKILL.md`. The warning-family checks in `test_ai_coding_implementation_gates.sh` reject deletion of inherited task authority, real-blocker handling or explicit user limits. Applied mutations failed their owning assertions and both controls passed; the current focused suite passes. Unknown completion, unavailable dependencies, unclear direction and missing authority remain actionable blockers after bounded remediation. Evidence is source-contract validation, not a host-enforced authorization mechanism. |
655
+ | Continuation recovers the current request and original authorized proposal, scopes blockers to dependent work, and uses a method checkpoint instead of renewing permission for necessary review | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#do not stop at a recommendation | updated | Owner key `product-rd-workflow/SKILL.md`. `test_ai_coding_implementation_gates.sh` binds the entry rules to `product-rd-workflow/references/pre-final-continuation-gate.md`; applied trigger and boundary removals fail their owning assertions and restored controls pass. The current suite passes. Classification fixtures cover inherited review authority, explicit review limits and out-of-scope review; their labels are not proof of tool execution or universal runtime improvement. |
656
+ | A complete checkpoint may bind source-refuted findings without rewriting external receipts or refreshing review authority | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/review_gate.py; bank-evidence: file:specs/continuation-control/routing-evidence.md#The code-review routing comparison must preserve | updated | Owner key `code-review/SKILL.md`. `code-review/scripts/test_review_client_compat.py` exercises `CompletionFindingDispositionTest`: the former passed-only predicate rejected complete same-candidate refutation evidence; the current 17 focused tests pass. Original ordered receipt hashes, canonical occurrence coverage, disposition evidence and candidate bindings remain checked; omitted or altered evidence, duplicate dispositions and unresolved findings are rejected. Validation establishes binding and coverage, not the truth of source reasoning. |
657
+ | Extraction reviewer limits bound each receipt sequence; source disposition, method changes and complete cumulative history govern necessary continuation under existing task authority | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The warning-family source check failed against the former mandatory-human-warning clause. Seven applied warning and delegation mutations failed their owning assertions with unchanged and restored controls passing; the current implementation-gate suite passes. `skill-extraction-workflow/references/dual-track-review-gate.md` preserves per-sequence bounds, source findings, cumulative spending and genuine decision boundaries. The new owner rows also repair a reproduced impact-chain failure for missing owner evidence; they do not turn source checks into runtime or external-review passes. |