@ccoalm/ccl-skills 0.7.0 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +8 -7
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +6 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +16 -17
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +195 -7
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +3 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +13 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +9 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +9 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/normalize_review_timeout.sh +22 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +9 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +1540 -129
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +8 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +76 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +1858 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +789 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/update_review_plan_intent.py +513 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +11 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +73 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/runtime-and-project-contract.md +58 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/verification-diagnostics-and-security.md +63 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +14 -16
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +10 -14
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +135 -86
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +66 -80
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/delivery-contract.md +275 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +88 -214
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +10 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +6 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +112 -95
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +30 -21
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +22 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +20 -17
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +7 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +14 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +3 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +9 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +37 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +8 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +16 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +16 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +4 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +8 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +104 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +11 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +102 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +103 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +6 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +4 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +69 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +22 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +49 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/obligation-ledger.py +2748 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +20 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/shared_git_surface_gate.py +1142 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +41 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh +82 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +336 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +82 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh +1416 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +141 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh +1696 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh +2117 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_loading_budget.sh +316 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +1176 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +31 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +9 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +980 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +8 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +16 -15
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +10 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +7 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
- package/dist/assets/release.json +215 -105
- package/package.json +1 -1
|
@@ -18,15 +18,15 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
18
18
|
├─ b. Draft skill / reference (new or update existing)
|
|
19
19
|
├─ c. Sibling-generalization mini-map (which sibling skills may share / route)
|
|
20
20
|
├─ d. Sanitization pass with checklist (cheap, seconds)
|
|
21
|
-
├─ e.
|
|
22
|
-
│ ├─
|
|
23
|
-
│ └─
|
|
21
|
+
├─ e. Owner review gate per mandatory table (deep, minutes)
|
|
22
|
+
│ ├─ Strict wording-only → one independent code-review pass
|
|
23
|
+
│ └─ Non-wording → extraction_review_gate review + at most two challenges
|
|
24
24
|
├─ f. Apply fixes, re-sanitize
|
|
25
25
|
├─ g. Commit per batch on a feature branch → MR pending review (never push to main)
|
|
26
26
|
└─ h. Update charter completion log
|
|
27
27
|
│
|
|
28
28
|
4. Closeout → ~/.<host>/skills/.extraction-work/<project>-completion.md
|
|
29
|
-
|
|
29
|
+
Non-wording terminal ledger validator + final state + deferred backlog
|
|
30
30
|
│
|
|
31
31
|
5. Provenance migration → ~/.<host>/.private-aliases/<project>.yaml
|
|
32
32
|
Move file keys / paths / counts / dates out of working files
|
|
@@ -91,9 +91,9 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
91
91
|
|
|
92
92
|
- When required: see `references/dual-track-review-gate.md` table.
|
|
93
93
|
- Choose the review tier from that table, not from intuition. Do not restate the rows locally; record the exact `dual-track-review-gate.md` table row used. Record `challenge: not-required` only when that row classifies the actual diff as challenge-not-required (for shared skills, this means strict wording-only with deterministic scope proof + independent review confirmation). Non-wording shared-skill changes cannot skip challenge.
|
|
94
|
-
- Run deterministic checks and implementer self-review first, and record what each proves before invoking review/challenge (this self-review-before-review ordering applies to every non-wording shared-skill change the dual-track table requires review for, not only the rows that look high-risk): `git diff --check` proves whitespace/conflict-marker hygiene only; validators prove schema/link/routing invariants; leakage/sanitization scans prove only their configured patterns; scope checks must name the changed files or expected file set; the self-review row is conclusive only when each required field is non-empty (acceptance criteria, changed-file scope, edge/failure paths, known residual risks) and the changed-file scope equals the candidate diff's changed-file set, or explicitly explains any excluded generated/irrelevant file.
|
|
95
|
-
- Review pass:
|
|
96
|
-
- Challenge pass: invoke
|
|
94
|
+
- Run deterministic checks and implementer self-review first, and record what each proves before invoking review/challenge (this self-review-before-review ordering applies to every non-wording shared-skill change the dual-track table requires review for, not only the rows that look high-risk): `git diff --check` proves whitespace/conflict-marker hygiene only; validators prove schema/link/routing invariants; leakage/sanitization scans prove only their configured patterns; scope checks must name the changed files or expected file set; the self-review row is conclusive only when each required field is non-empty (acceptance criteria, changed-file scope, edge/failure paths, known residual risks) and the changed-file scope equals the candidate diff's changed-file set, or explicitly explains any excluded generated/irrelevant file. Persist it before the review/challenge run in a fresh, non-overwritten task-evidence path outside the candidate diff, pass that exact file as the gate's review plan, and retain the gate result that binds its profile hash; do not edit the candidate merely to record self-review or review outcome, because that creates self-referential candidate churn. A candidate-local row is appropriate only when the row itself is a substantive deliverable under review. A plain in-place-editable MR description or scratch log is not ordering proof unless its edit history is retrievable and checked; a backfilled row is invalid and forces a rerun. If the candidate diff changes after the row is saved — a file added/removed OR the content of any listed file materially changed — refresh the row and rerun review/challenge against the new candidate. Changing only the external self-review record refreshes the profile binding; it does not by itself invalidate implementation tests or the candidate packet. A missing field, "ok" placeholder, mismatched scope, or unprovable ordering makes the row inconclusive. Do not spend LLM review rounds on issues a script or implementer-side checklist can decide. If the independent pass is the first place basic scope, contract, privacy, or test issues surface, apply those findings to the diff, close the self-review gap, and rerun the deterministic gates before rerunning review/challenge; the process-defect repair is in addition to resolving the findings, not a way to discard or downgrade them.
|
|
95
|
+
- Review pass: persist the complete self-review row and encode it in the review plan. For a **non-wording** lane, resolve the repository-owned `scripts/extraction_review_gate.sh` and use it from round 1; never substitute the generic controller, scan writable plugin roots, or supply a caller-selected budget. For a strictly proven **wording-only** lane, use the generic `code-review` proof-bound single-review recipe in `code-review/references/staged-review-contract.md` and record `challenge: not-required`; require its controller-derived wording scope plus the independent `wording_only_boundary` confirmation. This is the only extraction path that stays outside the multi-round wrapper and terminal ledger; the gate, not this page, decides whether a chainless review is legal, and it may still demand the tracked pair. Take all controller options from that runnable recipe, supplying the actual stage and exact candidate rather than an example default. The non-wording chain cannot be retrofitted, so a run started outside its owner wrapper is thrown away and restarted. Read the chain-opening and packet-composition rules in `references/dual-track-review-gate.md` first. Require conclusive JSON, selected-client attribution, packet/profile binding, family exclusion, and wrapper runtime evidence. When the host returns a live execution handle (`session_id`, `cell_id`, or equivalent), keep polling that exact handle until terminal exit; empty current output is progress, not a verdict, and no replacement/fallback reviewer may start while the original process is live. The result row records handle type, an opaque host transcript/tool-call reference and terminal exit status. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle. Never copy a credential-like raw handle into shared evidence. `findings` is not pass; inconclusive, malformed, or free-form output stays interim. Do not add a separate behavior probe.
|
|
96
|
+
- Challenge pass: for a non-wording lane, invoke `scripts/extraction_review_gate.sh` separately with the same plan, stage, candidate, family and tracked chain. Pass the next one-based index; later rounds include a distinct focus and all prior focuses. Preserve a separate result row with the same binding, exclusion, egress, attribution and conclusive checks. Review never satisfies challenge; missing or inconclusive required challenge keeps extraction interim. A wording-only lane has no challenge pass.
|
|
97
97
|
- Treat review/challenge as batch-level gates over the landing candidate, not as a per-bullet or per-line edit loop. Apply all findings from a round; when both lenses are required, re-run both on the updated candidate before landing.
|
|
98
98
|
- Skipping a required challenge = work can only land as interim, not complete.
|
|
99
99
|
|
|
@@ -119,6 +119,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
119
119
|
- File: `~/.<host>/skills/.extraction-work/<project>-completion.md`
|
|
120
120
|
- Final state: which batches done, which deferred, which sources unavailable.
|
|
121
121
|
- Lessons: what surprised; what would change in next extraction; what to add to skill-extraction-workflow.
|
|
122
|
+
- For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. A clean Round 2 plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 3 findings validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row but does not fabricate a multi-round ledger.
|
|
122
123
|
|
|
123
124
|
### 5. Provenance migration
|
|
124
125
|
|
|
@@ -147,8 +148,9 @@ Skip this step when nothing transferable surfaced.
|
|
|
147
148
|
|---|---|---|
|
|
148
149
|
| Anti-pattern grep panel | `references/recurring-anti-patterns-checklist.md` | Every commit; ~30s |
|
|
149
150
|
| `check-ccl-skills.sh` | `scripts/check-ccl-skills.sh` | Every commit; ~10s |
|
|
150
|
-
| `
|
|
151
|
-
| `
|
|
151
|
+
| Generic `code-review` gate | repository-owned skill | Strict wording-only independent review; ~5-10 min |
|
|
152
|
+
| `scripts/extraction_review_gate.sh` | this skill package | Non-wording review plus at most two challenges; ~5-15 min each |
|
|
153
|
+
| `scripts/validate_extraction_review_state.py <closeout.json>` | this skill package | Every non-wording terminal checkpoint |
|
|
152
154
|
| Source-read fallback ladder | `SKILL.md` Source-read remediation | When a source read fails or times out |
|
|
153
155
|
| Sibling mini-map | `SKILL.md` Step 4 stack-specific updates | Every stack-specific change |
|
|
154
156
|
| Private alias map `audit_cmd` | `~/.<host>/.private-aliases/<project>.yaml` or process-retro profile | Every commit's R0 audit |
|
|
@@ -42,6 +42,108 @@ The closeout row must name which private profile was used: `project-alias`, `pro
|
|
|
42
42
|
|
|
43
43
|
Pre-existing leakage from earlier extractions may be tracked as `known_debt` in the alias YAML and explicitly excluded from blocking the current change. New or modified content MUST remain zero-hit — `known_debt` cannot waive a newly introduced label.
|
|
44
44
|
|
|
45
|
+
### Shared Git and PR text
|
|
46
|
+
|
|
47
|
+
Git history and forge metadata are shared publication surfaces even when the
|
|
48
|
+
repository files are clean. The repository-owned
|
|
49
|
+
`skills/skill-extraction-workflow/scripts/shared_git_surface_gate.py` therefore
|
|
50
|
+
scans the exact candidate commit range, current branch name, and available PR
|
|
51
|
+
title/body text for AI session links or identifiers, model session/co-author
|
|
52
|
+
trailers, AI author/committer identities, generated footers,
|
|
53
|
+
conversation-process labels, and external-source provenance wording. Common
|
|
54
|
+
Markdown list (including GFM task lists), quote, heading, and nested wrappers
|
|
55
|
+
are presentation only and do not bypass any line-oriented class. The gate
|
|
56
|
+
reports only surface, locator, and category so neither a finding nor an error
|
|
57
|
+
diagnostic repeats the identifier, ref, path, or raw tool error it is blocking.
|
|
58
|
+
Candidate commit metadata is requested from Git as UTF-8 regardless of the
|
|
59
|
+
repository's ambient log-output encoding. Invalid UTF-8 or a Unicode
|
|
60
|
+
replacement character fails closed with only the commit locator and field
|
|
61
|
+
name, rather than silently erasing a CJK-only match. Each candidate's bounded
|
|
62
|
+
raw commit object is also checked before pretty formatting; an embedded NUL is
|
|
63
|
+
rejected because Git would otherwise truncate the visible message at that byte.
|
|
64
|
+
The batch response is length-parsed and bound one-for-one to the separately
|
|
65
|
+
enumerated full object IDs; abbreviated diagnostic locators are never used for
|
|
66
|
+
identity or ordering decisions.
|
|
67
|
+
`--repo` is the single repository identity. Every Git subprocess drops ambient
|
|
68
|
+
repository-routing variables that could replace its worktree, refs, object
|
|
69
|
+
store, ancestry, index, or namespace; a caller cannot point the gate at a clean
|
|
70
|
+
decoy with `GIT_DIR` while the prohibited candidate lives elsewhere. This gate
|
|
71
|
+
is a worktree pre-push/CI lane, not a receive-pack quarantine, so quarantine
|
|
72
|
+
object-store overrides are not accepted.
|
|
73
|
+
All object reads disable local replacement refs, which alter only the local
|
|
74
|
+
view and are not the bytes a push publishes. A non-empty legacy
|
|
75
|
+
`info/grafts` file likewise fails closed because it can rewrite the visible
|
|
76
|
+
parent chain without changing the commit objects sent by a push.
|
|
77
|
+
|
|
78
|
+
The base priority is explicit `--base-ref`, then the trusted PR event, then
|
|
79
|
+
`CCL_SKILL_BASE_REF`, then a repository-declared default; an unresolved selected
|
|
80
|
+
base is an error. The trusted event outranks ambient environment configuration
|
|
81
|
+
so a CI process variable cannot silently move the base forward and narrow the
|
|
82
|
+
event-defined candidate range. Only
|
|
83
|
+
history outside the selected `<merge-base>..<candidate-head>` range may be
|
|
84
|
+
`known_debt`; the merge-base itself is target-side history and is not a candidate
|
|
85
|
+
commit. Candidate commits, the current destination branch, and current PR text
|
|
86
|
+
have zero exceptions, including a violation hidden in an earlier candidate
|
|
87
|
+
commit behind a clean HEAD. CI runs the gate on direct pushes to `dev`/`main`
|
|
88
|
+
and on PR `edited` events because title/body changes do not require a new
|
|
89
|
+
commit. Before a PR is created or edited, pass its exact proposed title and
|
|
90
|
+
body through `--pr-text-file`; a local run with neither an event nor that file
|
|
91
|
+
has checked no PR text. The optional pre-push hook is early feedback; the
|
|
92
|
+
protected CI check is the merge boundary.
|
|
93
|
+
|
|
94
|
+
The repository command fallback `origin/dev` applies only to the normal
|
|
95
|
+
feature→dev lane. Promotion or any other target passes that target explicitly
|
|
96
|
+
(for example, `--base-ref origin/main` for dev→main); a default is not evidence
|
|
97
|
+
of the intended landing target. The pre-push hook uses an existing destination ref's remote SHA
|
|
98
|
+
as its base only when that destination is an actual landing target (`dev` or
|
|
99
|
+
`main`); ambient `CCL_SKILL_BASE_REF` cannot replace that SHA. The environment
|
|
100
|
+
base may delimit a new zero-OID landing target. Every non-delete pushed ref is
|
|
101
|
+
scanned at its own local object ID and destination name without checking it
|
|
102
|
+
out. An ordinary feature ref always uses an explicitly bound landing target or
|
|
103
|
+
the pushed remote's `dev` tracking ref; if that ref is absent, fetch it or set
|
|
104
|
+
`CCL_SKILL_BASE_REF`. Its existing remote feature tip is still candidate history
|
|
105
|
+
and must never become the base/`known_debt`. A new
|
|
106
|
+
`dev`/`main` destination without an explicit base fails closed. Provider aliases
|
|
107
|
+
are shared across the provider-shaped
|
|
108
|
+
surface patterns, but unknown future AI providers remain an enumerated-pattern
|
|
109
|
+
coverage risk; the shared prose prohibition is broader than the currently
|
|
110
|
+
recognized provider names.
|
|
111
|
+
|
|
112
|
+
Co-author detection also has an intentional precision boundary. An
|
|
113
|
+
unambiguous product display name such as `Claude Code` or `Codex`, a bounded
|
|
114
|
+
model/product-qualified form such as `Claude Sonnet`, `OpenAI Codex`, or
|
|
115
|
+
`ChatGPT-5`, or a known
|
|
116
|
+
provider GitHub App account carrying the `[bot]` suffix, is blocked with any
|
|
117
|
+
email. Account matching accepts GitHub slug separators, so `claude-code[bot]`
|
|
118
|
+
and `copilot-swe-agent[bot]` remain the same known-provider class. A known
|
|
119
|
+
provider `[bot]` account in the email local part is also sufficient when the
|
|
120
|
+
display name is neutral, with only the standard optional numeric GitHub ID
|
|
121
|
+
prefix accepted before that exact account. A provider token embedded as a
|
|
122
|
+
suffix inside another bot account is not treated as the provider. An exact
|
|
123
|
+
single-name alias that can also be a person's name needs an independent
|
|
124
|
+
`noreply`/`no-reply`/`bot` email signal. This keeps a human whose
|
|
125
|
+
real name matches an alias from being mechanically rejected, but it also means
|
|
126
|
+
an AI trailer that deliberately uses such an ambiguous name plus an ordinary
|
|
127
|
+
email is not detectable from the trailer alone. The root prohibition still
|
|
128
|
+
applies; proposed text with that ambiguity needs human readback rather than a
|
|
129
|
+
claim that the local gate proved every model co-author absent. Candidate author
|
|
130
|
+
and committer fields use a narrower rule: only unambiguous product names,
|
|
131
|
+
known-provider `[bot]` display names, or known-provider `[bot]` email accounts
|
|
132
|
+
with only an optional numeric GitHub ID prefix are blocked. A generic `noreply`
|
|
133
|
+
address is a normal human privacy setting there, so an ambiguous single-name
|
|
134
|
+
identity remains an explicit human-readback coverage gap; unrelated automation
|
|
135
|
+
bots and accounts that only end in a provider token are not reclassified as AI.
|
|
136
|
+
|
|
137
|
+
The current deterministic event surface does not fetch historical PR comments,
|
|
138
|
+
labels, or a platform-generated custom merge message. Those remain prohibited
|
|
139
|
+
by the root contract, but are an explicit coverage gap requiring forge-side
|
|
140
|
+
readback/enforcement; they are not reclassified as `known_debt` merely because
|
|
141
|
+
this local scanner cannot observe them. A pushed tag is scanned on every
|
|
142
|
+
surface the object graph carries: destination name, pointed-to commit range,
|
|
143
|
+
and each annotated tag-object layer's own message and tagger identity (nested
|
|
144
|
+
tags are peeled with a bounded depth, and an unresolvable or over-deep tag
|
|
145
|
+
chain fails closed).
|
|
146
|
+
|
|
45
147
|
## Fail-Closed Clause — Maintainer Discipline
|
|
46
148
|
|
|
47
149
|
Every new sanitized label introduced in this commit (e.g. an angle-bracket capability token like `<some-capability>`, a class tier like `<class-A1>`, or any other invented short name that stands for a real source artifact) MUST exist in the maintainer's alias YAML before the commit lands.
|
|
@@ -358,3 +358,106 @@ and inverted the sense (production, not product), and the coordinator now shares
|
|
|
358
358
|
| 发布工程技能有推进闸、回滚契约与审计日志,却没有发布过程自身的度量反馈环——DORA 词汇全仓三个平台技能零命中;feature flag 只被当成 dynamic config 的取值实例,缺生命周期纪律(deploy≠release 解耦、四分类、transient flag 的 owner+expiry、双侧测试、并发活跃数压低) | `platform-release-engineering` | references/promotion-gate-and-review.md 新增 Release-process metrics (DORA) 节(入口 SKILL.md 为历史超限零增长预算,正文落 reference、References 行做字数中性指名):从控制平面自身 audit log 机算 DORA 当前五因子(change lead time / deployment frequency / failed deployment recovery time〔MTTR 后继〕/ change fail rate / deployment rework rate),用作过程反馈、never 作个人或团队绩效评分(评分即污染信号:deploy 改标签、rollback 改叫 roll-forward)、算不出即 audit-log 缺口须修事件采集不得估算;references/secret-and-config-management.md dynamic 层补 flag 生命周期(Fowler/Hodgson 四分类按寿命分治、transient flag 创建即带 owner+expiry、过期即债务须浮出、双侧测试与 flag 组合空间不可测故压低并发活跃数、release toggle 的退休进 feature 的 definition of done); result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/platform-release-engineering/references/promotion-gate-and-review.md#never as individual or team performance scores | updated | owner key `platform-release-engineering/SKILL.md`。**RED-baseline 为覆盖差分 + 一手源核验**:改前 grep -riE "DORA\|deployment frequency\|lead time\|change fail" 于三个平台技能零命中(红),改后 R14 命中(绿)。DORA 一手源 dora.dev/guides/dora-metrics-four-keys 本轮 fetch:五因子引文逐条取得,含 MTTR→Failed Deployment Recovery Time 的演进说明与 deployment rework rate——按当前模型落地而非记忆里的旧四键。flag 一手源 martinfowler.com/articles/feature-toggles.html(Hodgson):Release/Experiment/Ops/Permissioning 四分类、carrying cost/inventory、expiration date for short-lived toggles 引文逐条取得。同轮顺带一手源复核未改动的既有主张:R13 gitlab-runner SIGQUIT/SIGTERM 语义与 killall/pkill 告诫和 docs.gitlab.com/runner/commands 逐字吻合(unchanged: already-covered,无观察失败)。deferred(记录不落地):SLSA/构建 provenance attestation 与镜像签名准入是 R2/deploy-pipeline 的候选补强——deploy-pipeline 已有 OpenGitOps/OCI/cosign-manifest 基底,image-build attestation 面待下轮按外部源核后落。 |
|
|
359
359
|
| 交付面只核「那一份变了没」核不出「放没放对地方」——挂错目录 / 空间 / 父容器的文档,内容标志核得再准也是错交付;且放置决策若依据过期快照或链接标题猜测,错位在发布前就已注定 | `tighten-doc` | delivery-face 新增 ①b 定位节(载体无关,适用任何多级容器承接面):放置前必须现读目标结构、不按链接标题猜父容器;落点选最具体的稳定容器、标题相似本身不构成父子依据(现读确认确为稳定容器的节点可作父节点——页面即容器的承接面适用,该限定由独立评审指出过宽后收窄);新建/移动后回读父容器/空间/完整路径并交付全路径;落点决定可见性——放置前记预期受众/访问边界、放置后回读生效受众/权限与 owner 并按方向分流:比预期更宽(已实际暴露)、无法排除更宽、或 owner 落到非预期主体(owner 自带控制与转授权能力)的按待安全处置转出安全/内容 owner 走收敛路径结案(closeout 触发枚举同步纳入、与「拿不准按命中」同则;预期 owner 随预期边界在放置前记录)、确知更窄(可修复交付缺陷)或边界符合但其他核验失败的按 blocked 修复后重投,两者均禁报已同步,且本节只核不改(结构移动与权限变更各需其自身授权)——该第四条由 review 与 challenge 两 lane 对 ACL 继承暴露面独立收敛后补,过宽/过窄分流由后续评审指出同判 blocked 会让真实泄露留在普通投放流后再收窄。SKILL.md 交付面 bullet 加最小 cue(①句「投放与定位都要回读远端」+ 指针「四态定义、定位判据」,后者同时修正此前登记的三态→四态指针 P2); result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/tighten-doc/references/delivery-face-closeout.md#放置前必须现读目标结构**:目录树 / 容器层级以本次实时读取为准 | `updated` | owner key `tighten-doc/SKILL.md`(轮起 base 53adde21,随并行轮 rebase 至当前 dev;派生与零损核对均以当前 origin/dev 为 base 复算)。源=对同事分享的中文技术写作技能包做 R0 复扫后剥出的载体无关方法形状(源包示例所在的具体场景域整体拒绝、仅抽象记录——域类别本身也不在共享树点名,名单在 per-host 私档;私有 denylist 审计 alias_audit_ok 由维护者侧持有、包内不可独立复核——名单本身就是需保密的标识符集合,随包公开名单即构成泄漏,故该不可复核性是脱敏设计的结构性边界而非可修缺口;包内可核的是 generic 扫描 token 与评审对全文的直接检视,此分级如实声明);三条规则不作业界 state-of-the-art 认领(generic framing,guidance stands alone)。**RED-baseline 为 headless 差分 applied differential**(claude-sonnet-5,n=3/臂,判分维度与通过线**各自冻结于其对应重跑之前**——初版三维冻结先于任何编辑,评审驱动的后增维(含两个安全维)各冻结先于其首跑、历次仪器修订全披露且时序锚于平台侧 SHA-keyed CI run 史(证据分级见包内 grading),臂快照与最终候选逐字节 diff 校验,隔离 cwd,污染 grep 零命中;判分语义为准、token grep 仅初筛代理,两例代理误报已改判并标注):行为学主张口径=规则可及性/引出差分(按清单作答能否产出义务;真实执行场景的决策行为未测、如实声明),唯一取自终态臂计分批(raw 全入包、臂逐字节校验;判分语义为准且覆盖全答;仪器修订全披露、定义冻结于对应重跑前):十一个差分维终批均 old 0/3 → new 3/3(D-verifyonly 曾有一批 old 读数 1/3——经既有删除条款类推达成,跨批漂移如实留痕于包内存档表)、C-ctl 两臂 3/3(后六维为评审指出安全关键条款需维覆盖后按冻结-先于-重跑纪律逐轮新增;D-acl-pre 初版误限 A 段判 1/3,评审对 raw 复核后按冻结全答判据改判,勘误留痕);D-acl-over 不作差分主张——校准表述:终态计分批 old 臂三跑均答清单未规定(两跑附条件性说明);此前一计分批曾观测 old 臂一跑经既有「拿不准按命中」兜底直接到达同一处置(观测保留于包内 grading 存档表与作者 charter)——该维对推理深度敏感且跨批不稳,故不以差分立论。该子句保留依据=双 lane 收敛暴露面 + closeout 触发枚举一致性修复(过宽暴露此前不在第 4 条枚举、操作者按编号流程进不了该状态,评审指出后补入并区分为收敛路径)。无角色存档 raw 按 convergence-by-deletion 移出树(三轮包内核验发现的源头;原因表与 charter 记录保留,ctl-invalid 组 raw 因自证保留)。**过程如实披露**:存档批(原因表保留、raw 除自证组外移出)(v1prompt-old、误吞②节标题的 arm-stale——由 governing-chain-diff 派生当场抓到并修复、v1prompt ctl-invalid、ACL 前 arm-stale、v1prompt-acl ctl-invalid、规则#2 收窄前的 old/new 两组)均不承载任何主张或依据,原因表在包内 grading、raw 除自证组外已按 convergence-by-deletion 移出树;终态派生 SKILL 5 行、reference 8 行(6 条改写行——含授权不投放行增「可回读授权来源」要件而原义务逐字保留——+ 2 条仅因父行改写而链变、自身文本逐字未动的子句;派生输出入包 chain-derivation.txt,base=当前 origin/dev)——closeout blocked/待安全处置行因新增触发就地改写,其既有删除类触发枚举逐字保留可 grep 复核,新增只扩触发集合不弱化既有强度,新 ①b 节为纯新增。候选绑定证据 `eval/evidence/placement-face-2026-08-28/`(REPLAY 零损对照含实测输出、臂快照、prompt 生成器、raw 12 份(计分 6 + ctl-invalid 自证 3——每个无效批保留省答那一跑 + 批 23 计分失败自证 3;判定不依赖的存档 raw 均按 convergence-by-deletion 移出树)、逐跑判分、SHA256SUMS、AGENTS 契约),除已按证据类分级如实声明者(私有 denylist 审计=维护者侧持有、冻结时序=平台侧 CI run 锚 + 包内部分自证,见包内 grading)外均包内可独立复算。 |
|
|
360
360
|
| 入口字数预算下,SKILL 交付面 bullet 的新增 cue 以收缩三处纯示例括注抵消——括注载体逐字/语义保留于 reference,义务本体句中未动 | `tighten-doc` | SKILL.md 交付面 bullet 三处括注收缩(面枚举括注、被要求删除触发枚举括注、指针三态→四态与加「定位判据」),义务零损;reference 侧为 ①b 新增 + closeout 触发扩展的就地改写(既有触发逐字保留); result-class: stable-success; behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/tighten-doc/SKILL.md#逐面判定,未判定不得报「已更新 / 已同步」:**①投放与定位都要回读远端 | `updated` | owner key `tighten-doc/SKILL.md`。零损对照表(每条被收缩片段 → reference 现行载体原文 → 复算命令与实测计数)在 `eval/evidence/placement-face-2026-08-28/REPLAY.md` §1;row set 由 governing-chain-diff 机械派生(SKILL 5 行=同一 bullet 改写行,reference 8 行=closeout 要件/触发扩展的 6 条就地改写行加 2 条仅因父行改写而链变、自身文本逐字未动的子句(派生输出入包 chain-derivation.txt,base=当前 origin/dev),既有触发逐字保留 grep 可核,①b 节纯新增);入口预算复测 entrypoint_word_budget_legacy_ok(9826≤9828)+ entrypoint_size_blocking_ok。保全清单三源走查:台账 firing-path 解析进本包锚点未触碰;校验脚本钉扎两遍扫无命中本轮改动短语;「三态定义」等旧短语的仓内残留仅在 062 冻结证据包(其 AGENTS 契约声明 concluded/candidate-bound,REPLAY §3 记交叉说明)。 |
|
|
361
|
+
| A Node.js implementation owner must resolve the repository's live runtime/module contract before code, keep asynchronous work bounded and cancellable, treat built-in TypeScript execution as distinct from type checking, and route architecture, diagnosis, test-layer policy, observability, connectivity, release, and terminal contracts to their existing owners | `nodejs-service-dev` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/nodejs-service-dev/SKILL.md#Do not swallow rejections or resume normal operation; result-class: insufficient-evidence; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-service-impl | updated | Owner key: `nodejs-service-dev/SKILL.md`. Initial RED: the routing-bank integrity gate rejected four fixtures because `nodejs-service-dev` did not exist. Candidate adds the owner, three task references, maintainer source map, catalog/README registration, and positive plus high-overlap negative routing fixtures. Primary technical claims were checked against live Node.js/npm documentation; OWASP/OpenSSF and two public skill corpora were used only to challenge coverage and skill shape. Deterministic GREEN and repository gates prove registration/conformance, not production behavior improvement; runtime-version and framework-specific claims remain live-check obligations. Sibling disposition: `testing-strategy`, `defect-diagnosis`, `product-rd-workflow`, `terminal-cli-dev`, `platform-observability`, `platform-service-connectivity`, `platform-release-engineering`, `web-react-dev`, and `llm-inference-integration` remain unchanged because the new skill routes to their existing contracts rather than copying them. |
|
|
362
|
+
| A catalog regression fixture that copies candidate catalog text must snapshot the candidate skill roots too; cloning committed HEAD while a new skill exists only in the working tree creates a false pristine mismatch and cannot validate the tree being landed | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh; result-class: failure; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-service-impl | updated | `make test` RED reproduced c6 as `extra_in_catalog=nodejs-service-dev`: the fixture cloned 32 committed skill roots but copied the 33-entry candidate catalog. `new_case` now replaces the temporary clone's `skills/` with the candidate working-tree snapshot before mutation; the full 12-case catalog suite turns GREEN. The gate assertions and failure tokens are unchanged, and all destructive fixture operations remain confined to the validated `mktemp` clone. |
|
|
363
|
+
| Trigger accuracy and body effectiveness are different claims for a new stack skill: route-bank examples can show owner selection while still proving nothing about the quality of the implementation advice, so both surfaces need frozen tasks and explicit failure criteria | `nodejs-service-dev` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/nodejs-service-dev/SKILL.md#Do not swallow rejections or resume normal operation; result-class: insufficient-evidence; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-test-strategy | updated | Owner key: `nodejs-service-dev/SKILL.md`. User correction required an explicit effectiveness surface after the initial candidate had only routing and repository conformance evidence. Four advisory human-judgment fixtures now cover runtime/module/TypeScript contracts, streaming cancellation/backpressure/shutdown, bounded CPU workers, and the testing-strategy-to-Node-mechanics handoff. They freeze prompts, rubrics, weak-answer criteria, and owner pointers; they do not claim a measured with/without improvement until repeated paired runs and transcript/outcome review exist. The shape follows a local skill repository's useful separation of full-catalog routing trials from blinded with/without body evaluation, while its product-specific names, infrastructure, stored outputs, and version policy remain out of this shared tree. |
|
|
364
|
+
| A candidate-tree regression fixture must snapshot the candidate task banks as well as skill roots: a new source-register row can legitimately point at an uncommitted routing or behavior fixture, and pairing that row with the committed eval tree makes the supposed pristine arm internally impossible | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh; result-class: failure; bank-evidence: file:eval/behavior-fixtures.jsonl#F31 | updated | Final `make test` reproduced c6 after the Node.js effect fixtures were added: the temporary clone received the candidate skills/source-register but retained HEAD's older eval banks. `new_case` now snapshots both `routing-tasks.jsonl` and `behavior-fixtures.jsonl`; c13 constructs candidate-only markers under `TEST_ROOT` so deleting either copy turns the regression suite RED even after today's fixtures are committed. The scope remains the two task banks consumed by shared-skill validation; no private outputs or unrelated eval artifacts are copied. |
|
|
365
|
+
| Before-review implementer closure for the new Node.js owner and its candidate-snapshot regression repair: acceptance is a concrete Node.js implementation owner with evidence-bounded runtime, async lifecycle, security, verification, testing-owner handoff, routing positives/negatives, and effect-evaluation fixtures, while preserving existing owner boundaries and making the catalog fixture validate the actual candidate tree | `nodejs-service-dev` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/nodejs-service-dev/SKILL.md#Do not swallow rejections or resume normal operation; result-class: stable-success; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-test-strategy | updated | Owner key: `nodejs-service-dev/SKILL.md`. **Changed-file scope, frozen before independent review/challenge:** `README.md`; `docs/SKILLS.md`; `eval/behavior-fixtures.jsonl`; `eval/routing-tasks.jsonl`; `packages/ccl-skills-npm/README.md`; `skills/nodejs-service-dev/SKILL.md`; `skills/nodejs-service-dev/agents/openai.yaml`; its four `references/*.md`; and this source register. **Load-bearing invariants and failure paths:** architecture/diagnosis/test-policy/CLI/observability/connectivity/release requests must redirect instead of being swallowed; Node test mechanics start only after `testing-strategy` chooses layers and coverage; version-sensitive claims require live repository/runtime or primary-doc evidence; TypeScript execution never substitutes for type checking; cancellation, backpressure, worker termination, fatal shutdown, dependency freezing, and secret handling stay explicit rather than happy-path examples; candidate regression clones must snapshot skills and both referenced task banks before mutations, remain confined to `mktemp`, and fail if either snapshot is omitted. **Self-check evidence:** the four Node routing rows and four F31-F34 body fixtures cover positive and high-overlap negative boundaries; catalog c6/c13 reproduced both snapshot defects before turning green; final `make test`, heavy regression lane, public sanitization, focused routing/fixture/catalog checks, and `git diff --check` were green on the then-current candidate. **Residual risks:** the body fixtures establish a falsifiable evaluation surface but no repeated blinded with/without trial has yet measured an effect delta; Node/runtime and ecosystem facts can drift and remain live-check obligations; framework-specific build conventions stay repository-local rather than being guessed into the shared owner. Dual-track classification: `dual-track-review-gate.md` first table row, non-wording new shared skill, so independent fact/consistency review plus adversarial challenge are both required; this row is the pre-review ordering record, not independent proof. |
|
|
366
|
+
| A self-review record written into the candidate solely to prove review readiness changes the packet it is meant to freeze, needlessly invalidates tests and review identity, and can recurse when outcome rows are added; ordering evidence and candidate evidence need separate lifecycles | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/extraction-quickstart.md#Run deterministic checks and implementer self-review first; result-class: failure | updated | The preceding Node.js closure row exposed the defect: adding it after a green exact-candidate run made the candidate dirty again and started another full suite even though no implementation, routing, runtime rule, or validator behavior had changed. That run was interrupted rather than credited. The quickstart now requires a fresh non-overwritten task-evidence path outside the candidate, passed verbatim as `--review-plan-file`; the gate result binds its profile hash and preserves the review ordering. Candidate-local rows remain valid only when the row is itself a substantive deliverable under review. Changing only external self-review evidence refreshes profile binding but does not invalidate implementation tests or candidate packet identity. The prior row remains append-only history and is superseded only for its storage mechanism; its substantive Node.js self-review claims are recopied into the external plan for this review. |
|
|
367
|
+
| Candidate skill-root snapshotting needs its own permanent candidate-only marker: relying on today's uncommitted new skill makes the regression turn inert as soon as that skill enters HEAD, so deleting the snapshot later can stay green until the next working-tree-only skill exposes the old false-pristine mismatch | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh; result-class: failure; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-service-impl | updated | The adversarial challenge found that c13 permanently mutated both task banks but did not create a candidate-only skill root. `new_case` now accepts separate candidate skill and eval roots; c13 builds all three marker surfaces under `TEST_ROOT`, passes them explicitly, and asserts the skill marker reached the case clone. Removing the skill-root copy therefore turns c13 RED independently of whether `nodejs-service-dev` is already committed. Default arguments preserve every existing case, and all writes/destructive cleanup remain inside the validated temporary clone or `TEST_ROOT`. |
|
|
368
|
+
| A Bash test label containing Markdown backticks must be passed as a literal; otherwise command substitution can emit stderr and erase label text while the suite still exits zero | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh; result-class: failure | updated | RED reproduced the existing A13 false green: suite rc=0, the literal `` `-` `` was absent from stdout, and stderr reported `-: command not found`. The durable A13 report oracle then made an applied restoration of the old double-quoted call fail with suite rc=1 and `A13: RED (summary lost literal `-`)`; the final single-quoted call is GREEN with rc=0, the complete A13 summary present, and empty stderr. Gate verdict and fixture semantics are unchanged. |
|
|
369
|
+
| A closure row with one owner-scoped firing path must not combine multiple owners: the Node rule anchor can adjudicate the Node owner, while workflow verifier changes retain their own rows and executable firing paths | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh; result-class: failure; bank-evidence: file:eval/behavior-fixtures.jsonl#F31 | updated | Owner key: `skill-extraction-workflow/SKILL.md`. The explicit-base repository gate rejected the combined closure row for `skill-extraction-workflow` because its only firing path resolved inside `nodejs-service-dev`. The closure row now binds only the Node owner; the candidate-snapshot workflow repairs remain covered by their dedicated RED-baseline rows and the changed catalog regression executable. No owner requirement or verifier result was downscoped. |
|
|
370
|
+
| A new implementation decision owner must be registered in the impact-chain owner set before its description-level bank evidence can be resolved; otherwise the gate can demand evidence from an owner while making that owner's row unreachable | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh; result-class: failure | `updated` | Owner key: `skill-extraction-workflow/SKILL.md`. A20 changes the fixture's committed `nodejs-service-dev` description, supplies an owner-scoped firing path and external bank-evidence locator, and reproduced `impact_chain_bank_evidence_missing` before registration. Adding this implementation owner to the existing curated set makes A20 GREEN without changing generic leaf resolution or curated-set semantics for other skills. |
|
|
371
|
+
| A new stack implementation owner joining an already-swept class must inherit the class-wide obligations its siblings carry — the language-stack CLI carve-out with reciprocal terminal-cli-dev skip legs, and the localized-refactor trigger pair — or requests in that language route to a skill that never claims the work | `nodejs-service-dev` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/nodejs-service-dev/SKILL.md#language-stack CLI implementation must not be routed back; result-class: stable-success; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-cli-not-terminal | updated | Owner key: `nodejs-service-dev/SKILL.md`. Independent review of the initial candidate flagged the asymmetry: python-service-dev and go-microservice-dev both advertise standalone CLI/tooling (with reciprocal terminal-cli-dev skip legs) and localized-refactor triggers, while the new Node owner claimed neither. RED-baseline is a coverage differential: before this change, CLI/命令行工具 and 重构/refactor wording had zero hits in the nodejs-service-dev description and routing rules; after it, the CLI description trigger, the owner-side routing rule, the reciprocal skip leg, two pointer-integrity anchors, the route-nodejs-cli-not-terminal fixture, and the localized-refactor trigger pair (重构 Node 服务里的某文件/某类(局部) / refactor a file/class within a Node.js service) are all present, mirroring the proven Python/Go legs. Per-obligation disposition for the remaining class obligations: defect-diagnosis-first, multi-stage→product-rd, and testing-strategy handoffs were already present in the initial candidate (`unchanged`); the bare service-wide refactor trigger is `not-applicable` — nodejs has no `*-architecture` sibling, and service redesigns already route to product-rd-workflow in the routing rules. The interface-contract split is preserved: the command/flag/help/exit/TTY contract stays with terminal-cli-dev. The routing fixture remains a frozen advisory input under the same evaluation boundary as the other Node fixtures in this round. |
|
|
372
|
+
| Extending a named skip-leg list is a routing-surface change on the skip-side owner too: the leg is only real when both sides land together, and the eval bank must assert the skip-side owner does not receive the redirected request | `terminal-cli-dev` | behavioral-evidence: semantic-control; observed-failure: no; bank-evidence: file:eval/routing-tasks.jsonl#route-nodejs-cli-not-terminal; result-class: stable-success | updated | Owner key: `terminal-cli-dev/SKILL.md`. Description-only change: the existing skip clause gains the `Node.js CLI → nodejs-service-dev` leg beside the Python and Go legs; the generic predicate ("a language whose dev skill owns it") and the no-owner fallback are unchanged, and the non-rendered command/flag/help contract stays owned here. Semantic-control: the change instantiates an already-pinned rule class for one more stack rather than introducing a new behavior claim — the Python and Go legs carry the RED history, and this leg reuses their oracle shape (reciprocal trigger anchor, skip-leg anchor, and a must-not-route fixture asserting terminal-cli-dev does not receive the Node CLI request). |
|
|
373
|
+
| A sibling-generalization map that only checks whether the new member copies sibling content misses the join-side dual of class-wide coverage: the obligations the class already landed on its members (triggers, skip-leg reciprocity, pinned anchors, fixtures) silently fail to transfer to the newcomer | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#must not be counted as obligation-inheritance coverage; result-class: failure | updated | Owner key: `skill-extraction-workflow/SKILL.md`. Observed failure: the initial nodejs-service-dev candidate shipped with a sibling map recording "siblings unchanged because the new skill routes to their existing contracts", yet the new member lacked two class obligations its siblings carry (the CLI carve-out reciprocity and the localized-refactor trigger pair) — the map answered the copy-content question and never asked the obligation-inheritance question, and only independent review caught it. RED-baseline: before this change the class-wide COMPLETE-set rule fired only on change-side sweeps ("every member of class C should carry X"), with no clause firing on member-join; the new Member-Join Inheritance section makes the join-side enumeration an explicit obligation with a per-obligation disposition requirement, and the SKILL.md class-wide bullet gains a word-neutral member-JOIN pointer (two pure example parentheticals moved verbatim into the section to offset the pointer under the entrypoint's zero-growth budget). |
|
|
374
|
+
| A created-routing-surface obligation must be bounded to entrypoints absent at the round's base: an existing non-curated skill editing its description otherwise incurs bank evidence that no ledger row can bind, because row-to-owner resolution spans only the curated name set — the gate demands evidence while making it unreachable | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb; result-class: failure | updated | Owner key: `skill-extraction-workflow/SKILL.md`. Observed: CI repository-gates went red on this branch with impact_chain_bank_evidence_missing owing the skip-side owner after the skip-leg round touched an existing non-curated description; the local run on the committed tree reproduced it, and a probe showed the owner's row set empty because name-level resolution iterates only resolvable curated names, so the appended row could never discharge the obligation. Fix bounds the created-surface pickup with an entrypoint-absent-at-base check via the existing regular-blob reader, exactly matching the block's stated brand-new-skill intent; the forcing function for genuinely new skills is preserved, since a base-absent entrypoint still joins the triggered set and still requires curation before its evidence resolves. New suite leg A21 holds an existing non-curated description edit green and turns red if the existence bound is removed. The skip-side owner's ledger row from the skip-leg round remains as unbound documentation by design. |
|
|
375
|
+
| Runtime-visible design work uses one context- and criterion-driven lifecycle rather than separate design checkpoints: design brief → test Phase 0 → producer/client execution records → test Phase 1/sufficiency → candidate-bound design verdict; each claim closes every required, non-substitutable evidence dimension, and heuristic review is risk discovery rather than acceptance proof | `product-ui-ux-design` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-ui-ux-design/SKILL.md#Automated accessibility checks do not replace | `updated` | Owner key `product-ui-ux-design/SKILL.md`. The pre-change focused contract test failed on the missing canonical record and reciprocal owners, four arbitrary five-second thresholds, and four eponym-only extraction anchors; later challenge also found duplicated client/test evidence ownership, reversed affected/unknown consumer status, and an over-broad rejected-surface rule. The owner entrypoint was compressed, `delivery-contract.md` added, the execution checklist converted to a profile router, theory rebuilt as an authority-classed claim ledger, source/code maps updated, and reader docs synchronized. Original-source classes include ISO human-centred/usability framing, W3C normative versus informative material, primary empirical papers with population/task limits, platform-scoped guidance, and the non-standard Design Tokens Community Group report. Current client-code extraction contributes only source-neutral static-source candidates, each limited to the implementation/test/script/CI subset actually observed; it makes no test-execution, render, runtime, or product-effect claim. The older platform-walkthrough model is explicitly superseded; its three immutable historical locators are row-digest-bound in `register-firing-path-resolution.rb`, not silently restored. Program record: [065 UI/UX evidence delivery](../../../specs/065-uiux-evidence-delivery/plan.md). |
|
|
376
|
+
| UI/UX testing no longer waits for a finalized design checkpoint that itself waits for test-layer selection: Phase 0 fills assertion/rendered layers and oracles from a testable design brief, while post-producer/client Phase 1 cites the complete design/test/producer/client record set, adds only test-owned evidence, and owns criterion results plus sufficiency without claiming the holistic design verdict | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#For every runtime-visible UI/UX slice, load the canonical sequence | `updated` | Owner key `testing-strategy/SKILL.md`. Baseline inspection found the circular checkpoint wording and no schema-level reciprocal contract; cross-owner challenge later found Phase 1 and the client owner duplicating the same evidence. The current sequence is Design brief → Phase 0 → producer/client execution → Phase 1/sufficiency → design verdict, and every design, test, producer, and client role writes its own facts once. A current bounded Design brief also satisfies testing's scope gate instead of forcing the user to restate it. The focused contract test requires both testing passes, their order, the shared-record handoff, single-write rule and bounded-scope reuse. Program record: [065 UI/UX evidence delivery](../../../specs/065-uiux-evidence-delivery/plan.md); [replayable validation](../../../specs/065-uiux-evidence-delivery/validation-evidence.md). |
|
|
377
|
+
| React/browser implementation consumes the shared design brief and Phase 0, writes candidate-bound Web runtime facts once into the shared client record, and lets testing Phase 1 own criterion interpretation/sufficiency before the design verdict | `web-react-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/web-react-dev/SKILL.md#For every visible UI change, load | `updated` | Owner key `web-react-dev/SKILL.md`. Baseline had one-way checkpoint consumption but no equally explicit closeout; a later challenge found direct-to-design return could skip testing sufficiency. The current Web return names route/server, viewport/container, artifacts, tested states/input, raw criterion-mapped observations, console/network observations, coverage and gaps exactly once in the shared record. Browser render remains scoped evidence rather than self-acceptance. |
|
|
378
|
+
| Cross-platform/native App implementation keeps its device evidence obligations, writes candidate-bound runtime facts once into the shared record, and leaves criterion sufficiency/verdict to testing/design owners instead of duplicating their state rules | `app-cross-platform-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/app-cross-platform-dev/SKILL.md#For every visible UI change, load | `updated` | Owner key `app-cross-platform-dev/SKILL.md`. Baseline App guidance already had strong post-render obligations, but cross-owner challenge showed its direct-to-design return could duplicate or bypass testing sufficiency. This round preserves device/form-factor, safe-area, keyboard, orientation, text-scale, lifecycle and artifact evidence, removes duplicated checkpoint/verdict schema, and routes the single client record through testing Phase 1 before verdict. |
|
|
379
|
+
| Mini-program implementation consumes the same design/Phase 0 record, translates it into each shipped host, and writes host/tool/device, capability/permission, state/dimension, artifact, coverage-boundary and gap facts once; browser/H5 preview cannot stand in for a shipped host | `miniapp-product-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/miniapp-product-dev/SKILL.md#For every visible UI change, load | `updated` | Owner key `miniapp-product-dev/SKILL.md`. Baseline repeated the design gate but lacked an equal closeout. The owner now keeps host-specific mechanics and runtime gates locally, writes raw criterion-mapped observations to the shared client record, and leaves sufficiency/verdict to testing/design owners. |
|
|
380
|
+
| Every user-facing terminal/CLI contract—including ordinary plain-text command trees, flags/defaults, help/output/exit behavior, confirmations, progress, and recovery—is a design surface: it consumes the shared design/Phase 0 record and writes the applicable command or PTY/runtime, dimension, fallback, interaction, artifact, coverage-boundary, and gap facts once | `terminal-cli-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; bank-evidence: file:specs/065-uiux-evidence-delivery/validation-evidence.md#基线与候选各完成 10 轮有效观测; firing-path: file:skills/terminal-cli-dev/SKILL.md#For every user-facing terminal/CLI contract change | `updated` | Owner key `terminal-cli-dev/SKILL.md`. Baseline had a one-way page-slice/checkpoint rule but no equivalent closeout. The owner now translates the shared contract to ordinary CLI semantics or cell-grid and terminal-lifecycle evidence as applicable, writes raw observations once, and leaves Phase 1 sufficiency and the design verdict to their owners. |
|
|
381
|
+
| Product readiness distinguishes parser/library-only CLI internals from every user-facing command/help/output/exit/confirmation/progress/recovery contract, and sequences one canonical design brief → Phase 0 → producer/client records → Phase 1/sufficiency → verdict record rather than maintaining a second checkpoint schema | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#Before coding any visible UI change, load | `updated` | Owner key `product-rd-workflow/SKILL.md` routes Web/App/mini-program/desktop/terminal visible delivery and canonical status vocabulary to the shared contract, preserves the full `visible surface: no` boundary, and permits only a labeled review-only draft MR to obtain a required independent design verdict. `design-routing-and-readiness.md` owns product sequencing, not a second evidence schema. Program record: [065 UI/UX evidence delivery](../../../specs/065-uiux-evidence-delivery/plan.md). |
|
|
382
|
+
| UI/UX extraction turns named theories and local source patterns into bounded observation → risk/mechanism → hypothesis → observable check → evidence-boundary records, and a deterministic regression gate protects the canonical design/testing/producer/client loop and retired folklore thresholds | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. `references/uiux-judgment-extraction.md` removes eponym-only instructions; the new executable test is registered in the fast regression lane and checks reciprocal owners, five top-level stages including both testing passes, verdict binding, profiles, evidence boundaries, primary-source ledger entries, retired five-second rules and resolvable references. The test was RED on the pre-change corpus and is GREEN on the current candidate. The firing-path resolver gains three explicit row-digest-bound waivers for the superseded platform-walkthrough locators so append-only history stays intact without pretending retired criteria still execute. |
|
|
383
|
+
| A Go service that emits or stores possibly client-rendered text closes only after authoritative consumer-universe classification; a known rendering client is `affected`, while only an incomplete/inaccessible universe or member is `unknown-consumers` | `go-microservice-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/go-microservice-dev/SKILL.md#create the applicable full or lightweight record in | `updated` | Owner key `go-microservice-dev/SKILL.md`. Cross-skill audit found a retired checkpoint alias and no canonical pointer; later challenge found the first rewrite incorrectly classified known visible consumers as unknown. The current rule proves an authoritative universe, records each member, routes `affected` clients into the full/lightweight contract, reserves `unknown-consumers` for incomplete/inaccessible evidence, and permits backend-only closure only when the complete inventory proves no client rendering. |
|
|
384
|
+
| A Python service that emits or stores possibly client-rendered text uses the same authoritative consumer-universe classification and canonical design/test/producer/client handoff: known rendering clients are `affected`; incomplete/inaccessible evidence is `unknown-consumers` | `python-service-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/python-service-dev/SKILL.md#create the applicable full or lightweight record in | `updated` | Owner key `python-service-dev/SKILL.md`. The baseline allowed a current-repo search to stand in for a complete inventory; the first rewrite also conflated affected and unknown. The current rule requires an authoritative universe plus per-member disposition before backend-only closure and routes affected/unknown states without turning known UI work into a false blocker. |
|
|
385
|
+
| An inference path that emits or persists possibly client-rendered copy, labels or generated status uses the canonical full/lightweight record when a known client is `affected`, reserves `unknown-consumers` for incomplete/inaccessible evidence, and cannot accept a surface from inference-side evidence alone | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/SKILL.md#you must load `../product-ui-ux-design/references/delivery-contract.md` | `updated` | Owner key `llm-inference-integration/SKILL.md`. The baseline allowed a bounded search over an author-chosen subset; the first rewrite conflated affected and unknown. The current rule requires an authoritative universe plus per-member disposition, routes known affected clients to their owner/testing stages, and leaves incomplete/inaccessible consumers blocked as unknown. |
|
|
386
|
+
| UI/UX contract and cross-owner RED-baseline claims remain independently replayable instead of living only in narrative source-register rows | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. [Replayable validation evidence](../../../specs/065-uiux-evidence-delivery/validation-evidence.md) binds the unchanged `origin/dev` corpus, candidate oracle, controlled RED exit, current GREEN result, built-in five-stage, entry-router, trigger-loss, role-member-drop, dirty-as-commit, mutable-external, owner-narrowing, ordinary-CLI, negated-return, affected-vs-unknown, and user-accepted-gap mutations, historical-locator waiver mutations, commands and evidence limits. The per-obligation carrier proof is kept separately in [obligation preservation](../../../specs/065-uiux-evidence-delivery/obligation-preservation.md). |
|
|
387
|
+
| A generic long-digit leak guard must distinguish opaque identifiers from canonical public identifiers at the matched URL span; exempting a whole line hides a second identifier, while rejecting every long digit blocks primary-source citations | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh | `updated` | The quick gate rejected three canonical DOI/W3C report URLs after the theory ledger restored original-source links. The checker now keeps the lexical domain scan unchanged and scans each long digit run separately. Only a run wholly inside a canonical `doi.org` or W3C Community Report URL without query/fragment is ignored; an unrelated long ID on the same line still fails. The regression covers allowed DOI/report fixtures, the same-line long-ID bypass, retained lexical terms, and end-to-end failure attribution. This is a narrow public-identifier classification, not a broad URL or line allowlist. |
|
|
388
|
+
| Runtime-visible delivery is a set problem, not a single-owner pipeline: producer changes can alter client state through API/event/schema/status/permission/result shape; source extraction and shared-system work compose with delivery depth; embedded content and host layers need separate owners, runtime records, and immutable binding members | `product-ui-ux-design` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-ui-ux-design/references/source-map.md#Implementation rules follow the complete affected client-owner set | `updated` | Independent cross-owner review reproduced five false-green paths: non-string producer changes bypassed the design contract; profile depth incorrectly cancelled orthogonal source/shared-system work; WebView/mini web-view/Electron collapsed content and host evidence into one owner; `base+dirty-sha256:` passed with an empty payload; and a backticked cross-skill pointer did not resolve. The contract now uses delivery-depth plus composable work-mode/risk unions, complete design/test/producer/client record and candidate-binding sets, closed non-empty binding kinds, and a dirty-bundle manifest over base, binary tracked diff, and sorted untracked path/mode/content. Go/Python/inference entries now require the applicable design record before Test Phase 0 and actual producer/client execution returns for affected consumers, or backend/product/testing handoff plus API/log/output evidence when the authoritative universe proves no client can react. Naming an owner or inventory alone is not closure. Focused mutations kill narrowed producer scope, missing pre-Phase-0 design records, name-only routing, single-choice profiles, dangling pointers, lost trigger classes, narrowed inner owner sets, dirty-as-commit or mutable-external bindings, ordinary-CLI escape, and missing design/test/producer/client members. |
|
|
389
|
+
| A contract oracle must parse the active contract structure it claims to protect; hard-coded test-local status sets, first-heading matches, and whole-file literals can certify extra contradictory rows, duplicate/out-of-order stages, or obligations moved into examples/comments | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh | `updated` | Independent gate review showed the former status fixture tested its own `case` statement instead of the contract table; the stage “swap” deleted a heading rather than swapping it; duplicate headings and fenced/commented carriers survived. The focused oracle now strips fenced blocks and HTML comments, requires each active stage heading exactly once and in order, derives the exact closed verdict/state and binding-kind sets from live Markdown tables, and includes killing mutations for real swaps, duplicates, added contradictory rows/kinds, fenced/commented obligations, and fenced tables. This proves those static contract properties only; actual Agent routing/effect remains a later task-outcome measurement, not inferred from green text gates. |
|
|
390
|
+
| Public-identifier exceptions in a privacy scanner require identifier grammar, not merely a trusted host prefix; otherwise arbitrary long IDs hidden under a DOI/W3C-looking path bypass the rule | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh | `updated` | Review probes showed `doi.org/not-a-doi/<opaque-id>` and `www.w3.org/community/reports/not-a-report/<opaque-id>` passed the first exception while separated IDs were blocked. The scanner now accepts only DOI paths beginning `10.<registrant>/...` and W3C Community Group `CG-FINAL-...<date>` report paths, still excluding query/fragment spans. One end-to-end fixture independently requires failures for same-line IDs, arbitrary DOI/report paths, query IDs, fragment IDs, noncanonical hosts, and IDs attached after a valid report URL; every marker must appear in the scan output so one caught class cannot mask another. |
|
|
391
|
+
| A shorter skill entry preserves capability only when each former task trigger still reaches its operative reference; a file that remains in the package but has no applicable inbound route is deleted behavior, not a successful rehost | `product-ui-ux-design` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-ui-ux-design/references/design-intake-and-acceptance.md#The changed producer owners, affected client owners, Test Phase 0 handoff | `updated` | Independent origin/dev-to-candidate trigger review found that the first compressed router left naming/version synchronization unreachable, limited the audit procedure to systemic redesign, and made multi-stack guidance reachable only through a token side path. A complete pass then found the same class for design-to-code primitives, narrow layout/state/screenshot work, local code-evidence audits, same-stack multi-subproject themes, and generic non-scenario surface/loop modeling. The router now chooses one delivery depth and unions every orthogonal source/code-evidence, design-to-code, audit, naming/version, same-stack, multi-stack, and risk lens whose trigger applies. Focused tests parse the live active-Markdown rows, require each old task class to resolve to its reference, and kill orphaned-pointer or narrowed-trigger mutations. The restored references were also reconciled to the canonical complete affected client-owner set and producer handoff so reactivation cannot hard-code React or mobile ownership for Vue/Svelte, Electron, terminal, desktop, or TV surfaces. This proves static reachability and owner consistency only; whether the shorter entry improves real Agent task outcomes still requires comparable task trials or longitudinal evidence. |
|
|
392
|
+
| Entry compression must be proved against immutable baseline obligations, not inferred from shorter files: every changed pre-existing skill obligation has one exact live carrier or a reviewed retirement, and unresolved rows block completion | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh | `updated` | The first compressed UI/UX candidate orphaned real routes, its first 544-row projection became stale after later semantic fixes, and the first parser skipped Markdown table cells. The canonical human-reviewed source is `specs/065-uiux-evidence-delivery/obligation-mapping.jsonl`; `obligation-ledger.py` derives the changed pre-existing row set from the immutable base, rejects fuzzy or provenance-only carriers, accepts a carrier bundle only when the old obligation is mechanically separable and every clause closes exactly once, and generates `obligation-preservation.md` as a reader projection. Thirty ledger mutations cover row-set drift, stale or duplicate carriers, qualifier weakening/reversal, table hiding, retirement misuse, and invalid bundles. The companion UI/UX oracle reproduces 236 baseline failures and kills 91 independent contract mutations (81 direct branches plus 10 parameterized owner/Phase-0 cases). These checks prove static preservation, reachability, and selected contract properties only; they do not prove lower task time, fewer corrections, better rendered design, or production effect. |
|
|
393
|
+
| Routing-bank evidence covers every skill description independently of the curated impact-chain owner set: a non-curated owner with valid evidence must pass, the same change with no row must fail, and one ambiguous row cannot discharge two owners | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`; implementation lives in `skills/skill-extraction-workflow/scripts/impact-chain-gate.rb`. Hosted CI exposed that `terminal-cli-dev` had a valid owner-scoped bank locator but was looked up only in the selected-owner map. Reproduction then isolated both directions: A20 failed with valid evidence, while A21 incorrectly passed with no ledger row because the outer gate never entered. The gate now has a bank-only resolver over base/head skill names and enters on any changed skill entrypoint, while selected owners keep the original impact-chain resolver. A20/A21 turn green, A22 rejects a two-owner row, A23 proves a stale exact owner path is rejected by the earlier missing-file gate, and round-attribution Leg M remains green for a non-curated body-only change. A separate lineage trace found no reachable difference in which the selected resolver loses a valid bank row: extra in-round owners remain ambiguous, absent exact paths fail missing-file, and absent package-prefix aliases fail the earlier ambiguity gate. Thus routing-bank coverage widens without widening impact-chain jurisdiction. |
|
|
394
|
+
| A preservation audit that lives only as a documented manual command drifts silently, and a Markdown obligation parser that ignores CommonMark container and code-span semantics lets inert example text become obligations or lets identifier-shaped URLs launder private numeric IDs; the real mapping/ledger must be audited by a registered gate against the base commit pinned in its own header | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. Independent review of the delivery found the committed obligation mapping stale at the PR head: the documented completion audit exited 1 while validation evidence recorded it as passing, and no CI lane ran the real-repository audit. The repaired mapping restores `audit_ok` over 1240 rows, and the new heavy-lane suite `test_obligation_ledger_repo_audit.sh` audits the real mapping/ledger against the base SHA pinned in the ledger header, so later carrier drift turns CI red instead of shipping silently. `obligation-ledger.py` parsers were aligned with CommonMark across eight adversarial review rounds: code-span pairing in table cells, section bodies sliced from masked visible text, clause splitting on code-span-masked structural text, and container-aware indented-code masking for top level, list items, and block quotes with quote-relative indentation; the parser-boundary leg in `test_obligation_ledger.sh` replays each defect RED on the pre-change tool. The long-digit scan in `check-ccl-skills.sh` now scopes its DOI/W3C exemption to the identifier payload with a punctuation-merged nine-digit cap, closing split, padded, and report-name laundering while keeping numeric-suffix citations green; `test_entrypoint_domain_scan_terms.sh` pins eight RED laundering classes and the green controls. The regenerated ledger stayed byte-stable across every parser change, and the round-by-round record with rejected-escalation rationale lives in the delivery spec's review-fix record. |
|
|
395
|
+
| A merge that combines a bank-only owner resolver with a base-absent created-surface bound must re-adjudicate the bound's premise: once rows bind over base plus head skill names, exempting existing non-curated description edits only reopens the evidence-free routing-surface hole the resolver closed | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. Integrating dev into this branch auto-merged two divergent gate changes: this branch's bank-only resolver and dev's base-absent bound on the created-surface pickup, whose stated premise was that non-curated obligations were undischargeable. Under the merged gate legs A21/A22 went RED (zero-row description edits passed). The bound is removed with its comment rewritten to record the supersession; A21/A22 return green, dev's Node owner leg is preserved as A24, and dev's exemption leg is rewritten as A25 to pin the dischargeable half: an existing non-curated owner's description edit owes bank evidence and a bank-only row discharges it. |
|
|
396
|
+
| Merging two branches that each rewrote one routing description resolves at the obligation level, not the text level: the union carries both rounds' obligations, and any budget trim must come out of non-carrier wording | `terminal-cli-dev` | result-class: stable-success; behavioral-evidence: semantic-control; observed-failure: no; bank-evidence: downscoped:PR76-INTEGRATION-TRIM-NO-BANK-RERUN | `updated` | Owner key: `terminal-cli-dev/SKILL.md`. Description-only integration: the compressed entry from this branch absorbs the `Node.js CLI → nodejs-service-dev` skip leg from dev beside the Python and Go legs, and the obligation mapping row for the skip sentence records the clause on both its before and carrier sides. The OpenCode budget overflow from the union is repaid inside the uncarried ownership sentence: the pinned short contract phrase and its parenthetical stay verbatim for the cross-skill pointer pair, and the fuller enumeration folds behind it as a with-clause; all four description carriers are byte-preserved and the re-pinned ledger audit stays `audit_ok` over 1240 rows. |
|
|
397
|
+
| A ledger effect vocabulary without a neutral deletion value forces every retired obligation to close as strengthened, so summary counts read as zero loss while carrier-less deletions ship; deletion closes as retired, live-carrier rows may not claim it, and a migration byte budget must carry real headroom with recorded rationale instead of freezing its documents at a reverse-fitted cap | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. The 065 ledger recorded 98 retired-dead and 25 partial-retirement rows as effect=strengthened because the vocabulary had no other closable value, so the rendered Effects line read as zero obligation loss. EFFECTS gains retired; retired-dead rows and retired-part partitions must close as retired; a retired effect on a live-carrier row is rejected; the rendered summary reports retired separately. The specs/065 mapping is relabeled (preserved=657, strengthened=460, retired=123) and re-rendered under the pinned-base audit. Three applied mutations, each attributed differentially with the unmutated suite green: strengthened on a retired-dead row fails RETIRED_EFFECT_INVALID where it previously survived validation to STALE_LEDGER; strengthened on a retired-part partition fails PARTITION_EFFECT_MISMATCH where it was previously the required baseline; retired on a rehosted carrier row fails RETIRED_EFFECT_INVALID where it previously fell to the generic INVALID_EFFECT. The uiux loading budget's direct-runtime cap sat at exactly its measured value (58860 = 90% of 65400, a zero-byte margin that froze the entry and contract files); the cap moves to 95% with the cap-setting rationale and re-derive-or-retire policy recorded in the script, and a router-shrink-plus-contract-pad mutation kills the 95% predicate specifically while the 110% specialized cap stays green. |
|
|
398
|
+
| A review round's fixes are their own gate round: a budget cap verified by one off-suite mutation can be silently weakened later, and a mixed partial retirement that closes row-level as retired must not let its strengthened surviving parts vanish from the rendered summary | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. Independent review and adversarial challenge of the retired-effect round produced two applied findings. First, the loading budget's caps move into pinned variables with an in-suite threshold probe: the 50/95/110 contract pin plus a cap-versus-cap+1 boundary walk per predicate; an applied 95-to-100 weakening mutation turned the probe RED and was restored. Second, 24 of the 25 real partial-retirement rows carry a strengthened surviving part, and the row-level retired close dropped that dimension from the summary; the renderer now emits part-level outcome counters, the fixture gains a mixed row pinning survived-preserved=3, survived-strengthened=1, retired=2, and a new mutant proves an unreviewed partition still fails MANUAL_REVIEW_REQUIRED, hardening the refuted manual-review finding into an invariant. The round-1 claim that retired rows escape manual review was refuted empirically: flipping one real partial-retirement row's manual_reviewed to false fails the pinned-base audit at the schema4 partition entry check. |
|
|
399
|
+
| Bound review-plan intent: overflow preserves bytes; compaction keeps only identified core+latest | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_update_review_plan_intent.sh; result-class: failure | updated | Helper bounds no-follow UTF-8 reads, checks optional opened digest, and atomically preserves mode. Compaction requires `review-plan-intent-stable-core-v1` character-count+SHA identity; legacy plans fail closed. Pre-rename failure leaves the plan unchanged; directory-sync failure after rename is committed/durability-unknown and MUST NOT be blindly retried. Tests pin boundaries, identity, hostile inputs, and both commit states. Callers own serialization and core semantics. |
|
|
400
|
+
| Repeated class: bind occurrences, sweep at three, stop at round three, and terminate after drift two | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh; result-class: failure | updated | Wrapper fixes review+two challenges at capability-only budget 2. V2 validator SHA-binds same-dir v3 receipts; ready needs exact review+challenge+`complete/self_reviewed`. Round-3 findings require `continuation_authorization_required`—never an automatic fourth round; a second recorded base drift terminates `baseline_race`. Findings classify once; unresolved stays open/human-decision and `source_refuted` requires caller proof. Occurrence three binds an authoritative sweep; zero unmatched is required only for ready, while continuation/race retain evidence. Ordered raw `ls-remote` catches A→B→A. Tests pin failures. Omitted history, live authority, CAS, and classification truth remain caller-owned. |
|
|
401
|
+
| Scan candidate commits, destination, and exact PR text; merge base excludes ancestor debt and PR edits rerun CI | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh; result-class: failure | updated | Bounded no-echo tests cover refs/ranges, provider metadata, a cross-product of every line-oriented metadata class with Markdown list/task-list/quote/heading/nested wrappers plus malformed-checkbox precision controls, CRLF, legacy commit encodings, malformed raw messages, malformed batch framing, colliding short locators, local replacement refs, legacy grafts, and controls. Candidate metadata is requested from Git with explicit UTF-8 output; invalid UTF-8 or a replacement character fails closed with only locator and field, so ambient `i18n.logOutputEncoding` cannot erase a CJK-only match. Candidate OIDs are enumerated separately; the bounded raw-object batch is length-parsed, type/size/trailer checked, and bound in order to each full OID before payload NUL rejection and pretty formatting. Twelve-character locators are diagnostic only. Every Git object read disables `refs/replace/*`, because a clean local replacement does not change the original object a push publishes; a non-empty `info/grafts` file fails before and after the scan because it can similarly hide the pushed parent chain. Provider-shaped unambiguous AI display names plus separator-normalized known-provider `[bot]` display or exact email accounts are blocked for coauthors, authors, and committers; email accounts allow only the standard optional numeric GitHub ID prefix, so a provider token embedded at the end of another bot account is not enough. Ambiguous single-name coauthor trailers need bot-like mail; author/committer headers never use generic email shape alone because a GitHub `noreply` privacy address is normal for humans, so those ambiguous identities stay an explicit human-readback gap and unrelated automation bots remain allowed. Base order is argument, trusted event, environment, default. A direct branch-push event with a missing or all-zero `before` may use a fallback only when it still yields a non-empty candidate range; a non-zero unreachable event base fails resolution, while local and PR-only empty ranges remain valid. Hook uses target/`origin/dev` for features and remote SHA for existing `dev`/`main`; it fails closed without a new-target base, skips deletes, scans all refs, and checks worktree only for pushed HEAD. Before PR create/edit, root policy requires exact draft title/body via `--pr-text-file`. CI direct-push triggers are limited to `dev`/`main`; PR triggers are `opened`/`synchronize`/`reopened`/`edited`, including text-only edits. Comments, labels, platform-generated merge messages, host-default footer conflicts, and ambiguous personal identities stay policy-owned; annotated-tag object messages are an explicit R0 coverage gap. |
|
|
402
|
+
| Supersedes `Bound review-plan intent...`: compaction retains every removed byte, review inputs are frozen and bounded, and the wording-only single-review exception is mechanically candidate-bound | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh; result-class: failure | updated | `code-review/SKILL.md` is the owner key. `test_update_review_plan_intent.sh` proves reversible suffix archival and fail-closed capacity; `test_review_gate.sh` proves open-once bounded files, Git routing/textconv isolation, frozen untracked bytes, and the one-skill full-context wording proof with independent `wording_only_boundary`. Missing/invalid proof leaves release/high-risk challenge-required; wording-only stays untracked and has no completion ledger. |
|
|
403
|
+
| Supersedes `Repeated class...`: the extraction wrapper applies only to non-wording lanes, and terminal readiness consumes the latest attested base while a second drift forbids later receipts | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh; result-class: failure | updated | `skill-extraction-workflow/SKILL.md` is the owner key. `test_extraction_review_gate.sh` pins owner-wrapper versus proof-bound wording-only routing; the validator accepts a later same-SHA recheck, rejects an unconsumed latest base, and terminates at the second ordered drift before any new receipt. Complete-history retention and live remote authority remain caller/platform boundaries. |
|
|
404
|
+
| Supersedes `Scan candidate commits...`: shared Git metadata scanning clears repository-routing state and closes wrapper, suffix, scheme-port, and branch-slug bypass classes in linear time | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh; result-class: failure | updated | `skill-extraction-workflow/SKILL.md` is the owner key. Tests cover hostile Git routing variables, scheme-correct ports, repeated-origin and separator performance, arbitrary paired Markdown wrappers, whole-link labels, model suffixes (including the Fable/Mythos model words the current harness signs with), emphasis wrapped around only the identity inside an anchored trailer, and English/Chinese branch continuation forms. Annotated-tag pushes are scanned per tag-object layer — message and tagger identity, nested tags peeled with a bounded depth — because head resolution otherwise peels straight to the commit. Comments, labels, custom merge messages, future provider aliases, and ambiguous human/provider names remain explicit platform/human-readback gaps rather than clean claims. |
|
|
405
|
+
| Supersedes the feature-base clause of `Scan candidate commits...`: pre-push derives the feature→dev base from the pushed remote instead of assuming `origin` | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh; result-class: failure | updated | The hook uses `refs/remotes/<pushed-remote>/dev`, fails with an explicit fetch-or-`CCL_SKILL_BASE_REF` remedy when it is absent, and keeps destination SHA precedence for existing `dev`/`main`. The focused suite reproduces a clone with only `upstream/dev`; the old literal blocks it and the corrected hook passes without narrowing to the remote feature tip. |
|
|
406
|
+
| Supersedes only the 600-second ceiling in the earlier Kimi packet-review row: a slow reviewer uses the existing generic `--timeout` instead of a client-specific option | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh; result-class: failure | updated | The 068 round plan freezes the scope. The controller now accepts 5..1200 seconds, rejects 1201 before client execution, and still defaults to 600; all four direct wrappers keep their default and clamp only above 1200. The cumulative `--total-timeout` deadline, per-mode allocation, client order, fallback rules, and provider/model selection are unchanged. RED on the prior candidate reported exactly two new failures: the controller rejected 1200, and the four-wrapper ceiling check found every old 600 clamp. The same full gate suite is GREEN after the change; `test_review_client_compat.py` also passes. |
|
|
407
|
+
| Supersedes the direct-wrapper input-domain and completion claims in the preceding timeout/plan-intent rows: bounded shell arithmetic is not a decimal parser, wording-only review is never completion evidence, and duplicate-key JSON is not loss-preserving input | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh; firing-path: command:skills/code-review/scripts/test_update_review_plan_intent.sh; result-class: failure | updated | The [068 timeout correction plan](../../../specs/068-review-loop-retro/plan.md) is an addendum to the existing retrospective candidate. One shared decimal-string normalizer now rejects values below 5 and clamps `1201` or a 50-digit decimal to 1200 before any raw-value arithmetic; all four wrappers bind to it, while Kimi inline mode's independent 120-second cap stays unchanged. Completion rejects any prior receipt carrying wording-only proof/scope. The plan updater rejects duplicate keys in every JSON object before mutation. Each newly added case was RED on the prior implementation and GREEN after the fix; defaults, client routing, fallback, provider/model selection, cumulative allocation, and ordinary plan updates remain controlled. |
|
|
408
|
+
| Model-qualified AI Git identities are the same prohibited attribution class as their unqualified product name | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh; result-class: failure | updated | The prior exact-name grammar allowed `Claude Sonnet`, `OpenAI Codex`, `ChatGPT-5`, `Qwen2.5-Coder`, and `Gemini 2.5 Pro` through co-author trailers, and corresponding author/committer fields also escaped. The bounded display-name grammar accepts only known product/version shapes after an AI provider, including attached Qwen versions and Gemini's versioned Pro form, without turning nearby human names such as Claude Monet, Qwen Li, or Gemini Proctor into AI identities. Trailer, author, and committer regressions were RED before the change and GREEN after it. |
|
|
409
|
+
| Committed-range Git-surface changes remain public-sanitized, including provider vocabulary and privacy-style email fixtures | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh; result-class: failure | updated | `skill-extraction-workflow/SKILL.md` is the owner key. The committed-range private-alias audit rejected a newly added provider token, and the public sanitizer rejected a non-example fixture address that pre-commit tracked-file scans had not seen. The registry now omits the colliding token, every literal fixture email uses an allowed example domain, and the focused behavior suite still passes. |
|
|
410
|
+
| Supersedes the wording-only clause of `Supersedes the direct-wrapper input-domain...`: a punctuation-only proof must preserve numeric tokens byte-for-byte and never certify code-container edits, and a committed plan update whose receipt is lost is a terminal nonzero state | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh; firing-path: command:skills/code-review/scripts/test_update_review_plan_intent.sh; result-class: failure | updated | `code-review/SKILL.md` is the owner key. Independent review of the 068 candidate found the punctuation-only skeleton certifying numeric-token merges (`5.5` to `55`) — the proof kind that waives the release challenge — and token replacement running inside fenced or HTML code containers; the pre-fix predicate reproducibly certified the merged-number edit (RED probe on the prior commit) and both classes now fail closed with numeric-merge and fenced-token regressions. The `wording_only_boundary` concern names the actually permitted scope: punctuation-only, whitespace edits rejected. `update_review_plan_intent.py` maps a stdout pipe closed before the success receipt to `plan_committed_receipt_lost` on stderr with a nonzero exit instead of rc 0, mirroring the durability-unknown contract; the already-correct `plan_core_identity_mismatch` and `intent_history_evidence_overflow` branches gained killing tests (coverage additions, not behavior fixes). |
|
|
411
|
+
| Supersedes the packet-fidelity and waiver clauses of `Supersedes the wording-only clause...`: a frozen base-mode packet must carry the exact changed bytes and the punctuation waiver must not flip assertive polarity | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh; firing-path: command:skills/code-review/scripts/test_update_review_plan_intent.sh; result-class: failure | updated | `code-review/SKILL.md` is the owner key. Cross-family external review (one review plus one challenge per partition over the four-packet partition of the exact candidate) surfaced: an in-tree `.gitattributes -diff` entry collapsed changed lines to a binary marker inside the frozen packet (now forced `--text`; a truly binary file fails the packet NUL check instead of passing unseen); repository-local `core.worktree` could redirect discovery to a decoy tree (now refused fail-closed, and the resolved root must contain `--cwd`); a punctuation proof certified `input.` to `input?` (question marks are now outside the certifiable set while exclamation marks remain pinned eligible by fixtures); the flushless success receipt let BrokenPipeError escape to interpreter shutdown, bypassing the receipt-lost handler and exiting 120 (now flushed in the guarded path with stdout parked on devnull before the terminal rc 2); and the packaged runtime closure did not require `update_review_plan_intent.py` (now in the verifier's required list with its own removal mutation test). A second adversarial pass on the corrected candidate closed the same-shape residue: `script`/`style`/`textarea` join the tracked code containers, untracked bytes split on LF only, and question/quote marks are compared with skeleton anchors across an extended Unicode question set so a mark cannot slide, appear, or vanish. Repository-local config attacks stay on the pinned neutralization posture (per-invocation `-c` disables with fixtures asserting the true packet survives a hostile include) rather than gaining refusal predicates; the clean-filter execution residue is a recorded disposition, not a silent claim. Each closed class carries a fixture that was RED against the prior behavior. |
|
|
412
|
+
| Supersedes the identity-grammar and tag-layer clauses of the shared-scan supersede row: the metadata scan recognizes current cross-provider model words, session-task URLs, and only canonically-shaped tag objects | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh; firing-path: command:skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh; result-class: failure | updated | `skill-extraction-workflow/SKILL.md` is the owner key. The same external round surfaced: bare `GPT-*` identities and `Gemini * Flash/Ultra` model forms escaped every surface (now in the registry and unambiguous grammar with trailer regressions); Codex task URLs on session origins (`.../codex/tasks/<id>`) matched no session-path grammar (now matched, with a product-page near-miss control); a literal tag object could duplicate its `object` header so the scanner followed a decoy chain while Git peels the first target (canonical header shape now required, duplicates fail closed); the closeout validator's completion-receipt `schema_version` guard lacked the exact-integer type check applied everywhere else (a float `3.0` validated; now rejected with a regression), and `bounded_text` accepted C1 controls and U+2028/U+2029 line separators into diagnostics (now rejected with per-codepoint regressions). The second pass additionally rejected default-ignorable format characters that could split one recurrence class into visually identical keys, and bound every counted chain round to the exact ledger candidate rather than only the final receipt. |
|
|
413
|
+
| Supersedes the fixture-portability posture of the plan-intent suite row: a mode assertion probes GNU stat before BSD stat | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_update_review_plan_intent.sh; result-class: failure | updated | `code-review/SKILL.md` is the owner key. GNU stat echoes unknown BSD directives (`%Lp`) verbatim with exit 0, so the suite's BSD-first probe never reached its GNU fallback and CI shard 2 failed `plan mode was not preserved` on every Linux run while macOS stayed green. The assertion now probes `stat -c` first and falls back to `stat -f '%Lp'`, mirroring the pattern `opencode_review.sh` already runs on both platforms. |
|
|
414
|
+
| Supersedes the comparison-domain clause of the obligation-preservation audit: a repository-frozen ledger pins BOTH ends of its domain, so unrelated later changes owe it nothing | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh; result-class: failure | updated | `skill-extraction-workflow/SKILL.md` is the owner key. The 065 obligation audit derived its row set from pinned-base..WORKING-TREE, so the first post-landing PR that rewrote any obligation line in any `skills/**/*.md` went red on preservation rows it never owed (observed: 12 phantom rows for this candidate's own code-review contract-paragraph rewrite). `obligation-ledger.py` now accepts `--head`, the ledger header pins `Head revision` beside the base, and the repo audit reads and requires it; carrier-drift detection still reads current files, so rewriting a BOUND carrier stays red (probed: the pinned audit still fails `CARRIER_COMPOSITE_NOT_UNIQUE` on a carrier rewrite). The synthetic suite adds the differential: a post-head non-carrier rewrite passes the pinned audit and fails the unpinned one with `ROW_SET_MISMATCH`. |
|
|
415
|
+
| Browser E2E business-success assertions must anchor on objective effects — weak proxy signals never suffice as the sole pass condition; billable metered-resource tests take an explicit lane with a named budget owner; the Playwright component-testing maturity claim is corrected against its primary source | `testing-strategy` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/testing-strategy/references/e2e-real-flow-testing.md#weak proxy signals are not business assertions | updated | Owner key `testing-strategy/SKILL.md` (unchanged this round; the changes are merges into two references, no new bullets). (1) Weak-proxy blacklist (page loaded / URL changed / non-empty text / generic element visible / success toast alone) merged into the existing click-through clause in `e2e-real-flow-testing.md` §E2E Scope Control, pairing the already-present objective-effect positive list in §Backend/API Real Flows with its negative list; (2) billable-resource lane semantics (paid model inference, per-call third-party APIs, real payment flows, cloud sandboxes/device farms → explicit marker/lane, out of default PR and scheduled-frequent lanes, named budget owner plus authorized test account) merged into the expensive-test clause in `test-topology-and-commands.md` §Command Tiers, with credential/provisioning discipline routed to `ci-fixtures-and-flake-control.md`; (3) factual correction in `e2e-real-flow-testing.md` §Browser E2E Rules item (d), scoped to exactly what the excerpt establishes: Playwright's current component-testing guide supersedes the experimental add-on packages — primary source `https://playwright.dev/docs/test-components` (fetched 2026-08-30): "This guide replaces the experimental `@playwright/experimental-ct-react` and `@playwright/experimental-ct-vue` packages"; the sibling Biome correction rests on `https://biomejs.dev/linter/` ("a total of 526 rules", "many of them inspired from other linters" — the latter covers the skill text's rule-origin parenthetical) plus `https://biomejs.dev/blog/biome-v1-9/` ("CSS formatter and linter are now considered stable"; the prior GraphQL-timeline clause was removed from the edited skill text rather than carried beyond its excerpt), and the S3 note on `https://docs.aws.amazon.com/AmazonS3/latest/userguide/checking-object-integrity.html` plus its upload subpage `.../checking-object-integrity-upload.html` (both fetched 2026-08-30: `CRC64NVME` "is the default checksum algorithm"; "A composite checksum is calculated based on the individual checksums of each part in a multipart upload" — establishing the composite branch the skill text describes; the per-algorithm support matrix lives on that subpage and is not restated here). RED-baseline (applied, differential; evidence scope: gate wiring and package integrity only, not per-clause semantics): on the committed candidate, `CCL_SKILL_BASE_REF=origin/dev check-ccl-skills.sh` ran green (`ccl_skill_check_clean_ok`, control); a probe commit deleting exactly this row turned the same run RED printing `impact_chain_gate_missing: upstream-owner skill changed without a matching source-register impact-chain row / missing evidence path: testing-strategy/SKILL.md` (attributable to the owning gate); restoring the row returned it to green — the oracle command is reproducible in-repo on any checkout of this candidate. Corpus-identifier scan over the five changed files (grep for the private source-vocabulary set) returned zero hits, and the repository R0 audit printed the clean private token in the same run. Clause semantics rest on this round's independent review and adversarial challenge rows. Scope of what THIS ROW asserts is deliberately narrow: the three inlined public excerpts above, the reproducible in-repo gate/scan runs, and nothing further — verification of other claims in the same round's diff is recorded in the maintainer's private round archive per the extraction-lifecycle handoff policy and is NOT certified by this row; a reviewer should judge those clauses against their own named public sources directly. Dual-track terminal record: multiple independent codex review rounds plus three full review+challenge×2 chains ran against successive candidates; every P0/P1 with a demonstrable failure path was applied (streaming retry/latch/resume/cancel semantics, canonical/hreflang interplay, submit latch, billable/payment-sandbox lanes, token-scan hardening, excerpt-scoping narrowings); the terminal chain's last three finding-fixes (cancel-vs-buffered-end, idle/heartbeat timeout, latch release on terminal state) were merged AFTER that chain per the repository's fixed-budget review-loop rule, so the exact landing candidate carries them un-rechallenged — listed for the merge decision-maker; the recurring bounded-packet self-certification findings (a packet reviewer cannot execute the in-repo oracle or see the private vocabulary) are dispositioned as verify-by-rerun: running the repository check script (`check-ccl-skills.sh` with `CCL_SKILL_BASE_REF=origin/dev`) on this candidate is the reviewer-executable oracle. Source class: local external engineering-skill corpus (sanitized to capability labels) plus public primary docs; sibling map: `web-react-dev` updated in the same round (its own package, not machine-gated), `product-ui-ux-design` unchanged (aesthetic-direction and anti-generic coverage already equal or stronger than the external source), `design-closed-contract-oracles.md` unchanged (criterion-to-proof already covered), `ci-fixtures-and-flake-control.md` unchanged (credential/provisioning face already covered). |
|
|
416
|
+
| Evidence-pipeline failure is a per-case third verdict: infra-error is never pass, never business-fail, never silent skip, and weaker surfaces cannot substitute for missing evidence | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#never a pass, never a business fail, and never a silent skip | `updated` | Owner key `testing-strategy/SKILL.md`. Benchmark round (private provenance alias: mediagen-platform) plus primary-source verification. RED baseline (replayed): `git show origin/dev:skills/testing-strategy/references/test-code-authoring-patterns.md` asserts the xUnit smell corpus as ~18 items while xunitpatterns.com lists 15 top-level (5/6/4), and the coverage floors carried no provenance — head corrects both and adds the infra-error verdict rule; baseline grep for the anchor line is zero-hit, head exactly one. |
|
|
417
|
+
| A continuation gate distinguishes multi-viable-approach and no-evidence-cause stops from a single dominant reversible path, which is carried through to a reviewable draft instead of stopping at a recommendation | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#do not stop at a recommendation | `updated` | Owner key `product-rd-workflow/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/product-rd-workflow/references/delivery-lifecycle.md` attributes DORA metrics' origin to the 2018 book while dora.dev/insights/dora-metrics-history dates the research line to 2014 — head corrects it; baseline zero-hit for the new stop/carry-through predicate, head exactly one. Review checklist gains guarantee grading and disabled-path walk (same package). |
|
|
418
|
+
| Query/lookup evidence reports 0, 1, or N matches distinctly; silently taking the first row of N is forbidden and an empty result over a named scope is itself evidence | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#never silently take the first row of N | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external production skill-pack doctrine (alias mediagen-platform), mechanism verified as stack-agnostic. RED baseline (replayed): baseline grep for the cardinality predicate is zero-hit across the package, head exactly one — the baseline tree gave no instruction against first-row-of-N, so an agent following it could silently mis-resolve identity lookups. |
|
|
419
|
+
| Risk classification runs on the change's objective shape: reporter tone or executive pressure never escalates tags and a small diff never de-escalates them | `feature-risk-router` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/feature-risk-router/SKILL.md#must never escalate a change's tags | `updated` | Owner key `feature-risk-router/SKILL.md`. Baseline covered only the small-diff de-escalation side (low-risk intuition clause); the tone/pressure escalation side was absent — RED baseline (replayed): baseline grep zero-hit for the anti-escalation predicate, head exactly one. Six objective evaluation axes recorded inline. |
|
|
420
|
+
| Reader annotations are evidence of reading breakdown: classify the root-cause class first, then sweep the whole document for the same class instead of patching only the flagged sentence | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/annotation-driven-revision.md#必须全文扫同类位置一起修,不得只改被标记的那一句 | `updated` | Owner key `tighten-doc/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/tighten-doc/references/figure-and-table-craft.md` lists the 25-word sentence limit as no-reliable-source while GOV.UK's writing guideline states it as its house style — head regrades it to a sourced single-institution style and records the 40-char claim's traceable community source; baseline zero-hit for the annotation-sweep predicate, head exactly one. |
|
|
421
|
+
| Review reuse depth is a deterministic function of the candidate delta with no manual downgrade, the reviewed-identity record is a freshness guard not cryptographic proof, and relayed blocking findings pass a false-positive check with auditable dispositions | `code-review` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/code-review/references/manual-invocation-and-prompts.md#deterministic function of the delta, never a manual downgrade | `updated` | Owner key `code-review/SKILL.md`. Benchmark adjudication recorded: the source pack's receipt survives rebase/amend; our stricter voids-on-any-edit line is deliberately KEPT and only the tiering-determinism and honest-trust-model clauses are absorbed (keep-stricter per Conflict Resolution). RED baseline (replayed): baseline grep zero-hit for the determinism predicate, head exactly one. |
|
|
422
|
+
| Observation code never intrudes on the observed path, cross-layer conclusions require aligned identifiers or time, existing signals are discovered before new ones are created, and explicit environment targets always win | `platform-observability` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/references/metrics-conventions.md#must not add retries, blocking waits, or business-logic branches | `updated` | Owner key `platform-observability/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/platform-observability/SKILL.md` attributes the good/valid SLI formula to Google SRE generically while the SRE Workbook's own text is good/total (the valid-events refinement is the Art of SLOs / GCP-blog line) — head corrects the attribution and self-labels the SLO ladder as team heuristic; baseline zero-hit for the non-intrusion predicate, head exactly one. |
|
|
423
|
+
| Benchmark rounds issue per-mechanism P/I/M/W verdicts with grep-anchored evidence, and an M verdict passes the functional-equivalent check before any borrow lands | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#requires the zero-hit grep recorded as replayable evidence | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Method absorbed from an external theory-audit corpus (alias mediagen-platform) whose own run found all 8 W-verdicts were internal drift rather than external knowledge gaps — the W-type internal-consistency sweep, disposition middle states (待实验/条件化采纳), and three conflict-synthesis shapes land in source-to-skill-extraction.md; keyword-activation evidence lands in description-authoring.md. RED baseline (replayed): baseline grep zero-hit for the functional-equivalent predicate, head exactly one. |
|
|
424
|
+
| A diff-scoped design review classifies every hard-coded visual-value hit into approved usage, pre-existing outside the change, or new violation — only new violations block, and pre-existing debt is routed, never blamed on the change | `product-ui-ux-design` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-ui-ux-design/references/ui-ux-audit.md#never reported as caused by this change | `updated` | Owner key `product-ui-ux-design/SKILL.md`. Changed refs: ui-ux-audit.md (diff-scoped three-bucket review), design-system-source-of-truth.md (definition-matrix completeness check), tokens-and-components.md (old code is not permission), platform-mobile-patterns.md (motion band self-labeled team heuristic per M3 tokens/Apple HIG verification; M3 Expressive research figures cited). RED baseline (replayed): baseline grep zero-hit for the three-bucket predicate, head exactly one. |
|
|
425
|
+
| Repo-pinned TC implementations pin the catalog revision, claim before implementing, and stop to fix a stale/contradictory source record before implementing against it | `test-artifact-management` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/test-artifact-management/references/update-lifecycle.md#停下先修源记录**,不得按旧版实现后事后补 | `updated` | Owner key `test-artifact-management/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/test-artifact-management/references/classical-test-design-techniques.md` states 2-way捕到 50–90% while NIST SP 800-142 Table 1 gives 53–97% with an explicit 10–40%+ miss warning, and tc-review-and-prioritization.md claimed a 30-year P×I consensus while ISTQB CTFL v4.0.1 §5.2 defines multiplication as the quantitative approach beside a qualitative matrix — head corrects both; baseline zero-hit for the revision-pin predicate, head exactly one. |
|
|
426
|
+
| Evaluation records keep human-review and machine fields with distinct writers, normalize cross-provider token accounting before comparison, and reconcile usage against billing where available | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#must never overwrite a human verdict | `updated` | Owner key `llm-inference-integration/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/llm-inference-integration/references/model-prompt-evaluation.md` states extended thinking is off by default per docs while the current platform docs deprecate the manual mode on 4.6, reject it on 4.7+, and default thinking on for the Claude 5 family — head rewrites the claim per model generation; prompt-caching modes and the provider-evaluation evidence ladder (spend contract, seven layers, open-loop load per NSDI'06) added in the same package. |
|
|
427
|
+
| Durable-state transitions capture the clock once and validate external-response structure with a typed error path before mapping | `nodejs-service-dev` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/nodejs-service-dev/references/async-lifecycle-and-performance.md#never a crash or a silently-defaulted field | `updated` | Owner key `nodejs-service-dev/SKILL.md`. Sibling-parity landing with the go/python state-machine references (same two predicates, stack-idiomatic wording). RED baseline (replayed): baseline grep zero-hit for the predicate, head exactly one. |
|
|
428
|
+
| Vendor/ecosystem status claims carry their verified level and date — the RPC-alternative entry records CNCF sandbox status and the accurate interop-test wording instead of an inflated maturity tier | `go-microservice-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/go-microservice-architecture/references/architecture-playbook.md#CNCF **sandbox** project (per connectrpc.com | `updated` | Owner key `go-microservice-architecture/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/go-microservice-architecture/references/architecture-playbook.md` asserts CNCF-incubated and Google-validated interop while connectrpc.com states sandbox level and self-run extended interop tests — head corrects both; crypto-erase citation upgraded to SP 800-88 r2 + EDPB 02/2025 in the same package. |
|
|
429
|
+
| Crypto-erase evidence conditions cite the current sanitization standard revision and regulator guidance rather than a withdrawn revision | `python-service-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/python-service-architecture/references/multi-tenant-isolation.md#do not use CE when data predates encryption enablement | `updated` | Owner key `python-service-architecture/SKILL.md`. RED baseline (replayed): baseline cites NIST SP 800-88 generically (Rev.1 withdrawn 2025) with no CE-condition specifics; head cites Rev.2's explicit do-not-use conditions and EDPB 02/2025's encrypted-data-is-still-personal-data holding — mirrored with the go sibling. |
|
|
430
|
+
| Gitflow-style release trains use direction-sensitive merges (squash only feature→develop), prompt dual back-merge, human-confirmed tags, and full-ladder hotfixes; canary template thresholds are labeled as tool example values, and user-bucketed ramps follow exposure/funnel/salt data-literacy rules | `platform-release-engineering` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-release-engineering/references/promotion-gate-and-review.md#must not squash** — prefer fast-forward/plain merge | `updated` | Owner key `platform-release-engineering/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/platform-release-engineering/references/canary-and-rollout-strategy.md` presents the 1%/500ms thresholds as bare template defaults with no provenance, while Flagger's builtin checks carry exactly those example values and Argo Rollouts' official example uses 95% — head attributes and scopes them; the merge-topology section is grounded in AWS Prescriptive Guidance's Gitflow pattern (verified 2026-08). |
|
|
431
|
+
| Verdict assignment is decided by fault origin with evaluated-false as business fail, and required infra-error cases block aggregate readiness | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#never by defaulting to whichever verdict looks better | `updated` | Owner key `testing-strategy/SKILL.md`. Fix-round rows for the review-chain findings: the pre-fix span (replayed via `git show` at the round base) carried the enumerated verdict mapping whose edges four review rounds broke in turn; the fault-origin predicate replaced the enumeration and the aggregate-readiness clause closed the clean-total false green. |
|
|
432
|
+
| Design-review exceptions are approved only by the design-system owner's recorded, unexpired, usage-covering approval; age is not approval; scan categories cover size/layout and motion | `product-ui-ux-design` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-ui-ux-design/references/ui-ux-audit.md#age is not approval: reusing or extending an old undocumented/expired exception | `updated` | Owner key `product-ui-ux-design/SKILL.md`. Fix-round: the pre-fix predates-arm allowed laundering an old undocumented/expired exception through a changed hunk (challenge-confirmed bypass, replayed at the round base); the arm was deleted (convergence-by-deletion) and the definition matrix aligned with the governed-category list. |
|
|
433
|
+
| TC claiming and write-back both ride atomic revision-conditioned updates; hash-pin mismatch always stops for source reconciliation | `test-artifact-management` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/test-artifact-management/references/update-lifecycle.md#停下改走人工/单线分配,不得按回读结果继续 | `updated` | Owner key `test-artifact-management/SKILL.md`. Fix-round: read-back-after-write was challenge-proven non-CAS (later writer silently replaces a verified claim, replayed at the round base); the landed rule requires platform CAS/optimistic-lock or a serialized coordinator, else concurrent claiming is unsupported and stops. |
|
|
434
|
+
| Relayed findings pass an existence-and-severity check, unverifiable stays blocking, and reuse voids on base/profile/lens change | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/code-review/references/manual-invocation-and-prompts.md#a confirmed-but-advisory or overstated-severity issue is relayed at its true severity | `updated` | Owner key `code-review/SKILL.md`. Fix-round: challenge rounds showed the fp-check validated existence only (severity inflation passed) and the unverifiable disposition could demote real defects; both closed, with the reuse boundary extended to base/profile/lens changes. |
|
|
435
|
+
| Auto-continue is subordinate to every stop condition, the metered-account carve-out keeps its existing/configured/self-use qualifiers, and stop wording preserves the pinned anchors | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#report interim or blocked with the next unblock step | `updated` | Owner key `product-rd-workflow/SKILL.md`. Fix-round: size-offset compression had silently widened the metered-account carve-out and detached carry-through from the stop list (review-confirmed, replayed at the round base); qualifiers restored, precedence bound, and checker-pinned anchor phrases restored after the pinned-phrase gate fired. |
|
|
436
|
+
| Vendor-standard citations are layered honestly: CE conditions cite Rev.1 §2.6 with Rev.2 superseding and continuing the framework | `go-microservice-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/go-microservice-architecture/references/multi-tenant-isolation.md#verify the corresponding Rev.2 section when citing it as the authority | `updated` | Owner key `go-microservice-architecture/SKILL.md`. Fix-round: the earlier candidate attributed CE do-not-use conditions to Rev.2 while the verification ledger located them in Rev.1 §2.6 (review-caught attribution mismatch); the landed text layers the citation and instructs Rev.2 section verification before citing it as authority. |
|
|
437
|
+
| The python sibling mirrors the layered CE citation byte-for-byte per the parity discipline | `python-service-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/python-service-architecture/references/multi-tenant-isolation.md#verify the corresponding Rev.2 section when citing it as the authority | `updated` | Owner key `python-service-architecture/SKILL.md`. Fix-round mirror of the go row above; parity gate keeps the mirrored region byte-identical. |
|
|
438
|
+
| Node's malformed-payload rule mirrors the go/python persisted-failure-transition semantics and single-now scopes to transition-stamped timestamps | `nodejs-service-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/nodejs-service-dev/references/async-lifecycle-and-performance.md#mirroring the go/python state-machine rendering | `updated` | Owner key `nodejs-service-dev/SKILL.md`. Fix-round: review caught the node wording drifting weaker than the go/python rendering (policy-handled vs persisted failure transition) and the single-now literalism re-stamping domain-provided times; both aligned. |
|
|
439
|
+
| Best-effort observation is scoped to diagnostic telemetry; audit/billing/deletion/release-gate records are business writes that never fail open | `platform-observability` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/references/metrics-conventions.md#their loss is a failure, never shrugged off as telemetry | `updated` | Owner key `platform-observability/SKILL.md`. Fix-round: review showed the unbounded best-effort license could excuse a mandatory record failing open; the boundary now names the mandatory-record classes and their durable-delivery semantics. |
|
|
440
|
+
| Annotation-driven revision scans the whole document but edits only within authorization, and factual/citation errors route to source verification | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/annotation-driven-revision.md#不得以「修同类」为名越权改动已定内容 | `updated` | Owner key `tighten-doc/SKILL.md`. Fix-round: review caught the unconditional same-class sweep expanding mutation past the approved scope and the root-cause classes omitting factual/citation errors; both landed with the scan-wide/edit-scoped split. |
|
|
441
|
+
| Keyword-activation guidance carries its sources and conditionality, and benchmark-figure reliance requires local reproduction | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/description-authoring.md#do not rely on the numbers without reproducing against your own catalog | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Fix-round: review flagged the activation claims as unsourced-in-ledger and unconditional; sources landed in the source-verification ledger and the conditionality bullet forbids relying on the numbers without local reproduction. This round also carries the crypto-erase supersede note above. |
|
|
442
|
+
| Hotfix back-merge covers every open release train and canary thresholds carry their tool-example provenance | `platform-release-engineering` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-release-engineering/references/promotion-gate-and-review.md#an active train missing the hotfix reverts it at that train's own merge | `updated` | Owner key `platform-release-engineering/SKILL.md`. Challenge-round: the earlier back-merge wording covered main and develop but not an open release branch (replayed at the round base), so a concurrent train could revert a shipped hotfix; the funnel invariant was also scoped as sanity-not-attribution and the error-budget comment renamed to a rolling error-rate threshold. |
|
|
443
|
+
| Benchmark M-verdict evidence pins command, scope, and baseline revision so replays run against the recorded baseline | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#replay runs against the recorded baseline | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Challenge-round: a NO_HITS record without its baseline revision self-hits once the borrow lands (review-caught, replayed at the round base); this round also relocates the ledger's supersede note below the row table so machine and human readers see one uninterrupted row stream. |
|
|
444
|
+
| Provider-evaluation spend is enforced at admission with spent-plus-in-flight headroom and billed attempts missing usage count against the provider | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#its cost enters via billing records, never silent exclusion | `updated` | Owner key `llm-inference-integration/SKILL.md`. Challenge-round: delayed usage/billing signals let an open-loop driver overshoot the cap and a provider omitting usage on billed failures flattered its ratios (replayed at the round base); the admission budget now subtracts recorded spend plus in-flight worst case, the stop latch halts admissions, and ratio exclusions are reported. |
|
|
445
|
+
|
|
446
|
+
| Spend enforcement is a reservation invariant (cap ≥ reconciled spend + outstanding reservations) and ratio exclusion never touches cost totals | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#total-cost accounting is the separate aggregate that never excludes | `updated` | Owner key `llm-inference-integration/SKILL.md`. Challenge-round: three successive edge findings on the admission-budget arithmetic converged by replacing the enumeration with the reservation invariant (reserve worst-case at admission, release only on billing reconciliation), and the ratio-exclusion clause was split from total-cost accounting so neither aggregate can be flattered by usage omission. |
|
|
447
|
+
|
|
448
|
+
| The billing-records anchor phrase is preserved inside the split-aggregate wording so prior ledger locators keep resolving | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#never silent exclusion from totals | `updated` | Owner key `llm-inference-integration/SKILL.md`. Locator-repair round: the split-aggregate rewrite had dropped the substring an earlier row anchors on (register_firing_path_unresolved, replayed at the round base); the phrase is restored within the new semantics so both locators resolve. |
|
|
449
|
+
|
|
450
|
+
| Spend reservations are atomic check-and-reserve on one ledger, closing the concurrent-headroom TOCTOU | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/inference-capacity-operations.md#read-headroom-then-reserve is a TOCTOU that lets concurrent workers jointly overshoot | `updated` | Owner key `llm-inference-integration/SKILL.md`. Challenge-round: two open-loop workers reading the same headroom could each reserve and jointly exceed the cap (replayed at the round base); the reservation is now an atomic check-and-reserve, completing the spend-invariant class alongside admission-halt and billing-release. |
|
|
451
|
+
|
|
452
|
+
| The hotfix full-ladder rule names the emergency-override section as its one sanctioned, loudly-logged exception | `platform-release-engineering` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-release-engineering/references/promotion-gate-and-review.md#the emergency-override section below is the one sanctioned, loudly-logged exception | `updated` | Owner key `platform-release-engineering/SKILL.md`. Review-round: the unconditional never-skips-a-gate wording contradicted the file's own emergency-override section during a production incident (replayed at the round base); the exception is now named inline so the two sections compose instead of conflicting. |
|
|
453
|
+
|
|
454
|
+
| Sub-threshold sample counts are reported as unreliable, never averaged into a verdict | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/inference-capacity-operations.md#must be reported as unreliable, never averaged into a verdict | `updated` | Owner key `llm-inference-integration/SKILL.md`. Review-round: the ~3-run floor carried no provenance label (replayed at the round base); it is now explicitly a team heuristic with a raise-per-variance instruction, and the rule moved to its own normative bullet. |
|
|
455
|
+
|
|
456
|
+
| The latency-SLI formula divides by the valid set, matching the good/valid definition in the same rule | `platform-observability` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/SKILL.md#the denominator is the availability SLI's valid set, never raw total | `updated` | Owner key `platform-observability/SKILL.md`. Challenge-round: the bullet preferred good/valid while its own latency formula divided by raw total (replayed at the round base); the denominator now names the valid set explicitly. |
|
|
457
|
+
| Verbatim-bound obligation carriers survive wording compression only verbatim: a size-budget compression must first check the sentence against the frozen preservation mapping, and byte offsets come from sentences the same round added | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#Do not describe such work as done, fixed, merge-ready, or release-ready | `updated` | Owner key `testing-strategy/SKILL.md`. Observed failure: two wording compressions in this round's size-budget offsets rewrote exact carrier sentences bound by the specs/065 obligation mapping (such-as→like; are-done→pass), and the heavy lane's real-repository obligation audit went red (CARRIER_COMPOSITE_NOT_UNIQUE count=0) while every entrypoint-scope gate stayed green — the preservation mapping is a verification surface the size-budget workflow did not consult. Fix: carrier sentences restored verbatim; equal-byte offsets taken from sentences this branch itself added (which the frozen mapping cannot bind); reader index regenerated at the mapping's pinned base/head. RED baseline (replayed): the repo-audit suite for the obligation ledger (`test_obligation_ledger_repo_audit.sh`) red before the restoration, `audit_ok domain=50 rows=1240 unresolved=0` plus `test_obligation_ledger_repo_audit_ok` after. |
|
|
458
|
+
| A wording compression that deletes a sentence boundary corrupts the hosting rule: the reinserted stop/continuation sentence must keep its full punctuation, and a fresh-eyes review of the reinsertion is what catches the truncation | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#explicit stop/pause needs no reconfirmation. A `continuing:` outcome | `updated` | Owner key `product-rd-workflow/SKILL.md`. Observed failure: the size-budget reinsertion of pinned continuation literals dropped a sentence boundary, leaving "needs no reconfirm A `continuing:` outcome" — two rules joined without punctuation on the entrypoint's stop/continue surface; caught by the supplementary post-delta review round (P1), invisible to every deterministic gate because pinned-phrase gates match their own literals only. Fix: boundary restored ("no reconfirmation. A"). Offset provenance, stated exactly: this entrypoint's continuation-gate block is wholly branch-rewritten (six modified base lines, no pure additions), so offsets necessarily live inside that rewritten block; a word-level diff against the base revision audited every token those compressions dropped — the two load-bearing drops it surfaced (metered model/tool scope; the missing-capability routing qualifier) are restored in their owners' rows, the ambiguity-or join was additionally reverted with its replacement byte taken from a demonstrably branch-added sentence (em-dash tightened to a colon in the dominant-approach clause), and the remaining list joins are recorded as audited-neutral. Where genuinely branch-added sentences exist, offsets come from them first. The truncation shape joins the round's carrier-restoration lesson: compression edits need a substring check against the frozen preservation mapping and plain sentence-boundary integrity. RED baseline (replayed): repo-wide grep for "needs no reconfirm A" one hit before the fix, zero after; the restored phrase greps exactly once. |
|
|
459
|
+
| A compression that drops a scope qualifier widens the rule it hosts: the external-pack routing clause routes only a MISSING method/tool-layer capability, and restoring the dropped word is the fix, not rewording around it | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#M needs the functional-equivalent check | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failure: a size-budget compression rewrote the reference-only routing clause from routing a missing method/tool-layer capability to routing that layer categorically, which would displace locally covered P-verdict capabilities and contradict the functional-equivalent check landed in the same round; caught by the supplementary post-delta challenge round enumerating compressed sentences (same class as the metered model/tool qualifier drop fixed in the sibling owner). Fix: the missing-capability qualifier restored; byte offsets from this branch's own pointer sentence, whose semantics live in the owning reference. RED baseline (replayed): word-level diff against the base revision showed the dropped qualifier before the fix and shows it restored after; the class sweep over all three owner entrypoints found no further load-bearing drops. |
|
|
460
|
+
| Stop reporting keeps its specificity qualifiers: the entrypoint demands the concrete stop reason and the exact evidence checked, and budget offsets come from relocating clauses whose semantics already live verbatim in the owning reference | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#state the concrete stop reason and the exact evidence checked | `updated` | Owner key `product-rd-workflow/SKILL.md`. Observed failure: a size-budget compression dropped the-concrete/the-exact from the stop-reporting sentence, licensing generic stop reports; the challenger graded it load-bearing, the implementer's word-sweep had graded it neutral, and the maintainer's standing delegation resolves such token disputes by the repository's fail-closed obligation standard, so the qualifiers are restored. Byte offset: the stale-source parenthetical is removed from the entrypoint because its full sentence lives verbatim in references/pre-final-continuation-gate.md (Status-source reconciliation) which the same sentence already cites — relocation, not compression. RED baseline (replayed): grep for the restored phrase zero-hit on the pre-fix entrypoint, exactly one hit after; the removed parenthetical greps once in the owning reference. |
|
|
461
|
+
| The dual-track reviewer's verification scope is a documented boundary: content semantics belong to the reviewer, deterministic-gate claims to CI, historical-process claims are testimony unless receipt-bound — ruled on once so packet-verifiability findings stop recurring per round | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#a finding that only restates this boundary is dispositioned against this rule, never re-litigated per round | `updated` | Owner key `skill-extraction-workflow/SKILL.md`; the change lands in references/dual-track-review-gate.md (new Reviewer verification scope section). Observed failure: the packet-verifiability finding class recurred across four supplementary review rounds and roughly a dozen occurrences in this round's chains — every reviewer independently rediscovered that the packet cannot carry the deterministic oracles, and every round paid the same finding again because the boundary was undocumented. The maintainer confirmed the operating reality (all consumers and reviewers are agents; the human role is authority, not readership), so the boundary is now standing text agents can disposition against, with the receipt-embedding backlog item named in place. RED baseline (replayed): grep for the boundary phrase zero-hit before this change, exactly one hit after, on an added normative list line. |
|
|
462
|
+
|
|
463
|
+
Supersede note (this round, before landing): the two crypto-erase rows above ("go-microservice-architecture" and "python-service-architecture") describe an earlier candidate state; the landed text attributes the CE do-not-use conditions to SP 800-88 Rev.1 §2.6 with Rev.2 (2025) superseding and continuing the framework — per the source-verification ledger row "CE conditions text location". The rows' RED-baseline probes and firing-path anchors are unaffected.
|