@ccoalm/ccl-skills 0.1.1 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +93 -16
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +249 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +394 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +50 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +59 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +16 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +391 -14
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +74 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +74 -33
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +11 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +421 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +576 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +176 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_lanes.sh +101 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +6 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +1 -1
- package/dist/assets/release.json +66 -21
- package/dist/cli.js +23 -2
- package/dist/update-notice.d.ts +71 -0
- package/dist/update-notice.js +173 -0
- package/dist/version-check.d.ts +2 -0
- package/dist/version-check.js +7 -0
- package/package.json +1 -1
|
@@ -46,7 +46,7 @@ Disposition:
|
|
|
46
46
|
- Supported: a rule's existence is not proof it executed; use behavioral assertions and coverage/firing evidence. The digest-binding practices these sources describe (complete subject sets, provenance, raw-result digests) are sound for supply-chain trust boundaries where authors and verifiers are distinct parties; the local policy below explains why this repository adopts the firing/coverage principle but not the digest binding.
|
|
47
47
|
- Local evidence policy: the impact-chain gate machine-verifies what is cheap and deterministic — an owner-scoped firing path that resolves to its round's added lines (a unique anchor on a changed normative numbered/list rule, or a changed owner executable), the letters/digits-free wording-only classification computed from that round's owner diff, and the owner-level floor that a non-wording package carries at least one `RED-baseline` row (a `semantic-control` label may supplement but never close a package alone, because an author-selected stable label cannot vouch for a different hidden delta). The `behavioral-evidence` and `observed-failure` fields themselves are required author declarations. A digest-bound attestation apparatus for these rows (in-toto-style subject digests, same-prompt model result pairs, command-result envelopes) was built, evaluated against real iteration, and deliberately removed: under the unsigned-repository-local trust model the author can regenerate every hash, so the apparatus only detected stale records — while costing a full-suite rerun and whole-evidence regeneration whenever any owner script changed by a single byte. That cost defeated normal multi-commit iteration (it broke its own author's branch twice), so behavior claims rest on the firing-path gate, honest authorship, and the mandatory independent review/challenge instead of hashes.
|
|
48
48
|
- Local trust model: register rows remain honest-but-fallible workflow evidence, not a hostile-author security boundary. The gate proves that a changed, normative, owner-scoped rule line (or changed owner executable) exists for every claimed firing path; it does not prove a model run occurred, that a named executable implements the claimed enforcement (a shebang stub passes the static check), that a mangled or ambiguous ledger row was honest (those are warned, not blocked, to avoid false positives on other table shapes), author identity, or non-tampering by an authorized contributor. Independent review/challenge and the fixed checker remain the assurance case.
|
|
49
|
-
- Machine format (relocated from the `SKILL.md` firing-mechanism rule; the local evidence policy above carries the rationale): every added source-register row must carry `behavioral-evidence: RED-baseline` (any observed delta — `observed-failure: yes` requires it) or `semantic-control` (only with `observed-failure: no`), an `observed-failure: yes/no` state, and an owner-scoped `firing-path` — each declaration in its own semicolon-delimited fragment of the cell (`…prose; behavioral-evidence: …; observed-failure: …; firing-path: …`), so a key embedded mid-prose never parses as a declaration. The firing-path anchor is at least 16 characters, occurs once in the file and once in its round's added lines, and lands on a numbered/list Markdown rule with a normative action.
|
|
49
|
+
- Machine format (relocated from the `SKILL.md` firing-mechanism rule; the local evidence policy above carries the rationale): every added source-register row must carry `behavioral-evidence: RED-baseline` (any observed delta — `observed-failure: yes` requires it) or `semantic-control` (only with `observed-failure: no`), an `observed-failure: yes/no` state, and an owner-scoped `firing-path` — each declaration in its own semicolon-delimited fragment of the cell (`…prose; behavioral-evidence: …; observed-failure: …; firing-path: …`), so a key embedded mid-prose never parses as a declaration. The firing-path anchor is at least 16 characters, occurs once in the file and once in its round's added lines, and lands on a numbered/list Markdown rule with a normative action. A row that survives at HEAD must also resolve to an owner this range actually changes — an owner reverted to its base bytes by a rebase or a base-side conflict resolution leaves the changed set while its row stays behind, and the row then vouches for a change the delivered diff does not contain. There is no author-declared escape from this: a corrective rewrite that back-fills a row for a round which merged red produces the same shape, and it is a deliberate, person-adjudicated repair that can adjudicate this refusal too.
|
|
50
50
|
|
|
51
51
|
- **Round scoping — a row is judged against the round it landed in, never the accumulating range.** A row is authored against one round's diff, so reading the whole `base..HEAD` range to classify it judges the row against work it never described. That mismatch produced both directions of the same defect: an already-gated row turned red once a LATER round touched the same owner (which is what the ledger's superseded-row notes were absorbing), and a description-only round lost its routing-surface locator because an EARLIER round had edited that owner's body. The gate cuts rounds at the commits that touch the ledger, walked first-parent so one merged worktree round is one boundary, and each round spans from the previous boundary so work commits sit in the round whose ledger append describes them. The partition is derived from git alone — an author cannot nominate, widen, or move their own scope.
|
|
52
52
|
- Both obligations move together, in opposite directions. **Classification** narrows to the round: whether a diff is wording-only, an identifier retarget, or description-only is asked of that round's bytes, which is what makes a verdict stable once it lands. **Presence** narrows to the round too: the round that changed an owner is the round that owes the row, so owner work committed after a ledger append can no longer ride on an earlier round's row. Narrowing classification without narrowing presence would have opened exactly that laundering route.
|
|
@@ -55,3 +55,61 @@ Disposition:
|
|
|
55
55
|
- Machine-verified no-behavior classes (each drops the firing-path requirement, each requires `observed-failure: no`, and neither is an author waiver — the gate recomputes the predicate from the owner's diff for that row's round and refuses on any mismatch): `not-required wording-only` when no changed line differs by a letter or digit and every changed owner file is Markdown prose; `not-required identifier-rename` when rewriting the base bytes with git-derived rename pairs reproduces the head bytes exactly. Both exist because the anchor uses the SHAPE of a changed line as a proxy for "an obligation changed", and both are diff shapes that carry no obligation at all.
|
|
56
56
|
- Canonical field locator for a routing-surface-only owner — NOT a third no-behavior class. When an owner's entire change is the `SKILL.md` frontmatter `description` entry (same body, same other top-level keys, the value still a YAML string, all checked byte-exact), its firing path is the constant `file:skills/<owner>/SKILL.md#description`. The row still declares `RED-baseline`, still owes its behavioural evidence — for a routing change that is the measured routing delta — and only the free-text anchor is replaced. It is deliberately not an exemption: a description edit decides which requests reach the skill, so exempting it would drop evidence from the class carrying the most behaviour. The locator is canonical rather than a substring because three successive substring designs were refuted — a numbered/list rule shape, then the `description:` line, then any line of the entry — the last because a substring that merely SURVIVES the edit identifies nothing. Same class three times is the signal to drop the proxy, not to patch it again; the predicate is what carries the proof, and the locator only names the field it proved. The predicate matters more here than in the two classes above because a wrong judgement is LOOSER rather than stricter, so it accounts for every changed byte and refuses whatever it cannot account for. A further shape that cannot be located is a signal to replace the proxy, not to widen again. The precondition judges BEHAVIOUR, not file count: the rest of the package — and the entrypoint's own body and other frontmatter keys — may differ by a rename retarget, because that is already a machine-proven no-behaviour class, and the two compose under one normalizer (apply git's rename pairs to the base bytes, then require the head bytes exactly, mode included). Without the composition an owner on an integration branch that accumulates rounds gets refused for retargets it already declared no-behaviour, while having no other rule to anchor on. The description entry is the one part exempt from the normalizer, because it is the change being evidenced; anything the normalizer cannot account for — one real byte — takes the locator away. Siblings are further restricted to regular-file Markdown: a script's bytes reproducing under a substitution says nothing about whether rewriting that identifier was safe there, and a `.md` SYMLINK's blob is its target, so a retargeted target reproduces while what the path resolves to changes. **Disclosed residual, unchanged from the rename class it composes with**: a Markdown sibling can still carry a fenced command or a persisted key whose old identifier is deliberately preserved — an installed artifact name that must not follow a source rename is the observed instance — and rewriting one of those IS a behaviour change the normalizer clears. The exposure is identical to the identifier-rename class, which already reproduces whole packages under the same pairs; this composition inherits that residual rather than widening it, and the pairs still come only from git's own tree state, so a slug is rewritten only when that skill was actually renamed in the same diff. Closing it means distinguishing prose mentions from executable/persisted ones inside Markdown, which needs its own evidence round. Two shapes are knowingly left refused, and refusal here costs nothing that was not already lost: before this locator existed no routing-surface owner could satisfy an anchor at all, so a refused shape simply keeps that prior behaviour rather than regressing from it. (1) A sibling frontmatter value whose YAML tag or class `safe_load` will not construct makes the parse fail, and a failed parse refuses rather than guesses. (2) The owner that ships `references/source-register.md` cannot use the class at all, because its own evidence row lands inside its package and so its diff is never the description alone — a consequence of not special-casing that file, which is what would otherwise let arbitrary ledger content ride along. Both are usability limits on a strictly additive path, not bypasses; closing either one means loosening a predicate whose wrong answers are LOOSER than the bar, so neither is worth doing without its own evidence.
|
|
57
57
|
- Boundary: OPA coverage supports the general distinction between policy presence and policy execution; the repository does not run OPA. in-toto/SLSA supply-chain formats were consulted for the evaluated-and-removed digest apparatus; the shipped gate deliberately does not bind digests and does not reuse or claim their security levels.
|
|
58
|
+
|
|
59
|
+
### `source-refuted` — withdrawing a claim a primary source refutes
|
|
60
|
+
|
|
61
|
+
A rule can fire perfectly and still be false. Removing such a claim usually produces **no measurable behavioural delta** — there is nothing for a `RED-baseline` to capture — while `semantic-control` is refused by the owner-level RED floor. Before this class existed the cheapest path was therefore to *leave the false claim in place*, which is the wrong incentive.
|
|
62
|
+
|
|
63
|
+
The class is not a waiver; it swaps one kind of evidence for another. All four bars are machine-checked and a row fails without any of them:
|
|
64
|
+
|
|
65
|
+
1. `observed-failure: no` — this gate asks whether the owning rule failed to *fire*; a withdrawn claim fired fine, it was simply wrong. Self-declaring `yes` cannot use this class.
|
|
66
|
+
2. The evidence cell cites a **primary source that refutes the claim** (a URL).
|
|
67
|
+
3. The evidence cell carries a pointer to a **zero-loss obligation map** that is a git-tracked regular `.md` inside the repository (no absolute path, no `..`, mode `100644`), whose **anchor resolves to an actual heading** — not merely to a substring, or `#a` would pass on any file containing the letter — and the map **quotes verbatim every substantive line the round deleted**. That last bar is what binds the pointer to the withdrawal itself: to remove a real obligation you must first copy it into the map, where a reviewer sees it.
|
|
68
|
+
4. The owner's round diff is a **pure deletion**, classified structurally rather than by line text: every changed path is a regular `.md` under that owner, every path's status is a modification, **both tree modes are `100644` and identical**, and the diff adds **zero** lines while deleting at least one. Read the modes from `--raw`, not `--name-status`: a `chmod` shows up as a plain modification there and contributes `0/0` to `--numstat`, so a mode change would otherwise ride along with a genuine deletion. Two earlier drafts used proxies — a net-byte floor, then "adds no normative rule line" — and adversarial review broke both: the first by offsetting a smuggled rule with unrelated deletions, the second because a script edit or an `Always`-phrased rule matches no prose predicate. A proxy for the invariant gets bypassed; the invariant itself does not.
|
|
69
|
+
5. **Every** row for that owner in the round is in this class. A lone compliant withdrawal must not lift the floor for a sibling row carrying a real change.
|
|
70
|
+
|
|
71
|
+
Because a withdrawal is a pure deletion it has no added line to anchor on, so this class drops the firing-path requirement exactly as the two `not-required` classes do. Write the explanation of *why* the claim was withdrawn into a reference, not beside the line being removed — an inline note is an added normative line and bar 4 refuses it.
|
|
72
|
+
|
|
73
|
+
These bars are **mechanical floors, not proof**: they cannot bind the cited source to the specific obligation withdrawn. The zero-loss review in the dual-track gate remains what catches that, and this gate does not pretend to replace it.
|
|
74
|
+
|
|
75
|
+
**This class does not lift the per-owner RED floor.** Seven independent review rounds broke every successive bar built to make an automatic lift safe — a net-byte floor, a normative-line heuristic, a status-letter check, a file-exists check, a substring anchor, and finally a well-formed-but-unbound pointer — and two lanes twice recommended not granting the lift until the pointer can be bound to the obligation actually withdrawn. The residual gap is not machine-checkable in principle: a pointer can resolve to a real heading and quote the deleted text verbatim while the cited source has nothing to do with the claim. So the class does what a machine can do — force an honest label, a real obligation map, and a genuinely pure deletion — and leaves the judgement where it belongs: a withdrawal still needs a `RED-baseline` row, or a named risk owner's waiver through the existing human channel.
|
|
76
|
+
|
|
77
|
+
## Designing a behavioral-evidence measurement
|
|
78
|
+
|
|
79
|
+
A `RED-baseline` row is only as good as the measurement behind it. The failure modes below were each observed while measuring one small skill change; **all of them biased the same way — toward "the change worked" and "the measurement is sound".** That is the tell: when the person designing the measurement is the author of the change being measured, design errors are not random.
|
|
80
|
+
|
|
81
|
+
**The one rule that matters: the grading standard must precede the change.** Not "write the rubric carefully afterwards" — afterwards you already know what the new text says, and the rubric grows into its shape. Either freeze the rubric before editing, or have a party that has not seen the candidate produce it. This repo already applies preregistration to *dispositions* (a preregistered reading rule committed before any run); apply it to the *instrument* too.
|
|
82
|
+
|
|
83
|
+
Everything else is hygiene, but each was observed failing:
|
|
84
|
+
|
|
85
|
+
| Failure | What it looked like | Rule |
|
|
86
|
+
| --- | --- | --- |
|
|
87
|
+
| Ceiling | The unchanged arm already passed 9/10, so no improvement was detectable | The arm you expect to fail must be *able* to fail; verify before comparing |
|
|
88
|
+
| Vocabulary inheritance | Regex written after the new text; the old arm said the same thing in other words and scored 0 (`剥离` vs a regex for `剥掉`) | Grade the **obligation**, paraphrase-tolerant, by a grader blind to which arm produced the answer |
|
|
89
|
+
| Tautology by construction | A "marker must be uniquely carried by this line" rule forced markers that only that line's vocabulary could satisfy — removing the line trivially removed the word | A validity constraint built for one polarity becomes a bias generator in the other: uniqueness is required of the **control**, never of the candidate (other carriers are the redundancy evidence, not an artifact) |
|
|
90
|
+
| Underpowered null | n=3 nulls recorded as findings; one flipped to a large effect at n=10 | Nulls need power and replication; large separations survive small n, nulls do not |
|
|
91
|
+
| Cross-run drift | Identical inputs, same n, unchanged arm scored 2/10 and 6/10 in two runs | Replicate; report the spread, never a point estimate |
|
|
92
|
+
| Transcribed arms | The "old" arm was hand-copied and did not match the real prior text; it inflated the effect | Extract both arms from version control |
|
|
93
|
+
| Slice inflation | Removing one line from a 22-line excerpt overstates its weight versus the 338-line artifact | Use the whole artifact, or report both readings |
|
|
94
|
+
| Post-hoc analysis change | Verdict logic was changed after an unwelcome result | Freeze thresholds and decision order with the rubric; if changed later, publish the full sensitivity grid, never the favourable cell |
|
|
95
|
+
|
|
96
|
+
**Proving a removal is harmless is not the mirror of proving an addition helps.** An addition is evidenced by a difference; a removal is evidenced by an *absence* of difference, which is indistinguishable from an instrument that detects nothing. So a deletion claim needs a **positive control** — remove something known to carry an obligation and show that obligation's satisfaction actually drops. Better still, remove two candidates in the same run so each is the other's control: a double dissociation (removing A drops only A's obligations, removing B only B's) validates the instrument and both verdicts at once.
|
|
97
|
+
|
|
98
|
+
## Instruction-following mechanisms
|
|
99
|
+
|
|
100
|
+
Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds the withdrawn predecessor, the sources, and their evidence grade; the entrypoint holds only the operative rules.
|
|
101
|
+
|
|
102
|
+
**Withdrawn (do not cite).** An earlier version said: *prose rules do not reliably co-fire — at any single decision point the agent tends to actively apply roughly the ONE most-salient rule*, with the corollary that *appending a bullet does not ADD compliance — it competes for, and can displace, the same salience slot*. Neither vendor's guidance nor any public benchmark supports a winner-take-all salience slot, and the corollary is contradicted outright: OpenAI's guidance is that a single unequivocal sentence is usually enough to steer the model. The claim rested on one session's self-observation.
|
|
103
|
+
|
|
104
|
+
**Sources and what each actually supports.**
|
|
105
|
+
|
|
106
|
+
| Source | What it gives | Evidence grade |
|
|
107
|
+
| --- | --- | --- |
|
|
108
|
+
| [Anthropic, *Effective context engineering for AI agents*](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) | finite attention budget; "context rot" reported as gentler in some models but emerging across those tested; smallest-set-of-high-signal-tokens; the right-altitude failure modes | vendor engineering post — no dataset, n, or error bars; version-bound; commercially aligned with context-management tooling |
|
|
109
|
+
| [OpenAI, *GPT-4.1 Prompting Guide*](https://developers.openai.com/cookbook/examples/gpt4-1_prompting_guide) | conflicting instructions tend to resolve to the one nearer the end; instructions at both ends of long context beat either alone; check-conflicts-first; a single clear sentence usually steers | same class; explicitly model-generation-bound ("GPT-4.1 tends to…") |
|
|
110
|
+
| [OpenAI, *GPT-5.1 Prompting Guide*](https://cookbook.openai.com/examples/gpt-5/gpt-5-1_prompting_guide) | check-conflicts-first; a published metaprompt recipe for finding contradictions in your own system prompt | same class |
|
|
111
|
+
| [RECAST](https://arxiv.org/html/2505.19030) | joint satisfaction degrades with constraint count; best model averaged 39.75% all-constraints-satisfied on their benchmark | benchmark paper proposing its own dataset and method — a low baseline flatters the contribution; that number is one hard benchmark's order of magnitude, not a usage failure rate |
|
|
112
|
+
|
|
113
|
+
**Why recency is a hazard, not a rule.** Vendor guidance reports that models *tend to follow* whichever instruction sits later — an observation about behaviour, not a licence to resolve conflicts by position. Two ways position becomes dangerous if read as a rule: a later permissive line beats an earlier stricter one (directly contradicting `Conflict Resolution`, which keeps the stricter data-loss/security/contract guard); and text embedded in **untrusted data** — a diff under review, a retrieved document, tool output — sits later within the same authority level and would win by placement alone, which is prompt injection with extra steps. Treat recency as a bias to design against: put the load-bearing rule where the decision happens, and never let placement confer authority.
|
|
114
|
+
|
|
115
|
+
**How to use them.** Vendor guidance and benchmarks are **hypotheses with good provenance** — they tell you what to test on your own corpus, they do not substitute for testing it. Citing them as settled is the same error as landing an unverified claim; so is overruling one with an underpowered probe.
|
|
@@ -36,7 +36,9 @@ Two wording rules for whichever form wins:
|
|
|
36
36
|
- **No nuance clauses.** "Don't X unless it matters" reopens the negotiation — appending a single nuance clause to a winning recipe degraded it from consistent to noisy in the same tests. Write a real exception as its own conditional on an observable predicate.
|
|
37
37
|
- **Exemption clauses don't scope.** "This limit doesn't apply to code blocks" still suppresses code blocks; if part of the output must be exempt, restructure the rule so it cannot reach that part.
|
|
38
38
|
|
|
39
|
-
Boundary vs the owning Core Rule's salience mechanism: that rule governs how JOINT requirements are structured (walked enumeration at the firing point; merge-over-append); this table governs the form of a SINGLE rule's text once its landing spot is chosen.
|
|
39
|
+
Boundary vs the owning Core Rule's salience mechanism: that rule governs how JOINT requirements are structured (walked enumeration at the firing point; merge-over-append); this table governs the form of a SINGLE rule's text once its landing spot is chosen.
|
|
40
|
+
|
|
41
|
+
- **When the chosen structure IS a walked pause-point checklist, its design follows checklist practice** (WHO Surgical Safety Checklist implementation manual, 2009 — the operational form of The Checklist Manifesto's killer items): an item earns its slot only by being both critical AND not adequately caught by another mechanism; a pause-point section stays roughly five to nine items, overflow routing to the owning rule text instead of more items; each item binds to one specific, unambiguous action — for a DO-CONFIRM item that action is the confirmation itself, and compound confirmed data is allowed only when ONE concrete confirmation event verifies a coherent tuple (WHO's identity-site-procedure-consent item is one verbal patient-verification event, not a container for unrelated checks); obligations that are independently executable are split into their own items or stay in the owning rule, the item names the specific predicate being confirmed, and a rule-index may serve only as the pointer to that predicate's owning rule, never as a substitute for naming it (the Pre-Completion DO-CONFIRM card's index-not-restatement form works because each indexed canonical rule names what is being confirmed); a new or reshaped gate checklist is exercised on a real task shape before rollout (the design-time operability check's author-dogfood leg); and an item is never deleted merely because it keeps failing or is inconvenient — WHO explicitly discourages removing safety steps because they cannot currently be accomplished; a chronically-failing item routes through the same-class-recurrence keep/delete/narrow/replace decision, not silent removal. Provenance is external — adopt the principle, but transferred evidence does NOT exempt the change from the dual-track gate's mandatory behavioral-evidence row: for a behavior-shaping rule, the `RED-baseline` (recorded incident, or a constructed scenario run against BOTH the unchanged baseline and the changed rule — `dual-track-review-gate.md`) is exactly where the form choice gets tested on YOUR case, with the instrument scaled per `harness-patterns-and-eval.md` §3 (before-after / golden-trace for a single skill edit; the system-wide `skill-behavior-eval` fixture battery is ONLY for always-on layer changes, not routine rule edits). A prohibition chosen for an output-shaping failure, or a recipe replacing a prohibition, is a form choice the source measured as capable of backfiring — it must not land on transferred deltas alone.
|
|
40
42
|
|
|
41
43
|
## Obligation-preservation table (required for any rewrite / retire)
|
|
42
44
|
|
|
@@ -261,3 +261,37 @@ and inverted the sense (production, not product), and the coordinator now shares
|
|
|
261
261
|
| The Python side of the same architecture-pair registration: the package-owned runtime-invariant registry candidate parks as a backlog entry beside the runtime-readiness rules it would extend, with the evaluation deferred to the owner's next runtime-readiness/diagnostics round and the outcome recorded against the entry | `python-service-architecture` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/python-service-architecture/references/packaging-runtime-readiness.md#Evaluate at the next Python architecture round touching runtime readiness | updated | `python-service-architecture/SKILL.md` is the owner key and is unchanged in this round; the change is a new `## Topic-extension backlog` tail entry in `python-service-architecture/references/packaging-runtime-readiness.md` (registration only; packaging/runtime rules untouched), cross-linking the Go-side sibling entry. Candidate evidence status: weak, single source, sanitized as the same agent-native product repository (evolving portfolio). RED-baseline (applied, differential; evidence scope: registration↔ledger binding, not candidate semantics): same throwaway-copy mutation protocol — deleting this entry turns the impact-chain gate red on exactly this row's anchor; restore returns green; control green before mutation. Terminal dispositions of the remaining three candidates (discard / no-new-lesson / confirm-only, with reasons and deciding authority) are recorded in `specs/033-runtime-face-borrowing/plan.md` 第三档处置登记 rather than as owner rows, because they change no owner package; dual-track: 033 tier-3 registration round. |
|
|
262
262
|
| A product-agnostic architecture skill's boundary-and-contract main line carries its external grounding inside the package, scoped as adopted-in-part — bounded-context model-boundary criteria (as boundary inputs, not a service-per-context mandate) and Conway's-law team-structure criteria ground the split rules without claiming the full DDD strategic-design method or an inverse-Conway process — so the rules' authority class is verifiable from the package alone rather than from a repo-level doc | `go-microservice-architecture` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/go-microservice-architecture/references/architecture-playbook.md#External grounding (adopted in part) | updated | `go-microservice-architecture/SKILL.md` is the owner key and is unchanged in this round; the change is one grounding paragraph in `go-microservice-architecture/references/architecture-playbook.md` §Service Boundary Rules (citations plus a borrowed-scope statement only; the boundary rules themselves are untouched). Source verification per the theory doc's verify-title-not-just-HTTP-200 rule: both cited pages fetched this round with titles confirmed — "Bounded Context" (Martin Fowler, 2014-01-15, central DDD strategic-design pattern) and "Conway's Law" (Martin Fowler, 2022-10-20, law credited to Melvin Conway's 1968 Datamation article, named by Brooks); primary sources also verified — Conway's 1968 paper at the author's own page ("Committees Paper", melconway.com, thesis text confirmed) is cited in both bullets, and the Evans origin is verified via the Fowler page's in-text footnote ("Eric Evans in Domain-Driven Design"). RED-baseline (applied, differential; evidence scope: grounding↔ledger binding, not rule semantics): in a throwaway copy of the landing candidate, rewording the grounding bullet's anchor phrase away turns the impact-chain gate red naming exactly this owner's evidence path missing, the Python-side row unaffected; restore returns green; control green before mutation. A bare deletion probe is degenerate for this one-line change shape — deleting the added line can restore the owner file to base so the owner leaves the changed set and the row is rightly unchecked — which is why the applied mutation rewords rather than deletes. Program record: `specs/034-theory-debt-repayment/plan.md` 批次一, repaying the self-named debt row in `docs/skills-theory-foundations.md`; dual-track: 034 batch-1 round. |
|
|
263
263
|
| The Python side of the same grounding pair: the boundary rule in the architecture playbook carries the identical adopted-in-part citations with the same borrowed-scope statement and a cross-link to the Go sibling, so the pair cannot drift to one stack only | `python-service-architecture` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/python-service-architecture/references/architecture-playbook.md#External grounding for the boundary rule above | updated | `python-service-architecture/SKILL.md` is the owner key and is unchanged in this round; the change is one grounding bullet in `python-service-architecture/references/architecture-playbook.md` §Architecture Decisions beside the boundary rule it grounds (citations plus borrowed-scope statement only; the rules themselves are untouched). Same source verification as the Go-side row. RED-baseline (applied, differential; evidence scope: grounding↔ledger binding, not rule semantics): same throwaway-copy reword protocol — rewording this bullet's anchor phrase away reds the gate naming exactly this owner's evidence path missing, the Go-side row unaffected; restore returns green; control green before mutation. The deletion probe was run first and observed GREEN: deleting the single added bullet restores this file to base, the owner leaves the changed set, and the row is unchecked — recorded as the degenerate-probe arm that motivated the reword mutation. Program record: `specs/034-theory-debt-repayment/plan.md` 批次一; dual-track: 034 batch-1 round. |
|
|
264
|
+
| Platform-spec conformance is a walkthrough lens with per-criterion single-source provenance: first-party HIG/Material rules that intersect the skill's existing main lines (state completeness, accessibility, platform conventions) land only as pass/fail walkthrough criteria labeled by source, and spec content that cannot form a criterion is recorded as deliberately-not-absorbed instead of being silently dropped or link-dumped | `product-ui-ux-design` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md#Disabled semantics are real, not painted | updated | `product-ui-ux-design/SKILL.md` is the owner key and is unchanged in this round; the change lands in `product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md` (new Platform Convention Walkthrough section: 25 criteria each labeled (HIG) or (Material) single-source, 5 deliberately-not-absorbed records, Source Anchors upgraded to name the first-party specs) and `product-ui-ux-design/references/ui-ux-audit.md` (Audit Procedure step 10 entry), with `docs/skills-theory-foundations.md` UI/UX row upgraded to dual-lens 🔗 in the same landing. Source verification per the theory doc's verify-title-not-just-HTTP-200 rule: four HIG pages (Feedback / Accessibility / Loading / Designing for iOS) fetched full-text via Apple's documentation JSON endpoint with titles confirmed; two Material 3 pages (States, slug interaction-states; Designing, slug designing — Color contrast / Structure / Flow / Elements) fetched full-text via the site's content JSON endpoint after headless-browser endpoint discovery (SPA body unreachable by plain HTTP — blocked-source remediation, not a skipped read). The 44×44pt minimum hit region verified against Apple's own HIG Buttons page text, not Material's paraphrase of it. RED-baseline (applied, differential, throwaway copy; evidence scope: criterion↔ledger binding, not criterion semantics): rewording the anchored criterion line inside the committed round without a superseding row turns `impact-chain-gate.rb` red naming exactly this owner's evidence path missing; restoring returns green; control green before mutation — the reword-not-delete protocol inherited from 批次一's degenerate-deletion-probe record; the command/exit transcript, bound to the probe commit, gate-script sha256, and candidate-diff sha256, is retained in the per-host round evidence (per the extraction-lifecycle rule that keeps provenance artifacts out of the shared tree). Criterion semantics are reference prose reviewed by this round's dual-track rows. Program record: `specs/034-theory-debt-repayment/plan.md` 批次二; dual-track: 034 batch-2 round. |
|
|
265
|
+
| The platform walkthrough binds design acceptance only; implementation mechanics for the same platform rules (Dynamic Type adoption, native reduce-motion detection, predictive-back integration) stay with the stack owner via the routes the criteria themselves name, and this round adds no new implementation obligation | `app-cross-platform-dev` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md#predictive-back geometry routes to | unchanged | `app-cross-platform-dev` package byte-identical this round (paired control); each mechanic named here is routed in the criterion that raises it — the text-scaling criterion assigns Dynamic Type adoption mechanics to the stack implementation owner directly, and the reduce-motion and predictive-back criteria route through `platform-mobile-patterns.md` sections that already hand mechanics to the stack owner. Program record: `specs/034-theory-debt-repayment/plan.md` 批次二. |
|
|
266
|
+
| Mini-program surfaces stay outside the HIG/Material walkthrough scope: their platform spec is the host platform's own operating rules (already cited in the theory doc's platform row), and importing Apple/Google conventions into mini-program acceptance would be exactly the false cross-platform standard claim the walkthrough's per-criterion source labels exist to prevent | `miniapp-product-dev` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/product-ui-ux-design/references/ui-ux-audit.md#Check platform-convention conformance | unchanged | `miniapp-product-dev` package byte-identical this round (paired control); the audit step scopes the walkthrough to iOS/Android app surfaces by name. Program record: `specs/034-theory-debt-repayment/plan.md` 批次二. |
|
|
267
|
+
| Walkthrough criteria are design-acceptance checks, not test-layer selections: which assertion layer or rendered-evidence form proves a criterion in CI stays owned by the testing skill, and this round changes no layer-selection rule | `testing-strategy` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md#Interaction-state matrix is complete | unchanged | `testing-strategy` package byte-identical this round (paired control); the walkthrough section intro frames every criterion as a rendered-surface acceptance check, naming no assertion layer. Program record: `specs/034-theory-debt-repayment/plan.md` 批次二. |
|
|
268
|
+
| An LLM product eval names the decision it supports before choosing datasets or scorers and carries an explicit run protocol; a selection benchmark's average is not by itself a release verdict | `llm-inference-integration` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#Conclusions do not transfer automatically across decision classes | updated | Expected executable behavior: the Evaluation design field list opens with the supported-decision field and includes a run-protocol field, so an eval designed from this reference cannot silently reuse a selection-style average as a release gate. `llm-inference-integration/SKILL.md` is the owner key and is unchanged this round; the change is two added fields in `llm-inference-integration/references/model-prompt-evaluation.md` §Evaluation (supported-decision field with the cross-class non-transfer clause; run-protocol field). External grounding verified first-hand this round: `docs.anthropic.com/en/docs/test-and-evaluate/define-success` (success-criteria-first), `developers.openai.com/api/docs/guides/evaluation-best-practices` ("Define eval objective" as design step 1; nondeterminism framing), `docs.langchain.com/langsmith/evaluation-types` (when/why eval-type taxonomy separated from evaluator implementations), `arxiv.org/abs/2411.00640` Recommendation 3 (resample per task, score question-level averages); the decision-class list is this reference's working framework informed by those public taxonomies, not claimed as source-established, and a secondary popularization article served as the discovery lead only with no clause resting on it as evidence. The clause set states non-transfer as not-automatic with an explicit combined-predeclared-design allowance, scopes fixed-conditions to everything except the declared variable under test with both values recorded, scopes question-level averaging to aggregate quality metrics with red-line checks aggregating any-hit, and matches the repetition protocol to the decision and observed nondeterminism (single runs only for end-to-end-deterministic paths, generation included, or with recorded variance justification — a deterministic grader over stochastic output does not qualify; resampling scoped to offline/replay comparisons, online production-monitoring decisions using predeclared windowed operational thresholds, failure diagnosis using sanitized isolated replay of recorded inputs, and live side-effecting requests never re-executed). RED-baseline (constructed scenario, both arms actually run, same prompt except the guide excerpt, headless same-family model; changed arm run against the exact landed text): baseline arm — a field list built from the unchanged §Evaluation names no decision class, no red-line blockers, and no fixed-conditions protocol; changed arm — names the release/regression-gate decision as field 1, treats the prompt version as the declared variable under test, requires any-hit red-line aggregation, and requires recorded justification for single-run subjective scoring. Scope honesty: both arms reached a do-not-ship-yet verdict (the baseline model reached it through its own variance reasoning), so the RED delta is design-field coverage and the named rationale, not a flipped final verdict; run artifacts archived in the maintainer-local extraction workdir. Downstream dispositions: `testing-strategy` unchanged (layer/CI-gate policy untouched; LLM eval mechanics remain routed to this owner), `product-rd-workflow` unchanged (launch-gate ownership already routed in the owner SKILL.md routing section), implementation landing is this owner's own reference. Dual-track: this round's review and challenge rows recorded in the round validation log. |
|
|
269
|
+
| The regression runner's execution-surface contract gains a `--heavy-only` mode so CI runs the fast and heavy lanes as parallel jobs instead of serially inside one `--full` job, while `--full` stays the local aggregate and the registration self-audit keeps every sibling suite in exactly one lane — lane parallelization therefore cannot drop a suite from CI | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh | updated | `skill-extraction-workflow/SKILL.md` is unchanged; the change is the mode dispatch, usage text, lane-conditional execution, and a fixture-redirect for `SCRIPTS_DIR` via `REGRESSION_SCRIPTS_DIR` in `scripts/test_check_ccl_regressions.sh`, plus the new sibling `scripts/test_regression_runner_lanes.sh` registered in fast_tests (the registration self-audit is untouched; heavy_tests is untouched). RED-baseline (temporal, applied, differential): on the pre-change script, `--heavy-only` exits 2 with usage (observed, recorded in the round plan), while `--fast`, `--full`, and `--list-unregistered` keep rc-parity pre- and post-change (control arms); post-change `--heavy-only` runs exactly the heavy_tests entries with one `regression_test_timing` line each, ends `test_check_ccl_regressions_heavy_only_ok`, exit 0. The lane semantics are held mechanically by `test_regression_runner_lanes.sh` (per-mode exact-lane multiset against a stub fixture, red-heavy propagation, per-mode tokens), added on the 036 challenge's P2 finding; the same round's second challenge hardened the registration self-audit from an advisory tail to a hard failure in every execution mode (an unregistered sibling test now reds `--fast`/`--full`/`--heavy-only` directly, so enforcement no longer depends on the registration guard test staying registered), with the RED path held by the lanes test's unregistered-sibling case. CI consumption moves to two jobs (regression-fast `--fast`, regression-heavy `--heavy-only`) with the same base-ref guard on both — the fast lane's clone-driving catalog suite needs `CCL_SKILL_BASE_REF` (the 022 round's observed failure class). Round plan and decision table: `specs/036-ci-parallel-shards/plan.md`; dual-track: 036 round. |
|
|
270
|
+
| Recognized theory carries a second harvest with a named in-workflow trigger (recognition-dependent, honestly bounded) and a legal empty-handed outcome: the Deep RCA gains a bounded second-harvest check fetching the theory doc's uncovered-parts and failure-boundary questions, external attributions in RCA carry the same falsifiable-evidence bar as self-causes with restating-louder-under-challenge named as the self-sealing signature, formal process descriptions are work-as-imagined requiring a work-as-done trace before their rules count as practice-derived, bounded retros weigh frequency alongside severity when selecting events to dig into, the but-for test gets an emergent-outcome boundary (reconstructed factors stay probabilistic and the control shifts to condition-constraint form), and walked pause-point checklists follow the killer-item selection, five-to-nine size, one-action binding (compound canonical obligations allowed as WHO's own items are), tested-before-rollout, and no-deletion-for-inconvenience bounds | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#The mirror trap is the untested | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round (entrypoint is severe-debt, capacity recorded as a map input; every landing merges into reference sections the entrypoint already routes to: Deep RCA move 3, the Documents source-type row, the new Second-harvest check paragraph and two method-provenance lines in `references/source-to-skill-extraction.md`, and one checklist-design bullet in `references/rule-consolidation.md`), with `docs/skills-theory-foundations.md` gaining the reciprocal firing-point note in the same landing. Sources fetched primary and title-verified: Argyris HBR 1991 full text via academic mirror of the original reprint after the origin paywall (blocked-source remediation), the EUROCONTROL Safety-II white paper PDF, and the WHO Surgical Safety Checklist implementation manual 2009 (ISBN-verified); the unfetchable Gawande book contributes only the killer-items term via the theory doc's existing row, with every executable bound re-anchored to WHO primary text. Pilot outcome 3/3 landed, so the plan's empty-handed downgrade clause was not triggered. Behavioral evidence, scoped per clause: (a) the external-attribution clause got a guidance-as-prompt base/head differential (provider `codex`, codex-cli 0.148.0, 4 rounds/arm, arms = the pre/post move-3 text with recorded sha256s, one neutral-domain RCA scenario, prompts and all eight raw answers archived in the per-host round evidence) — the evidence-bar half scored **base 4/4 vs head 4/4** (no delta: the unchanged hypothesis-grade language already pushes there; recorded honestly rather than dressed up), while the self-sealing-signal half scored **base 0/4 vs head 4/4** (no base answer classified restating-louder as a defensiveness signal on the analysis; every head answer did); (b) the other five clauses each got a bounded 2-round/arm guidance-as-prompt scenario differential (same provider and archival discipline; per-clause scenarios and all twenty raw answers in the per-host round evidence), scored honestly per half: WAI/WAD **base 0/2 vs head 2/2** (only head answers demand a work-as-done trace and the WAI-only label); second-harvest **base 1/2 vs head 2/2** (one base answer improvised a similar bounded pass from workflow priors but without the two-question bound or the terminal-outcome contract); checklist-design **base 0/2 vs head 2/2** on selection/size/rollout specifics (base answers reached the deletion guard from existing blocked-verification priors — recorded as base-covered for that half); frequency-weighting **outcome no-delta (both arms 2/2 chose the frequent event)** with a rationale delta only (head cites frequency-weighting explicitly; base reasoned from sustain/non-luck priors); emergent-boundary **verdict half no-delta** (base already reaches `probabilistic` from the existing one-trace rule) with the control-form half showing the delta (only head answers produce condition-constraint control rows). A constructed absence check additionally distinguishes the arms textually (base blob greps 0 for the new obligations' key phrases, head ≥1); (c) row↔rule-line binding: rewording the anchored move-3 sentence turns `impact-chain-gate.rb` red naming exactly this owner; restore green; control green before mutation; hash-bound transcript regenerated on the final landing candidate. Anchor-coverage disclosure (per this ledger's narrower-than-the-rule precedent): this row's locator binds the move-3 sentence only; the emergent-boundary and checklist-bounds landings carry their own anchored rows below with their own applied probes, while the WAI/WAD (table row), second-harvest (titled paragraph), and frequency (paragraph) landings sit on line shapes the gate's list-line anchor cannot bind — deleting them while keeping the anchored sentences leaves the gate green, so their protection rests on the dual-track rows and the obligation-preservation machinery, not on this gate, and this row does not claim otherwise. Program record: `specs/034-theory-debt-repayment/plan.md` 批次三; dual-track: 034 batch-3 round. |
|
|
271
|
+
| The but-for test carries an emergent-outcome boundary: when a counterfactual will not stabilize in an intractable interaction, reconstructed factors stay `probabilistic` — no forced necessary/sufficient verdict — and the control row shifts to move 5's condition-constraint form, constraining the conditions that let the outcome emerge instead of eliminating a reconstructed cause | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#Emergent-outcome boundary (Safety-II) | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round; the clause lands inside Deep RCA move 4 (`references/source-to-skill-extraction.md`), sourced from the EUROCONTROL Safety-II white paper's emergence chapter (causes reconstructed rather than found; control the conditions). Scenario differential (2 rounds/arm, archived per-host): verdict half no-delta — base already reaches `probabilistic` via the one-trace rule — control-form half base 0/2 vs head 2/2 (only head answers produce condition-constraint control rows). RED-baseline (applied, differential, throwaway copy): rewording this row's anchored sentence turns the gate red naming exactly this owner; restore green. Program record: `specs/034-theory-debt-repayment/plan.md` 批次三. |
|
|
272
|
+
| A walked pause-point checklist is designed by checklist practice: items earn slots by the killer-item test (critical AND not adequately caught elsewhere), a section stays roughly five to nine items with overflow routed to the owning rule text, each item binds one confirmation event to a named predicate (compound data only as one coherent tuple), the gate is exercised on a real task shape before rollout, and no item is deleted merely for failing or being inconvenient — chronic failure routes through the keep/delete/narrow/replace decision | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/rule-consolidation.md#When the chosen structure IS a walked pause-point checklist | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round; the bullet lands in `references/rule-consolidation.md`, sourced from the WHO Surgical Safety Checklist implementation manual 2009 (Focused/Brief/Actionable/Tested principles; removing-safety-steps discouraged), with the killer-items term attributed via the theory doc's existing Checklist Manifesto row. Scenario differential (2 rounds/arm, archived per-host): selection/size/rollout specifics base 0/2 vs head 2/2; the deletion-guard half recorded honestly as base-covered via existing blocked-verification priors. RED-baseline (applied, differential, throwaway copy): rewording this row's anchored bullet start turns the gate red naming exactly this owner; restore green. Program record: `specs/034-theory-debt-repayment/plan.md` 批次三. |
|
|
273
|
+
| Attribution testing for defect causes stays with the diagnosis skill's existing first-hand-evidence discipline — AI-proposed or externally-blamed causes are already hypothesis-grade there until reproduced; what this round adds is retro-process-specific (the self-sealing signature on the ANALYSIS itself), owned by the extraction workflow, so the diagnosis skill carries no new rule | `defect-diagnosis` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#state it so others can test it | unchanged | `defect-diagnosis` package byte-identical this round (paired control); its description and phase discipline already demand first-hand failure evidence before causes are accepted. Program record: `specs/034-theory-debt-repayment/plan.md` 批次三. |
|
|
274
|
+
| Checklist-design bounds are rule-form guidance for authoring walked gates, not test-layer or coverage policy: which assertion layer proves a behavior stays owned by the testing skill, and the mutation-walk/coverage rules it owns are untouched by this round | `testing-strategy` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/rule-consolidation.md#an item earns its slot only by being both critical | unchanged | `testing-strategy` package byte-identical this round (paired control); the new bullet governs how a pause-point checklist is designed inside rule text, naming no assertion layer or coverage threshold. Program record: `specs/034-theory-debt-repayment/plan.md` 批次三. |
|
|
275
|
+
| A lane whose cost is dominated by WAITING (process timeouts, stub sleeps) or by independent subprocesses is scheduled, not rewritten: a shared bounded-concurrency suite runner owns the missing-file precondition, per-suite timing, input-order output replay, and failure propagation, so a lane's cost drops from the sum of its suites to its slowest suite while every suite's assertions stay untouched — and the precondition that makes it sound (no mutable out-of-repo state shared between suites in one lane) is a per-lane audit the runner cannot verify for you | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh | updated | `skill-extraction-workflow/SKILL.md` is unchanged; the change is that `scripts/test_check_ccl_regressions.sh` replaces its serial `run_test` loop with `run_lane`, delegating both lanes to the new repo-root `scripts/run-parallel-suites.sh` (fast_tests/heavy_tests arrays, mode dispatch, and the registration hard-fail are untouched). RED-baseline (applied, differential, throwaway copies): three mutations of the runner each red its own assertion in `scripts/test_run_parallel_suites.sh` and nothing else — reverting the semaphore to `jobs -r` (line-counting, which reads one multi-line background job as several) reds `concurrency is not happening`; neutering the failure exit reds `did not make the runner exit nonzero`; reversing launch+replay order reds `replayed in completion order, not input order`; the unmutated control is green each time. That first mutation is not hypothetical: it was the runner's actual first-draft defect, caught by the concurrency assertion before landing. Lane semantics under concurrency are unchanged — `test_regression_runner_lanes.sh` and the registration guard both stay green. Sibling landing in the same round: `code-review`'s `test_init_policy_matrix.sh` registers its 16 mutants and dispatches them concurrently against the same unmodified assertion function (239s to 46s locally, 16/16 still asserted); making one registered mutation a no-op reds the walk with `flipped NOTHING`, proving the parallel walk kept its sensitivity. Round plan, the per-lane shared-state audit, and the corrected floor claim: `specs/037-ci-intra-job-parallel/plan.md`; dual-track: 037 round. |
|
|
276
|
+
| A mutation-walk suite is sped up by rescheduling it, never by thinning it: the registered mutants are dispatched concurrently against the SAME unmodified assertion helper, each still writing its own disposable parser copy, and every registered index must report an EXPLICIT terminal status — a worker that dies without reporting is a failure, because treating a missing marker as success would bank a mutant that never ran as proof of sensitivity | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/test_init_policy_matrix.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged in this round; the change is confined to `skills/code-review/scripts/test_init_policy_matrix.sh`, which now registers its 16 mutants and dispatches them with bounded concurrency instead of calling them inline. `mutate_and_expect_mismatch` — the helper holding every assertion, including the guard self-check that proves a broken mutant is rejected for the right reason — is byte-identical, and the guard self-check still runs serially ahead of the walk. RED-baseline (applied, differential, throwaway copy): making one registered mutation a no-op (replacement identical to its find string) turns the walk red with `mutation drop-host-vocabulary-breach-guard flipped NOTHING`, so the parallel walk kept the sensitivity the serial one had; the unmutated control is green. Measured 239s -> 46s locally with 16/16 `sensitive to` lines still emitted. The missing-terminal-status rule was added on the 037 challenge's finding that a killed worker would otherwise read as a pass. Round plan and the lane-isolation gate that now enforces the concurrency precondition mechanically: `specs/037-ci-intra-job-parallel/plan.md`; dual-track: 037 round. |
|
|
277
|
+
| Capturing a worker's outcome must preserve errexit: running the assertion helper as the condition of an `if` disables errexit for its entire body, so an unguarded failure partway through can still be followed by the helper's final successful command and be recorded as a pass — a mutant that never completed banked as proof of sensitivity. The worker runs the helper as a simple command and records the real exit status from an EXIT trap | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/test_init_policy_matrix.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged; the change is confined to the concurrent dispatch in `skills/code-review/scripts/test_init_policy_matrix.sh`, and `mutate_and_expect_mismatch` itself stays byte-identical. Found by the 037 review chain's eighth round against the dispatch this same round introduced. RED-baseline (applied, differential): with the fix in place, making one registered mutation a no-op still reds the walk with `flipped NOTHING`, and the full walk reports 16/16 `sensitive to` at exit 0 in 46.8s; the unmutated control is green. Round plan: `specs/037-ci-intra-job-parallel/plan.md`; dual-track: 037 round. |
|
|
278
|
+
| A deferred fix registered against our own mechanism decays on both halves — the behavior it recorded and the remedy it proposed — so the fix round reproduces each registered form against the current baseline with a control leg before implementing, closes (not fixes) forms that no longer reproduce, and re-derives the remedy from the current code instead of applying it as written | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#reproduce each registered form against the current baseline | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round: the entrypoint is a severe size-debt surface and its own no-monotonic-growth gate refused the addition, so the clause lands in `skill-extraction-workflow/references/source-to-skill-extraction.md` as a charter-adjacent subsection, firing at Step 0 — the moment a fix round for a deferred item opens — as the inward arm of the entrypoint's existing hypothesis-grade rule. Observed failure: three gate forms deferred across three batches, and on the fix round one had a precondition the registration never mentioned, one's registered remedy was unimplementable as written (measured against this register's own rows it would have rejected a large minority of correctly-authored ones), and the control leg caught two probe verdicts that were red for a fixture defect rather than for the form. RED-baseline (applied, differential, throwaway copy): rewording this row's anchored sentence turns the gate red naming exactly this owner; restore green. Program record: `specs/038-impact-chain-attribution/plan.md`. |
|
|
279
|
+
| The impact-chain gate resolves a row's owner by owner-package prefix — conditioned, for the widened form, on the row carrying a behavioral-evidence declaration so the register's other tables keep their prior behavior — and asks changed-ness as a separate question afterwards, so an owner leaving the changed set no longer switches off both the row's evaluation and its demand at once; it refuses a surviving row that vouches for an owner this diff does not change, blocks (no longer merely warns) a row resolving to more than one selected owner, and evaluates rows when the ledger alone changed | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key; the change lands in `skill-extraction-workflow/scripts/impact-chain-gate.rb` with probes in `skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh` (a control leg, the two false-green forms, and an unrelated-five-column-table regression guard). Observed failure: an owner reverted to base bytes kept a green ledger row; a row citing a package path other than SKILL.md was never evaluated. Both reproduce against the base gate and are closed against this one; the regression guard reproduces against the intermediate unconditioned widening. Before/after differential over EVERY merge commit reachable from the baseline — 64 points, the list asserted against git's own enumeration — plus two synthetic replay cases, one baseline-red so the loosening direction is reachable: 60 points agree in both directions, and 4 diverge as the intended refusal of the corrective-rewrite back-fill shape, each named individually and constrained to `newly refused` with that specific diagnostic. Both an always-accept and an always-refuse gate are caught by it. Program record: `specs/038-impact-chain-attribution/plan.md`. |
|
|
280
|
+
| Probe-control discipline (a differential fixture needs a control leg proving its own shape passes, so a verdict red for a fixture defect is not read as the defect reproducing) is already owned by the testing side of the dual-track gate's RED-baseline attribution rule, which requires the owning assertion to pass in the unmutated control and fail under the mutant; this round exercises that rule rather than adding one | `testing-strategy` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#reproduce each registered form against the current baseline | `unchanged` | `testing-strategy` package byte-identical this round (paired control); its mutation-walk and coverage rules are untouched, and the new clause governs when a deferred registration may be implemented, naming no assertion layer or coverage threshold. Program record: `specs/038-impact-chain-attribution/plan.md`. |
|
|
281
|
+
| A gate change owes evidence in BOTH directions: probes for every acceptance requirement it states — not only for the defect forms that motivated it — and a re-runnable no-verdict-regression differential, because an asserted \|I checked the old history\| is the half a reviewer or reverter cannot reproduce | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key; the suites land in `skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh` (now six legs — a control, the two forms, and one guard each for the unrelated-table, multi-owner-blocking and non-curated-owner acceptance requirements) and `skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh` (twelve SHA-pinned integration points, baseline gate versus candidate gate, failing on regression and on a missing ref rather than skipping), both registered in `skill-extraction-workflow/scripts/test_check_ccl_regressions.sh`. Observed failure: the suite reported \|all forms closed\| while two stated acceptance requirements had no leg at all, and the regression claim rested on a scratch script. RED-baseline (applied, differential): the multi-owner leg passes against the base gate until the round's two owners are validly declared first, at which point it correctly reproduces the warn-and-drop it exists to detect. Program record: `specs/038-impact-chain-attribution/plan.md`. |
|
|
282
|
+
| A behavior that a control's own docstring names as its reason to exist must have a probe that reds when that branch is short-circuited: `host_entry_is_whole` rejects non-string host-vocabulary entries because a dict can hide a path in a sibling key or under an allowed built-in name, yet that branch could be flipped to always-whole with the dedicated suite still green — a lossy-read control whose own anti-lossiness branch was unprobed | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_parse_probe_result.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the assertion lands in `code-review/scripts/test_parse_probe_result.sh`. Applied mutation with differential attribution: unmutated tree + new case = rc 0 (`parse_probe_result_tests_ok`); non-string branch short-circuited to whole + new case = rc 1 (`expected native-skill probe parser failure`); implementation restored with `git diff` empty against it. The case is an owner-aware init whose `slash_commands` carries a dict whose `name` is an allowed built-in while a sibling key holds the discarded proof |
|
|
283
|
+
| An anchor that RESOLVES is not an anchor that FIRES: a `file:` locator into an executable artifact was graded solely on its anchor text being present, so once the mechanism it belonged to retired and nothing ran that file, the rows kept asserting firing evidence while the gate stayed green for the whole life of the retirement | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the check lands in `skill-extraction-workflow/scripts/register-firing-path-resolution.rb`. A `file:` locator whose target carries an executable extension is additionally checked for a runner: with no entry point outside `specs/` naming that file it is reported as an unrunnable target. Advisory by construction — the exit status never changes — because the failure it names is a decayed claim, not an unsafe landing; prose (`.md`) anchors stay out of scope since they are read, not run. Live-corpus premise check, mandatory for a tightening: the run is NOT clean — five rows anchored into a retired gate's evidence script are named, prose-anchor false positives are zero, and `register_firing_path_resolution_ok (176 locators resolved)` still prints with exit 0 |
|
|
284
|
+
| An advisory added to a gate is itself a gate change and owes the round that added it a row of its own: the round landed the unrunnable-target check in the resolver while its only ledger row named the sibling owner, so the diff carried a behaviour change for this package with nothing vouching for it | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round; the change lands in `skill-extraction-workflow/scripts/register-firing-path-resolution.rb`, which gains an advisory reporting `file:` locators whose executable target no runner names. RED-baseline (applied, differential): the advisory names ten locators across nine ledger rows on the live corpus and none once their target is reachable, while `register_firing_path_resolution_ok` and exit 0 are unchanged in both runs — the check adds a report, never a verdict |
|
|
285
|
+
| An anchor that RESOLVES is not an anchor that FIRES: a `file:` locator into an executable artifact was graded solely on its anchor text being present, so once the mechanism it belonged to retired and nothing ran that file, the rows kept asserting firing evidence while the gate stayed green for the whole life of the retirement | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the check lands in `skill-extraction-workflow/scripts/register-firing-path-resolution.rb`. A `file:` locator whose target carries an executable extension is additionally checked for a runner: with no entry point outside `specs/` naming that file it is reported as an unrunnable target. Advisory by construction — the exit status never changes — because the failure it names is a decayed claim, not an unsafe landing; prose (`.md`) anchors stay out of scope since they are read, not run. Live-corpus premise check, mandatory for a tightening: the run is NOT clean — ten locators across nine ledger rows, all anchored into a retired gate's evidence script, are named, prose-anchor false positives are zero, and `register_firing_path_resolution_ok (176 locators resolved)` still prints with exit 0 |
|
|
286
|
+
| An advisory that promises never to change the exit status must survive the tree it walks: hoisting the candidate reads out of the per-locator loop to kill an O(locators x files) cost also removed the short-circuit that had been hiding them, so one unreadable file anywhere under the root took the whole gate red — the claimed behaviour and the actual behaviour diverged, which is the exact defect class this advisory exists to report | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the change lands in `skill-extraction-workflow/scripts/register-firing-path-resolution.rb`. Applied mutation with differential attribution: a `chmod 000` file placed at the repo root took the gate to exit 1 with `Errno::EACCES` before the fix and leaves it at exit 0 after, with the advisory still naming the same ten locators across nine ledger rows (one row carries two locators) in both the with-probe and without-probe runs. The reachability corpus is now built once behind a lambda memoized in a local, so a repo with no executable locators never walks the tree at all; unreadable candidates are skipped rather than raised, which biases toward silence on that candidate — the safe direction for a report that must not block |
|
|
287
|
+
| A regression case earns its place by failing for the reason it names, and its fixture must not buy that failure with a side effect the repo forbids: the first version pinned the dict-shaped host entry with a literal `/tmp` path, which the lane-isolation gate flags, and the gate's own escape hatch — declare the string inert and register it in ALLOWED — would have let the author self-certify the very exemption the case exists to deny elsewhere | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_parse_probe_result.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/test_parse_probe_result.sh`. The fixture value was rewritten to carry no fixed temporary path instead of being exempted, so `lane_isolation_ok` reports zero unreviewed members with no new entry in ALLOWED. Discrimination re-verified after the rewrite: unmutated tree = rc 0 (`parse_probe_result_tests_ok`), non-string branch of `host_entry_is_whole` short-circuited to whole = rc 1 (`expected native-skill probe parser failure`), implementation restored with `git diff` empty against it |
|
|
288
|
+
| A report states only counts it can recompute from its own output, and a count read through a filter that sees half the output is not one: `grep -c` over the path lines counted neither the anchor lines beneath them nor the two locators sharing a row, so three artifacts carried three different wrong totals and an external reviewer arrived at a fourth by arithmetic from the wrong shape | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the change lands in `skill-extraction-workflow/scripts/register-firing-path-resolution.rb`. The advisory now declares its covered runner shapes and its blind spots (Rakefiles, extensionless bin scripts, `find -name` invocation) in its own output rather than leaving them for a reader to discover, and reports how many candidates were unreadable and skipped so a lone-unreadable-runner false alarm is distinguishable from a decayed anchor. RED-baseline (applied, differential): a `chmod 000` file at the repo root makes the skip line appear while the exit status stays 0, and removing it makes the line disappear; the named set is ten locators across nine ledger rows in both runs, taken from the block's raw output rather than from a filter over it |
|
|
289
|
+
| A suite that starts session-detached, signal-immune helper processes owns both their lifetime bound and their reaping on abort: the case's own postcondition can be green while the process table is not, because the component under test is the only reaper on the happy path and an abort is exactly what removes it. Bound the fixture's hang, extend the cleanup trap to the abort signals, and assert at suite exit that nothing the run started is still alive — the acceptance object is a clean process table, not a green case. The probe for such a property must prove its own preconditions before its verdict is readable: selecting the target by a timing heuristic, or reading the verdict without showing the reaper was actually removed, produces a probe that passes against the very defect it exists to catch | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate_abort_leak.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `skills/code-review/scripts/test_review_gate.sh` (300s hang bound in both stubs, path-scoped reaper, INT/TERM/HUP traps, exit-time leak assertion), the new `skills/code-review/scripts/test_review_gate_abort_leak.sh`, `Makefile` and `.github/workflows/ci.yml` (own job, never a shard member). Round plan: `specs/040-review-suite-process-leak/plan.md`. Observed failure: a `claude_review.sh --timeout 5` fixture wrapper found alive for 39 hours with ppid=1, ignoring SIGTERM. Proven cause: the controller starts each wrapper with `start_new_session=True`, so no group- or terminal-directed signal reaches it, and the `hang` fixture makes it TERM-immune — leaving the controller's own timeout path as the sole reaper, while the fixture's loop was unbounded. Reproduced by killing the wrapper's controller and then aborting the suite; three earlier abort shapes did NOT leak, which is what narrowed the cause to controller-death-before-timeout rather than suite-death. RED-baseline (applied, differential): the probe reds against the pre-fix suite with the wrapper alive at 30s after all three preconditions are shown green, and greens against the candidate with zero residue. F3 sensitivity, applied on the controller: mutating `signal_reviewer_process_group` to a no-op reds the exit-time assertion naming three leaked pids, and the trap still reaps them. Two drafting defects were caught the same way and are recorded in the plan: a pid ledger keyed on the wrapper's own pid missed the `( ... ) &` child that shares its command line; an `awk -v` needle carrying the harness path made the scanner match itself; an argv-path-only predicate was a proxy that a bare-argv descendant walks through, closed by recording each wrapper's process GROUP and auditing the union of group and path — shown by a mutation whose leaked `sleep 999` is named only by the group side; and a `2>/dev/null >&2` diagnostic sent itself to /dev/null. Lane placement is mechanical, not stylistic: as a shard-1 member the probe burned 408s without reaching its target case, so it runs in its own CI job |
|
|
290
|
+
| A fixture helper's lifetime must be pushed far past any plausible run WITHOUT becoming unbounded, and the suite — not the component under test — owns reaping it: the reaper the helper normally relies on is exactly what an abort removes, so an infinite loop is how a test helper outlives the machine rather than the run | `testing-strategy` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#far past, but never UNBOUNDED | updated | `testing-strategy/SKILL.md` is the owner key and is unchanged this round; the change is a merge into the existing sleep-replacement clause in `skills/testing-strategy/references/ci-fixtures-and-flake-control.md` (no new bullet), whose prior wording — "push that lifetime far past any plausible run so it stops being a deadline" — is what licensed the unbounded form. Observed failure: a signal-immune fixture wrapper alive for 39 hours with ppid=1 in this repository's own `code-review` suite, whose sibling fixture in the SAME file already carried a bounded lifetime; round record `specs/040-review-suite-process-leak/plan.md`. Downstream executable owner is `code-review`, whose row in this round carries the mechanism; the merged clause's content is load-bearing there by applied differential mutation: removing only the lifetime bound reds exactly the probe leg that watches an unreaped helper self-terminate while the trap-reaped leg stays green, and removing only the abort traps reverses which leg reds. The ownership half is likewise mutation-backed — a reaper that never signals a process group leaves a bare-argv group member alive and reds both legs. Zero-loss: every obligation of the prior clause (oracle-as-state, wall-clock as coarse backstop only, zombie-vs-liveness sampling, bounded reap grace) survives verbatim; the merge only adds the bounded-lifetime and suite-owns-reaping halves |
|
|
291
|
+
| An evidence taxonomy with no class for a claim that fires correctly and is false makes leaving the false claim cheaper than withdrawing it; the missing class is not a waiver but a different evidence kind — primary-source refutation plus obligation preservation — and it must carry machine-checked bars so it cannot become a general exemption | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; implementation in `skill-extraction-workflow/scripts/impact-chain-gate.rb`, six frozen synthetic cases in `skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh`, acceptance frozen before the gate was touched in `specs/043-evidence-class-source-refuted/frozen-acceptance.md`. RED baseline: the compliant-withdrawal case fails on the unmodified gate and passes after the applied change, while the four abuse cases stay failing and the existing RED-baseline path stays passing; the repo's four existing impact-chain tests stay green |
|
|
292
|
+
| A gate bar that approximates its invariant with a wording heuristic is bypassed by whatever the heuristic does not recognise, so a class whose whole meaning is "this change only deletes" must classify the diff structurally — path kind, change status, and added/deleted line counts — rather than inspect the text of added lines | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; implementation in `skill-extraction-workflow/scripts/impact-chain-gate.rb`, regression cases in `skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh`. RED baseline: reverting this classifier turns the suite red on the two cases that name the bypass — an owner-script edit with nothing deleted, and an `Always`-phrased rule the previous predicate deliberately excluded — both of which passed under the heuristic and fail under the structural check; the other eight cases and the repo's four existing impact-chain tests are unchanged either way. Observed failure: two independent review lanes each reached this bypass on the candidate before it landed |
|
|
293
|
+
| A diff classifier that reads change status but not file mode admits the one edit git reports as an unremarkable modification while contributing nothing to the line counts, so a pure-deletion class must compare both tree modes rather than trust the status letter | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; implementation in `skill-extraction-workflow/scripts/impact-chain-gate.rb`, regression in `skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh`. RED baseline: reverting the mode comparison turns the suite red on the case that pairs a `chmod` with a genuine deletion, which passed while the classifier read only `--name-status`; the other ten cases and the repo's four existing impact-chain tests are unchanged either way |
|
|
294
|
+
| A pointer check that only proves the pointer is well formed never shows that it accounts for what was removed, so a withdrawal class must require the obligation map to quote verbatim every substantive line the round deleted, and must resolve that pointer through git rather than the filesystem so it cannot escape the repository or match an arbitrary substring | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; implementation and regressions in `skill-extraction-workflow/scripts/impact-chain-gate.rb` and `skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh`. RED baseline: reverting the verbatim-accounting bar turns the suite red on the case whose map omits the deleted text; reverting the git-resolved pointer turns it red on the traversal case and on both existing-substring anchor cases; the other twelve cases and the repo's four existing impact-chain tests are unchanged either way |
|
|
295
|
+
| When every successive bar built to make a mechanical exemption safe is broken by independent review, the signal is that the exempting capability should not exist rather than that another bar is owed; an evidence class may still force an honest label and a real obligation map while leaving the unverifiable judgement to a named risk owner | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; implementation and regressions in `skill-extraction-workflow/scripts/impact-chain-gate.rb` and `skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh`. RED baseline: restoring the automatic floor lift turns the suite red on the compliant-withdrawal case, which must now stay blocked; the other fifteen cases and the repo's four existing impact-chain tests are unchanged either way. Seven review rounds and ten findings are recorded in `specs/043-evidence-class-source-refuted/frozen-acceptance.md` |
|
|
296
|
+
| A rule may fire perfectly and still be false: an entrypoint claim asserting that only the single most-salient prose rule applies while co-resident ones stay dormant is unsupported by either major vendor's published guidance, and the corollary it generated — that appending a clear rule does not add compliance — is contradicted outright, so it must be withdrawn rather than re-explained. Withdrawing an unsupported rationale is a `semantic-control` change, not a behaviour delta: every obligation the bullet carried is preserved verbatim and the entrypoint shrinks. **The `observed-failure: no` classification is a judgement worth challenging**: the withdrawn claim did mislead a reader into building a redirection on it, but this field asks whether the owning rule failed to *fire*, and it fired — it was simply wrong. The repo's evidence taxonomy has no class for 「一条正确触发但内容为假的规则」; recording it here rather than mislabelling it as a RED-baseline | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/SKILL.md#Descriptive, not permissive | updated | `skill-extraction-workflow/SKILL.md` 的机制 bullet(1128 → 825 字节,入口净 −304);出处、证据分级与被撤回的原文保留在 `skill-extraction-workflow/references/external-practice-controls.md#instruction-following-mechanisms`;一手源为 Anthropic 的 context-engineering 文与 OpenAI 的 GPT-4.1 / GPT-5.1 prompting guide;本轮对该撤回做过的三次行为测量全部作废并记账于 `specs/042-skill-corpus-optimization/batch-1-result.md` 与 `evidence/AGENTS.md` |
|
|
297
|
+
| A suite that adds a sibling test without registering it in a lane ships a false green: the file exists, reads as covered, and never runs, so the registration self-audit must be treated as a merge-blocking check rather than a lint nicety — and a clone-based multi-case suite belongs in the heavy lane, not the pre-commit one | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the lane registry is in `skill-extraction-workflow/scripts/test_check_ccl_regressions.sh` and the audit in `skill-extraction-workflow/scripts/test_regression_runner_registration.sh`. RED baseline: CI demonstrated it — `regression-fast` and `regression-heavy` both failed with `test_*.sh not registered in fast_tests/heavy_tests` naming `test_impact_chain_source_refuted.sh`, which the previous round had added and never registered; both lanes pass once the suite is registered in the heavy lane, and reverting the registration turns them red again |
|
|
@@ -56,6 +56,15 @@ Use this before reading deeply or editing any skill. The charter is the guardrai
|
|
|
56
56
|
| Evidence plan | Which source categories must be inspected, routed, discarded, or marked unavailable? For a task/session retrospective over a session that produced artifacts, the FIRST source class listed MUST be those produced artifacts (deliverables, reports, scripts, datasets — the a0 enumeration, owed at charter time, not only before an exhaustion claim); a session with genuinely no produced artifacts records an explicit `produced artifacts: not-applicable` entry carrying the reason, the minimum checked surfaces (deliverable directories, script/output locations, dataset paths), and a resolvable inventory-check locator — the command or listing that establishes absence — instead of the class. Either way, correction turns and the agent's own summaries are friction-biased digests — they record only what rubbed, so what went RIGHT is structurally invisible in them — and cannot substitute for the artifact class or excuse skipping the check. |
|
|
57
57
|
| Completion standard | What pressure scenario, independent review, command, install check, or source-map evidence proves done? |
|
|
58
58
|
|
|
59
|
+
### A deferred registration is a hypothesis source class, not a finding
|
|
60
|
+
|
|
61
|
+
Fires when the round's job is to implement something a PREVIOUS round deferred against our own mechanism — a gate/validator/checker hardening item parked in a ledger row, a spec's follow-up section, a plan's unpaid-item list, or memory, because the round that found it was scoped elsewhere. This is the inward face of the `SKILL.md` rule that an inherited reading of an external convention is hypothesis-grade: here the inherited reading is of *our own code*, and the primary source is that mechanism's CURRENT behavior.
|
|
62
|
+
|
|
63
|
+
- Such a registration records two things and **both decay**: how the mechanism behaved when the gap was seen, and what remedy looked right against *that* code. The mechanism keeps moving between the deferral and the fix round — often through several unrelated rounds that hardened the same file — so by the time the round opens neither half is a finding. List the registration in the charter's Evidence plan as a hypothesis-grade source class, never as the statement of work.
|
|
64
|
+
- Before implementing, you must reproduce each registered form against the current baseline, and **every probe is paired with a control leg** that runs the probe's own shape with the form absent. Without the control a probe that fails for a fixture defect reads as the form reproducing, and the same defect elsewhere reads as the form being closed — both directions have been observed in one sitting.
|
|
65
|
+
- A form that no longer reproduces is **closed, not fixed**: record it closed, keep the probe as its regression guard, and do not write code for it. The registered remedy is re-derived from the current code and re-measured against the real corpus it would judge; it must never be applied as written.
|
|
66
|
+
- Never treat the count of registered forms as the count of live defects, or a registered remedy as a specification. Failure shape: three gate forms deferred across three batches; on the fix round one turned out to need a precondition the registration never mentioned (so its recorded trigger was far broader than the real one), one's registered remedy was unimplementable as written — measured against the shipped ledger it would have rejected a large minority of real, correctly-authored rows — and the control legs caught two probe verdicts that were red for a fixture defect rather than for the form.
|
|
67
|
+
|
|
59
68
|
Depth rules:
|
|
60
69
|
|
|
61
70
|
- Wording cleanup/no new source read: no new source-derived rule; only clarify trigger, routing, naming, or validation text.
|
|
@@ -111,8 +120,8 @@ Do not force exactly five questions, and do not accept a single straight chain.
|
|
|
111
120
|
- **Detection gap** — why did no validator, review, or closeout catch it after the act?
|
|
112
121
|
A straight line of causes with no branches means you stopped early.
|
|
113
122
|
2. **Separate active failure from latent conditions.** The visible act (the agent skipped a gate) is the sharp-end *active failure*; the blunt-end *latent conditions* were authored long before (the gate was advisory, the closeout didn't check it, the rule lived only in a dropped description). Deep RCA fixes the latent conditions — you change the conditions agents work under, not their fallibility. This does NOT replace the delivery-chain pass: fixing an upstream latent condition (a trigger/description) and skipping the downstream detection/validation/closeout control is exactly the miss delivery-chain RCA exists to catch — keep both layers (defence-in-depth).
|
|
114
|
-
3. **Local rationality, not hindsight.** Knowing the outcome makes the right path look obvious in retrospect — hindsight is the primary obstacle to honest analysis. Do **not** write "the agent should have known / should have run X"; that describes a world that didn't exist and explains nothing. Ask: *given what the agent could see and was optimizing for, why did this action make sense at the time?* The answer points at a missing constraint or feedback, not at diligence. (This is the deep form of the "discipline gap is a non-cause" Core Rule.)
|
|
115
|
-
4. **Counterfactually test each candidate cause — to rank causal weight, not to prune defence-in-depth.** A why-chain never tests its links. For each contributing factor run the but-for test: *if this had been removed or changed, would the failure still have happened?* Use it to separate genuine coincidence from causes and to rank levers — never to delete redundant safeguards:
|
|
123
|
+
3. **Local rationality, not hindsight.** Knowing the outcome makes the right path look obvious in retrospect — hindsight is the primary obstacle to honest analysis. Do **not** write "the agent should have known / should have run X"; that describes a world that didn't exist and explains nothing. Ask: *given what the agent could see and was optimizing for, why did this action make sense at the time?* The answer points at a missing constraint or feedback, not at diligence. (This is the deep form of the "discipline gap is a non-cause" Core Rule.) The mirror trap is the untested **external** attribution (Argyris's defensive reasoning): a cause assigned to the host, tool, reviewer, environment, or user without evidence someone else could check is hypothesis-grade exactly like a one-trace factor — state it so others can test it, and test it before it drives a control row (the "it's a host bug" bar in `skill-listing-budget.md` is this rule's narrow instance). The self-sealing signature — answering a challenge to an attribution by restating it more forcefully instead of producing the test — marks the RCA itself as defensive: treat it as a repeated-correction signal on the analysis, never as confirmation of the attribution.
|
|
124
|
+
4. **Counterfactually test each candidate cause — to rank causal weight, not to prune defence-in-depth.** A why-chain never tests its links. For each contributing factor run the but-for test: *if this had been removed or changed, would the failure still have happened?* Use it to separate genuine coincidence from causes and to rank levers — never to delete redundant safeguards. Emergent-outcome boundary (Safety-II): in intractable interactions the but-for test may have no stable answer — the "causes" are reconstructed, not found. Invoking this boundary requires a RECORDED counterfactual attempt showing the instability under named interacting conditions (which factors were varied, why no stable necessary/sufficient verdict emerged); an unattempted but-for test keeps the ordinary ranking and unsupported-factor rules — "intractable" is never a label of convenience. Once so evidenced, do not force a verdict — mark the reconstructed factor `probabilistic`, and shift the control row to move 5's condition-constraint form (constrain the conditions that let the outcome emerge, rather than eliminating a reconstructed cause):
|
|
116
125
|
- Removing it prevents the outcome → **necessary** (a primary lever).
|
|
117
126
|
- It alone would produce the outcome → **sufficient**.
|
|
118
127
|
- Not individually necessary *only because another failed layer would also have allowed the outcome* → this is a **redundant safeguard / secondary control**: KEEP it (its failure is real defence-in-depth erosion — list it as a secondary control exactly as the delivery-chain and incident-postmortem methods require), do not discard it as correlation.
|
|
@@ -120,6 +129,8 @@ Do not force exactly five questions, and do not accept a single straight chain.
|
|
|
120
129
|
A factor asserted from a single trace is a hypothesis, not a proven cause — mark it `probabilistic`/`unsupported` and keep it as a candidate until a second observation or a mechanism confirms or refutes it; do not hard-delete a plausible latent or secondary factor on one trace.
|
|
121
130
|
5. **Site the fix as a mechanical control on the failure CLASS — provably firing, practically bounded.** The fix is never "remember next time" (sharp-end diligence — the dodge the Core Rules already forbid). Restore or add a *control*: a constraint the workflow enforces mechanically **plus** the *feedback* that confirms it fired. It must pass the **firing-path proof** (name the trigger, validator, closeout row, hook, or merged clause that mechanically catches the case — AND the **reachable surface** where the next agent actually encounters it in time: bootstrap/description/closeout/hook/CI. A clause landed only in a deep reference that nothing routes the agent to read is NOT a firing control — that is the "rule exists but did not trigger" gate. "The final answer mentions it" / "the agent will consider it" is not feedback and is still a diligence patch in disguise). Place it so the whole failure class is caught, not just this instance (defence-in-depth). Where several branches each have a controllable point, prefer the **highest-leverage practical control** (one upstream control often closes several branches) over the first patchable point — but "practical" is bounded exactly as the incident method's *earliest practical prevention point*: a named artifact, an owner who can change it today, and an observable check. Preferring leverage does NOT license deleting a useful narrow gate, removing a wording-only escape hatch, or inflating one miss into an over-broad global hook; keep the failed secondary controls as defence-in-depth. When the same failure class recurs across rounds, question whether the risky capability should exist at all, not just patch it again.
|
|
122
131
|
|
|
132
|
+
**Second-harvest check (bounded, recognition-dependent).** When a contributing factor, control, or lesson in this RCA maps onto a theory the theory doc already recognizes (`docs/skills-theory-foundations.md` per-skill rows; its 〔认出或借来之后〕 section is this check's fetch point), spend one bounded read answering that section's two questions — what does the theory cover that we have not adopted, and what are its known failure boundaries — one source pass, never a literature review. The outcome is one of three: a candidate rule routed through the normal extraction ledger; a recorded `no-new-lesson` whose reason cites the exact existing rule(s) covering the source principle — a vague "already covered" is the dodge the firing-path-proof gate forbids; or an evidence-backed `discarded`/`not-applicable` for a genuinely uncovered principle that is irrelevant, unsafe, or outside the skill's scope, recording the principle, the rejection reason, and the owner/scope decision in the ledger. All three are complete answers, and manufacturing a lesson to justify the read is the same padding the RCA-depth rule above forbids. Honesty bound (same bar as the self-detect firing point): recognizing that a factor maps onto a recognized theory is author-dependent — this check binds what happens AFTER recognition, and no validator detects an unrecognized mapping; do not report it as a mechanical gate.
|
|
133
|
+
|
|
123
134
|
Good stopping points (per branch):
|
|
124
135
|
|
|
125
136
|
- A missing source row, coverage depth, branch check, Figma node/page register, document register, or evidence grade.
|
|
@@ -147,7 +158,7 @@ Source-specific prompts (where each branch's RCA should resolve):
|
|
|
147
158
|
| --- | --- | --- |
|
|
148
159
|
| Code | A pattern was missed, overgeneralized, or routed to the wrong skill | Required code inventory depth, branch/module coverage, runtime/test evidence, sibling skill owner, or validator/test gate |
|
|
149
160
|
| Figma/design | A visual/interaction/behavior/psychology rule was shallow or unsourced | Required page/node/state-family coverage, Figma-plus-code boundary, judgment layer delta, pressure scenario, or source-neutral design/app/test routing |
|
|
150
|
-
| Documents | A workflow rule was copied from wording instead of extracted from durable practice | Document set boundary, source authority, contradiction handling, stale-doc exclusion, owner artifact, or no-skill decision |
|
|
161
|
+
| Documents | A workflow rule was copied from wording instead of extracted from durable practice | Document set boundary, source authority, contradiction handling, stale-doc exclusion, owner artifact, or no-skill decision. For a formal process description (SOP, runbook, playbook, template) the formal text is work-as-imagined (Hollnagel): include at least one work-as-done trace — a real session, execution log, or produced artifact of the process actually running — to confirm, narrow, or contradict the imagined flow, and mark rules landed from the formal description alone as WAI-only rather than practice-derived |
|
|
151
162
|
| Task/session retrospective | A lesson was summarized but not durably landed | Correction RCA, lesson classification, shared-skill vs memory-only boundary, durable owner, validation command, or final-response limit |
|
|
152
163
|
|
|
153
164
|
**Method provenance** (the moves are standard root-cause / safety-science practice, checked against primary or authoritative-secondary sources — for audit, not required reading; some origin *dates* are approximate or contested and flagged inline, so do not treat every bundled attribution below as equally primary-verified):
|
|
@@ -155,6 +166,7 @@ Source-specific prompts (where each branch's RCA should resolve):
|
|
|
155
166
|
- Single-chain / single-root-cause is the documented failure of 5 Whys (named and operationalized within the Toyota Production System by Taiichi Ohno, 1950s; the earlier "Sakichi Toyoda, 1930s" origin is widely repeated but weakly sourced — do not state it as settled) — the standard 5-Whys critique (named e.g. by ex-Toyota MD Teruyuki Minoura): stops at symptoms, knowledge-bound, non-reproducible, isolates one cause; and John Allspaw, *The Infinite Hows* (2014, kitchensoap.com): ask *how/what conditions*, not *why*, because "why" drifts to "who". Widening across parallel categories = Ishikawa/fishbone (Kaoru Ishikawa, 1960s); combination of causes = Fault Tree Analysis (H.A. Watson / Bell Labs, ~1962); "every effect has ≥2 causes, a straight line means you stopped early" = Apollo RCA (Dean Gano).
|
|
156
167
|
- "Overt failure requires multiple faults; post-accident attribution to a single root cause is fundamentally wrong" = Richard Cook, *How Complex Systems Fail* (how.complexsystems.fail). "Root cause is an arbitrary stopping rule; accidents are inadequate enforcement of constraints across a control structure" = Nancy Leveson, STAMP/CAST (*Engineering a Safer World*, MIT, 2011) — the source of the constraint + feedback / control framing in moves 1 and 5.
|
|
157
168
|
- Active vs latent conditions and defence-in-depth = James Reason, "Human error: models and management", *BMJ* 2000 (Swiss-cheese model). Local rationality / hindsight & counterfactual trap = Sidney Dekker, *The Field Guide to Understanding Human Error*. "Human error is a symptom, fix systems not people" and "contributing factors, not one root cause" = Google SRE *Postmortem Culture* (Lunney & Lueder) and Etsy blameless-postmortem practice. Necessity / sufficiency counterfactual test = Judea Pearl, *Causality* (probability of necessity / sufficiency).
|
|
169
|
+
- Untested external attribution and the self-sealing loop in move 3 = Chris Argyris, "Teaching Smart People How to Learn", *HBR* May–June 1991 (full text re-verified for this rule: defensive reasoning keeps premises private and untested, attributions "never really tested … a closed loop, remarkably impervious to conflicting points of view", and criticism of the analysis met by repeating the claim more vehemently). Work-as-imagined vs work-as-done in the Documents row = Erik Hollnagel et al., *From Safety-I to Safety-II: A White Paper* (EUROCONTROL, 2013), re-verified for this rule.
|
|
158
170
|
|
|
159
171
|
## Task Retrospective Extraction
|
|
160
172
|
|
|
@@ -189,7 +201,7 @@ Minimum retrospective table:
|
|
|
189
201
|
|
|
190
202
|
### LARGE-Session Lesson Axes And The Delivery-State Axis
|
|
191
203
|
|
|
192
|
-
A retrospective over a LARGE multi-batch / multi-phase session owes per-axis coverage (the `SKILL.md` Core Rule names the four lesson-TYPE axes — content/craft, program/process, workflow/meta, sustain — and the LARGE trigger; the sustain-entry bar of mechanism + non-luck evidence + owner routing lives there too): each axis is covered or explicitly marked `no-new-lesson`. Sustain and content/craft both read primarily from the produced-artifact source class, not from correction turns — a retro whose only evidence is corrections structurally cannot fill them (the Evidence plan hard-data-first rule). The user **re-asking the same session retrospective after a "done" claim** (especially the SAME wording) is a **same-scope correction signal, not proof an axis was dropped**: first classify it — a plain rerun, a genuinely new scope to honor, or a missed-axis re-ask. Only when it is same-scope AND the prior retro lacks *explicit* per-axis coverage do you owe a full four-axis enumeration (content/craft, program/process, workflow/meta, sustain); if scope actually changed, honor the new scope instead. Either way, never manufacture a lesson to satisfy the re-ask — an evidence-backed `no-new-lesson` mark on an axis is a complete answer for that axis.
|
|
204
|
+
A retrospective over a LARGE multi-batch / multi-phase session owes per-axis coverage (the `SKILL.md` Core Rule names the four lesson-TYPE axes — content/craft, program/process, workflow/meta, sustain — and the LARGE trigger; the sustain-entry bar of mechanism + non-luck evidence + owner routing lives there too): each axis is covered or explicitly marked `no-new-lesson`. Sustain and content/craft both read primarily from the produced-artifact source class, not from correction turns — a retro whose only evidence is corrections structurally cannot fill them (the Evidence plan hard-data-first rule). The user **re-asking the same session retrospective after a "done" claim** (especially the SAME wording) is a **same-scope correction signal, not proof an axis was dropped**: first classify it — a plain rerun, a genuinely new scope to honor, or a missed-axis re-ask. Only when it is same-scope AND the prior retro lacks *explicit* per-axis coverage do you owe a full four-axis enumeration (content/craft, program/process, workflow/meta, sustain); if scope actually changed, honor the new scope instead. Either way, never manufacture a lesson to satisfy the re-ask — an evidence-backed `no-new-lesson` mark on an axis is a complete answer for that axis. When choosing WHICH events a bounded retro digs into, weigh frequency as well as severity (Safety-II): learning potential is NOT proportional to severity, the accumulated cost of frequent small-scale events may easily be larger than a rare severe incident's, and frequent small events are easier to understand and easier to manage — while deferring them to "some other time" is the documented failure: that time never comes. A recurring minor friction is therefore a first-class extraction candidate, not noise to defer (this weighs the two, it does not deprioritize investigating a severe event).
|
|
193
205
|
|
|
194
206
|
For the delivery-state axis (a long operational delivery session's non-lesson axis, additional to the lesson axes): the source register must include a delivery-state row family before "whole-session retro complete" is claimed — changed artifact set, branch/worktree state, remote/MR state, CI or local verification state, cancelled/retried pipeline state, unresolved risks, and the next concrete action. This is still provenance, not shared-skill content: use sanitized labels in the shared landing, but keep the real per-repo evidence in scratch/private archive. Closeout evidence must record either `delivery-state rows: <scratch/private/source-register locator>` or `artifact/status axis: not-applicable, reason=<...>`. Pure doc/skill-text retros with no deployable artifact or remote delivery state record the not-applicable row instead of manufacturing empty rows. If required rows are absent, the retro can be reported only as `interim`, even when the extracted lesson text is correct.
|
|
195
207
|
|