@ccoalm/ccl-skills 0.13.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (24) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +14 -41
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +11 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/correction-routing-map.md +22 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +7 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +26 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +2 -2
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +2 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +6 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +4 -2
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +8 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/incident-postmortem-extraction.md +8 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +1 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +34 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +11 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +11 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/entrypoint_form_census.py +169 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/reference-access-census.sh +157 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +454 -103
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +8 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_form_census.sh +174 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_reference_access_census.sh +209 -0
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +331 -0
  23. package/dist/assets/release.json +45 -20
  24. package/package.json +1 -1
@@ -558,3 +558,37 @@ Round 073-receipt-bundling rows (new table so the entry renders as a table row a
558
558
  | Fan-out width is counted in workers that each own one bounded slice — tightly related items may sit inside that one slice — and every brief must carry an effort budget, so a width rule can never produce multi-task workers whose ownership, deadline, and failed-return classification cannot be attributed to one bounded unit | `multi-agent-delegation` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/multi-agent-delegation/SKILL.md#each owning one bounded slice | `updated` | Owner key `multi-agent-delegation/SKILL.md`. Lane-5 review (codex) on candidate 56342b9: gate item 4's 'several tasks each' contradicted execution step 3's one bounded task per agent (P1); fixed with zero-loss trims so the entrypoint stays within the 5000-word gate. RED baseline is the reviewed wording in the lane-5 round-1 receipt. |
559
559
  | Revocation inside a tool-bearing loop is stated as one fencing invariant rather than accumulated ordering patches: the generation advance and the call's final generation check at the irreversible handler's commit boundary are serialized by the same lock or fence, a call rechecks immediately before crossing that boundary, and the lease and fencing-token mechanics already required for stale agents are reused rather than re-derived | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/agent-tool-dispatch.md#revocation is a fencing problem, not a prompt problem | `updated` | Owner key `llm-inference-integration/SKILL.md`. Four consecutive same-class findings (lanes 1, 3, 4, 5: append-only vs invalidation, eviction vs invalidation, in-flight revocation, admission ordering) showed the class was a concurrency protocol being specified one patch at a time; the lane-5 succession's 'started is ambiguous' finding (P1) is fixed by the invariant form and by routing the mechanics to the existing fencing-token rule, per the same-class-recurrence design rule. RED baseline is the reviewed wording in the lane-5 succession receipt. |
560
560
  | The code-then-execute pattern cuts the trifecta edge only when the privileged code, its allowed sinks, and its permitted data flows are generated and frozen before any untrusted content is read, untrusted input entering afterwards only as non-instruction typed data with the negative test restoring that ordering; and the serving-lever table states prefill/decode disaggregation's throughput effect as engine- and workload-dependent to be measured on the target engine, not as a categorical claim | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/retrieval-agent-safety.md#generated and frozen before any untrusted content is read | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-6 review (codex) on candidate 09a040c: the code-then-execute option did not require code and sinks to be frozen before exposure (P1) and the disaggregation caveat was categorical (P2); the packet-evidence finding is accepted_tradeoff as before; the missing lane-5 authorization record and the authorization-chain wording were corrected in the evidence directory. RED baseline is the reviewed wording in the lane-6 round-1 receipt. |
561
+ | The merge gate binds an integration branch that accumulated several already-bound rounds and is promoted as one pull request by walking HEAD's first-parent chain down to the first commit already on the target: each round merge is rebound by the same gate in a detached checkout of its second parent against its first parent with that checkout's own controller and validator, a merge whose second parent is already on the target is a sync merge that owes nothing, every step must equal the automatic merge of its parents, a non-merge commit on the chain is refused, and the chain is consulted only after the single-ledger and manifest paths and only for the default path set | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, its suite, references/dual-track-review-gate.md, references/extraction-quickstart.md, and a .github/workflows/ci.yml comment). Observed failure: a promotion of three stacked rounds could be bound neither by one ledger nor by a path partition, because two of the rounds appended to this register and no round ever froze the sum of both appends, so the release had to land as three separate pull requests each pointing at one round's merge commit. Every such round had already been bound at its own base when it merged; the gate simply discarded that evidence at promotion time. Chain cases were written first and observed failing on the previous gate (10 failing), then green; a mutation walk reds each load-bearing predicate exactly on its own cases; the real three-round promotion binds as three rounds. Merge-queue aggregation of unmerged pull requests remains unsolved and is now stated as such in the ci.yml comment. |
562
+ | A historical round on the first-parent chain is judged with the landing tree's controller and validator, never the round's own, and the detached checkout it is rebound in must be verifiably released (removal result checked, registry read back, no repository-wide prune) or the verdict is an error; the walk accepts a chain of exactly the bound's length and refuses one longer — superseding the previous row's clause that a round is rebound with its own checkout's tools | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, its suite, and references/dual-track-review-gate.md). The round's own dual-track lane found all three: the adversarial challenge (P1) showed a round could install a hollow validator in its own branch, forge a ledger the real validator rejects, and stay bound after a later round restored the real validator, because history was being judged with history's tools; the independent review (P1) showed the checkout release discarded the removal result and ran a repository-wide prune that also drops registrations the run did not create; both lanes found the off-by-one that refused a chain of exactly the bound's length (P2). Fixes held until after the challenge; three cases were written first and observed failing on the pre-fix gate, then green. RED baseline is the pre-fix gate accepting the forged round, refusing the 64-step chain, and reporting ok past a failed release. |
563
+ | A detached checkout whose creation fails after git registered or populated it is released through the same verified path as a successful one, so a failed rebind leaves no repository worktree state behind and the error names any release problem alongside the creation failure | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and its suite). The round's succession challenge left this open at the exhausted lane budget: `git worktree add --detach` can return nonzero after creating its registration, and the failure branch only deleted the directory. Closed in an authorized continuation lane: the case (a git shim that runs the real add and then fails) was written first and observed failing on the pre-fix gate, then green. RED baseline is the pre-fix gate leaving the registration behind. |
564
+ | Releasing a detached checkout distinguishes a registration that survives removal (an error) from an unregistered directory this run created before git registered anything (deleted by the run itself, reported only if it survives), so a creation that fails before registering errors cleanly without a release complaint | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and its suite). Found while pre-covering the previous row's fix before review: routing a failed creation through the verified release path made an add that never registered anything report a spurious release problem and leave its directory behind. The benign before-registration sibling case sits beside the registered-failure case as the precision row; the registered-failure case is the RED baseline for the class and this row's case pins that the benign shape neither complains nor leaks. |
565
+ | The worktree registry is read back NUL-terminated when a detached checkout is released, because line-oriented porcelain cannot carry every path the gate may compare against — a newline splits the line, and quoting of other unusual bytes depends on git version and core.quotePath — so a surviving registration under such a temporary directory would match neither spelling and read as gone | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and its suite). The continuation lane's adversarial challenge (P2) named a tab-bearing TMPDIR; on the git in use a tab is printed raw and did not reproduce, while a newline-bearing path did split the line, so the case uses a newline: an add that registers and returns nonzero plus a refused remove left the registration in place while the release reported clean. Fixed by reading `git worktree list --porcelain -z` and comparing exact NUL-delimited fields; the case was written first and observed failing on the pre-fix gate, then green. RED baseline is the pre-fix gate reporting no release problem with the registration surviving. |
566
+ | Retirement or relocation of an over-budget entrypoint or reference cites the host-local usage census (per-file session counts derived from the agent's own transcripts, counts only) instead of the author's opinion, and the census must report `unevaluated` rather than a zero table whenever it could not evaluate | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_reference_access_census.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Source: external primaries — per-bullet helpful/harmful usage counters in itemized-context evolution (ACE, arXiv 2510.04618 §3.1–3.2), root-vs-auxiliary placement by frequency (Trace2Skill, arXiv 2603.25158 §2.1), and the official skill-authoring guidance that a bundled file the agent never accesses is unnecessary or poorly signaled. Observed failure this round: the first census version printed `sessions_touching_package=0` on a 60-day window (the candidate list exceeded one exec batch and the batch error was swallowed) while a 14-day window on the same host counted touching sessions normally — an instrument that reads zero when it could not evaluate is the false-green shape ratchet invariant 2 forbids. Fixed by batching; the regression test's ARG_MAX leg (450 synthetic transcripts, 3 touching one reference) and its unevaluated-sentinel and privacy-contract legs are the RED-to-GREEN evidence. RED (measured on fa0a7de): `NO_HITS: never accessed\|ignored content\|access log\|usage count` across SKILL.md and references (ledger excluded). Supporting paths: scripts/reference-access-census.sh, scripts/test_check_ccl_regressions.sh (fast lane registration), references/attention-budget-ratchet.md §Retirement and relocation signal, references/extraction-quickstart.md tool inventory. Host census figures stay in the private charter. |
567
+ | Low-frequency entrypoint detail relocates verbatim into the reference the entrypoint already points at (UI/UX judgment obligations, correction-type routing for test/verifier/deferred-evidence corrections, read-modify-write example-code obligations), the entrypoint keeps a one-bullet summary carrying the load-bearing obligations and every check-ccl phrase pin, and the entrypoint's body must shrink, never grow, in a round that touches it while it is over budget | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#never close a judgment-delta row whose visual direction/tokens fields were not inspected | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failure: the entrypoint that enforces a 5,000-body-word cap on every other skill carried 16,748 body words at base (size gate `entrypoint_size_severe_debt`, 3.3× the cap; 90 top-level bullets, 12 over 300 words) — the placement rule it states was violated by its own file, and Codex consumers cannot read it in one pass. RED (measured on fa0a7de): `check-size-budget.sh` base_body_words=16748; GREEN at head: head_body_words=16426, `changed_entrypoint_word_delta … delta_body_words=-322`, `entrypoint_size_blocking_ok`. Zero-loss obligation map: six UI/UX bullets → references/uiux-judgment-extraction.md §Entrypoint obligations (byte-identical), DFE-CONT + tests-before-test-cases rules → references/correction-routing-map.md (byte-identical, new file linked from Reference Loading), five RMW obligations → references/validation-and-landing.md §Read-modify-write example code (byte-identical); the four `skill_extraction_test_case_first_gate` phrases stay inline in the summary bullet (the first draft dropped one and the gate went RED, then GREEN after restoring it); the five ledger `file:` anchors and eight script pins into SKILL.md were enumerated first and none sits in the moved text (`register_firing_path_resolution_ok`). Remaining debt after this batch: 16,426 body words; the census and cost row exist so the next batch is chosen from evidence. |
568
+ | Content placement sorts by firing frequency as well as by kind: a rule that fires only on a narrow source class or correction type is low-frequency detail even when it is non-negotiable and must live verbatim in the reference the entrypoint already points at, behind a one-bullet summary that keeps the load-bearing obligations inline; the census is advisory and never a gate | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#must live verbatim in the reference the entrypoint already points at | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Source: Trace2Skill §2.1 (root document stores broadly applicable procedural knowledge, auxiliary files provide lower-frequency details) and §2.4/App. B (joint consolidation over many traces beats order-dependent one-lesson-at-a-time editing — the regime this skill's per-round appends have been in); Anthropic skill-creator progressive disclosure and the best-practices navigation observations (overreliance → move into the entrypoint; ignored → remove); Goodhart for the never-a-gate clause (same anchor as the health roll-up). RED (measured on fa0a7de): `NO_HITS: lower-frequency\|low-frequency\|firing frequency` in SKILL.md and references (ledger excluded); the placement rule sorted by kind only, so every non-negotiable read as entrypoint material. Head carries exactly one anchored list line. |
569
+ | A skill-effect comparison defines its baseline arm by change type (no-skill for a new skill, a frozen pre-change snapshot for an existing one), starts both arms in the same turn, records tokens and duration per arm, and reads results per assertion (pass-both, fail-both, one-sided, high-variance) before trusting an aggregate pass rate | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/harness-patterns-and-eval.md#the with-skill and baseline arms must start in the same turn | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Source: Anthropic skill-creator SKILL.md (head 2026-04-20; Step 1 spawn with-skill and baseline runs in the same turn with the snapshot rule for existing skills, Step 3 capture `total_tokens`/`duration_ms`, Step 4 analyst pass) and agents/analyzer.md §Analyzing Benchmark Results step 2 (per-assertion patterns). RED (measured on fa0a7de): `NO_HITS: non-discriminating\|passes in both\|without-skill\|duration_ms\|total_tokens` in harness and validation references; §3.1 compared before/after arms only, had no cost column, and read only aggregate deltas. Supporting path: references/external-practice-controls.md gains the IFScale row (density degradation, primacy bias; evidence grade recorded), which is a table row and carries no anchor. Concept-adjacent coverage is not functional equivalence; that bar makes this row I→updated, not P. |
570
+ | Every non-wording closeout records a cost row (review and challenge rounds run, findings fixed, accepted, or deferred, wall-clock from charter to PR, and net body-word delta per touched entrypoint) and the next round must read it before choosing its batch shape | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/extraction-quickstart.md#must read this row before deciding its batch shape | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Source: Google SRE Postmortem Culture (regularly survey whether writing a postmortem entails too much toil and improve the process from the answers) and the skill-creator per-run cost capture. RED (measured on fa0a7de): `NO_HITS: toil\|wall-clock` in extraction-quickstart.md; the closeout template recorded final state and lessons only, so process cost (one recent round ran 18 review rounds across 7 lanes; another re-ran full verification 8 times) lived only in private notes and never fed the next round's batch decision. Head carries exactly one anchored list line. |
571
+ | The census CLI treats a flag without its value as a usage error (exit 2, stderr) rather than an unbound-variable abort, and its regression suite's large-window leg must exceed every common ARG_MAX so that the first version's single-exec expansion cannot pass it | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_reference_access_census.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Self-review under the terminal owner and the testing owner (controller-derived owners for the new CLI and its test) found two defects before external review: a flag without a value aborted under `set -u` with exit 1, and the ARG_MAX leg used 450 short paths (~54 KB), which no ARG_MAX rejects — a mutation never applied, so a hypothesis. Observed RED: with 11,000 transcripts of ~230-character paths (~2.5 MiB) the exact first-version mutation applied to a disposable copy failed the leg with `sessions_touching_package=0` (mutant rc=1); the restored suite is GREEN in 3.9 s. Supporting paths: scripts/reference-access-census.sh, scripts/test_reference_access_census.sh. |
572
+ | The census withholds its table on any input error (unreadable, vanished, or unexecutable transcript inputs: exit 2 with an `unevaluated` count-only message) instead of printing zeros or the ok token, and every sentinel and usage error is path-free — no log root, caller path, or host detail reaches stdout or stderr | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#withholds the table (exit 2 on input errors) | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Round-1 review (codex, candidate c2c4d377) findings 0 and 1 and the round-2 challenge finding 0: the empty-log sentinel interpolated the log-root path (an absolute private path) and the test asserted that leak; find/grep/xargs errors were collapsed by `\|\| true` into zero counts followed by the ok token. Fixed: path-free sentinels (test asserts the log root and any absolute path are absent), a stderr-collecting error log that withholds the table with exit 2 (new test leg: a chmod-000 transcript → exit 2, `1 input error(s)`, no table, no ok token, no path), and the ratchet bullet now states the contract. RED is the reviewed candidate; GREEN is `test_reference_access_census_ok` on the fixed tree. Correction to two earlier rows of this round: their `NO_HITS` evidence for the terms `passes in both` and `wall-clock` overstated the grep — the base has one unrelated hit each (`passes in both directions of the wrong answer` in rule-consolidation.md; `wall-clock deadline` in the harness MAST table); the mechanisms had no prior carrier, which is the consolidation question those rows answer; recorded here because rows are append-only. Review findings 2 (ledger over 500 lines) and 3 (two rows sharing a command locator) are `accepted_tradeoff` with evidence in the round's disposition files. Supporting paths: scripts/reference-access-census.sh, scripts/test_reference_access_census.sh, specs/111-extraction-skill-benchmark/evidence/primary-source-excerpts.md (finding 4, fixed). |
573
+ | The census treats an explicitly supplied log root that does not exist as an input error (an absent default root stays normal), never echoes a caller argument in a usage error, and guards every stat substitution so an input error always reaches the withholding path instead of aborting under set -e | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#must be treated as an input error | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Continuation lane (scope: the defects named in the preceding closeout). Observed failures: succession challenge (codex, candidate 282c6e14) P2 ×2 and the mis-packeted first succession attempt's P1 (explicit missing root skipped silently). Each new test leg went RED under its applied mutation on a disposable copy (explicit-root check removed → leg 8 red; argument echo restored → leg 7 red; stat guard removed → leg 9 red with the pipeline status) and GREEN on the fixed tree. Supporting paths: scripts/reference-access-census.sh, scripts/test_reference_access_census.sh. |
574
+ | From the second review round on, the packet's exclusion list must exclude every evidence JSON the round has already added — receipts, dispositions, closeout files — so that the reviewed candidate equals the candidate the merge-side binder computes, and bound evidence (base attestations, excerpt files) is committed before the round | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/extraction-quickstart.md#must exclude every evidence JSON the round has already added | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed twice in consecutive rounds: the previous round's private notes recorded three receipts binding three candidates, and this round's first succession attempt excluded only the two receipts and bound candidate 0b91f263 while the binder computed 282c6e14 (`review_ledger_binding.py --print-candidate` before and after; the mis-packeted receipt is kept as history in the evidence directory). RED is that receipt; GREEN is the retry with the eight-file exclusion set whose receipt binds 282c6e14 and the validator's `extraction_review_state_ok`. |
575
+ | A stable-success row records its mechanism as work-as-done (the adjustment that matched the actual conditions) before any rule-fired claim, and a sustained practice names the condition under which it is re-tried against an alternative | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/incident-postmortem-extraction.md#Record the success mechanism as work-as-done | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Source: external primaries read this round to check learning-from-success theory — Hollnagel & Leonhardt 2013 (Safety-II white paper, PDF read: performance adjustment as the reason things go right), Levitt & March 1988 (competency trap, superstitious learning), Ellis & Davidi 2005 (successes+failures review beats failures-only, PubMed abstract), TC 25-20 (AAR sustain/improve lists, PDF read). Functional-equivalent check at fa0a7de: the classification rule's mechanism / non-luck / reuse-conditions / disconfirming-observation fields and the LARGE-session sustain axis exist (P for Ellis & Davidi, AAR, superstitious learning); `NO_HITS: work-as-done\|competency trap\|re-examination trigger` — the compliance-vs-adjustment distinction and the re-exploration trigger had no carrier. Head carries exactly one anchored list line; entrypoint gains a 20-word pointer while its round delta stays negative. |
576
+ | A description trigger evaluation grades each bank utterance at least three times per description version, keeps negatives as near-misses that share vocabulary with the skill, and selects an auto-rewritten description on a held-out split, never on the training split | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/eval-routing.md#each bank utterance must be graded at least three times per description version | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Source: Anthropic skill-creator Description Optimization (20 realistic queries, near-miss negatives, 3 runs per query, 60/40 train/held-out, best by test score). Functional-equivalent check at fa0a7de: the bank already carries `expected_skill: none` sentinels and `must_not_route_to` decoys and the runner has `--replicas` (P for negatives and repetition capability); `NO_HITS: held-out\|train/test\|overfit` — repetition was optional and no held-out discipline existed. Landed because an identified gap is a fix item, not a record. |
577
+ | The census keeps a NUL-delimited, pathname-deduplicated transcript inventory so overlapping or repeated log roots and newline-bearing names count a session once, tolerates an unset HOME (no default roots, never an unbound-variable abort), never echoes an invalid `--skill` value, and adds a shim-forced transcript-read error leg that does not depend on file permissions; the success-review reference requires a sustain row only when stable-success evidence meets the classification entry bar | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#overlapping or repeated log roots must count a transcript once | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Continuation lane (chain r3, codex): review finding 2 and challenge finding 0 (overlap double-count, newline-split names), review finding 3 (`--skill` echoed), challenge finding 1 (HOME unset abort; root-dependent unreadable leg), challenge finding 2 (the success-review pairing sentence universalized a sustain row beyond the entry bar — reworded to require `no-new-lesson`/`unstable` otherwise). Applied mutations on disposable copies: dedupe removed → leg 10 RED; `--skill` echo restored → leg 11 RED; HOME guard removed → leg 13 RED; restored suite GREEN in 3.6 s. The `seq` portability finding is refuted on this host (`/usr/bin/seq` present on macOS) and recorded as such in its disposition. Review findings 0/1 (ledger over 500 lines; command locators) repeat the earlier accepted tradeoffs; challenge finding 3 (packet evidence) repeats the packet-evidence class. |
578
+ | The census maps grep's no-match status to success inside each scan batch and treats any other batch failure — with or without a stderr line — as an input error that withholds the table, and its ARG_MAX regression fixture builds its padding without `seq` and asserts the inventory exceeds the largest common ARG_MAX before the mutation-sensitive check | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#a scan batch that fails without writing stderr must still withhold the table | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Second continuation lane (scope: the defects named in the preceding closeout). Observed failures: continuation-lane succession (codex, candidate c949ad35) findings 1 and 2. Applied mutation on a disposable copy: restoring the plain `xargs grep` pipeline makes the new silent-exit-2 shim leg (15) fail with a table and the ok token; the fixture-size assertion is checked at run time (inventory > 2 MiB). Restored suite GREEN in 6.5 s. Supporting paths: scripts/reference-access-census.sh, scripts/test_reference_access_census.sh. |
579
+ | The census batch wrapper survives an inherited errexit (a caller-exported SHELLOPTS=errexit no longer turns an all-no-match batch into an input error), and the ARG_MAX regression leg proves its fixture exceeds the running host's exec limit by executing the single-exec shape and requiring E2BIG instead of trusting a fixed byte figure | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#must not misreport an all-no-match batch under an inherited errexit | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Third continuation lane (scope: the defects named in the preceding closeout). Observed failures: second continuation lane's challenge (codex, candidate ef0010e3) P2 findings 1 and 2. Applied mutations on disposable copies: the errexit-unsafe wrapper (`grep; rc=$?`) fails leg 16 under `env SHELLOPTS=errexit`; restoring the single-exec `$(cat …)` scan makes the fixture leg fail with E2BIG on this host (rc 126). Restored suite GREEN in 7.8 s. Supporting paths: scripts/reference-access-census.sh, scripts/test_reference_access_census.sh. |
580
+ | The entrypoint's relocation summary bullets keep every load-bearing obligation of the moved text inline (mini-program owner split, execution-only disclosure, the pending/out-of-scope disposition for uninspected token fields, the DFE-CONT validation and rename-sync clauses); shared-tree evidence carries neutral scope facts only, never a conversation transcript or adjudication narrative; and a transcript listing that fails without stderr is an input error that withholds the table | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#do not claim design-judgment extraction | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Third continuation lane (chain r6, codex, candidate 420ab84e): review findings 0 (summary bullets dropped obligations), 3 and the challenge finding (three evidence notes and three ledger rows carried conversation-level content on a shared surface — removed and reworded to scope facts), 4 (silent find failure). Applied mutation on a disposable copy: ignoring find's status makes the new leg 17 fail with exit 0 and the no-transcript sentinel where an input error is required; restored suite GREEN in 5.4 s. The entrypoint grows by the retained obligations (+80 body words) while its round delta stays negative. |
581
+ | The ARG_MAX regression leg derives its fixture size from the running host's exec limit (getconf ARG_MAX) and skips its E2BIG probe with a printed note when the limit exceeds the fixture bound, never asserting a fixed byte figure; the entrypoint summary bullets carry the relocated observable-proxy checklist, breakpoint scope, the deferred-evidence trigger set, and the bare-mention qualifier inline; and shared-tree evidence and ledger rows state scope facts without host census figures or conversation-level narrative | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#a bare mention of runtime or external access is not enough | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failures: the CI fast lane on the Linux runner failed the previous E2BIG probe (2.5 MB inventory below that host's exec limit while it exceeds macOS's 1 MiB) — the host-derived-threshold class the prior succession named; the prior succession's summary-bullet and shared-surface findings. Applied mutation on a disposable copy: restoring the single-exec scan fails the leg on this host (ARG_MAX 1 MiB, E2BIG); restored suite GREEN in 5.6 s. Rows citing host figures or narrative were corrected at their origin commits (the branch is unmerged; rows are never edited in place). Supporting paths: scripts/test_reference_access_census.sh, specs/111-extraction-skill-benchmark/evidence/continuation-lanes.md, specs/111-extraction-skill-benchmark/evidence/primary-source-excerpts.md. |
582
+ | The census last-touched column is the newest touching transcript's date and the counting leg asserts that exact date (and the older date on the other file) from deterministic mtimes, so selecting the oldest timestamp or emitting an unparsed value fails the suite | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#the last-touched column must be the newest touching transcript's date | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Fourth continuation lane review (codex) P2: successful-row assertions accepted any non-empty date. Fixed with `touch -t` mtimes on two transcripts and exact `alpha.md \| 2 \| <newest> \| 66%` / `beta.md \| 1 \| <older> \| 33%` assertions. Applied mutation on a disposable copy: selecting the oldest date (`sort \| head -1`) fails the leg; restored suite GREEN in 5.9 s. Supporting paths: scripts/test_reference_access_census.sh. |
583
+ | A transcript older than the census window moves no count, share, or last-touched date, and the suite proves it with a stale fixture: removing the mtime filter turns the denominator assertion red | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_reference_access_census.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Registered gap from the previous round's terminal succession (no test leg for out-of-window transcripts), reproduced against the current script with a control leg: with the filter present the stale fixture leaves every existing assertion green; with the filter removed the suite fails at the denominator assertion; restored tree green. Supporting evidence: `skill-extraction-workflow/scripts/test_reference_access_census.sh`. |
584
+ | Relocating entrypoint detail to a reference keeps every obligation under exactly one carrier: the row set is derived by the governing-chain tool, each derived obligation has one grep carrier under a named chain, every ledger anchor and script-pinned phrase survives verbatim, and no over-cap reference grows | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/SKILL.md#must cover three axes, not only the executor's canonical phrase | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Zero-loss obligation map: `specs/113-extraction-entry-slim/obligation-preservation.md` (row set from `governing-chain-diff.py`; every survivor phrase counted once across the package with this ledger excluded). Supporting evidence: `skill-extraction-workflow/references/coverage-exhaustion-traps.md`, `skill-extraction-workflow/references/description-authoring.md`, `skill-extraction-workflow/references/external-practice-controls.md`. Paired control: the pinned-phrase, sync-pointer, contract-anchor, register firing-path resolution, and size gates report the same ok status on the base tree and on the landing tree, with the entrypoint body-word delta negative. |
585
+ | A usage-census mention share is not an open count: before a census figure justifies relocating, splitting, promoting, or retiring a file, the read shape in the same window is counted and the whole-read count carries the read-side cost argument | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#A mention count is not an open count | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Baseline (recorded incident): the previous round's closeout registered a ledger split as follow-up work on the ledger's mention share alone. With the rule applied: the same window's read-shape count showed whole-file loads in a small minority of the mentioning transcripts and bounded reads or gate echo in the rest, so the split was withdrawn as a read-side measure; the counts stay in the private charter. Supporting evidence: `skill-extraction-workflow/references/attention-budget-ratchet.md`. |
586
+ | Shared-tree guidance states a usage-census conclusion qualitatively; every measured ratio or count from a host census stays in the private charter, and a clause carried by a reference has exactly one carrier in that file | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#the read shape must be counted in the same window | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failure: the first review round and the same-candidate challenge both found a measured ratio in the new read-shape bullet; the fix batch replaced it with the qualitative conclusion and reworded one cross-reference lead-in in the dual-track reference so the clean-only-oracle clause has a single carrier (obligation table row 74). Supporting evidence: `skill-extraction-workflow/references/attention-budget-ratchet.md`, `skill-extraction-workflow/references/dual-track-review-gate.md`, `specs/113-extraction-entry-slim/obligation-preservation.md`. |
587
+ | A suite case that deliberately leaves shared fixture state mutated must restore it inside the case and assert the restoration: a registry prune only drops registrations whose directory is already gone, so deleting that directory after the prune leaves a stale registration every following case inherits | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Registered gap carried open in a prior round's frozen closeout, treated as hypothesis and reproduced against the current suite with a control leg. RED baseline: with the case's own cleanup ordering unchanged, the added restoration assertion is the single failing case in the suite and every other case stays green, so the mutated registry was invisible to the existing assertions -- including the passing-state case that runs immediately after it. GREEN: deleting the directory before the prune turns the same assertion green with no other case changed. Sibling-form evidence: the failed-add case earlier in the same file already captures the pre-case registry and asserts equality, so the corrected case matches a form the suite already contains. Supporting evidence: `skill-extraction-workflow/scripts/test_review_ledger_binding.sh`. |
588
+ | The other finding class carried open in the same frozen closeout is recorded closed rather than fixed: a failed detached-checkout creation is released through the verified path and a suite case already asserts the registry is restored | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `unchanged` | Owner key `skill-extraction-workflow/SKILL.md`. The registration's predicate (a failed add only removes the directory and never unregisters or verifies) does not reproduce against the current baseline, so no code lands for it. Oracle sensitivity proven rather than assumed: replacing the failure branch's verified release with the predicate's own described behavior turns the owning failed-add case red, together with the newline-path case that also reads that release path, and no unrelated case fails; the tree was restored and re-verified green. Supporting evidence: `skill-extraction-workflow/scripts/review_ledger_binding.py`, `skill-extraction-workflow/scripts/test_review_ledger_binding.sh`. |
589
+ | A claim about a whole file's guidance form is a measurement, so it carries a runnable instrument that states its own definitions and is committed with its baseline reading before the edit it will judge; a figure recorded without its method supports no trend claim in either direction | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/rule-consolidation.md#A whole-file form claim carries a recorded ruler | `updated` | Owner key `skill-extraction-workflow/SKILL.md`; lands in references/rule-consolidation.md (form-by-failure wording rules) plus scripts/entrypoint_form_census.py and scripts/test_entrypoint_form_census.sh, registered in the lane every run exercises. Observed failure (recorded incident): an earlier round recorded this entrypoint's prohibitive-token count with no recorded counting method; recounting the same file this round by an independently written method produced a different figure over a different scope, so no trend could be claimed from the pair and the earlier number could not be used as a baseline. RED baseline (applied mutation, differential): making a blank line close a rule turns the continuation case red and no other case, restored green. The first draft of that case anchored on the rule count, which does not move under the defect because a dropped continuation is still not a new top-level bullet -- the mutation ran green against it, and the case was re-anchored on the token total, which does move. The instrument encodes no threshold. Supporting evidence: `skill-extraction-workflow/scripts/entrypoint_form_census.py`, `skill-extraction-workflow/scripts/test_entrypoint_form_census.sh`. |
590
+ | The density of imperative-negative vocabulary in a rule set is not a measure of its guidance form: a required-slot rule, a conditional keyed to an observable predicate, and a positive recipe all carry that vocabulary inside them, so the count names a set to classify one rule at a time and never a set to rewrite | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/rule-consolidation.md#The density of imperative-negative vocabulary is not itself a measure of form | `updated` | Owner key `skill-extraction-workflow/SKILL.md`; merged into the form-by-failure wording rule that already owns the ruler requirement rather than appended as a second bullet. Observed failure (recorded incident): an earlier round inferred from this entrypoint's prohibitive-token-to-rationale ratio that the entrypoint did not apply its own form table, and registered a rewrite programme on that inference. Walked evidence: every rule the census flags was classified against the form-by-failure rows in `specs/114-entry-form-and-routing-baselines/form-classification.md`; the flagged set is overwhelmingly already in a form the table endorses, one row is mixed, none is in a form the table calls wrong, so the inference is withdrawn and no rewrite lands. Residual gap recorded rather than closed by writing: the discipline-slip rows need rationalization-vs-reality pairs quoted verbatim from baseline or pressure runs, a form this package prescribes in two places and realizes in none, with no capture channel operated -- the missing input is evidence, not authoring effort. Supporting evidence: `skill-extraction-workflow/references/rule-consolidation.md`, `specs/114-entry-form-and-routing-baselines/form-classification.md`. |
591
+ | A delivery that adds or removes a routing-bank row rebuilds the baseline in the same round or records that every prior baseline is now orphaned: the runner treats a differing bank fingerprint as a different ruler and suppresses the diff, so a bank edit does not age a baseline, it makes it permanently incomparable while it still reads as a baseline on disk | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/eval-routing.md#改 bank 的那一轮必须同轮重建基线 | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failure (recorded incident): the only full-bank report on disk was taken against a 136-row bank at one replica; the bank has since grown to 158 rows and the routing surface moved by six changed descriptions and one added skill, and no run in between could have detected it. RED baseline, produced by the runner itself rather than asserted: passing that report as `--baseline` to a fresh full-bank run emitted `baseline not compared — different ruler: bank content differs`, with both `newly_failed` and `newly_passed` empty because the comparison never ran. The new report carries `replicas: 3` and its `bank_sha256`, so a later round using the same bank and replica count is comparable to it; that is the property the rule exists to preserve. Second measured result from the same run, recorded because it is what the three-replica discipline buys: 13 of 14 failures carry `ownership_split` (the replicas disagreed and conservative consensus reports FAIL) and exactly one fails consistently across all three gradings, a distinction a single-replica report cannot express. Routing dispositions are deliberately not taken here -- editing a description while holding a routing measurement in the same round is the co-change shape the runner warns about. Supporting evidence: `eval/evidence/routing-baseline-replicas3-2026-09-03/`. |
592
+
593
+ Supersede note (round 114, ledger correction with no rule change): the row above beginning "Shared-tree guidance states a usage-census conclusion qualitatively" carries, in its evidence cell, an account of which review round and which challenge produced which finding. That is conversation-level process narrative on a shared surface, and the register is append-only, so the row stays byte-identical and is corrected here by pointer rather than edited. The obligation it records, restated at artifact level: `attention-budget-ratchet.md` carries the read-shape conclusion in qualitative form with no ratio or count; the measured census figures live only in the private charter; the clean-only-oracle clause has exactly one carrier in `dual-track-review-gate.md`, recorded as row 74 of `specs/113-extraction-entry-slim/obligation-preservation.md`. The superseded row's own behavioral-evidence declaration and firing-path anchor are unaffected. This is a note rather than a table row because it changes no rule and therefore has no owner-scoped anchor of its own to declare.
594
+ | An instrument whose stated definitions ARE its contract owes a case per definition, and the counting rule it documents must be the rule its figures were produced by: a regular-expression alternation matches leftmost and non-overlapping, so a listed phrase absorbs the words inside it and a comment promising independent counts describes a different ruler than the one that ran | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_entrypoint_form_census.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failure: two independent reviewer lenses on one candidate both read the census script's comment as promising that a phrase and the words inside it are counted separately, and no fixture exercised the overlap, so the documented rule and the reported totals were different rulers while the suite stayed green. The figures are unchanged by the correction because the behaviour was always leftmost non-overlapping and only the comment was wrong: the same run reports the same totals before and after. Two cases added, each with an applied mutation: a phrase-absorption case pins `must not` at one token, and an empty-Core-Rules case pins the false-empty guard -- disabling that guard turns exactly that case red with no other case failing, restored green. The classification artifact gains a stable per-rule identifier so each verdict maps to a rule without the generated JSON. Supporting evidence: `skill-extraction-workflow/scripts/entrypoint_form_census.py`, `skill-extraction-workflow/scripts/test_entrypoint_form_census.sh`, `specs/114-entry-form-and-routing-baselines/form-classification.md`. |
@@ -177,3 +177,14 @@ Use this when validating that an extraction skill or design/client skill actuall
177
177
  3. Record what changed in the extraction method before the test, then run the source evidence through the required chain: observation, judgment, rule, acceptance.
178
178
  4. Compare the result against current skills. If the source reveals a gap, update the smallest owning skill/reference. If it confirms existing guidance, record that no new rule was needed.
179
179
  5. Mark the coverage honestly: targeted pressure test, file-level refresh, node/artifact inventory, or full workflow extraction. Never upgrade a targeted pressure test into a full-source claim.
180
+
181
+ ## Entrypoint obligations
182
+
183
+ The six rules below are the entrypoint-level obligations for UI/UX, Figma, frontend, app, miniapp, and client sources. They were relocated verbatim from `SKILL.md`'s `What to extract` Core Rules group (low-frequency detail per the entrypoint's content-placement rule); the entrypoint keeps a one-bullet summary that points here, and Step 6's UI/UX validation rows resolve against this section. Wording changes here go through the same shared-skill gates as an entrypoint edit.
184
+
185
+ - Design/client extraction must cover the judgment layer, not only the engineering layer. For UI/UX, extract aesthetic logic, interaction logic, behavioral logic, and user psychology from source evidence before landing rules about layout, components, breakpoints, or tests.
186
+ - UI/UX judgment extraction must use observable proxies, not adjectives. Read state families, navigation/entry/return paths, disabled reasons, recovery controls, timing/feedback, accessibility, responsive/device variants, and code state machines before claiming behavioral or psychology rules. Use `references/uiux-judgment-extraction.md` for the required method.
187
+ - UI/UX lessons usually route to multiple owners. Before editing, map each candidate to design, web, app, miniapp, testing, product workflow, or this extraction workflow using `references/uiux-routing-map.md`; do not land only the design rule when implementation or scenario testing is required. For mini-program lessons, `testing-strategy` owns layer/scenario selection, while `miniapp-product-dev` owns host-platform implementation, developer-tool or real-device evidence, review/release mechanics, and miniapp runtime constraints.
188
+ - Judgment-layer extraction must name what changed. For UI/UX/client sources, record whether each judgment layer produced a new rule, confirmed an existing rule, narrowed an existing rule, or found no new evidence. If the pass only improves execution/validation, say so instead of implying new aesthetic, behavioral, psychology, or interaction knowledge.
189
+ - For UI/UX/client extraction, the judgment-dimension axis enumeration lives in `references/uiux-judgment-extraction.md`. When the adjacency-scan rule fires on a UI/UX source, walk that enumeration — do not re-derive the axis list from memory.
190
+ - A UI/UX judgment-delta row is not complete with labels such as `confirmed`, `narrowed`, or `no new evidence` alone. Each visual direction/tokens row must satisfy the field list in `references/uiux-judgment-extraction.md`; if those fields were not inspected, mark the row `pending` or `out of scope` and do not claim design-judgment extraction.
@@ -54,6 +54,17 @@ When the skill is installed outside the source repo, resolve the installed `skil
54
54
 
55
55
  If the validator reports `missing_required_command`, keep the failure visible and complete the static validation bullets manually.
56
56
 
57
+ ## Read-modify-write example code
58
+
59
+ Relocated verbatim from `SKILL.md`'s `Validation & the dual-track gate` Core Rules group (the entrypoint keeps a one-bullet summary that points here); it applies to every reference file or skill example that ships runnable code against external mutable state, and wording changes here go through the same shared-skill gates as an entrypoint edit.
60
+
61
+ - Reference example code that performs read-modify-write on external mutable state (Bitable records, database rows, file contents, API state) must:
62
+ - Read failure: raise explicitly; never return `{}`, `""`, `None`, or any empty-success value that silently drops the prior state.
63
+ - Write failure: raise or skip, never silently continue.
64
+ - Uniqueness invariant: when the example declares, assumes, or depends on one, detect violations such as duplicate unique keys and raise before propagating bad state.
65
+ - Lost-update control: use optimistic concurrency controls (ETag, version field, CAS, transaction, or compare-and-swap), append-only API semantics, or an explicitly declared single-writer precondition — RMW examples that silently assume no concurrent writers will produce lost-update bugs under normal conditions.
66
+ - Data-loss anti-pattern: log-and-continue after a read failure on an append-only field.
67
+
57
68
  ## Behavioral Validation
58
69
 
59
70
  - For new skills or major workflow changes, use `writing-skills` for RED-baseline/test-first methodology — **and the firing point is BEFORE drafting the body, not only before finalizing**. Eval-first authoring for a NEW skill (or a new hard-rule section): (1) write the evaluation scenarios first — **at least three** for a new skill (both the vendor's published authoring guide and the high-star practice pack converge on three-plus scenarios before body text; a single-rule edit may scope down to that rule's own scenario); (2) run them WITHOUT the skill and record the observed failures verbatim — and for a discipline-slip failure (the agent knows the rule and skips it under pressure), capture the agent's rationalizations word-for-word: each verbatim excuse is the raw material for one rationalization-vs-reality row and one red-flag line in the skill text (the discipline-slip form in `rule-consolidation.md`'s form-by-failure table); an invented hypothetical excuse does not qualify — counter only what a run actually said, and don't add rows for excuses no run produced; **a no-skill control that does not exhibit the failure is a stop signal — do not author guidance for a failure you cannot observe** (record the null finding instead; this is the pre-draft face of "Evidence must come before new rules"); (3) draft the **minimal** content that addresses the observed failures, then re-run the same scenarios WITH the skill; (4) when a later run, review round, or live miss surfaces a NEW rationalization for an existing discipline gate, add its explicit counter row to that gate's table and re-run the tempting scenario — counter tables accrete from observed excuses across rounds, never from imagination. The code-level RED-GREEN-REFACTOR method (write the failing case first, watch a fresh agent violate the rule WITHOUT the skill, then add the skill and watch it comply) is owned by `superpowers:writing-skills` + `superpowers:test-driven-development` — **if installed, route there; otherwise apply the RED-baseline rule inline** (manually record the without-change failure and the with-change compliance). This is the skill-authoring face of **eval-driven development** (for a behavior/routing change, run the scenario before you finalize; never special-case the scenario just to make it pass) — borrow the *principle*, not a claim of production-grade eval rigor.
@@ -0,0 +1,169 @@
1
+ #!/usr/bin/env python3
2
+ """Count the guidance FORM of an entrypoint's rules, so form claims carry a ruler.
3
+
4
+ `references/rule-consolidation.md`'s form-by-failure table picks the guidance form
5
+ from the baseline failure a rule answers. Judging whether an entrypoint follows
6
+ its own table needs a number, and the number is worthless unless the next round
7
+ can recompute it: an earlier round recorded a prohibitive-token count with no
8
+ recorded method, and a later round counting the same file by a different method
9
+ got a different figure, so no trend could be claimed in either direction. That is
10
+ the failure this script exists to prevent -- not the counting itself, which is
11
+ easy, but the counting being reproducible.
12
+
13
+ What is counted, stated here because the definition IS the instrument:
14
+
15
+ - A `rule` is one top-level `- ` bullet inside `## Core Rules`, together with
16
+ every continuation and sub-bullet line up to the next top-level bullet or the
17
+ next heading. Sub-bullets are not separate rules; they are part of the rule
18
+ whose form is being judged.
19
+ - A `prohibitive token` is a match of PROHIBITIVE_RE: the imperative-negative
20
+ vocabulary the form table calls the right form for a discipline slip and the
21
+ wrong form for every other baseline failure. Case is significant only where
22
+ the capitalised spelling is itself the emphasis (MUST / NEVER / ALWAYS).
23
+ - A `named baseline failure` is a match of FAILURE_SHAPE_RE: the rule states the
24
+ observed failure it answers, rather than only the prohibition. The form table
25
+ needs the baseline failure to pick a form, so a prohibition with no named
26
+ failure is a rule whose form was never derived from anything.
27
+
28
+ The reported diagnostic is `unanchored_prohibition_rules`: rules carrying at
29
+ least one prohibitive token and no named baseline failure. That is the set the
30
+ form table has something to say about; it is not a defect count, because a
31
+ discipline-slip rule is legitimately a prohibition -- it is the set a form pass
32
+ must classify one by one.
33
+
34
+ Exit status is 0 whenever the file parses; this is an instrument, not a gate.
35
+ Nothing here decides whether a form is right, and no threshold is encoded: a
36
+ threshold would make the ruler an argument for its own reading.
37
+ """
38
+
39
+ from __future__ import annotations
40
+
41
+ import argparse
42
+ import json
43
+ import re
44
+ import statistics
45
+ import sys
46
+ from pathlib import Path
47
+
48
+ # The imperative-negative vocabulary. Alternation is leftmost and matches do not
49
+ # overlap, so the LONGEST listed spelling wins at a position and the words inside
50
+ # it are not counted again: "must not" is one token, not "must not" plus "must".
51
+ # The phrase alternatives are therefore listed before the words they contain, and
52
+ # that ordering is load-bearing rather than cosmetic. The figure counts token
53
+ # occurrences, never distinct rules.
54
+ PROHIBITIVE_RE = re.compile(
55
+ r"\bMUST NOT\b|\bMUST\b|\bNEVER\b|\bALWAYS\b"
56
+ r"|\bmust not\b|\bmust\b|\bnever\b|\bcannot\b|\bcan not\b"
57
+ r"|\bdo not\b|\bdon't\b|\bdoes not\b|\bmay not\b|\bshall not\b"
58
+ r"|\bforbid(?:s|den)?\b|\bprohibit(?:s|ed)?\b|\bno[tn]-negotiable\b"
59
+ )
60
+
61
+ # The rule states the failure it answers. These are the phrasings this package
62
+ # already uses for that job; a rule that names its baseline failure some other
63
+ # way reads as unanchored here, which biases the diagnostic toward over-reporting
64
+ # rather than under-reporting -- the safe direction for a set meant to be walked.
65
+ FAILURE_SHAPE_RE = re.compile(
66
+ r"failure shape|failure-shape|failure mode|the failure it prevents"
67
+ r"|recurring shape|observed failure|the exact .{0,40}failure"
68
+ r"|recurrence signal|the tell that|the defect this prevents"
69
+ r"|failure it prevents|the dodge this prevents",
70
+ re.IGNORECASE,
71
+ )
72
+
73
+ BOLD_RE = re.compile(r"\*\*[^*]+\*\*")
74
+ CORE_RULES_HEADING = "## Core Rules"
75
+
76
+
77
+ def parse_rules(text: str) -> list[dict]:
78
+ """Return one record per top-level bullet inside `## Core Rules`."""
79
+ lines = text.splitlines()
80
+ try:
81
+ start = next(i for i, l in enumerate(lines) if l.strip() == CORE_RULES_HEADING)
82
+ except StopIteration:
83
+ raise SystemExit(f"entrypoint_form_census_error: no {CORE_RULES_HEADING!r} heading")
84
+ # The section ends at the next same-level heading.
85
+ end = len(lines)
86
+ for i in range(start + 1, len(lines)):
87
+ if lines[i].startswith("## "):
88
+ end = i
89
+ break
90
+
91
+ rules: list[dict] = []
92
+ current: dict | None = None
93
+ group = None
94
+ for line in lines[start + 1 : end]:
95
+ if line.startswith("### "):
96
+ group = line[4:].strip()
97
+ current = None
98
+ continue
99
+ if line.startswith("- "):
100
+ current = {"group": group, "lines": [line]}
101
+ rules.append(current)
102
+ continue
103
+ if current is not None:
104
+ # A blank line does not close a rule: sub-bullets and continuations
105
+ # are separated by blanks in this file, and treating a blank as a
106
+ # terminator would split rules and inflate the rule count.
107
+ current["lines"].append(line)
108
+ return rules
109
+
110
+
111
+ def measure(rule: dict) -> dict:
112
+ body = "\n".join(rule["lines"])
113
+ prohibitions = PROHIBITIVE_RE.findall(body)
114
+ named_failure = bool(FAILURE_SHAPE_RE.search(body))
115
+ return {
116
+ "group": rule["group"],
117
+ "head": rule["lines"][0][2:][:80],
118
+ "words": len(body.split()),
119
+ "lines": len(rule["lines"]),
120
+ "prohibitive_tokens": len(prohibitions),
121
+ "bold_spans": len(BOLD_RE.findall(body)),
122
+ "names_baseline_failure": named_failure,
123
+ "unanchored_prohibition": bool(prohibitions) and not named_failure,
124
+ }
125
+
126
+
127
+ def main() -> int:
128
+ ap = argparse.ArgumentParser(description=__doc__)
129
+ ap.add_argument(
130
+ "path",
131
+ nargs="?",
132
+ default=str(Path(__file__).resolve().parent.parent / "SKILL.md"),
133
+ help="entrypoint to measure (default: this package's own SKILL.md)",
134
+ )
135
+ ap.add_argument("--json", dest="json_path", help="write the full per-rule table here")
136
+ args = ap.parse_args()
137
+
138
+ text = Path(args.path).read_text(encoding="utf-8")
139
+ records = [measure(r) for r in parse_rules(text)]
140
+ if not records:
141
+ raise SystemExit("entrypoint_form_census_error: no rules parsed")
142
+
143
+ words = [r["words"] for r in records]
144
+ unanchored = [r for r in records if r["unanchored_prohibition"]]
145
+ summary = {
146
+ "path": args.path,
147
+ "rules": len(records),
148
+ "rule_words_median": int(statistics.median(words)),
149
+ "rule_words_max": max(words),
150
+ "rules_over_300_words": sum(1 for w in words if w > 300),
151
+ "prohibitive_tokens": sum(r["prohibitive_tokens"] for r in records),
152
+ "bold_spans": sum(r["bold_spans"] for r in records),
153
+ "rules_naming_baseline_failure": sum(1 for r in records if r["names_baseline_failure"]),
154
+ "unanchored_prohibition_rules": len(unanchored),
155
+ }
156
+ for key, value in summary.items():
157
+ print(f"{key}={value}")
158
+ print("entrypoint_form_census_ok")
159
+
160
+ if args.json_path:
161
+ Path(args.json_path).write_text(
162
+ json.dumps({"summary": summary, "rules": records}, ensure_ascii=False, indent=2) + "\n",
163
+ encoding="utf-8",
164
+ )
165
+ return 0
166
+
167
+
168
+ if __name__ == "__main__":
169
+ sys.exit(main())
@@ -0,0 +1,157 @@
1
+ #!/usr/bin/env bash
2
+ # Reference-access census: which of a skill package's files were actually
3
+ # touched by real agent sessions on THIS host, over a lookback window.
4
+ #
5
+ # Why this exists: a rule set must not grow monotonically, but "retire what is
6
+ # not pulling its weight" needs a signal that is not the author's opinion. The
7
+ # closest observable proxy on a local host is the per-file access record in the
8
+ # agent's own session transcripts (Claude Code `~/.claude/projects/**/*.jsonl`,
9
+ # Codex `~/.codex/sessions/**/*.jsonl`): a reference that no session mentioned
10
+ # in N days is a relocation/retirement candidate; one mentioned in most sessions
11
+ # that loaded the skill is a candidate for promotion into the entrypoint. This
12
+ # is the local instantiation of the usage counters that context-evolution
13
+ # methods keep per bullet (helpful/harmful counts) and of the "ignored content
14
+ # is unnecessary or poorly signaled" observation in the official skill-authoring
15
+ # guidance — advisory only, never a gate (Goodhart: a count that becomes a
16
+ # target gets gamed by mentioning files).
17
+ #
18
+ # Privacy contract: the transcripts are private per-host data. This script
19
+ # prints ONLY repo-relative skill file paths, per-file session counts, and
20
+ # last-touched dates. It never prints transcript text, prompts, absolute paths
21
+ # outside the skill tree, or session ids. Its output is safe to paste into a
22
+ # private charter; it is still not shared-tree content by itself.
23
+ #
24
+ # Counting unit: a SESSION (one transcript file) counts once per skill file it
25
+ # mentions, whatever the tool (Read, sed, grep, Skill load). "Mentioned" is a
26
+ # superset of "read to depth" — treat a count as an upper bound on real use.
27
+ #
28
+ # Usage:
29
+ # reference-access-census.sh [--skill <name>] [--days <n>] [--repo-root <dir>]
30
+ # [--logs <dir>[,<dir>...]]
31
+ # Defaults: skill=skill-extraction-workflow, days=60, repo-root=cwd-derived,
32
+ # logs=$HOME/.claude/projects,$HOME/.codex/sessions
33
+ # Exit 0 on a completed census or an honest unevaluated result (no transcripts);
34
+ # exit 2 on usage errors and on input errors (unreadable/vanished/unexecutable
35
+ # inputs) — counts are withheld rather than printed as zeros.
36
+ set -euo pipefail
37
+
38
+ skill="skill-extraction-workflow"
39
+ days=60
40
+ repo_root=""
41
+ logs="${HOME:+${HOME}/.claude/projects,${HOME}/.codex/sessions}"
42
+ logs_explicit=0
43
+ while [ $# -gt 0 ]; do
44
+ case "$1" in
45
+ --skill|--days|--repo-root|--logs)
46
+ if [ $# -lt 2 ]; then echo "reference_access_census_usage_error: $1 needs a value" >&2; exit 2; fi ;;
47
+ esac
48
+ case "$1" in
49
+ --skill) skill="$2"; shift 2 ;;
50
+ --days) days="$2"; shift 2 ;;
51
+ --repo-root) repo_root="$2"; shift 2 ;;
52
+ --logs) logs="$2"; logs_explicit=1; shift 2 ;;
53
+ -h|--help) sed -n '2,32p' "$0"; exit 0 ;;
54
+ *) echo "reference_access_census_usage_error: unknown argument (see --help)" >&2; exit 2 ;;
55
+ esac
56
+ done
57
+ if [ -z "$repo_root" ]; then
58
+ repo_root="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
59
+ fi
60
+ skill_dir="$repo_root/skills/$skill"
61
+ if [ ! -d "$skill_dir" ]; then
62
+ echo "reference_access_census_usage_error: --skill names no tracked skill package under the repo root" >&2; exit 2
63
+ fi
64
+ case "$days" in ''|*[!0-9]*) echo "reference_access_census_usage_error: --days must be an integer" >&2; exit 2 ;; esac
65
+
66
+ # Inventory: every tracked markdown file in the package (entrypoint + references).
67
+ files=()
68
+ while IFS= read -r f; do files+=("$f"); done < <(cd "$repo_root" && git ls-files "skills/$skill/SKILL.md" "skills/$skill/references/*.md" 2>/dev/null | sort)
69
+ if [ "${#files[@]}" -eq 0 ]; then
70
+ echo "reference_access_census_usage_error: --skill names a package with no tracked SKILL.md or references" >&2; exit 2
71
+ fi
72
+
73
+ # Candidate transcripts: any .jsonl under the log roots modified within the window.
74
+ tmp="$(mktemp -d)"; trap 'rm -rf "$tmp"' EXIT
75
+ roots=(); if [ -n "$logs" ]; then IFS=',' read -r -a roots <<<"$logs"; fi
76
+ : > "$tmp/candidates0"; : > "$tmp/errors"
77
+ # Input errors are never zeros: every scan phase collects stderr, and a non-empty
78
+ # error log withholds the table (exit 2) instead of printing counts that would
79
+ # read as evidence. grep's no-match status writes nothing to stderr, so it never
80
+ # trips this; unreadable, vanished, or unexecutable inputs do. The error log is
81
+ # summarized as a count only — file names would be private paths.
82
+ withhold_if_errors() {
83
+ if [ -s "$tmp/errors" ]; then
84
+ n="$(wc -l < "$tmp/errors" | tr -d ' ')"
85
+ echo "reference_access_census_unevaluated: $n input error(s) while scanning transcripts (unreadable, vanished, or unexecutable inputs); counts withheld"
86
+ exit 2
87
+ fi
88
+ }
89
+ # A default root that is absent is normal (the host may run only one agent); an
90
+ # explicitly supplied root that is absent is an input error — the caller asked
91
+ # for a scan that cannot happen, and skipping it would read as a real count.
92
+ for r in "${roots[@]}"; do
93
+ [ -n "$r" ] || continue
94
+ if [ ! -d "$r" ]; then
95
+ if [ "$logs_explicit" -eq 1 ]; then echo "missing explicit log root" >> "$tmp/errors"; fi
96
+ continue
97
+ fi
98
+ find "$r" -type f -name '*.jsonl' -mtime "-$days" -print0 >> "$tmp/candidates0" 2>> "$tmp/errors" || echo "transcript listing failed under one log root" >> "$tmp/errors"
99
+ done
100
+ withhold_if_errors
101
+ # The inventory is NUL-delimited (a pathname may contain a newline) and
102
+ # deduplicated by pathname: overlapping roots (a root and its own subtree, or
103
+ # the same root twice) list a transcript more than once, and a session must
104
+ # count once. Two names for one file (symlinked roots) are not canonicalized.
105
+ sort -zu "$tmp/candidates0" -o "$tmp/candidates0"
106
+ total_sessions="$(tr -cd '\0' < "$tmp/candidates0" | wc -c | tr -d ' ')"
107
+ if [ "$total_sessions" -eq 0 ]; then
108
+ echo "reference_access_census_unevaluated: no transcripts within ${days}d under ${#roots[@]} log root(s); counts withheld"
109
+ exit 0
110
+ fi
111
+
112
+ # Sessions that mention the package at all (fixed-string prefilter keeps this fast).
113
+ # xargs batches keep this under ARG_MAX; grep exit 1 (no match in a batch) is not an error.
114
+ # grep exit 1 (no match in a batch) is success; any other non-zero status —
115
+ # with or without a stderr line — is an input error, so each batch runs under a
116
+ # small wrapper that normalizes 1 to 0 and lets xargs report anything else.
117
+ # `|| rc=$?` keeps the batch alive when the caller exported SHELLOPTS=errexit:
118
+ # a bare `grep; rc=$?` would exit the child at status 1 before normalizing it.
119
+ grep_batch() { local rc=0; grep -lF --null "$@" || rc=$?; [ "$rc" -le 1 ] && return 0; return "$rc"; }
120
+ export -f grep_batch 2>/dev/null || true
121
+ if ! xargs -0 -n 200 bash -c 'grep_batch "$@"' _ "skills/$skill/" < "$tmp/candidates0" > "$tmp/touching0" 2>> "$tmp/errors"; then
122
+ echo "transcript scan failed in at least one batch" >> "$tmp/errors"
123
+ fi
124
+ withhold_if_errors
125
+ touching="$(tr -cd '\0' < "$tmp/touching0" | wc -c | tr -d ' ')"
126
+
127
+ # Per-file session count + last-touched date (mtime of the newest touching transcript).
128
+ : > "$tmp/rows"
129
+ for f in "${files[@]}"; do
130
+ rel="${f#skills/$skill/}"
131
+ if [ "$touching" -eq 0 ]; then n=0; last="-"; else
132
+ if ! xargs -0 -n 200 bash -c 'grep_batch "$@"' _ "$f" < "$tmp/touching0" > "$tmp/hits0" 2>> "$tmp/errors"; then
133
+ echo "transcript scan failed in at least one batch" >> "$tmp/errors"
134
+ fi
135
+ withhold_if_errors
136
+ n="$(tr -cd '\0' < "$tmp/hits0" | wc -c | tr -d ' ')"
137
+ if [ "$n" -gt 0 ]; then
138
+ # BSD stat first, GNU stat as the fallback; only the fallback's failure is an
139
+ # error. Every substitution is guarded so a failing pipeline cannot abort the
140
+ # script under set -e before withhold_if_errors runs.
141
+ last="$(xargs -0 stat -f '%Sm' -t '%Y-%m-%d' < "$tmp/hits0" 2>/dev/null | sort | tail -1)" || last=""
142
+ if [ -z "$last" ]; then
143
+ last="$(xargs -0 stat -c '%y' < "$tmp/hits0" 2>> "$tmp/errors" | cut -c1-10 | sort | tail -1)" || last=""
144
+ fi
145
+ if [ -z "$last" ]; then
146
+ [ -s "$tmp/errors" ] || echo "stat produced no timestamp for a touching transcript" >> "$tmp/errors"
147
+ withhold_if_errors
148
+ fi
149
+ else last="-"; fi
150
+ fi
151
+ if [ "$touching" -gt 0 ]; then share="$(( n * 100 / touching ))%"; else share="-"; fi
152
+ printf '%s | %s | %s | %s\n' "$rel" "$n" "$last" "$share" >> "$tmp/rows"
153
+ done
154
+ printf '%s\n' "reference_access_census: skill=$skill window=${days}d transcripts=$total_sessions sessions_touching_package=$touching"
155
+ printf '%s\n' "file | sessions | last_touched | share_of_touching"
156
+ sort -t'|' -k2,2nr "$tmp/rows"
157
+ echo "reference_access_census_ok"