@ccoalm/ccl-skills 0.15.5 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +21 -19
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +113 -123
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/wording-only-review.md +136 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +99 -11
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +297 -100
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +162 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +17 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_order.sh +30 -15
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +445 -13
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +14 -7
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +22 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +56 -18
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +41 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +68 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/closeout-reread.md +40 -0
- package/dist/assets/release.json +31 -21
- package/package.json +1 -1
|
@@ -25,6 +25,7 @@ The read side already defends against oversized files (chunked reads under ~200
|
|
|
25
25
|
- An existing over-limit reference is frozen per invariant 4: shrink or stay level; growth blocks. Additions to a frozen reference are funded by consolidating existing text in the same file.
|
|
26
26
|
- Append-only ledgers are structurally excluded: `references/source-register.md` grows by contract (append-only, supersede-by-pointer, rows never edited), so a line cap would block the ledger discipline itself; the gate skips it and prints a visibility token when it is over the figure. Residual risk, accepted under the same trusted-contributor model as the entrypoint gate: a prose file named `source-register.md` would dodge the cap — review owns that shape.
|
|
27
27
|
- A new reference over 100 lines must be structured with `##` sections so chunked reads and greps can navigate it; a heading-less long file draws an advisory token (never a block). A table-of-contents list is optional — section structure is the invariant, not a TOC block.
|
|
28
|
+
- **Funding an addition by trimming prose means editing text that may be pinned — resolve the pins before rewriting, not after.** The ratchet's per-file freeze makes every addition to a legacy surface a rewrite of something else in the same file, and load-bearing sentences are pinned in two places: declaratively in `../../skill-extraction-workflow/scripts/contract-anchors.tsv`, which the fast repo gate checks, and as `grep -Fq` assertions inside owner suites, whose break a full lane run reports half an hour later. Read BOTH for the file you are about to trim — `awk -F'\t' '$2 == "<path>"' skills/skill-extraction-workflow/scripts/contract-anchors.tsv` lists the registry rows pinning it, and `grep -rn 'grep -Fq' skills/*/scripts/*.sh` finds the suite assertions — because a funded trim that silently retired two pinned wait-contract obligations was reported by the slow lane only, long after the edit. A pinned sentence may be reworded only together with whatever pins it, in the same landing.
|
|
28
29
|
- Authoring anti-patterns (verified against the official skill-authoring checklist, see verdicts below): time-sensitive facts outside an explicit old-patterns section; inconsistent terminology for one concept; abstract examples where a concrete input/output pair fits; Windows-style paths; unexplained constants; scripts that defer error handling to the model instead of solving it.
|
|
29
30
|
|
|
30
31
|
## Retirement and relocation signal (usage census)
|
|
@@ -89,3 +89,25 @@ Firing point: **producing or first-publishing a reader-facing deliverable is an
|
|
|
89
89
|
但按该规则自己的定义,**针对外部源的缺口清单就是它所说的 findings 回合**:一旦产出,charter 就只能事后补写。
|
|
90
90
|
|
|
91
91
|
观测实例:一轮里先产出四条「外部有我们没有」的缺口,之后才 invoke 提炼工作流;改前的触发词表逐字检索该轮实际措辞得零命中。
|
|
92
|
+
|
|
93
|
+
## The review-chain case — the owning skill is not loaded where the situation arises
|
|
94
|
+
|
|
95
|
+
The same-class-recurrence rule is owned by this workflow, but the situation it governs —
|
|
96
|
+
findings coming back round after round — arises inside a `code-review` chain, where this
|
|
97
|
+
skill is typically never loaded. Naming the owner in prose therefore never made it fire.
|
|
98
|
+
The trigger sits at the transition instead:
|
|
99
|
+
|
|
100
|
+
- **The controller must raise it, not the reader:** when a round returns findings and the
|
|
101
|
+
history it carries already holds one, the gate adds `recurring_findings_design_check` to
|
|
102
|
+
that round's required self-review triggers and `decide_keep_delete_narrow_replace` to its
|
|
103
|
+
allowed actions, in the round's own envelope
|
|
104
|
+
(`../../code-review/references/staged-review-contract.md`). What discharges it is the
|
|
105
|
+
`keep` / `delete` / `narrow` / `replace` decision the owning rule defines, ratified by a
|
|
106
|
+
risk owner other than the one proposing it.
|
|
107
|
+
- **What it counts is bounded by what a receipt carries:** the chain the controller is
|
|
108
|
+
handed, plus the predecessor a succession names. Succession does not compose, so a third
|
|
109
|
+
chain opened fresh carries no history and the controller claims none; from there the
|
|
110
|
+
recurrence is the round's own record to keep.
|
|
111
|
+
- **It over-fires by design:** two findings rounds need not share a risk class, so the
|
|
112
|
+
question is sometimes inapplicable. Answering an inapplicable question costs a line; the
|
|
113
|
+
round a missed design question costs does not.
|
|
@@ -664,3 +664,28 @@ The pending classification above is superseded by the executed source comparison
|
|
|
664
664
|
| The shared implementation-gates fixture pins the continuation gate's non-stop clauses, so a later edit cannot silently delete them | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The sibling row for `product-rd-workflow` records the failure itself; this row records why the fix cannot regress silently. Four assertions were added to the fixture -- the reference's non-stop clause, its turn-end firing check, its outcome-contract line, and the entrypoint's own clause. The fourth was added after independent review observed that the outcome-contract line could be deleted with every assertion still green, which is the same false-green shape the pins exist to prevent. RED-baseline (applied, differential): deleting each protected sentence reds only its owning assertion, with every other assertion passing and the unmutated control clean, so a partial deletion is attributable rather than lost in an aggregate failure. The fixture was chosen over a new suite because it already owns cross-owner rule-retention pins; no new registration surface is introduced. |
|
|
665
665
|
| The landing binder names the ordering cause at the failure point: evidence a round adds that stays inside the candidate is listed when nothing binds | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/review_ledger_binding.py | updated | Owner key `skill-extraction-workflow/SKILL.md`. Third occurrence of one class. The rule that bound evidence is committed before the review rounds already exists verbatim in the quickstart and already carries a register row marked observed twice in consecutive rounds; this round hit it again because the round was driven from the delivery and review owners and never opened that quickstart. Two prior landings answered the recurrence with more prose, so this one changes the mechanism instead: when nothing binds, the binder enumerates the added evidence that is NOT excluded -- the complement of the receipt exclusion it already computes -- and states that only added JSON carrying a candidate_sha256 is excluded, so committing a base attestation or excerpt after the rounds moves the candidate out from under their receipts. RED-baseline (applied): on this round's own failing candidate the pre-change binder reported only that nothing bound it, naming neither the file nor the ordering; the changed binder lists `landing-base.txt` and the round's markdown dispositions and states the ordering. The five binding suites pass unchanged. The diagnosis now reaches an agent at the moment it fails rather than requiring it to know which document to open. |
|
|
666
666
|
| A harness whose RECORDS are the evidence — an evaluation or benchmark runner, a conformance suite feeding a comparison, an A/B or regression rig — can be corrupted by the data it produces in three ways that all read green: absence stored as a bare null cannot separate confirmed-absent from never-observed, planned units and retries sharing one counter let a retry move the denominator, and a later attempt overwrites an earlier failure. Its own record layer is a high-risk failure class of the same standing as the canonical list, and is built against these before the happy path | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#Absence carries a coded reason beside the value | updated | `testing-strategy/SKILL.md` is the owner key and is unchanged this round: the entrypoint is over its size budget and the growth gate blocks it, so the rule lands in `testing-strategy/references/ci-fixtures-and-flake-control.md` and is reached as a failure class from the canonical high-risk list in `testing-strategy/references/scenario-testing.md`, an enumeration the entrypoint already tells readers to walk. RED baseline: a paired walk over real artifacts — a held-out harness that contributed nothing to deriving the rules fails all three rows, each defect named by exactly one row while the other two do not mention it, while the control harness passes two and partially satisfies the first, so the check discriminates rather than accepting whatever is put to it. `observed-failure` is `no` deliberately: no malfunction of an existing repository rule was recorded this round, and the delta is measured against the held-out artifact rather than against a regression this repository observed; `result-class` is `failure` because that held-out artifact does exhibit all three defects the rule names. Sources read this round: the health-interchange data-absent-reason code system, a monitoring query language's absent-vector operators, and the controlled-trial reporting guidance for the flow diagram and per-group denominators. Known limit, stated in the landed text itself: the assembled rule has no located prior name, and measurement system analysis is the adjacent established field covering instrument accuracy and repeatability rather than record integrity. |
|
|
667
|
+
| The reviewer's packet and the landing candidate are two objects: `--diff-file` widens what the reviewer reads while `--base` keeps `candidate_sha256` the base-derived identity the merge-side binder recomputes, and the gate accepts the combination only when the packet BEGINS with that candidate, byte for byte | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/review_gate.py | `updated` | Owner key `code-review/SKILL.md`. Observed failure: the previous round's first review chain returned five findings of one class -- the code a containment claim depends on is not in the packet -- and answering them was impossible, because the skill tells the author that insufficient input is an input defect to be answered by widening the packet and rerunning the lane, while `freeze_packet` refused `--diff-file` together with `--base` and recorded `candidate_sha256` as the packet's own hash, so a widened rerun produced a receipt `review_ledger_binding.py` can never match. Following the contract produced evidence the merge side rejects. The controller now computes the candidate from the base independently of the packet, accepts the combination only when the packet begins with that candidate, and derives `candidate_paths` from the candidate so owner selection and the wording-only `changed_files` comparison cannot be widened by appended context; the in-chain checks that asserted the two hashes were equal now compare candidate to candidate, which is what lets a later round in one chain read more than an earlier one. Backward compatibility is byte-exact: with no `--diff-file` both hashes keep today's value, and `--diff-file` alone keeps today's meaning. The anchor is a prefix rather than a bare containment test because the round's adversarial challenge showed what containment alone buys: a packet PRECEDING the candidate with a sanitized decoy diff passes, and the reviewer then reads the decoy as the change and the real candidate as trailing context -- the repository's authoring rule already said context sits on top of the candidate, and until this round it was documented and unenforced. RED-baseline (applied, differential): collapsing the candidate identity back into the packet hash reds only the new acceptance case in `test_review_gate.sh` while its other 268 cases pass; removing the anchor check reds that same case alone; the unmutated control is green. Recorded because it cost a red rather than being reasoned out: the packet file must live outside the repository, because the base-derived candidate includes untracked files and a packet written into the worktree becomes part of the candidate it has to contain. |
|
|
668
|
+
| The binding gate's cross-side agreement is asserted rather than assumed: a widened packet's recorded candidate identity must equal the one the gate recomputes, and the base and paths for that comparison come from the gate's own scope resolution rather than being rebuilt by hand | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. `review_ledger_binding.py` has no behavior delta of its own -- it never passes `--diff-file`, so its candidate and its packet stay the same bytes -- and the honest record of that is a mutation which does NOT discriminate: making it return the packet hash instead of the candidate hash leaves all 85 binder cases green. So what lands here is the cross-side assertion, because the claim that motivated the change spans both sides and neither suite alone can hold it. The binder suite now freezes a widened packet through the controller and requires the identity recorded for it to equal the one this gate recomputes. It derives the base and the paths from `candidate_scope` instead of hand-building them, which is itself the finding: the base is a fork point and the paths carry the receipt exclusions the round has already added, so `--print-candidate` is the authority an author reads rather than a value an author reconstructs. RED-baseline (applied, differential): collapsing the controller's candidate identity into the packet hash reds this case alone while the other 85 pass, and the unmutated control is green. |
|
|
669
|
+
| A reviewer wrapper classifies a transport failure from the transport's own error channel, not from whichever stream is habitual: a CLI run under a structured-output flag reports supply and credential failures as events on stdout, so a classifier reading stderr alone reports a routine quota exhaustion as an unclassifiable client fault and stops the lane instead of cascading; and because only the transport's TOP-LEVEL error events count, model-authored text quoting the reviewed packet cannot steer that decision. Every transport failure additionally carries a bounded, redacted excerpt of what the transport said, because a failure whose captured streams are deleted leaves nothing that can contradict a wrong hypothesis | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/codex_review.sh | `updated` | Owner key `code-review/SKILL.md`. Observed failure: the codex lane returned `codex_run_failed` / `unknown_client_failure` three rounds running, and the recorded suspicion was packet size. Both halves were wrong. A 20KB packet reproduces the failure identically, and the preserved run directory shows the account's usage limit announced on the event stream while stderr carried only an unrelated models-manager message -- the quota regex matches the first file and not the second, and the wrapper greps only the second. Invoking the CLI directly against the user's own home returns the same message, so the condition is account-level rather than wrapper-induced. The cost is not cosmetic: `review_gate.py` admits a cascade only for a reason code in `CANDIDATE_LOCAL_CODES` carrying `cascade_eligible` true, `quota` is in that set and `unknown_client_failure` is not, so the misclassification produced `stop_reviewer_lane` and the recovery was an operator reordering the clients by hand. Why the wrong hypothesis survived three rounds is the second half of the rule: the EXIT trap removes the run directory with both captured streams, and the receipt held only an exit code, so rounds 122 and 123 left six receipts with no codex record between them and nothing in the repository could contradict the size theory. RED-baseline (applied): eight assertions written against the unchanged wrapper each fail on the row they name; four single-predicate mutations then turn exactly their predicted rows red and no others -- removing the top-level restriction reds the two rows proving packet-derived text cannot classify, removing redaction reds the two redaction rows, removing the bound reds the truncation row, and restoring the stderr-only grep reds the two event-stream rows. Two rows are green on the unchanged baseline by construction and it is the mutations, not the rows, that establish their meaning. Recorded because each cost a red rather than being reasoned out: a heredoc nested in a command substitution is scanned by Bash 3.2 for shell quoting, so an apostrophe in a comment inside it ends the parse of the whole script; and a fixture carrying a literal credential-shaped value is refused by this repository's own credential scanner, which is the scanner behaving correctly. |
|
|
670
|
+
| A redactor guarding an evidence excerpt keys on the SHAPE of an assignment, not on a list of credential-sounding key names: the value of every `key=value` and quoted `"key": "value"` pair is removed whatever the key is called, while the key and any prose that is not assignment-shaped survive so the excerpt still says what failed | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/codex_review.sh | `updated` | Owner key `code-review/SKILL.md`. Observed failure, twice in one round on the same predicate: the first review chain found that anchoring a word boundary before the keyword made `access_token=` and `client_secret=` unmatchable, because `_` is itself a word character, and the list was widened; the succeeding chain's challenge then found `session=`, `cookie=`, `auth=`, `code=` and `bearer=` still missing, plus quoted JSON forms. Widening a third time was the obvious move and is the one this row rejects: a list of credential-sounding names has no state in which it is finished, so the recurrence is evidence that the predicate is a proxy rather than the invariant. The list is deleted. What replaces it does not read the key at all. The cost is real and accepted: an informative `error=timeout` loses its value too, which is why the key is preserved rather than the whole pair, and why prose is left alone -- the message this round exists to classify contains no assignment and passes through whole, verified against the live condition rather than argued. RED-baseline (applied): a fixture carrying six unlisted key names and two quoted forms fails against the widened list and passes against the shape rule, and it is the only row that moves. Recorded because it cost a red rather than being reasoned out: this gate judges per commit, so a register row for an owner package must land in the same commit as the package change, not in a later one. |
|
|
671
|
+
| When a failure destroys its own evidence, the fix is to stop destroying it, not to copy an excerpt somewhere durable: the run directory holding the captured streams survives a failed run and the receipt names it, while the receipt itself carries only the transport's own error message. Arbitrary process output never reaches a committed artifact, so no filter over it has to be complete | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/codex_review.sh | `updated` | Owner key `code-review/SKILL.md`. This supersedes the excerpt-plus-redactor shape two rows above, and the reason is measured rather than argued: three review chains each found a different escape from the same filter -- a key name the list lacked (`access_token`), an assignment form the shape rule lacked (`session=`, quoted JSON), and URL userinfo, which is not an assignment at all. Replacing the name list with a shape rule was recorded as an invariant change and was not one: names and shapes are both enumerations of how a secret might look, so the recurrence continued. The predicate was never completable, because the input was arbitrary process output and the property wanted of it -- that nothing secret-shaped survives -- is not decidable. What ends it is removing the input: raw stderr has no path into the receipt. The original defect was also misread on the way in. `RUN_ROOT` is deleted by the EXIT trap, so the streams that named the cause were gone at the moment the cause became interesting; the minimal repair for that is to keep the directory on failure, which also leaves the whole streams for diagnosis rather than 600 redacted bytes. The directory is mode 0700 under TMPDIR and holds exactly what it held while the run was in flight, so nothing is exposed that was not already, and reclaiming it stays the platform's temp-directory lifetime. Stated limit, not claimed away: a CLI-authored error message can still echo a credential, so the redaction stays as defence in depth -- it is no longer the control the safety rests on. RED-baseline (applied): restoring the stderr fallback reds the row asserting stderr is not quoted; deleting the directory on failure reds the two preservation rows; dropping the userinfo strip or the quoted-value alternation reds the shape row; each mutation reds only its own rows. The redaction fixtures were moved onto the event stream in the same change, because on stderr they would have asserted nothing under the new design. |
|
|
672
|
+
| A test fixture that only LOOKS like a credential is still a credential to every scanner that guards a boundary, so a fixture exercising a redactor assembles its credential-shaped values at runtime rather than writing them into the repository | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_cli_review_wrappers.sh | `updated` | Owner key `code-review/SKILL.md`. Observed twice in one round against two different guards. A literal `sk-` token in the fixture was refused by this repository's own credential scanner in `validate-skill.sh`, turning `make test` red. A URL written with credentials in its userinfo component, in the same fixture, was later refused by the review gate's egress tripwire, which returned `egress_denied` and stopped the review lane before any reviewer ran -- and would have done so on every later review of this repository, because the fixture lives in the tracked tree that every packet carries. Both guards were behaving correctly; the fixture was the defect. Approving the egress by flag was available and rejected: it would spend a real control on synthetic data and teach the flag as routine. Assembling the value inside the fixture at runtime keeps the test at full strength -- the wrapper still sees the whole shape -- while the shape never exists in a tracked file. RED-baseline (applied): the suite is green with the assembled value and the redaction rows still red under their mutations, and the literal form is absent from the tree. |
|
|
673
|
+
| A test that reads a path out of a receipt expands it before touching the filesystem: a path recorded with the home directory elided is correct in the receipt and meaningless to a filesystem check, so the check false-REDs on exactly the hosts where the eliding fires | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_cli_review_wrappers.sh | `updated` | Owner key `code-review/SKILL.md`. Raised by the landing review and confirmed by hand rather than reasoned about: with TMPDIR under the home directory the wrapper records `transport_run_dir` as `~/.cache/.../codex-review.9YWG3E`, and the assertion's `[ -d ... ]` then tests a literal tilde path. The failure mode is a false RED, never a false GREEN, which is why it is a portability defect rather than a correctness one -- but a shared suite that reds on someone else's machine costs them the same time it would cost here. The suite's own TMPDIR is not under the home directory, so no row exercises the expansion; the tilde form was reproduced directly against the wrapper instead, and that is the evidence, not a suite row. |
|
|
674
|
+
| Eliding a home directory out of a persisted path compares PHYSICAL paths on both sides, not the literal environment variable: a home spelled with a trailing slash, or reached through a symlink, is the same directory, and a literal comparison leaves the username in the artifact on exactly the hosts that spell it differently | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/codex_review.sh | `updated` | Owner key `code-review/SKILL.md`. Raised by the landing challenge against the eliding this round had just added, and it is the same class the round spent three chains on in a different place: matching one spelling of a thing is not matching the thing. `cd \| pwd -P` on both sides normalizes trailing slashes and symlinks in one move rather than enumerating the ways a path can be written, and the redactor gets the same treatment through a set of home spellings applied longest-first. RED-baseline (applied): a row exporting a home with a trailing slash and a TMPDIR beneath it reds against the literal comparison and greens against the physical one, and reverting only the physical resolution reds that row alone. Recorded because it cost a red: the suite's `run_codex` helper pins TMPDIR, so a row needing a TMPDIR under a specific home has to invoke the wrapper directly instead of through the helper -- the helper silently won, and the row failed for a reason that had nothing to do with the fix. |
|
|
675
|
+
| A guard's fixture must be inert to every OTHER rule in the same pipeline, or the guard cannot fail when its own rule is deleted: a secret shaped like a neighbouring rule's input is removed by that neighbour, and the row stays green over the hole it was written to watch | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_cli_review_wrappers.sh | `updated` | Owner key `code-review/SKILL.md`. Found by the landing challenge, not by the mutation pass that was supposed to catch exactly this: the row guarding the URL query strip used `?token=abc123`, which the assignment redactor rewrites on its own, so deleting the query strip left the row green and a bare non-assignment query secret would have reached a committed receipt with no red row anywhere. The earlier mutation run did remove the query strip and did report the row red -- but that run predated the assignment redactor, so the coverage it proved expired when a later rule was added and nothing re-established it. That is the reusable part: a mutation result is evidence about the pipeline as it stood, and adding a rule can silently subsume a neighbour's fixture. RED-baseline (applied, after the fixture was changed to a non-assignment shape): deleting the query strip reds that row alone, and the unmutated control is green. |
|
|
676
|
+
| A rule that strips a delimited region consumes to the LAST delimiter the region can legally contain, not the first: URL userinfo may itself contain the separator, so a non-greedy strip leaves the tail of the secret behind | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/codex_review.sh | `updated` | Owner key `code-review/SKILL.md`. Fourth consecutive chain to find something in the same defence-in-depth redactor, and the count is the point rather than the individual defect: the class is bounded only because the redactor is no longer what the safety rests on -- raw process output has no path into a committed receipt, so what remains is a narrowing series of improvements to a secondary control rather than an open hole. This one: excluding the separator from the consumed class stopped the strip at the first one, so a password containing a literal separator left its tail. Consuming greedily to the last separator before the path boundary closes the whole shape rather than the one spelling that was reported. RED-baseline (applied): a fixture whose password contains the separator reds against the non-greedy rule and greens against the greedy one, and reverting only that character class reds that row alone. |
|
|
677
|
+
| A fixture built to exercise a redactor is chosen so that no PART of it reads as a different sensitive shape: a reviewer quoting the finding puts the fixture into a committed receipt, where every other content gate then reads it | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_cli_review_wrappers.sh | `updated` | Owner key `code-review/SKILL.md`. Third distinct guard this round has tripped on its own test data, after the credential scanner and the egress tripwire. The userinfo fixture needed a password containing the separator; with a dotted host, the tail of that password plus the host reads as an email address, and the public-sanitization gate refused the receipt a reviewer wrote it into. A host without a dot exercises the same wrapper behaviour and forms no such shape. The reusable part is the indirection: the fixture is not what the gate scans -- the receipt is, and its content is chosen by a reviewer quoting the fixture, so the fixture has to be clean under every gate rather than under the one it was written for. Also recorded here because it cost a full CI round: `check-public-sanitization.py` and `review_ledger_binding.py` run in CI's repository-gates job and are absent from `make test`, so a green local run is not evidence about those two. |
|
|
678
|
+
| A fixture that FAILS TO PRODUCE its input turns every assertion reading that input green, so a change to a fixture is verified by observing the input it emits, not by the suite's exit status: a green suite is exactly what a dead fixture produces | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_cli_review_wrappers.sh | `updated` | Owner key `code-review/SKILL.md`. Landed and caught by the next review round, which is the honest record: an edit to the fixture indented one statement inside an embedded interpreter block, the interpreter died before emitting anything, the wrapper fell through to its no-output placeholder, and five redaction assertions passed over a placeholder while the suite reported its success token. The mutation evidence recorded for those rows was true when taken and had silently expired. Two rules follow, and the second is the one that would have caught it: after changing a fixture, observe the input it now emits; and re-run the mutation for the rows that fixture feeds rather than citing the earlier run. RED-baseline (applied, after the indentation was corrected): removing the redactor reds all five rows and the unmutated control is green -- which is the check that was missing, because the same removal against the broken fixture left all five green. |
|
|
679
|
+
| A value used as a redaction NEEDLE is rejected when it is only structure: a home directory of `/` is a legitimate environment and a catastrophic needle, because replacing it rewrites every separator in the text and disables every rule that runs after it | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/codex_review.sh | `updated` | Owner key `code-review/SKILL.md`. Raised by the landing review against the normalization the previous round added: gathering every spelling of the home directory is right, but a spelling that carries no content is not a path to elide. With `HOME=/` -- root, or an arbitrary-uid container -- the needle set contained `/`, the replacement ran before the URL rules, and the excerpt came out mangled with its credentials intact. The general shape is that a needle derived from the environment needs a content test, not only a presence test. RED-baseline (applied): a row invoking the wrapper with `HOME=/` reds without the content test and greens with it, and removing only that test reds that row alone. |
|
|
680
|
+
| A filter over free text in a persisted artifact is replaced by having no free text: "nothing secret-shaped survives" is not decidable over arbitrary text, so an adversarial reviewer can always spell one more escape, and the terminal state is a constant the input cannot influence rather than a filter that keeps growing | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/codex_review.sh | `updated` | Owner key `code-review/SKILL.md`. The measurement is the row: eight review chains after the input was already narrowed to CLI-authored error messages, each found a different escape -- an unlisted key name, an assignment form, URL userinfo, a password containing the separator, a fixture whose own shape tripped a neighbouring gate, an escaped quote closing a quoted value early, an uppercase scheme, a separator-only home used as a needle. Every one was real and none was derivable from the previous one. Two intermediate diagnoses were wrong on the way and are recorded above: swapping a key-name list for an assignment-shape rule was called an invariant change and was another enumeration, and narrowing the input was called sufficient when it only slowed the rate. What ends the class is that the receipt now carries a constant and the transport's output stays in the preserved run directory. The property is stated as equality with that constant, which a test can hold, instead of the absence of a list of shapes, which no test can. RED-baseline (applied): echoing the extracted message into the receipt reds the invariant row, and the unmutated control is green. Cost, recorded because the next round should be able to weigh it: each chain was roughly twelve minutes of wall clock, and the merge gate accepts no open challenge finding, so there was no landing state that carried the residue. |
|
|
681
|
+
| A fixture string is read by every scanner in the repository, not only by the suite it belongs to, so it is chosen to be inert under all of them: a host:port that exists only inside a quoted test payload still reads as a listening port to a lane-isolation scanner | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_cli_review_wrappers.sh | `updated` | Owner key `code-review/SKILL.md`. Fourth guard this round to reject the round's own test data, after the credential scanner, the egress tripwire and the public-sanitization gate. The userinfo fixture carried a port it never needed, and the parallel-lane isolation scanner reads any host:port in a lane member as evidence that concurrent suites could race on it. Dropping the port exercises the same wrapper behaviour. Recorded as one rule with the three before it: the cost of learning this one guard at a time was a full verification cycle each, and the cheaper order is to sweep every local gate after touching a fixture, before spending a review chain on the candidate. RED-baseline (applied): `test_lane_isolation.py` reds on the ported form and greens on the bare host, with the wrapper suite green either way -- which is why the suite alone was not evidence. |
|
|
682
|
+
| A rule that lives in a skill the situation never loads does not fire, however well it is written: the controller that runs the review chain raises the design question itself, inside that round's own envelope, and counts only the history a receipt carries | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md`. Both halves were registered deferrals, reproduced against the current controller before any code was written, each probe paired with a control leg. `recurring_findings_design_check` joins the required self-review triggers, with `decide_keep_delete_narrow_replace` among the allowed actions, whenever a findings round's own history already holds one — an earlier round of this chain, or the predecessor a succession names. `claim_strength` becomes a required self-review concern at build and release depth, owed before round 1 because that is the only point in a round where correcting an overstated claim is free. Applied-mutation RED baseline, differential: control 274 ok / 0 FAIL; removing the succession carry fails exactly `findings after a predecessor chain that also returned findings raise the design check` (273 ok / 1 FAIL); exempting `claim_strength` from the plan-coverage check fails exactly `a plan that skips the claim-strength walk fails before any provider runs` (273 ok / 1 FAIL); no non-owning assertion moves in either mutant. Two control legs ship inside the suite so a FIRST findings round is proved not to raise the trigger. The additions crossed the 500-line reference cap, so the wording-only exception moved verbatim into `code-review/references/wording-only-review.md` (the ratchet's own split-by-subtopic remedy) and the entrypoint's cumulative-budget paragraph became a pointer whose every number already lives in `code-review/references/timeout-auth-and-capabilities.md`; entrypoint body words end 8 below base. Round-1 review (kimi) returned one P2 on exactly that relocation — the numbers were delegated to a file outside the packet with nothing pinning them there — so the three literals are now contract anchors; the applied deletion mutation on one of them turns the anchor gate red (1 of 13) and the unmutated control runs green. Adding a required concern is a repository-wide compatibility event: five suites carried review-plan fixtures that omitted it and only the full lane found them, so the lane is run to green before a review chain opens rather than after. Supporting evidence: `code-review/scripts/review_gate.py`, `code-review/scripts/test_review_gate.sh`, `code-review/references/staged-review-contract.md`. |
|
|
683
|
+
| The same-class-recurrence rule states where it fires and what the firing controller is allowed to count, because the owner skill is not loaded at the transition where the situation arises | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/firing-point-placement.md#The controller must raise it, not the reader | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. The rule is owned here, but the situation it governs arises inside a review chain where this skill is never loaded, so naming the owner in prose never made it fire; the firing point moves onto the transition and the statement of it lands in `references/firing-point-placement.md`, the reference the entrypoint already points at for firing-point mechanics, because the attention-budget ratchet holds this entrypoint at its base measure. The reference records what discharges the trigger (the keep/delete/narrow/replace decision, ratified by a risk owner other than the one proposing it), the bound on what the controller may count (this chain plus the predecessor a succession names; succession does not compose, so a third chain opened fresh carries none), and that the trigger over-fires by design. RED baseline is the controller's own suite: with the succession carry removed, the chain this text describes stops raising the check and exactly that assertion fails (273 ok / 1 FAIL against a 274 ok / 0 FAIL control), with no other assertion moving. |
|
|
684
|
+
| A gate that forces evidence to be bought on an artifact that will never land is charging for the wrong thing: a chain ends where the candidate moves, so the round it ended on is whichever round came last -- and requiring that round to be a challenge only moved the spend, never the proof | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md`. A succession may now carry a chain whose terminal receipt is its REVIEW, not only its challenge. Cost of the old shape, recorded as the author's measurement of a prior round rather than as anything this packet can reproduce: a fix applied straight after a review ended the chain by moving the owner digest, and reaching the post-fix candidate then required spending the chain's challenge on the pre-fix candidate first, so challenges were bought on candidates that never landed. Exempted class, stated behaviourally: a challenge on a candidate that will never land, which carries no evidence about the one that does. Every other binding holds -- the candidate must still have moved, succession still does not compose, the per-chain budget is untouched, and this path spends fewer rounds than the old one. What bounds it is the receipt's own arithmetic, and review named that for what it is: a FORGERY guard, not a history check. A genuine round-1 review reads the same whether its chain later ran a challenge or not, so a caller who spent the challenge and presents only the review is accepted, and the successor inherits no challenge focuses -- one that chain did spend can be spent again. The round had claimed that refusal in its acceptance and 'proved' it with a receipt shape the controller never emits; the claim is withdrawn rather than mechanised, because no check at the succession call site can close an omitted-history gap that the rest of this contract already declares. Controller comment, contract text, probe name and acceptance criterion all now say only what is enforced. Applied-mutation RED baseline, differential and attributed to the guard rather than to a diagnostic string: reverting the whole change reds the new probes only because the base rejects EVERY review predecessor, which proves nothing about the arithmetic guard, so the recorded mutation disables that guard ALONE -- review predecessors still admitted, their chain-ended arithmetic no longer checked. Under it exactly one assertion fails, `a succession rejects a forged review receipt whose own arithmetic says its chain is spent`, and it fails because the succession was ACCEPTED; the unmutated control runs the suite green. The same round adds `--print-required-concerns`. Parity is proved over ANSWERS and over REFUSALS, the second only after challenge found the export answering for inputs the enforcer rejects -- a risk tag containing whitespace, an empty tag, an over-long tag -- which is the divergence the export exists to remove, reproduced against the shipped binary and now refused identically on both sides -- asserted on BOTH, after challenge observed the first parity probes invoking only the printer, so removing the enforcer's own rejection would have left them green. Removing it in an isolated clone now reds exactly those three assertions and nothing else. The clone matters: three earlier attempts mutated the worktree and restored it at the end of the same command, and one such restore -- from a backup another still-running task had taken while the file was already mutated -- put a controller with its tag validation stripped back into the tree, caught only because the task's output was shorter than expected. Destructive probes run on a copy. Challenge then found the parity checks themselves discarding exit status: this suite runs without errexit, so a command substitution swallows it and a printer that emitted the right concerns before failing would still have satisfied a non-empty check. The valid calls now assert their status, and a mutant that prints correctly then returns non-zero reds exactly those two assertions. The same failure-propagation blind spot appeared twice in one round -- the first parity probe had instead enabled errexit, aborting the suite at its first non-zero command with zero FAIL lines -- so both directions are now covered by assertions rather than by shell defaults. Asked to sweep the class rather than patch the instance, the next challenge found the remaining one -- a printer call NESTED inside a comparison, whose status no assignment could capture. All four printer invocations in the suite now capture status; a mutant failing only for explicit release depth with no risk tag reds exactly the assertion that covers it, and nothing else. The answer direction is proved in both senses: a plan built from it is accepted, and dropping each printed concern in turn turns the gate red -- at BOTH depths, after the challenge observed that proving it only at build left the release/high-risk branch free to diverge with every test green. |
|
|
685
|
+
| A consumer that keeps its own copy of a set the controller owns drifts the moment that set changes, and the suite runner aborts at its first failing target so the drift surfaces rounds later -- while relocating a working pin mechanism to buy faster feedback costs more than the latency it buys | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. This owner's wrapper suite derives the required concern set from the controller that enforces it instead of holding a copy. RED-baseline, paired control differing in exactly one variable: with one concern added to the controller's build stage, the base fixture fails `self_review_incomplete` on the no-independent-reviewer assertion (rc 1) while the derived fixture on this candidate passes (rc 0); with the mutation removed the derived fixture passes again. Second thread, withdrawn rather than landed: the round first relocated twenty prose pins from a slow suite into the fast registry, and five consecutive challenge findings landed inside the matching normalisation that relocation required -- this repo's own cue to question the capability instead of patching it again -- while the change additionally applied looser whitespace-insensitive matching to fourteen pre-existing anchors that never asked for it. The four mechanism files are byte-identical to base. What lands from that thread is the write-side norm that sends an author to BOTH pin surfaces, with a command for each that was run before it was written down -- the first draft shipped a registry lookup that scanned the wrong scripts directory and returned nothing, which challenge caught; the replacement lists a file's anchor ids and was verified against a file that has them -- before a budget-funded trim; the latency that motivated the relocation is left to its own change, where wiring the owner suite into the fast gate buys the same feedback with no new matching semantics. |
|
|
686
|
+
| 改写既有文档时新写的句子不得把用词水位抬到原文之上——病根是编辑落笔用的是自己的词库而不是宿主文档的;查法是把本轮新句单独拎出、逐个术语查它在原文里出现过没有、没出现的换成原文说法或当场白话解释 | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/SKILL.md#改写时新句用词不得高出原文水位、抬高读者门槛 | `updated` | Owner key `tighten-doc/SKILL.md`。Observed failure:一篇面向业务读者的大白话协作文档被逐句就地修订约 25 次,每一次替换单独看都正确;收尾回读的既有触发词是「分享 / 发布之前」,而就地编辑的每一刀落地即发布,永远到不了那个时刻,整体回读因此一次也没跑;新写的句子同时带进了编辑自己的行话,文档 owner 的反应是「改成看不懂的了」,返工重写了全部新句。两个缺陷都只在跨刀整体读时显形。证据边界照实说明:prose 收尾规则在本仓没有可执行的行为 oracle,锚点钉住的是规则在场与措辞,不是模型行为差分,不按行为实测记。owner-generalization map 在 `specs/125-doc-closeout-and-register-drift/evidence/owner-map.md`(十个 owner 逐条 updated/unchanged/routed/not-applicable)。RED-baseline(applied,differential):把该锚定句改成非规范措辞,`register-firing-path-resolution.rb` rc=1 并点名 `source-register.md` 的这一行与该 locator;控制组与恢复后均 rc=0(落地时重跑为 527 locators resolved,与同轮提交的 mutation-walk.txt 一致;518 是本行初稿时的捕获值,账本此后被追加过),同一 diff 的其余检查两侧不变。 |
|
|
687
|
+
| 就地编辑一篇已发布文档时不存在「分享前」这一刻——每一刀落地即发布——所以挂在分享前的收尾回读永远不触发;触发点顺延到本轮最后一次写操作之后,收工前必须整体回读一遍 | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/closeout-reread.md#没有「发布前」这一刻**:每一刀落地即发布,读者随时可能正在读。触发点顺延到**最后一次写操作之后**:收工前必须把整篇(或受影响那一面)整体回读一遍 | `updated` | Owner key `tighten-doc/SKILL.md`。Observed failure:一篇面向业务读者的大白话协作文档被逐句就地修订约 25 次,每一次替换单独看都正确;收尾回读的既有触发词是「分享 / 发布之前」,而就地编辑的每一刀落地即发布,永远到不了那个时刻,整体回读因此一次也没跑;新写的句子同时带进了编辑自己的行话,文档 owner 的反应是「改成看不懂的了」,返工重写了全部新句。两个缺陷都只在跨刀整体读时显形。证据边界照实说明:prose 收尾规则在本仓没有可执行的行为 oracle,锚点钉住的是规则在场与措辞,不是模型行为差分,不按行为实测记。owner-generalization map 在 `specs/125-doc-closeout-and-register-drift/evidence/owner-map.md`(十个 owner 逐条 updated/unchanged/routed/not-applicable)。细节落在同包的 `tighten-doc/references/closeout-reread.md`(触发点、这一遍要拿出的证据、三类不能顶替它的东西、语域漂移查法),entrypoint 只留触发与硬规则并因此净缩小(bytes 49905→49454,body words 9822→9806)。RED-baseline(applied,differential):把该锚定的规范列表行改写掉,`register-firing-path-resolution.rb` rc=1 并点名该 locator;控制组与恢复后 rc=0。这一行与上一行分别钉住本轮两条规则,任何一半**被锚定的那段字面**被删或被改写都会红,不靠同一个锚代管两件事;但改动规则的**适用条件**(例如给它加一个前置条件)不动锚内任何字,闸检不出——这一条与本表下方那行同口径,不作更强声称。锚点特意跨过承重从句——「没有发布前这一刻 / 每一刀落地即发布 / 顺延到最后一次写操作之后 / 收工前必须」连成一条字面量,删掉其中任一从句而只留末句都会让 locator 失配;这是评审指出的绕过(只锚末句时,删掉前面的适用条件仍能过闸)。覆盖边界照实说明:entrypoint 里那句一行复述不单独钉锚,本行不声称闸能护住它。 |
|
|
688
|
+
| 语域漂移的查法本身是规则的一部分:新句里原文没出现过的术语是候选,每个候选必须换成原文已有的说法或当场用一句白话解释,原样留着不算处理 | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/closeout-reread.md#每个候选必须二选一:换成原文已经在用的说法,或当场用一句白话把它解释掉 | `updated` | Owner key `tighten-doc/SKILL.md`。查法与触发点同属本轮那条收尾规则,分行钉锚是**归属选择**不是解析器限制——守卫会把逗号分隔的多个 locator 拆开各自解析(`register-firing-path-resolution.rb` 的 locator 拆分与 multi_bad/multi_ok 两条用例);分行是为了让红的时候直接指到是哪一步被动了。促成拆分的实测是:只钉规则句时,把查法三步删掉仍能过闸。RED-baseline(applied,differential):删掉该规范列表行,`register-firing-path-resolution.rb` rc=1 并点名该 locator;控制组与恢复后 rc=0。覆盖边界:本行只钉「候选术语必须处理」这一步;隔离新句与对照改前文本两步由下两行分别钉住。 |
|
|
689
|
+
| 语域漂移的查法要先隔离本轮新句:必须把新写的句子单独拎出来单独看,混在整篇里读就看不出用词是谁的 | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/closeout-reread.md#必须把本轮**新写的句子**单独拎出来单独看 | `updated` | Owner key `tighten-doc/SKILL.md`。与上三行同属本轮那条收尾规则,分行是归属选择——守卫支持一条 firing-path 里放多个 locator 并各自解析,分行只是为了让红时能指到具体那一步。RED-baseline(applied,differential):删掉或改写该规范列表行,`register-firing-path-resolution.rb` rc=1 并点名该 locator;控制组与恢复后 rc=0。这一行来自人工授权轮的评审:三个 locator 都不覆盖「隔离新句」,删掉它仍能过闸。 |
|
|
690
|
+
| 语域漂移的候选判定必须以改动之前的文本为对照:不得拿改完的文档做对照,否则新词已经在里面,永远查不出候选 | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/closeout-reread.md#逐个术语查它**在改动之前的文本里**出现过没有;不得拿改完的文档做对照,否则新词自己就在里面,永远查不出候选。没出现过的就是候选 | `updated` | Owner key `tighten-doc/SKILL.md`。与上三行同属本轮那条收尾规则,分行是归属选择——守卫支持一条 firing-path 里放多个 locator 并各自解析,分行只是为了让红时能指到具体那一步。RED-baseline(applied,differential):删掉或改写该规范列表行,`register-firing-path-resolution.rb` rc=1 并点名该 locator;控制组与恢复后 rc=0。这一行同样来自人工授权轮。锚点从「逐个术语」起,覆盖遍历范围、对照对象、禁令与候选定义四段连写:上一轮 challenge 实测出,只钉禁令时把「在改动之前的文本里」换成「在术语表里」,五条 locator 全部照旧匹配而查法已废;扩锚后该替换直接失配。同类 finding 已连出三轮(只护一半规则 / 锚点截断 / 换掉正面对照对象),据此在账本里把结论写死:substring 锚钉的是字面不是语义,它保证的只是**被锚定的那段字面**被删或被改写时会红;它不保证规则被改写时会红——实测:把查法的引导词改成「以下查法仅在用户明确要求时执行」,五条 locator 全部照常匹配而整条查法已成可选。适用条件与语义完整性由评审与人读负责,本行不作此声称。 |
|
|
691
|
+
| 钉住散文规则的锚点闸必须自带能失败的测试:删掉被锚定的规则行必须让闸变红并点名该 locator,而锚点之外的文字被掏空时闸不得报红——后一条把「锚钉字面不钉语义」这条边界写成被执行的事实,而不是账本里的一句声明 | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`(本轮未改,改动落在同包的该测试脚本)。Observed failure:本轮连续三轮 challenge 都指出同一件事——变异记录依赖的守卫不在按 diff 装的评审包里,包内无法核验;先后用「路径+blob 哈希」和「候选内的风险责任人接受书」作答都被驳回,后者尤其是错的:被评审的东西不能自己给自己发授权。RED-baseline(applied,differential,隔离路径归因):完整套件 fail-fast,任何守卫变异都先撞红最早受影响的既有用例,新增两条根本跑不到——本轮 challenge 正是据此推翻了先前那条「改诊断串」的证据,那次变异撞的是既有的 reworded-anchor 用例,什么也没归因到。改用**每条用例各一份单例副本**:把 `next if body.include?(anchor)` 换成整行相等时,边界用例红、删除用例绿;换成从不报缺失锚点时,删除用例红、边界用例绿;控制组与恢复四格全绿。两次不翻转的格子都是实跑观测;探针已提交为 `specs/125-doc-closeout-and-register-drift/evidence/attribution-probe.sh`,一条命令重跑整张矩阵;它在被测提交的一次性 detached worktree 里执行,调用者的检出只读不写(首版写进活动检出,被本轮评审判为 P1 并已重写)。三次被推翻的归因尝试(改诊断串撞到既有用例/两条用例同副本致后一条不执行/副本临时不可复核)连同原因一并留在走查里。两次变异与原始输出见 `specs/125-doc-closeout-and-register-drift/evidence/mutation-walk.txt` 的第二段。 |
|
|
@@ -14,3 +14,7 @@ coverage-tier-provenance skills/testing-strategy/references/test-code-authoring-
|
|
|
14
14
|
pairwise-trigger-range skills/test-artifact-management/references/classical-test-design-techniques.md 2-way 累计触发 53–97% 071-chainC-r1f4: NIST SP 800-142 empirical range, externally verified (specs/071 source-verification)
|
|
15
15
|
bva-two-vs-three-value skills/test-artifact-management/references/classical-test-design-techniques.md 2-value(边界 + 下一格)和 3-value(边界 + 两侧) 071-chainC-r1f4: ISTQB v4 BVA variant definitions, externally verified (specs/071 source-verification)
|
|
16
16
|
merge-side-ledger-binding-failclosed skills/skill-extraction-workflow/scripts/review_ledger_binding.py return 0 if args.allow_unevaluated else 2 084-r4f1: the no-base fail-closed branch; flipping it to an unconditional 0 restores a gate that passes having checked nothing
|
|
17
|
+
lane-budget-default-range skills/code-review/references/timeout-auth-and-capabilities.md defaults to 2400 seconds and accepts 5 to 3600 125-r1f1: the entrypoint now points here for the cumulative lane budget instead of restating it; dropping or drifting the default/range would silently strip the bound from both surfaces
|
|
18
|
+
lane-budget-mode-minimums skills/code-review/references/timeout-auth-and-capabilities.md 21 total seconds for review and 16 125-r1f1: the per-mode fail-closed minimums the entrypoint delegates here; without the pin the fail-closed limit can vanish with no suite failing
|
|
19
|
+
lane-budget-reserved-seconds skills/code-review/references/timeout-auth-and-capabilities.md while reserving ten controller 125-r1f1: the reserved controller seconds in the per-invocation division the entrypoint delegates here
|
|
20
|
+
recurrence-rule-owner-pointer skills/code-review/references/staged-review-contract.md `../../skill-extraction-workflow/SKILL.md` owns that rule 125-binding-r2f1: the chain-bound reader cannot load the owning skill, so this pointer is the only route to the rule that discharges the trigger; a wrong depth resolves inside code-review and no link checker sees a backticked path
|
|
@@ -443,14 +443,14 @@ def candidate_hash(module: types.ModuleType, repo_root: Path, base: str, paths:
|
|
|
443
443
|
paths=list(paths),
|
|
444
444
|
wording_only_proof_file=None,
|
|
445
445
|
)
|
|
446
|
-
packet_path,
|
|
446
|
+
packet_path, _packet_sha256, candidate_sha256, _bytes, _paths, _secrets = module.freeze_packet(
|
|
447
447
|
args, time.monotonic() + 120
|
|
448
448
|
)
|
|
449
449
|
try:
|
|
450
450
|
Path(packet_path).unlink(missing_ok=True)
|
|
451
451
|
except OSError:
|
|
452
452
|
pass
|
|
453
|
-
return
|
|
453
|
+
return candidate_sha256
|
|
454
454
|
|
|
455
455
|
|
|
456
456
|
def canonical_digest(value: object) -> str:
|
|
@@ -55,12 +55,14 @@ PY
|
|
|
55
55
|
# Exercise the installed wrapper/controller pair without invoking a model. The
|
|
56
56
|
# same real controller defaults to budget 0 when called directly, while the
|
|
57
57
|
# extraction wrapper must make the emitted receipt report budget 1.
|
|
58
|
-
python3 - "$TMP/real-controller.diff" "$TMP/real-controller-plan.json" <<'PY'
|
|
58
|
+
python3 - "$TMP/real-controller.diff" "$TMP/real-controller-plan.json" "$ROOT" <<'PY'
|
|
59
59
|
import json
|
|
60
|
+
import subprocess
|
|
60
61
|
import sys
|
|
61
62
|
from pathlib import Path
|
|
62
63
|
|
|
63
|
-
diff_path, plan_path = map(Path, sys.argv[1:])
|
|
64
|
+
diff_path, plan_path = map(Path, sys.argv[1:3])
|
|
65
|
+
root = Path(sys.argv[3])
|
|
64
66
|
diff_path.write_text(
|
|
65
67
|
"diff --git a/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh "
|
|
66
68
|
"b/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh\n"
|
|
@@ -69,20 +71,41 @@ diff_path.write_text(
|
|
|
69
71
|
"@@ -1 +1 @@\n-old wrapper\n+new wrapper\n",
|
|
70
72
|
encoding="utf-8",
|
|
71
73
|
)
|
|
74
|
+
# The required set has ONE owner. A fixture that keeps its own copy stops
|
|
75
|
+
# satisfying the gate the moment that set changes, and the runner aborts at its
|
|
76
|
+
# first failing target so the later shards never report it -- five suites drifted
|
|
77
|
+
# that way in one round. Derive it instead.
|
|
78
|
+
required = subprocess.run(
|
|
79
|
+
[
|
|
80
|
+
sys.executable,
|
|
81
|
+
str(root / "skills/code-review/scripts/review_gate.py"),
|
|
82
|
+
"--print-required-concerns",
|
|
83
|
+
"--stage",
|
|
84
|
+
"build",
|
|
85
|
+
],
|
|
86
|
+
capture_output=True,
|
|
87
|
+
text=True,
|
|
88
|
+
check=True,
|
|
89
|
+
).stdout.split()
|
|
90
|
+
assert required, "the controller printed no required concerns"
|
|
72
91
|
conclusions = {
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
"failure_paths": "The no-independent-reviewer boundary remains structured and fail closed.",
|
|
76
|
-
"tests_evidence": "Direct and wrapped calls provide a differential budget assertion.",
|
|
77
|
-
"compatibility": "The generic controller default remains zero while extraction fixes one.",
|
|
78
|
-
}
|
|
79
|
-
skills = {
|
|
80
|
-
"correctness": "skill-extraction-workflow",
|
|
81
|
-
"safety": "code-review",
|
|
82
|
-
"failure_paths": "python-service-dev",
|
|
83
|
-
"tests_evidence": "testing-strategy",
|
|
84
|
-
"compatibility": "terminal-cli-dev",
|
|
92
|
+
concern: f"The differential budget probe covers {concern} within the fixture it was observed on."
|
|
93
|
+
for concern in required
|
|
85
94
|
}
|
|
95
|
+
# Owners are a separate obligation: the controller derives them from the candidate's
|
|
96
|
+
# own paths, so they do not drift when the concern set does. Spread the required
|
|
97
|
+
# concerns over them, then give any owner the spread missed a row of its own.
|
|
98
|
+
owners = [
|
|
99
|
+
"skill-extraction-workflow",
|
|
100
|
+
"code-review",
|
|
101
|
+
"python-service-dev",
|
|
102
|
+
"testing-strategy",
|
|
103
|
+
"terminal-cli-dev",
|
|
104
|
+
]
|
|
105
|
+
skills = {concern: owners[index % len(owners)] for index, concern in enumerate(required)}
|
|
106
|
+
extra_rows = [
|
|
107
|
+
(required[0], owner) for owner in owners if owner not in set(skills.values())
|
|
108
|
+
]
|
|
86
109
|
plan = {
|
|
87
110
|
"intent": "Prove the extraction wrapper and real review controller agree on budget one.",
|
|
88
111
|
"acceptance": ["The wrapped real-controller receipt reports challenge_budget one."],
|
|
@@ -94,6 +117,15 @@ plan = {
|
|
|
94
117
|
"evidence_refs": ["real-controller-differential"],
|
|
95
118
|
}
|
|
96
119
|
for concern, conclusion in conclusions.items()
|
|
120
|
+
]
|
|
121
|
+
+ [
|
|
122
|
+
{
|
|
123
|
+
"concern": concern,
|
|
124
|
+
"skill": owner,
|
|
125
|
+
"conclusion": f"{owner} is covered for {concern} by the same differential probe.",
|
|
126
|
+
"evidence_refs": ["real-controller-differential"],
|
|
127
|
+
}
|
|
128
|
+
for concern, owner in extra_rows
|
|
97
129
|
],
|
|
98
130
|
"evidence": [
|
|
99
131
|
{
|
|
@@ -274,6 +306,11 @@ landing = (
|
|
|
274
306
|
staged = (
|
|
275
307
|
root / "skills/code-review/references/staged-review-contract.md"
|
|
276
308
|
).read_text(encoding="utf-8")
|
|
309
|
+
# The wording-only exception is specified in its own reference; the contract must
|
|
310
|
+
# still name it, so the pin proves both the specification and its reachability.
|
|
311
|
+
wording_only = (
|
|
312
|
+
root / "skills/code-review/references/wording-only-review.md"
|
|
313
|
+
).read_text(encoding="utf-8")
|
|
277
314
|
code_review = (root / "skills/code-review/SKILL.md").read_text(encoding="utf-8")
|
|
278
315
|
|
|
279
316
|
for label, text in {
|
|
@@ -310,10 +347,11 @@ for label, text in {
|
|
|
310
347
|
assert "wording_only_boundary" in text, (
|
|
311
348
|
f"{label} does not require independent wording-only semantic confirmation"
|
|
312
349
|
)
|
|
313
|
-
assert "
|
|
314
|
-
assert "--
|
|
315
|
-
assert "
|
|
316
|
-
assert "markdown-
|
|
350
|
+
assert "wording-only-review.md" in staged
|
|
351
|
+
assert "--wording-only-proof-file" in wording_only
|
|
352
|
+
assert "--challenge-budget 0" in wording_only
|
|
353
|
+
assert "markdown-punctuation-only" in wording_only
|
|
354
|
+
assert "markdown-token-replacement" in wording_only
|
|
317
355
|
assert "opens no challenge chain or `complete` checkpoint" in dual
|
|
318
356
|
assert "codex review --base" not in quickstart
|
|
319
357
|
assert "codex exec adversarial" not in quickstart
|
|
@@ -720,5 +720,45 @@ run_gate "$FIX"
|
|
|
720
720
|
assert_rc "$rc" 0 "a fenced example row is an illustration, not a live declaration"
|
|
721
721
|
pass "fenced declaration-looking example does not red a register with no live declarations"
|
|
722
722
|
|
|
723
|
-
|
|
723
|
+
# ── N. RED: the anchored list rule DELETED outright ──────────────────────────
|
|
724
|
+
# A prose rule pinned for a documentation round is normally removed, not
|
|
725
|
+
# reworded: the round that added this case pinned five rules whose only
|
|
726
|
+
# mechanical protection is this gate, and its walk deleted each in turn.
|
|
727
|
+
new_fixture deleted_rule "$LITERAL"
|
|
728
|
+
# A mutation that does not apply proves nothing, so each edit below is checked
|
|
729
|
+
# both before and after: the target must be present first, and the edit must
|
|
730
|
+
# have changed the file. Without this a fixture drift turns these cases into
|
|
731
|
+
# assertions about an unmutated fixture that still pass.
|
|
732
|
+
grep -qF 'Never bypass the demo isolation boundary' "$FIX/skills/demo-skill/SKILL.md" \
|
|
733
|
+
|| fail "fixture drift: the rule this case deletes is not in the fixture"
|
|
734
|
+
grep -v 'Never bypass the demo isolation boundary' "$FIX/skills/demo-skill/SKILL.md" \
|
|
735
|
+
> "$FIX/skills/demo-skill/SKILL.md.tmp"
|
|
736
|
+
mv "$FIX/skills/demo-skill/SKILL.md.tmp" "$FIX/skills/demo-skill/SKILL.md"
|
|
737
|
+
grep -qF 'Never bypass the demo isolation boundary' "$FIX/skills/demo-skill/SKILL.md" \
|
|
738
|
+
&& fail "mutation did not apply: the rule is still present"
|
|
739
|
+
run_gate "$FIX"
|
|
740
|
+
assert_rc "$rc" 1 "deleting the anchored rule outright must be caught"
|
|
741
|
+
assert_contains "anchor text absent from target" "$out" "deleted anchored rule"
|
|
742
|
+
assert_contains "Never bypass the demo isolation boundary" "$out" "names the dead locator"
|
|
743
|
+
pass "deleting an anchored rule turns the gate RED"
|
|
744
|
+
|
|
745
|
+
# ── N+1. GREEN by design: text OUTSIDE the anchor may be gutted ──────────────
|
|
746
|
+
# The stated boundary, asserted rather than promised: a substring locator binds
|
|
747
|
+
# the letters it names and nothing else. Here the anchored clause survives while
|
|
748
|
+
# the rest of its line is replaced, and the gate is silent — which is why the
|
|
749
|
+
# ledger rows that rely on it must not claim semantic protection.
|
|
750
|
+
new_fixture unanchored_clause_gutted "$LITERAL"
|
|
751
|
+
grep -qF ' when dispatching work.' "$FIX/skills/demo-skill/SKILL.md" \
|
|
752
|
+
|| fail "fixture drift: the clause this case guts is not in the fixture"
|
|
753
|
+
sed -i.bak 's/ when dispatching work\./ — every safeguard around it removed./' \
|
|
754
|
+
"$FIX/skills/demo-skill/SKILL.md"
|
|
755
|
+
grep -qF ' — every safeguard around it removed.' "$FIX/skills/demo-skill/SKILL.md" \
|
|
756
|
+
|| fail "mutation did not apply: the unanchored clause is unchanged"
|
|
757
|
+
run_gate "$FIX"
|
|
758
|
+
assert_rc "$rc" 0 "gutting text outside the anchor is invisible to this gate"
|
|
759
|
+
assert_contains "register_firing_path_resolution_ok" "$out" "boundary control"
|
|
760
|
+
assert_not_contains "anchor text absent" "$out" "no false RED on unanchored text"
|
|
761
|
+
pass "text outside the anchor can be gutted while the gate stays green (declared boundary)"
|
|
762
|
+
|
|
763
|
+
[ "$passed" -eq 59 ] || fail "expected 59 assertions, saw $passed (a case was skipped or misplaced)"
|
|
724
764
|
echo "register_firing_path_resolution_tests_ok ($passed assertions)"
|
|
@@ -903,6 +903,74 @@ else
|
|
|
903
903
|
'{ [ "$real_rc" = 0 ] && { case "$real_out" in [0-9a-f]*) [ ${#real_out} = 64 ];; *) false;; esac || case "$real_out" in *review_ledger_binding_no_change*) true;; *) false;; esac; }; } || { [ "$real_rc" = 1 ] && case "$real_out" in *"cannot freeze the candidate packet"*"review packet exceeds 200000 bytes"*) true;; *) false;; esac; }'
|
|
904
904
|
fi
|
|
905
905
|
|
|
906
|
+
# The point of splitting the candidate from the packet is that a round which
|
|
907
|
+
# widened its packet to answer an evidence-gap finding still produces a receipt
|
|
908
|
+
# this gate can accept. That claim spans both sides, so it is asserted across
|
|
909
|
+
# both: the identity the controller records for a widened packet must be the
|
|
910
|
+
# identity this gate recomputes from the repository. Mutate the gate to return
|
|
911
|
+
# the packet hash instead and this goes red while nothing else does.
|
|
912
|
+
printf 'widened-candidate\n' >>"$REPO/skills/skill-extraction-workflow/SKILL.md"
|
|
913
|
+
git -C "$REPO" add -A
|
|
914
|
+
git -C "$REPO" commit -qm widened
|
|
915
|
+
WIDENED_BASE="$BASE"
|
|
916
|
+
GATE_CANDIDATE="$(run_gate --base "$WIDENED_BASE" --print-candidate)"
|
|
917
|
+
controller_candidate="$(
|
|
918
|
+
CONTROLLER_DIR="$(dirname "$CONTROLLER")" REPO="$REPO" BASE="$WIDENED_BASE" \
|
|
919
|
+
GATE="$GATE" WORK="$WORK" python3 - <<'PY' 2>&1
|
|
920
|
+
import importlib.util
|
|
921
|
+
import os
|
|
922
|
+
import sys
|
|
923
|
+
import time
|
|
924
|
+
from pathlib import Path
|
|
925
|
+
from types import SimpleNamespace
|
|
926
|
+
|
|
927
|
+
sys.path.insert(0, os.environ["CONTROLLER_DIR"])
|
|
928
|
+
import review_gate
|
|
929
|
+
|
|
930
|
+
spec = importlib.util.spec_from_file_location("binder", os.environ["GATE"])
|
|
931
|
+
binder = importlib.util.module_from_spec(spec)
|
|
932
|
+
spec.loader.exec_module(binder)
|
|
933
|
+
|
|
934
|
+
repo = Path(os.environ["REPO"])
|
|
935
|
+
# An author cannot hand-build these: the base is a fork point and the paths
|
|
936
|
+
# carry the receipt exclusions this round already added. Both come from the gate
|
|
937
|
+
# that will judge the receipt, which is what makes it one identity.
|
|
938
|
+
base, _excludes, paths, _changed = binder.candidate_scope(
|
|
939
|
+
repo, os.environ["BASE"], (".",)
|
|
940
|
+
)
|
|
941
|
+
|
|
942
|
+
|
|
943
|
+
def freeze(**overrides):
|
|
944
|
+
args = SimpleNamespace(
|
|
945
|
+
cwd=str(repo),
|
|
946
|
+
diff_file=None,
|
|
947
|
+
base=base,
|
|
948
|
+
paths=list(paths),
|
|
949
|
+
wording_only_proof_file=None,
|
|
950
|
+
)
|
|
951
|
+
for key, value in overrides.items():
|
|
952
|
+
setattr(args, key, value)
|
|
953
|
+
return review_gate.freeze_packet(args, time.monotonic() + 60)
|
|
954
|
+
|
|
955
|
+
|
|
956
|
+
narrow_path, _narrow_packet, narrow_candidate, _n, _p, _s = freeze()
|
|
957
|
+
subject = narrow_path.read_bytes()
|
|
958
|
+
narrow_path.unlink()
|
|
959
|
+
|
|
960
|
+
# Outside the repository: the candidate includes untracked files.
|
|
961
|
+
widened = Path(os.environ["WORK"]) / "widened-for-binder.patch"
|
|
962
|
+
widened.write_bytes(subject + b"\n--- appended context for the reviewer ---\n")
|
|
963
|
+
wide_path, wide_packet, wide_candidate, _wn, _p, _s = freeze(diff_file=str(widened))
|
|
964
|
+
wide_path.unlink()
|
|
965
|
+
|
|
966
|
+
assert wide_packet != wide_candidate, "the widened packet must not be its own candidate"
|
|
967
|
+
assert wide_candidate == narrow_candidate, "widening moved the candidate"
|
|
968
|
+
print(wide_candidate)
|
|
969
|
+
PY
|
|
970
|
+
)"
|
|
971
|
+
check "a widened packet records the candidate identity this gate recomputes" \
|
|
972
|
+
'[ ${#GATE_CANDIDATE} = 64 ] && [ "$controller_candidate" = "$GATE_CANDIDATE" ]'
|
|
973
|
+
|
|
906
974
|
if [ "$fails" -gt 0 ]; then
|
|
907
975
|
echo "test_review_ledger_binding: $fails failing case(s)" >&2
|
|
908
976
|
exit 1
|
|
@@ -23,7 +23,7 @@ For spec/standard/guideline artifacts, the `pre-owner blocked` marker rule appli
|
|
|
23
23
|
|
|
24
24
|
Before sharing, syncing, committing, or publishing a concrete deliverable artifact, run Tighten mode by default. Do not ask the user whether to optimize unless the edit would change a substantive decision, remove required detail, or risk collaborative-comment loss. A request to write a template, SOP, report, checklist, Feishu/Lark doc, task card, or launch material already includes the optimization pass after the owner skill has settled the substance.
|
|
25
25
|
|
|
26
|
-
After multi-round edits that added, restructured, or changed reader-facing prose or meaning, always run the Tighten mode before sharing or publishing; rounds of only exempt trivial edits (typo/link/format) may record that exemption instead.
|
|
26
|
+
After multi-round edits that added, restructured, or changed reader-facing prose or meaning, always run the Tighten mode before sharing or publishing; rounds of only exempt trivial edits (typo/link/format) may record that exemption instead. **就地编辑线上文档没有「发布前」这一刻**,触发点顺延到最后一次写操作之后。这一遍的证据标准、扫描范围与语域漂移查法见 `references/closeout-reread.md`。
|
|
27
27
|
|
|
28
28
|
**技术 / 参考文档:form 优化 ≠ 内容已核。** 当本轮**新增 / 改写 / 发布**了可对源码核验的承载性技术断言(API 签名、env 行为 / 优先级、字段 / label 名、默认值——这类文档常自述真值源),**且一手源用户已给、在当前工作区可得、或用户明确授权查证**时,form 清理之外要把这些断言对一手源码(SDK / specs / fixtures / 可跑命令)核一遍。核出不符:**仅 source-literal 的 typo / 名称 / 默认值笔误**可直接照证据改;**会改动 substantive contract / decision 的**只标 discrepancy / pending 交 owner(stack / review skill),tighten-doc 不接管纠错。源不可得、超出本轮范围、或只是轻量润色:只做 form 并声明「形式已优化、内容未核」,不得裸称「文档已优化 / 准确」(别把每次 tighten 滚成一次代码审计)。
|
|
29
29
|
|
|
@@ -63,7 +63,7 @@ owner · 硬规则 · 完成标准/DoD · 里程碑 · 数值阈值 · the real
|
|
|
63
63
|
- **代码进代码块,不进段落**:**多行 / 独立执行步骤 / 长 flag 串命令 / 多命令序列**放带 lang 的代码块。**随文的短 one-liner / 表达式、表格单元格、`func()` 式符号引用可留 inline**,以不妨碍扫读为限。多语言对照两端形态对齐;随复制必需的说明:支持注释则放块内,否则随附;长篇原理放正文。
|
|
64
64
|
- **Enumeration sections (依赖/兜底/分工/里程碑 子项) = multi-line sub-bullets, NOT a `;`-collapsed single line.** Readability beats compactness here; a `- 依赖:A;B;C;D` run is hard to scan — split to `- 依赖:` + one `- A` sub-bullet per item. Do not collapse to one `;` line just for parity with another card; parity is not a reason to reduce scanability. Single-line `;` is only for a true 2-item short pointer where sub-bullets would be heavier than the content.
|
|
65
65
|
- Table cells that list multiple skills, owners, checks, environments, or evidence items should be split into multiple lines or shorter rows. A readable table beats a compressed cell when the cell is used as an execution checklist.
|
|
66
|
-
- Terms unified and glossed once in a 白话 section (e.g. 红灯 = 卡住/NO-GO 到点必升级; 排障手册 = 排障 SOP). Also catch **intra-doc term drift**: the same concept written two different ways in one doc → align to that doc's prevailing term. Drift includes **unit drift in a sequenced ladder** (a milestone list mixing 第N周 and N天 — align the lone odd unit to the ladder's prevailing one). 陌生或自造缩写仅在明显缩短且反复使用时引入;正式名称按下条保留。
|
|
66
|
+
- Terms unified and glossed once in a 白话 section (e.g. 红灯 = 卡住/NO-GO 到点必升级; 排障手册 = 排障 SOP). Also catch **intra-doc term drift**: the same concept written two different ways in one doc → align to that doc's prevailing term. Drift includes **unit drift in a sequenced ladder** (a milestone list mixing 第N周 and N天 — align the lone odd unit to the ladder's prevailing one). Drift 还包括**语域漂移**:改写时新句用词不得高出原文水位、抬高读者门槛(查法同上引用)。 陌生或自造缩写仅在明显缩短且反复使用时引入;正式名称按下条保留。
|
|
67
67
|
- A column/section header must match what its cells actually hold (a "文档化进度" header over cells that hold 现状 is a defect — rename the header to the truth). An editorial paren in a header/heading that restates an intro rule is the same 编辑性括号 as DELETE #9 — cut it.
|
|
68
68
|
- **Reader-facing published docs: the problem is unexplained or non-navigable internal references, not the names themselves**. Fix three recurring reader-blockers: ① internal repo paths used as navigation (`see README.md`) a non-author can't follow → name the human destination or link the published doc; ② opaque internal gate/code labels (`R0` / `F4`-style) → plain-name or drop the code; ③ unglossed in-house English / abbreviations (`mTLS` / `PTY` / `SLO`) → 中文化 or gloss at first use. **Keep** anything the reader actually operates on or that is a public convention / protocol / API / field / contract / standard name (`AGENTS.md`, `CODEOWNERS`, `package.json`, well-known abbrevs) — gloss if unfamiliar, don't delete.
|
|
69
69
|
- **A cell must fit its column's semantic role.** A 负责人/owner column entry must be a who (person/role), a 事项/规则 column a what (a parseable clause). Over-terse text — including a value the user dictated in an earlier pass — that no longer parses as that column's type ("业务真值+ 误差" in a 负责人 column; "…必需的指标建立" as a 规则 clause) is a 病句 (DELETE #7). On re-review, read each dictated/compressed value back **in its column context**, not in isolation; flag it with the rule even if the user set it (don't silently override, but don't pass it as clean either).
|
|
@@ -181,6 +181,7 @@ When the user probes sentence-by-sentence ("这是废话么 / 什么意思 / 能
|
|
|
181
181
|
- For Feishu/Lark comment preservation, range update mechanics, overwrite pitfalls, and cross-doc rename safety, read `references/comment-safe-feishu.md`.
|
|
182
182
|
- 「外部基线」条一触发就必读 `references/self-benchmark-baseline.md`:抽样、分布定位输出、可发现性语料与词频读法、两条边界。
|
|
183
183
|
- 交付面收尾(投放回读、三态定义与绕过路径、旧版退役与「被要求删除」的转出、静态导出全展开)展开在 `references/delivery-face-closeout.md`;closeout 第 7 项与「交付面」条都指向它。
|
|
184
|
+
- 收尾回读(触发点、证据标准、语域漂移查法)在 `references/closeout-reread.md`。
|
|
184
185
|
|
|
185
186
|
## CROSS-MODEL / CODEX CO-REVIEW CAVEAT
|
|
186
187
|
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/closeout-reread.md
ADDED
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# 收尾回读:触发点、证据、语域漂移
|
|
2
|
+
|
|
3
|
+
`SKILL.md`「After multi-round edits…」那条的展开。触发与硬判据在那条,这里放执行细节;两处冲突以 `SKILL.md` 为准。
|
|
4
|
+
|
|
5
|
+
## 触发点
|
|
6
|
+
|
|
7
|
+
- 默认触发点是**分享 / 同步 / 提交 / 发布之前**。只含 typo / 链接 / 格式这类豁免小改的轮次,记豁免即可。
|
|
8
|
+
- **就地编辑线上文档(协作文档、wiki、已发布页面)没有「发布前」这一刻**:每一刀落地即发布,读者随时可能正在读。触发点顺延到**最后一次写操作之后**:收工前必须把整篇(或受影响那一面)整体回读一遍。
|
|
9
|
+
- 「逐刀都很小」不是跳过的理由。**逐刀都对 ≠ 整篇读着对**:重复、术语漂移、语域漂移都是跨刀才显形的缺陷,只有整体回读能看见。
|
|
10
|
+
|
|
11
|
+
## 这一遍要拿出的证据
|
|
12
|
+
|
|
13
|
+
按 rubric 回读受影响的交付面:重复(含你自己本轮早些时候改出来的**同内容换面孔**)、一坨、术语漂移、语域漂移、cell-fits-column。
|
|
14
|
+
|
|
15
|
+
回读后记下五项:
|
|
16
|
+
|
|
17
|
+
- 回读的是哪一面。
|
|
18
|
+
- 自上次整体回读以来动过的块——这是**扫描范围的下限不是上限**,含义被这轮改动改变的小节也要进来。
|
|
19
|
+
- 是否命中 `SKILL.md` 里的整卡 / 全家族扫描触发条件。
|
|
20
|
+
- 查出的缺陷,或明确写「无」。
|
|
21
|
+
- 范围如果停在动过的块,写清为什么。
|
|
22
|
+
|
|
23
|
+
## 不能顶替它的三类东西
|
|
24
|
+
|
|
25
|
+
写的时候逐条守 rubric、收尾只跑 grep(半角 / 残留词扫描)、以及实质性内容评审——这三样**单独或加在一起都不算**这一遍的证据。
|
|
26
|
+
|
|
27
|
+
反复踩的就是把它们当等价证据,然后把自己改出来的重复发出去——那正是整体回读一眼能看见的东西。
|
|
28
|
+
|
|
29
|
+
## 语域漂移:改写场景里最高频的自引入缺陷
|
|
30
|
+
|
|
31
|
+
给已有文档做修订时,新写的句子用了比原文更专业的词,读者门槛被你抬高了。病根是编辑落笔用的是**自己的词库**,不是宿主文档的词库;原文越是刻意写成大白话,越容易被抬高。
|
|
32
|
+
|
|
33
|
+
查法:
|
|
34
|
+
|
|
35
|
+
- 必须把本轮**新写的句子**单独拎出来单独看,不混在整篇里读——混着读就看不出是谁的词。
|
|
36
|
+
- 逐个术语查它**在改动之前的文本里**出现过没有;不得拿改完的文档做对照,否则新词自己就在里面,永远查不出候选。没出现过的就是候选。
|
|
37
|
+
- 每个候选必须二选一:换成原文已经在用的说法,或当场用一句白话把它解释掉;原样留着不算处理。
|
|
38
|
+
- 收尾只读新句:读着**不像原文同一个人写的**,就是漂了。
|
|
39
|
+
|
|
40
|
+
指针型改写同样受这条管——写「按 X 执行」而不抄数值时,称呼那份 X 要用原文已有的说法,别换成你自己的行话。
|