@ccoalm/ccl-skills 0.16.0 → 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (18) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +6 -4
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +73 -123
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/wording-only-review.md +136 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +120 -11
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +17 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_order.sh +30 -15
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +209 -12
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +14 -7
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +1 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +22 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +10 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +4 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +56 -18
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +41 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +3 -2
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/closeout-reread.md +40 -0
  17. package/dist/assets/release.json +27 -17
  18. package/package.json +1 -1
@@ -74,7 +74,7 @@ Positive challenge capacity opens it at index 1; budget zero is untracked.
74
74
  The sole release/high-risk budget-zero exception is a controller-proved
75
75
  `markdown-punctuation-only` review: it requires `wording_only_boundary`, permits
76
76
  no `complete`, and rejects an author assertion alone (recipe:
77
- `references/staged-review-contract.md`). After a clean/source-refuted tracked
77
+ `references/wording-only-review.md`). After a clean/source-refuted tracked
78
78
  challenge, `complete` may close early and preserve unused rounds. Every result
79
79
  exposes controller-owned `self_review_gate`; an outstanding checkpoint blocks
80
80
  only external review or completion, not implementation or tests. Even a passed
@@ -280,7 +280,7 @@ For diffs over roughly 2,000 changed lines or 50 files, split review/challenge b
280
280
 
281
281
  Never treat a timeout, silence, or empty output as approval or "no findings": any timeout is inconclusive, and the final status must say `inconclusive` with the timeout reason so downstream review or merge state cannot treat the missing lane as approval. A live host execution handle such as `session_id` or `cell_id` means the same command is still running; poll that exact handle to terminal exit, and never start a replacement/fallback while its process is live. Do not use a 30 second silence as a review failure — narrow diff reviews legitimately take 2-3 minutes and broad reviews about 5. Challenge makes one formal Claude invocation; review and consult may make at most two only for their existing bounded result-recovery paths. After that, mark the Claude lane inconclusive and apply fallback only if the owning gate allows it. The wrapper traps TERM/INT/HUP and emits terminal `operator_interrupt`; the gate never starts another client after an operator interruption. SIGKILL and host crashes cannot be trapped, so non-zero exit without valid JSON remains inconclusive/manual-review-required, never as success. The timeout bound is per formal invocation, not per wrapper run; outer timeouts must cover the mode's worst case and must never kill the wrapper and then treat the killed output as success. For a yielded run, the caller's lane evidence records handle type, an opaque host transcript/tool-call reference rather than a raw credential-like handle, and terminal exit status. If that handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; a `ps`/process-tree capture and wrapper artifacts are diagnostic only and cannot reconstruct the missing terminal result. The outer host assigns this handle after launch, so this is a host-workflow obligation rather than a controller-owned field; exact enforcement requires a trusted host adapter. Recovery detail and timing formulas live in `references/timeout-auth-and-capabilities.md`.
282
282
 
283
- `review_gate.sh` also enforces a cumulative reviewer-lane budget: `--total-timeout` defaults to 2400 seconds, accepts 5 to 3600, reserves ten controller seconds, and divides the remainder by the mode's maximum invocation count. Starting a lane requires at least 21 seconds for review or 16 for challenge, plus setup overhead; smaller accepted values fail closed. Git preflight subprocesses share this deadline; direct filesystem reads still need the host's outer timeout. A timed-out client process group gets bounded cleanup and may cascade only while enough total budget remains. Total exhaustion returns terminal inconclusive `gate_timeout`; killed output is never a verdict. Timing details live in `references/timeout-auth-and-capabilities.md`.
283
+ `review_gate.sh` also enforces a cumulative reviewer-lane budget (`--total-timeout`), shared by git preflight but not by direct filesystem reads, which still need the host's outer timeout. Its default, range, reserved controller seconds, per-invocation division, and per-mode fail-closed minimums are in `references/timeout-auth-and-capabilities.md`. A timed-out client process group gets bounded cleanup and may cascade only while enough total budget remains. Total exhaustion returns terminal inconclusive `gate_timeout`; killed output is never a verdict.
284
284
 
285
285
  ## Auth And CLI Pitfalls
286
286
 
@@ -311,7 +311,8 @@ activation was observed; only a public event/export may populate
311
311
 
312
312
  ## Reference Loading
313
313
 
314
- - `references/staged-review-contract.md` — required review-plan schema, stage concerns, high-risk depth, prompt layers, and challenge budget. Load before review/challenge.
314
+ - `references/staged-review-contract.md` — review-plan schema, stage concerns, high-risk depth, prompt layers, challenge budget. Load before review/challenge.
315
+ - `references/wording-only-review.md` — the proof-bound wording-only single review. Load when claiming it.
315
316
  - `references/manual-invocation-and-prompts.md` — manual command shape, filesystem-boundary text, and the review/challenge prompt templates. Load only when debugging or patching the wrapper or its prompt construction.
316
317
  - `references/timeout-auth-and-capabilities.md` — wait-policy timing tables, the numbered auth-recovery procedure, per-mode tool-flag matrix, and CLI capability adoption notes. Load on timeout/auth failures or when maintaining wrapper flag adoption.
317
318
  - `references/client-routing.md` — `review_gate.sh` client order, family exclusion, egress, Kimi/Codex boundaries, OpenCode user-model binding, and concurrency rollback. Load when running or diagnosing review/challenge routing.
@@ -341,8 +342,9 @@ In the final work summary, include:
341
342
  - mode: review, challenge, complete, or consult
342
343
  - command scope, not the full prompt unless useful
343
344
  - result: blocking findings, no blocking findings, or inconclusive
344
- - the reviewed identity — a hash of the exact diff packet reviewed, **required** whenever the reviewed content includes staged, unstaged, untracked, or generated files (later worktree edits keep the same base/head SHA, so SHA alone cannot detect the change); the base/head commit SHA alone suffices only for a clean, fully-committed candidate tree. A review/challenge `no blocking findings` is valid **only** for that exact reviewed content: any later edit, rebase, amend, or new commit voids it and requires a fresh run (mirrors the agentic candidate-SHA binding). A caller — especially one invoking this skill standalone, outside a controller that already tracks the head SHA — must not reuse a prior pass as approval for changed content.
345
+ - the reviewed identity — a hash of the exact diff packet reviewed, **required** whenever the reviewed content includes staged, unstaged, untracked, or generated files (later worktree edits keep the same base/head SHA, so SHA alone cannot detect the change); the base/head commit SHA alone suffices only for a clean, fully-committed candidate tree. A `no blocking findings` result is valid **only** for that exact content: any later edit, rebase, amend, or new commit voids it and requires a fresh run (mirrors the agentic candidate-SHA binding), and no caller — least of all one invoking this skill standalone, outside a controller tracking the head SHA — may reuse a prior pass as approval for changed content.
345
346
  - any follow-up fixes made because of the review
347
+ - if `recurring_findings_design_check` fired, the `keep`/`delete`/`narrow`/`replace` decision, what it recurred across, and who ratified it
346
348
  - if skipped or inconclusive, the exact reason
347
349
  - for a host-yielded execution, the handle type, opaque host transcript/tool-call reference, and terminal exit status; if the handle was lost, record that infrastructure-inconclusive state, diagnostic artifacts, and that fallback was unavailable; never persist a credential-like raw handle in shared evidence
348
350
 
@@ -58,7 +58,12 @@ hand-attested plan (`review_plan_source=implementer-supplied` otherwise).
58
58
  Self-review accumulates stage concerns: explore covers correctness and
59
59
  safety; build adds failure paths, tests, and compatibility; release adds rollout
60
60
  and operations. High-risk input raises depth to release and adds
61
- `high_risk_boundary`.
61
+ `high_risk_boundary`. That set has one owner, and
62
+ `review_gate.sh --print-required-concerns --stage <stage> [--risk-tag <tag>]`
63
+ prints it, so a caller building a plan derives the list instead of keeping a copy
64
+ that silently stops satisfying the gate when the set changes. It prints what the
65
+ PLAN owes: the synthetic challenge slot and the wording-only boundary, which the
66
+ controller adds for the reviewer and never for the plan, are absent.
62
67
 
63
68
  The serialized plan is at most 32,000 bytes and `intent` is 8..4,000
64
69
  characters. Those are validation limits, not permission for a caller to slice a
@@ -168,6 +173,28 @@ top-level agents, commands, hooks, or MCP servers is not loaded.
168
173
  Wrappers keep an explicit selected-owner count instead of testing empty Bash
169
174
  arrays under `set -u`, preserving the no-owner lane on Bash 3.2.
170
175
 
176
+ ### The claim-strength walk, and why the late correction is not cheaper
177
+
178
+ `claim_strength` is a required self-review concern at build and release depth. The
179
+ plan walks the candidate's load-bearing claims — absolutes, universals, causal
180
+ statements, exhaustiveness — and for each one either names evidence that would
181
+ survive a challenge or weakens the claim on the spot. It is owed before round 1
182
+ because that is the only point in a round where correcting a claim is free: once a
183
+ round binds the candidate, an edit inside a selected owner package voids every
184
+ receipt bound to it, so a sentence that claims too much costs exactly what a changed
185
+ predicate costs. The concern also reaches the reviewer, so a claim that survives the
186
+ walk comes back as a round-1 finding — inside the fix batch the round was going to
187
+ pay for anyway — rather than at closeout, where the remaining moves are a fresh chain
188
+ or leaving it standing.
189
+
190
+ There is deliberately no cheap late path. The proof-bound single review
191
+ (`wording-only-review.md`) refuses any changed non-punctuation character, which is
192
+ exactly what weakening a claim is, and an exception keyed on the author's own "this
193
+ edit only weakens a claim" is an assertion the controller cannot re-derive — a waiver
194
+ of that shape was carried here once and removed, because a predicate that approximates
195
+ meaning keeps admitting shapes it did not anticipate. The price stays uniform in both
196
+ directions; the walk is what moves the correction to where the price is zero.
197
+
171
198
  ## Base-derived packet input boundary
172
199
 
173
200
  `--base` freezes the tracked diff plus every non-ignored untracked path in
@@ -226,126 +253,9 @@ later round read more than an earlier one.
226
253
 
227
254
  ## Proof-bound wording-only single review
228
255
 
229
- The wording-only exception is one untracked `review` with
230
- `challenge_budget=0`; it is not a chain, challenge, or `complete` checkpoint.
231
- Supply `--wording-only-proof-file` to bind the exception to the exact packet.
232
- Without that proof, an explore/build budget-zero review remains an ordinary
233
- single review and cannot be recorded as the wording-only exception;
234
- release/high-risk budget zero fails before inference.
235
-
236
- At release depth, including depth raised by a high-risk tag, only a
237
- controller-proved `markdown-punctuation-only` check may use this exception.
238
- `markdown-token-replacement` remains available for explore/build budget-zero
239
- review, but it cannot waive the release/high-risk challenge: byte-exact token
240
- replacement does not prove that the old and new tokens have the same meaning.
241
-
242
- The proof is a single-link regular UTF-8 JSON file of at most 16,000 bytes:
243
-
244
- ```json
245
- {"schema_version":1,"candidate_sha256":"<packet-sha256>","check":{"kind":"markdown-punctuation-only"}}
246
- ```
247
-
248
- The other fixed check is
249
- `markdown-token-replacement`, whose `check` also contains `old_token`,
250
- `new_token`, and integer `expected_count` (1..100). The controller never trusts
251
- a caller-supplied pass result. It reparses the frozen packet and derives the
252
- status, files, changed-line count, replacement count, and scope SHA-256.
253
-
254
- The accepted packet is deliberately narrow: a canonical full-context unified
255
- Git diff, LF-terminated, at most 200,000 bytes, changing existing regular
256
- Markdown files inside exactly one existing non-linked skill package. Every
257
- file's first hunk starts at line 1 so frontmatter is inspectable. Adds,
258
- deletes, renames, multi-skill changes, frontmatter or `description` edits,
259
- non-regular Git modes, custom/compact packets, extra context outside the diff,
260
- and files without a final newline fail closed. `markdown-punctuation-only`
261
- accepts only one-for-one plain-prose line replacements whose non-punctuation
262
- characters remain identical; numeric tokens must additionally survive
263
- byte-for-byte (deleting the dot in `5.5` is a threshold change, not
264
- punctuation), and a question mark may not be added or removed (a statement
265
- turned into a question is a meaning change). Lines must start at column zero
266
- and contain prose; line adds/deletes, Markdown headings, lists, block quotes,
267
- links, tables, inline code, fenced or indented code, and raw HTML `pre`/`code`
268
- containers fail closed. `markdown-token-replacement` requires every changed
269
- line pair to differ only by the named whole-token replacement, with the exact
270
- total count, and rejects packets whose changed lines touch a Markdown or HTML
271
- code container.
272
-
273
- This recipe produces the exact packet and proof without a second parser or a
274
- pretend verifier command. Set `WORDING_KIND=markdown-punctuation-only`, or set
275
- `WORDING_KIND=markdown-token-replacement` plus `WORDING_OLD`, `WORDING_NEW`, and
276
- `WORDING_COUNT`:
277
-
278
- ```bash
279
- : "${CODE_REVIEW_SKILL_DIR:?set the installed code-review skill directory}"
280
- : "${REPO_ROOT:?set the absolute repository root}"
281
- : "${REVIEW_BASE:?set the exact base ref}"
282
- : "${SKILL_NAME:?set the one existing skill package name}"
283
- : "${REVIEW_STAGE:?set explore, build, or release}"
284
- : "${IMPLEMENTER_FAMILY:?set the implementer model family}"
285
- : "${REVIEW_PLAN_FILE:?set the absolute review-plan JSON path}"
286
- : "${REVIEW_EVIDENCE_DIR:?set an existing durable private evidence directory}"
287
- : "${WORDING_KIND:?set one supported wording-only check kind}"
288
-
289
- umask 077
290
- WORDING_RUN_DIR="$(mktemp -d "$REVIEW_EVIDENCE_DIR/wording-review.XXXXXX")" || exit 1
291
- WORDING_DIFF="$WORDING_RUN_DIR/candidate.diff"
292
- WORDING_PROOF="$WORDING_RUN_DIR/proof.json"
293
- WORDING_RESULT="$WORDING_RUN_DIR/review.json"
294
-
295
- git -C "$REPO_ROOT" diff --no-color --no-ext-diff --no-textconv --full-index \
296
- --src-prefix=a/ --dst-prefix=b/ --unified=1000000 \
297
- "$REVIEW_BASE" -- "skills/$SKILL_NAME" >"$WORDING_DIFF" || exit 1
298
-
299
- python3 - "$WORDING_DIFF" "$WORDING_PROOF" "$WORDING_KIND" \
300
- "${WORDING_OLD:-}" "${WORDING_NEW:-}" "${WORDING_COUNT:-0}" <<'PY'
301
- import hashlib
302
- import json
303
- import sys
304
- from pathlib import Path
305
-
306
- diff_path, proof_path = map(Path, sys.argv[1:3])
307
- kind, old, new, count = sys.argv[3:]
308
- check = {"kind": kind}
309
- if kind == "markdown-token-replacement":
310
- check.update(old_token=old, new_token=new, expected_count=int(count))
311
- elif kind != "markdown-punctuation-only":
312
- raise SystemExit("unsupported WORDING_KIND")
313
- payload = {
314
- "schema_version": 1,
315
- "candidate_sha256": hashlib.sha256(diff_path.read_bytes()).hexdigest(),
316
- "check": check,
317
- }
318
- proof_path.write_text(
319
- json.dumps(payload, ensure_ascii=False, separators=(",", ":")) + "\n",
320
- encoding="utf-8",
321
- )
322
- PY
323
-
324
- WORDING_RISK_ARGS=()
325
- for tag in ${REVIEW_RISK_TAGS:-}; do WORDING_RISK_ARGS+=(--risk-tag "$tag"); done
326
- if ! bash "$CODE_REVIEW_SKILL_DIR/scripts/review_gate.sh" \
327
- --mode review --stage "$REVIEW_STAGE" --challenge-budget 0 \
328
- --cwd "$REPO_ROOT" --diff-file "$WORDING_DIFF" \
329
- --review-plan-file "$REVIEW_PLAN_FILE" \
330
- --wording-only-proof-file "$WORDING_PROOF" \
331
- ${WORDING_RISK_ARGS[@]+"${WORDING_RISK_ARGS[@]}"} \
332
- --implementer-family "$IMPLEMENTER_FAMILY" >"$WORDING_RESULT"; then
333
- cat "$WORDING_RESULT" >&2
334
- exit 1
335
- fi
336
- cat "$WORDING_RESULT"
337
- ```
338
-
339
- A valid result carries `wording_only_proof_sha256`, controller-derived
340
- `wording_only_scope.status=passed`, and a reviewed
341
- `wording_only_boundary` concern. That concern independently confirms the edit
342
- changes no trigger, scope, routing, validation, acceptance, rule, threshold,
343
- boundary, frontmatter, description, or other meaning. If it is missing,
344
- inconclusive, or reports a possible semantic change, the wording-only exception
345
- does not apply: use the normal challenge and behavioral-evidence path. Any
346
- candidate edit regenerates the packet and proof and requires a new review.
347
- Keep the diff, proof, and result together; a digest whose source artifact was
348
- deleted is not independently auditable evidence.
256
+ The wording-only exception — its depth limits, the proof schema, the accepted
257
+ packet, the recipe that produces both, and how a valid result is read — is
258
+ specified in `wording-only-review.md`.
349
259
 
350
260
  ## Agent review chain
351
261
 
@@ -362,8 +272,10 @@ index 1; an untracked initial review is single-round and therefore uses budget 0
362
272
  candidate can never be challenged inside it. One succeeding chain may open at
363
273
  index 1 in `challenge` mode by supplying `--predecessor-chain-result-file` — the
364
274
  ended chain's terminal receipt — instead of an in-chain prior result. The
365
- controller accepts it only when that receipt is a tracked challenge at its own
366
- chain's terminal index, carries this chain's `review_scope_sha256` and matching
275
+ controller accepts it only when that receipt is the tracked round its chain ended
276
+ on — a terminal challenge, or a round-1 review whose own
277
+ arithmetic still reports its challenge unspent — carrying this chain's
278
+ `review_scope_sha256` and matching
367
279
  stage/depth/risk-tags/budget, preserves the controller digest, owner-selection
368
280
  source, and selected owner names, and binds a candidate that DIFFERS from this
369
281
  packet: the owner-package digest is the one binding allowed to move, because its
@@ -375,6 +287,24 @@ differ from every focus the ended chain spent. The result records
375
287
  satisfied self-review trigger. Succession carries history rather than resetting
376
288
  it: consumers still sum rounds across both chains.
377
289
 
290
+ A chain ends where the candidate moves, and a fix applied straight after the review
291
+ moves the owner digest exactly as one applied after the challenge does. Requiring a
292
+ challenge receipt here never protected the landing candidate — the succession
293
+ challenge binds that either way — it only forced the challenge to be spent on a
294
+ candidate the author had already decided to replace. The single class that stops
295
+ being owed is a challenge on a candidate that will never land, which carries no
296
+ evidence about the one that does; every other binding is unchanged, the candidate
297
+ must still have moved, succession still does not compose, and this path spends
298
+ fewer rounds than the old one, never more. What bounds it is the receipt's own arithmetic, and that is a
299
+ forgery guard rather than a history check: a genuine round-1 review reads the same
300
+ whether its chain later ran a challenge or not, so a caller who spent the challenge
301
+ and presents only the review is accepted, and the successor inherits no challenge
302
+ focuses — a focus that chain did spend can be spent again. This is the same
303
+ omitted-history boundary the rest of this contract states rather than a new one, and
304
+ the closeout validator's ordered receipt set is where a retained challenge receipt
305
+ would show it; no check at the succession call site can close it, and none is
306
+ claimed.
307
+
378
308
  The chain binds task scope, candidate identity per round, result hashes, mode,
379
309
  status, challenge focus, controller, and selected owners. The opaque
380
310
  `review_scope_sha256` always hashes normalized intent, acceptance, stage/depth,
@@ -430,6 +360,26 @@ checkpoint, and before a completion claim. Findings never produce a blind
430
360
  review-fix-review loop: they block another reviewer call, return to implementer
431
361
  triage, and still allow implementation, tests, and independent runnable work.
432
362
 
363
+ **Findings that come back are a design question.** When a round returns findings and
364
+ the history it carries already holds one — an earlier round of this chain, or the
365
+ predecessor chain a succession names — the gate adds
366
+ `recurring_findings_design_check` to the required triggers and
367
+ `decide_keep_delete_narrow_replace` to the allowed actions. It blocks nothing that
368
+ `findings_returned` does not already block; what it adds is the question the next
369
+ patch would walk past: whether the reviewed surface should exist in this shape at
370
+ all, answered as `keep`, `delete`, `narrow`, or `replace`, resting on the rounds and
371
+ findings it recurred across, and ratified by a risk owner other than the one
372
+ proposing it. `../../skill-extraction-workflow/SKILL.md` owns that rule; this is where
373
+ it fires, because the situation arises inside a chain and that skill is usually not
374
+ loaded there. Two findings rounds need not share a class, so the trigger over-fires
375
+ by design — answering an inapplicable question costs a line, and the round it saves
376
+ does not.
377
+
378
+ The count is what the controller can prove, and no more: the rounds of this chain plus
379
+ the predecessor a succession names, which is why the trigger reaches across a chain
380
+ break at all (Chain succession, below). Succession does not compose, so a third chain
381
+ opened fresh carries no history and the recurrence becomes the round's own record.
382
+
433
383
  A passed final external round returns
434
384
  `next_action=deep_self_review_before_completion` and remains
435
385
  `completion_gated=true`. `--mode complete` accepts one exact-candidate passed
@@ -0,0 +1,136 @@
1
+ # Proof-bound wording-only single review
2
+
3
+ The wording-only exception to `staged-review-contract.md`: when one review may
4
+ stand in for the review-plus-challenge pair, what the controller re-derives
5
+ before it will say so, and how to produce the packet and proof it accepts.
6
+
7
+ ## The exception and its depth limits
8
+
9
+ The wording-only exception is one untracked `review` with
10
+ `challenge_budget=0`; it is not a chain, challenge, or `complete` checkpoint.
11
+ Supply `--wording-only-proof-file` to bind the exception to the exact packet.
12
+ Without that proof, an explore/build budget-zero review remains an ordinary
13
+ single review and cannot be recorded as the wording-only exception;
14
+ release/high-risk budget zero fails before inference.
15
+
16
+ At release depth, including depth raised by a high-risk tag, only a
17
+ controller-proved `markdown-punctuation-only` check may use this exception.
18
+ `markdown-token-replacement` remains available for explore/build budget-zero
19
+ review, but it cannot waive the release/high-risk challenge: byte-exact token
20
+ replacement does not prove that the old and new tokens have the same meaning.
21
+
22
+ ## The proof
23
+
24
+ The proof is a single-link regular UTF-8 JSON file of at most 16,000 bytes:
25
+
26
+ ```json
27
+ {"schema_version":1,"candidate_sha256":"<packet-sha256>","check":{"kind":"markdown-punctuation-only"}}
28
+ ```
29
+
30
+ The other fixed check is
31
+ `markdown-token-replacement`, whose `check` also contains `old_token`,
32
+ `new_token`, and integer `expected_count` (1..100). The controller never trusts
33
+ a caller-supplied pass result. It reparses the frozen packet and derives the
34
+ status, files, changed-line count, replacement count, and scope SHA-256.
35
+
36
+ ## The accepted packet
37
+
38
+ The accepted packet is deliberately narrow: a canonical full-context unified
39
+ Git diff, LF-terminated, at most 200,000 bytes, changing existing regular
40
+ Markdown files inside exactly one existing non-linked skill package. Every
41
+ file's first hunk starts at line 1 so frontmatter is inspectable. Adds,
42
+ deletes, renames, multi-skill changes, frontmatter or `description` edits,
43
+ non-regular Git modes, custom/compact packets, extra context outside the diff,
44
+ and files without a final newline fail closed. `markdown-punctuation-only`
45
+ accepts only one-for-one plain-prose line replacements whose non-punctuation
46
+ characters remain identical; numeric tokens must additionally survive
47
+ byte-for-byte (deleting the dot in `5.5` is a threshold change, not
48
+ punctuation), and a question mark may not be added or removed (a statement
49
+ turned into a question is a meaning change). Lines must start at column zero
50
+ and contain prose; line adds/deletes, Markdown headings, lists, block quotes,
51
+ links, tables, inline code, fenced or indented code, and raw HTML `pre`/`code`
52
+ containers fail closed. `markdown-token-replacement` requires every changed
53
+ line pair to differ only by the named whole-token replacement, with the exact
54
+ total count, and rejects packets whose changed lines touch a Markdown or HTML
55
+ code container.
56
+
57
+ ## Producing the packet and proof
58
+
59
+ This recipe produces the exact packet and proof without a second parser or a
60
+ pretend verifier command. Set `WORDING_KIND=markdown-punctuation-only`, or set
61
+ `WORDING_KIND=markdown-token-replacement` plus `WORDING_OLD`, `WORDING_NEW`, and
62
+ `WORDING_COUNT`:
63
+
64
+ ```bash
65
+ : "${CODE_REVIEW_SKILL_DIR:?set the installed code-review skill directory}"
66
+ : "${REPO_ROOT:?set the absolute repository root}"
67
+ : "${REVIEW_BASE:?set the exact base ref}"
68
+ : "${SKILL_NAME:?set the one existing skill package name}"
69
+ : "${REVIEW_STAGE:?set explore, build, or release}"
70
+ : "${IMPLEMENTER_FAMILY:?set the implementer model family}"
71
+ : "${REVIEW_PLAN_FILE:?set the absolute review-plan JSON path}"
72
+ : "${REVIEW_EVIDENCE_DIR:?set an existing durable private evidence directory}"
73
+ : "${WORDING_KIND:?set one supported wording-only check kind}"
74
+
75
+ umask 077
76
+ WORDING_RUN_DIR="$(mktemp -d "$REVIEW_EVIDENCE_DIR/wording-review.XXXXXX")" || exit 1
77
+ WORDING_DIFF="$WORDING_RUN_DIR/candidate.diff"
78
+ WORDING_PROOF="$WORDING_RUN_DIR/proof.json"
79
+ WORDING_RESULT="$WORDING_RUN_DIR/review.json"
80
+
81
+ git -C "$REPO_ROOT" diff --no-color --no-ext-diff --no-textconv --full-index \
82
+ --src-prefix=a/ --dst-prefix=b/ --unified=1000000 \
83
+ "$REVIEW_BASE" -- "skills/$SKILL_NAME" >"$WORDING_DIFF" || exit 1
84
+
85
+ python3 - "$WORDING_DIFF" "$WORDING_PROOF" "$WORDING_KIND" \
86
+ "${WORDING_OLD:-}" "${WORDING_NEW:-}" "${WORDING_COUNT:-0}" <<'PY'
87
+ import hashlib
88
+ import json
89
+ import sys
90
+ from pathlib import Path
91
+
92
+ diff_path, proof_path = map(Path, sys.argv[1:3])
93
+ kind, old, new, count = sys.argv[3:]
94
+ check = {"kind": kind}
95
+ if kind == "markdown-token-replacement":
96
+ check.update(old_token=old, new_token=new, expected_count=int(count))
97
+ elif kind != "markdown-punctuation-only":
98
+ raise SystemExit("unsupported WORDING_KIND")
99
+ payload = {
100
+ "schema_version": 1,
101
+ "candidate_sha256": hashlib.sha256(diff_path.read_bytes()).hexdigest(),
102
+ "check": check,
103
+ }
104
+ proof_path.write_text(
105
+ json.dumps(payload, ensure_ascii=False, separators=(",", ":")) + "\n",
106
+ encoding="utf-8",
107
+ )
108
+ PY
109
+
110
+ WORDING_RISK_ARGS=()
111
+ for tag in ${REVIEW_RISK_TAGS:-}; do WORDING_RISK_ARGS+=(--risk-tag "$tag"); done
112
+ if ! bash "$CODE_REVIEW_SKILL_DIR/scripts/review_gate.sh" \
113
+ --mode review --stage "$REVIEW_STAGE" --challenge-budget 0 \
114
+ --cwd "$REPO_ROOT" --diff-file "$WORDING_DIFF" \
115
+ --review-plan-file "$REVIEW_PLAN_FILE" \
116
+ --wording-only-proof-file "$WORDING_PROOF" \
117
+ ${WORDING_RISK_ARGS[@]+"${WORDING_RISK_ARGS[@]}"} \
118
+ --implementer-family "$IMPLEMENTER_FAMILY" >"$WORDING_RESULT"; then
119
+ cat "$WORDING_RESULT" >&2
120
+ exit 1
121
+ fi
122
+ cat "$WORDING_RESULT"
123
+ ```
124
+
125
+ ## Validating the result
126
+
127
+ A valid result carries `wording_only_proof_sha256`, controller-derived
128
+ `wording_only_scope.status=passed`, and a reviewed
129
+ `wording_only_boundary` concern. That concern independently confirms the edit
130
+ changes no trigger, scope, routing, validation, acceptance, rule, threshold,
131
+ boundary, frontmatter, description, or other meaning. If it is missing,
132
+ inconclusive, or reports a possible semantic change, the wording-only exception
133
+ does not apply: use the normal challenge and behavioral-evidence path. Any
134
+ candidate edit regenerates the packet and proof and requires a new review.
135
+ Keep the diff, proof, and result together; a digest whose source artifact was
136
+ deleted is not independently auditable evidence.
@@ -14,6 +14,7 @@ from pathlib import Path, PurePosixPath
14
14
  import signal
15
15
  import stat
16
16
  import subprocess
17
+ import sys
17
18
  import tempfile
18
19
  import time
19
20
  from typing import Any
@@ -370,6 +371,10 @@ STAGE_CONCERNS = {
370
371
  "compatibility",
371
372
  "Compatibility, maintainability, and unnecessary-complexity regressions.",
372
373
  ),
374
+ (
375
+ "claim_strength",
376
+ "Claims the cited evidence does not carry: absolutes, universals, causal statements, exhaustiveness.",
377
+ ),
373
378
  ),
374
379
  "release": (
375
380
  ("correctness", "Functional correctness and acceptance coverage."),
@@ -397,6 +402,10 @@ STAGE_CONCERNS = {
397
402
  "observability_operations",
398
403
  "Operational visibility, diagnosis, support, and recovery evidence.",
399
404
  ),
405
+ (
406
+ "claim_strength",
407
+ "Claims the cited evidence does not carry: absolutes, universals, causal statements, exhaustiveness.",
408
+ ),
400
409
  ),
401
410
  }
402
411
  HIGH_RISK_TAGS = {
@@ -1730,25 +1739,56 @@ def _validate_chain_succession(
1730
1739
  if prior.get("predecessor_chain_id") is not None:
1731
1740
  reject("predecessor is itself a succession round; succession does not compose")
1732
1741
  prior_budget = prior.get("challenge_budget")
1742
+ prior_mode = prior.get("mode")
1733
1743
  if (
1734
1744
  prior.get("schema_version") != 3
1735
- or prior.get("mode") != "challenge"
1745
+ or prior_mode not in ("review", "challenge")
1736
1746
  or prior.get("status") not in ("passed", "findings")
1737
1747
  or prior.get("review_chain_tracked") is not True
1738
1748
  or not isinstance(prior_budget, int)
1739
1749
  or isinstance(prior_budget, bool)
1740
1750
  or prior_budget < 1
1741
1751
  ):
1742
- reject("predecessor is not a tracked challenge receipt")
1743
- if (
1744
- prior.get("autonomous_review_index") != prior_budget + 1
1745
- or prior.get("challenge_index") != prior_budget
1746
- or prior.get("autonomous_reviews_remaining") != 0
1747
- or prior.get("autonomous_review_allowed") is not False
1748
- ):
1749
- # Terminality is the receipt's own arithmetic, not just its index: a
1750
- # forged receipt can carry a terminal index while every other field still
1751
- # says the chain has rounds left.
1752
+ reject("predecessor is not a tracked review or challenge receipt")
1753
+ # Terminality is the receipt's own arithmetic, not just its index: a forged
1754
+ # receipt can carry a terminal index while every other field still says the
1755
+ # chain has rounds left.
1756
+ #
1757
+ # A chain ends where the candidate moves, and a fix applied straight after the
1758
+ # REVIEW moves the owner digest exactly as one applied after the challenge
1759
+ # does, so a review can be the last round its chain ever had. Requiring a
1760
+ # challenge receipt here did not protect the landing candidate -- the
1761
+ # succession challenge binds that either way -- it only forced the challenge to
1762
+ # be spent on a candidate the author had already decided to replace. The one
1763
+ # class this stops owing is a challenge on a candidate that will never land,
1764
+ # which carries no evidence about what does. Everything else is unchanged: the
1765
+ # candidate must still have moved, succession still does not compose, and the
1766
+ # per-chain budget is untouched (this path spends fewer rounds, never more).
1767
+ if prior_mode == "challenge":
1768
+ chain_ended = (
1769
+ prior.get("autonomous_review_index") == prior_budget + 1
1770
+ and prior.get("challenge_index") == prior_budget
1771
+ and prior.get("autonomous_reviews_remaining") == 0
1772
+ and prior.get("autonomous_review_allowed") is False
1773
+ )
1774
+ else:
1775
+ # The review is round 1 with its chain's challenge still unspent -- as the
1776
+ # receipt itself reports it. This is a FORGERY guard, not a history check:
1777
+ # a genuine round-1 review reads the same whether its chain later ran a
1778
+ # challenge or not, because a stateless controller sees only the receipt it
1779
+ # is handed. A caller who spent the challenge and presents only the review
1780
+ # is therefore accepted here, and the successor inherits no challenge
1781
+ # focuses, so a focus that chain really did spend can be spent again. That
1782
+ # is the same omitted-history boundary the rest of this contract states,
1783
+ # and the closeout validator's ordered receipt set is where a retained
1784
+ # challenge receipt would show it; nothing at this call site can close it.
1785
+ chain_ended = (
1786
+ prior.get("autonomous_review_index") == 1
1787
+ and prior.get("challenge_index") == 0
1788
+ and prior.get("autonomous_reviews_remaining") == prior_budget
1789
+ and prior.get("autonomous_review_allowed") is True
1790
+ )
1791
+ if not chain_ended:
1752
1792
  reject("predecessor is not its chain's terminal round")
1753
1793
  predecessor_chain_id = prior.get("review_chain_id")
1754
1794
  if not isinstance(predecessor_chain_id, str) or not predecessor_chain_id.strip():
@@ -1800,6 +1840,10 @@ def _validate_chain_succession(
1800
1840
  "result_sha256": result_hash,
1801
1841
  "candidate_sha256": prior_candidate_hash,
1802
1842
  "focuses": focuses,
1843
+ # The budget is one review plus one challenge, so a fix ends the chain and the
1844
+ # second findings round lands HERE rather than in-chain. Carrying the ended
1845
+ # chain's verdict is what lets the recurrence be counted at all.
1846
+ "returned_findings": prior.get("status") == "findings",
1803
1847
  }
1804
1848
 
1805
1849
 
@@ -3161,6 +3205,7 @@ def freeze_review_profile(
3161
3205
  previous_challenge_focuses: list[str] = []
3162
3206
  prior_review_result_hashes: list[str] = []
3163
3207
  prior_review_candidate_hashes: list[str] = []
3208
+ prior_findings_rounds = 0
3164
3209
  succession: dict[str, Any] | None = None
3165
3210
  inherited_challenge_focuses: list[str] = []
3166
3211
  if review_chain_tracked:
@@ -3309,6 +3354,8 @@ def freeze_review_profile(
3309
3354
  previous_challenge_focuses.append(focus)
3310
3355
  prior_review_result_hashes.append(result_hash)
3311
3356
  prior_review_candidate_hashes.append(prior_candidate_hash)
3357
+ if prior.get("status") == "findings":
3358
+ prior_findings_rounds += 1
3312
3359
  if challenge_focus and challenge_focus in (
3313
3360
  previous_challenge_focuses + inherited_challenge_focuses
3314
3361
  ):
@@ -3329,6 +3376,9 @@ def freeze_review_profile(
3329
3376
  "review_chain_required",
3330
3377
  )
3331
3378
 
3379
+ if succession is not None and succession["returned_findings"]:
3380
+ prior_findings_rounds += 1
3381
+
3332
3382
  self_review_satisfied_triggers: list[str] = []
3333
3383
  if args.mode in ("review", "challenge"):
3334
3384
  self_review_satisfied_triggers.append("before_external_review")
@@ -3401,6 +3451,7 @@ def freeze_review_profile(
3401
3451
  succession["candidate_sha256"] if succession else None
3402
3452
  ),
3403
3453
  "self_review_satisfied_triggers": self_review_satisfied_triggers,
3454
+ "prior_findings_rounds": prior_findings_rounds,
3404
3455
  "required_concerns": [
3405
3456
  {"id": concern_id, "description": description}
3406
3457
  for concern_id, description in reviewer_concern_pairs
@@ -4168,7 +4219,52 @@ def build_parser() -> argparse.ArgumentParser:
4168
4219
  return parser
4169
4220
 
4170
4221
 
4222
+ PRINT_REQUIRED_CONCERNS_FLAG = "--print-required-concerns"
4223
+
4224
+
4225
+ def _print_required_concerns(argv: list[str]) -> int:
4226
+ """Print the concern ids a review plan must cover, one per line.
4227
+
4228
+ The plan's required set is derived from the stage and the risk tags, and a
4229
+ caller that hardcodes its own copy of that list drifts the moment the set
4230
+ changes -- measured as five suites whose fixtures stopped satisfying the gate
4231
+ when one concern was added, none of which the fast lane could report because
4232
+ the runner aborts at its first failing target. The list has exactly one owner;
4233
+ this prints it so callers derive instead of duplicating.
4234
+
4235
+ Deliberately narrower than the reviewer's concern set: this answers what the
4236
+ PLAN owes, so the synthetic challenge slot and the wording-only boundary --
4237
+ which the controller adds for the reviewer, never for the plan -- are absent.
4238
+ """
4239
+ parser = argparse.ArgumentParser(prog="review_gate.py", add_help=True)
4240
+ parser.add_argument(PRINT_REQUIRED_CONCERNS_FLAG, action="store_true", required=True)
4241
+ parser.add_argument("--stage", choices=("explore", "build", "release"), default="build")
4242
+ parser.add_argument("--risk-tag", action="append", default=[])
4243
+ args = parser.parse_args(argv)
4244
+ # The same tag validation the run path applies. Without it the printer answers for
4245
+ # inputs the enforcer refuses, which is the printer/enforcer divergence this export
4246
+ # exists to remove -- a caller deriving from a malformed tag would get a list where
4247
+ # the real round fails closed.
4248
+ for index, tag in enumerate(args.risk_tag):
4249
+ if not tag or len(tag) > 80 or any(ch.isspace() for ch in tag):
4250
+ parser.error(f"invalid risk tag at index {index}")
4251
+ stage_rank = {"explore": 0, "build": 1, "release": 2}
4252
+ depth = "release" if HIGH_RISK_TAGS.intersection(args.risk_tag) else args.stage
4253
+ if stage_rank[depth] < stage_rank[args.stage]:
4254
+ depth = args.stage
4255
+ ids = [concern_id for concern_id, _ in STAGE_CONCERNS[depth]]
4256
+ if HIGH_RISK_TAGS.intersection(args.risk_tag):
4257
+ ids.append("high_risk_boundary")
4258
+ for concern_id in ids:
4259
+ print(concern_id)
4260
+ return 0
4261
+
4262
+
4171
4263
  def main(argv: list[str] | None = None) -> int:
4264
+ if PRINT_REQUIRED_CONCERNS_FLAG in (sys.argv[1:] if argv is None else argv):
4265
+ # Answered before the run parser, which requires --mode/--cwd/--implementer-family
4266
+ # for an actual review; asking what a plan owes needs none of them.
4267
+ return _print_required_concerns(sys.argv[1:] if argv is None else argv)
4172
4268
  script_dir = Path(__file__).resolve().parent
4173
4269
  packet_path: Path | None = None
4174
4270
  profile_path: Path | None = None
@@ -4560,6 +4656,7 @@ def main(argv: list[str] | None = None) -> int:
4560
4656
  "deep_self_review",
4561
4657
  "continue_implementation",
4562
4658
  ]
4659
+ recurring_findings = profile["prior_findings_rounds"] > 0
4563
4660
  if result["autonomous_review_allowed"]:
4564
4661
  next_action = "implementer_self_review"
4565
4662
  review_state = "findings_pending"
@@ -4573,6 +4670,18 @@ def main(argv: list[str] | None = None) -> int:
4573
4670
  )
4574
4671
  allowed_self_review_actions.append("continue_independent_work")
4575
4672
  allowed_self_review_actions.append("resolve_review_findings")
4673
+ if recurring_findings:
4674
+ # Findings have now come back across rounds. The next patch is
4675
+ # not the default move: decide whether the reviewed surface
4676
+ # should exist in this shape at all. Two rounds of findings need
4677
+ # not share a class, so this over-fires by design -- answering an
4678
+ # inapplicable question is cheap, and the miss it prevents is not.
4679
+ required_self_review_triggers.append(
4680
+ "recurring_findings_design_check"
4681
+ )
4682
+ allowed_self_review_actions.append(
4683
+ "decide_keep_delete_narrow_replace"
4684
+ )
4576
4685
  current_self_review_gate = self_review_gate(
4577
4686
  required_triggers=required_self_review_triggers,
4578
4687
  satisfied_triggers=profile["self_review_satisfied_triggers"],