@ccoalm/ccl-skills 0.11.0 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +21 -12
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +42 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +7 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-context-freshness.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +20 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/retrieval-agent-safety.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +13 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/references/multi-agent-delegation-playbook.md +12 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +364 -23
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +185 -0
- package/dist/assets/release.json +20 -20
- package/package.json +1 -1
|
@@ -524,7 +524,7 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
|
|
|
524
524
|
|
|
525
525
|
The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
|
|
526
526
|
|
|
527
|
-
Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **The lane spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There the autonomous lane ends. An authenticated human may request later review, but that is separately attributed human-requested evidence outside this chain/budget, never an additional Agent round. Unused generic capacity never authorizes automatic continuation. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
|
|
527
|
+
Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **The lane spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There the autonomous lane ends. An authenticated human may request later review, but that is separately attributed human-requested evidence outside this chain/budget, never an additional Agent round. Unused generic capacity never authorizes automatic continuation. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
|
|
528
528
|
|
|
529
529
|
**Self-hosted chains break on every fix; the budget is summed across chains, never per chain.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break spends the SAME Agent-autonomous budget. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
|
|
530
530
|
|
|
@@ -119,7 +119,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
119
119
|
- File: `~/.<host>/skills/.extraction-work/<project>-completion.md`
|
|
120
120
|
- Final state: which batches done, which deferred, which sources unavailable.
|
|
121
121
|
- Lessons: what surprised; what would change in next extraction; what to add to skill-extraction-workflow.
|
|
122
|
-
- For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. Ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate hashes to first: if the held fix batch moved it, the ledger owes the succession challenge bound to that hash, and the same script is the merge-side gate that refuses a landing whose evidence binds a different candidate. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row and does not fabricate a multi-round ledger — but note the cost the merge-side gate imposes on it: that gate accepts only a validator-checked ledger, because it cannot authenticate a hand-writable receipt, so a wording-only change that touches the bound paths still owes the two-round chain before it can land.
|
|
122
|
+
- For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. Ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate hashes to first: if the held fix batch moved it, the ledger owes the succession challenge bound to that hash, and the same script is the merge-side gate that refuses a landing whose evidence binds a different candidate. When the whole candidate exceeds one packet, split the review rather than the pull request: `--print-manifest --partition <paths> ...` renders a landing partition manifest, and one validated ledger per partition plus the committed manifest is what the gate binds. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row and does not fabricate a multi-round ledger — but note the cost the merge-side gate imposes on it: that gate accepts only a validator-checked ledger, because it cannot authenticate a hand-writable receipt, so a wording-only change that touches the bound paths still owes the two-round chain before it can land.
|
|
123
123
|
|
|
124
124
|
### 5. Provenance migration
|
|
125
125
|
|
|
@@ -12,6 +12,10 @@
|
|
|
12
12
|
- **Anthropic, "Effective context engineering for AI agents"** (2025, anthropic.com/engineering/effective-context-engineering-for-ai-agents) — context 即稀缺资源 / compaction / 结构化笔记 / JIT 检索 / context rot(提炼入 §5)
|
|
13
13
|
- **Anthropic, "Demystifying evals for AI agents"** (Jan 09 2026) — agent eval 方法(realistic tasks / robust criteria / multiple graders / transcripts)
|
|
14
14
|
- **Anthropic Claude Code sub-agent docs**(Claude Code 官方文档 sub-agents 段)— sub-agent 隔离 / 独立 context window / 独立 permission
|
|
15
|
+
- **Anthropic, "Building multi-agent systems: When and how to use them"** (claude.com/blog, Jan 23 2026) — single-agent-first 三问(真实约束 / 按 context 而非按角色拆 / 有清晰验证点)、3–10× token 溢价、verification-subagent 模式(提炼入 `multi-agent-delegation` 的 fan-out gate)
|
|
16
|
+
- **Google Research, "Towards a science of scaling agent systems"** (arXiv 2512.08296, Dec 2025) — 并行可分解任务下集中式协调收益大;严格顺序推理任务所有多 agent 变体退化;独立并行 agent 的错误放大远高于带 orchestrator 的 hub;单 agent 基线已强时协调收益递减或为负
|
|
17
|
+
- **Cemri et al., "Why Do Multi-Agent LLM Systems Fail?"** (arXiv 2503.13657, NeurIPS 2025) — MAST:3 类 14 种失败模式(§2 映射表);多数失败源于系统设计而非模型
|
|
18
|
+
- **Cognition, "Don't Build Multi-Agents"** (Jun 2025) — 两原则:共享完整 trace 而非摘要;动作携带隐式决策、冲突决策产坏结果(提炼入 fan-out gate 的 shared-decisions 项)
|
|
15
19
|
- **SWE-bench** (Jimenez et al., arXiv 2023, ICLR 2024, Princeton + UChicago) — agent 在真 GitHub issues 上的 task replay eval
|
|
16
20
|
- **Aider benchmarks / leaderboard** (aider.chat/docs/leaderboards/) — code editing/refactoring 固定 task set + pass rate
|
|
17
21
|
- **AutoGen** (Microsoft Research, 2023) — multi-agent conversation framework
|
|
@@ -75,6 +79,16 @@ Workflow 更适合可预测 / 可调试 / 可控成本的任务;agent 更适
|
|
|
75
79
|
- `multi-agent-delegation` skill 主体覆盖 isolation 决策;本 ref 补充失败模式 checklist
|
|
76
80
|
- 用 sub-agent 后**必须独立核对结果**(reading diff / grep specific changes),不依赖 sub-agent self-report
|
|
77
81
|
|
|
82
|
+
- **与文献标准命名(MAST)的对应**——上表是自用名;评审一次失败的 worker 返回时必须先按 MAST 类别定位该修哪一层(MAST 的结论:多数失败源于系统设计而非模型,先改 brief / 拓扑 / 验证,别先换模型):
|
|
83
|
+
|
|
84
|
+
| MAST 类 | 失败模式 | 对应上表 / 我们的规则 | 修哪层 |
|
|
85
|
+
|---|---|---|---|
|
|
86
|
+
| FC1 系统设计 / 规格 | 1.1 违背任务规格;1.2 违背角色规格;1.3 步骤重复;1.4 丢失对话历史;1.5 不知终止条件 | Context starvation;`multi-agent-delegation` 的 spec-compliance review、同错 ~3 次升级、wall-clock deadline、stop line | brief(目标 / 边界 / 输出形状 / effort budget)或拓扑 |
|
|
87
|
+
| FC2 agent 间失配 | 2.1 对话重置;2.2 该问不问;2.3 任务跑偏;2.4 扣留信息;2.5 忽略他方输入;2.6 推理-动作不一致 | Hidden dependency;brief 逐字携带全局约束与邻接契约、5 字段 escalation、owned-path manifest 核对、integrate 步查隐式决策分歧 | brief 契约 / escalation / 集成检查 |
|
|
88
|
+
| FC3 任务验证 | 3.1 过早终止;3.2 无 / 不完整验证;3.3 错误验证 | Trust drift;不信 success report、blocked 声明先补救、reviewer 的 verdict_scope / cannot_verify 槽位 | 控制器侧验证 |
|
|
89
|
+
|
|
90
|
+
Result inflation 没有 MAST 对应——它是 context / 成本问题,不是任务失败。
|
|
91
|
+
|
|
78
92
|
---
|
|
79
93
|
|
|
80
94
|
## §3 Skill effectiveness eval(task replay / before-after / golden trace)
|
|
@@ -520,3 +520,41 @@ Round 073-receipt-bundling rows (new table so the entry renders as a table row a
|
|
|
520
520
|
| A probe that cannot tell "this checkout cannot be measured" from "the thing being measured is broken" is deleted, not patched again: the repository checker runs against synthetic fixtures inside other suites, and a smoke that reds there fails suites that have nothing to do with it | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/check-ccl-skills.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/check-ccl-skills.sh). Supersedes by pointer the row above that added this smoke. Observed failure, three times in one round and each time in a suite that does not own the gate: the catalog suite, then route-drift and sync-pointers together, then the source-register lifecycle suite — every one of them runs the checker against a synthetic repository where the smoke legitimately cannot operate, and the third failure additionally exposed that its capture was not `set -e` safe, so the checker died silently mid-run before printing the verdicts those suites read. Two patches had already narrowed the predicate (skip without a controller, skip without a parent commit) and a third would have narrowed it again, which is the signal this repository already records: when the same class recurs, question whether the capability should exist. It should not. The gate's own suite owns the real-checkout path with a case that runs it against the checkout it ships in, and the CI step is the enforcement point — verified green on this candidate's own pull request. What is lost is stated rather than glossed: nothing else runs the gate during `make test-repo-gates`, so a break in it surfaces at the CI step rather than locally. RED-baseline (applied): the lifecycle suite reds against the smoke-bearing checker and is green after its removal, with the gate's own eighteen-case suite unchanged in both runs. |
|
|
521
521
|
| The merge gate binds every tracked path minus exactly what this round ADDS under a round's evidence directory, and refuses a candidate tree that is not committed: a whitelist binds only the paths some round happened to review, so unreviewed executable content rode along on a valid ledger; a written-down `specs/` exclusion would additionally hide edits to the committed review history itself; and a packet frozen from a dirty working tree produces a hash no clean checkout recomputes | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, its suite, and .github/workflows/ci.yml). Observed failure: the succession round of the previous round's own lane raised the whitelist half and it was deferred with its reason -- every fix moves the candidate and voids the ledger, and that lane's budget was spent -- so it was recorded for the next round that touches this gate. This round's own review and challenge then each found that the first inversion traded one hole for another, and both land here rather than as further deferrals. The review found that the frozen packet includes untracked files, so a scratch file inside the widened set produces a candidate hash only that working copy can reproduce: the author records it in the ledger and the merge-side run then reports that nothing binds the landing candidate, refusing valid work. The challenge found that excluding all of `specs/` excludes the committed review history, so a pull request could delete or rewrite an earlier round's plan and receipts with no evidence required. The exclusion is therefore computed from the round's own diff rather than written down: only added paths under a round's evidence directory stay outside, because a receipt inside the bound set would move the hash it records, while every modification and deletion under `specs/` is bound like any other file. The succession round then broke that shape too -- an arbitrary added file under an evidence directory, a script included, was excluded for the same reason -- which is the third occurrence of one class: the rule kept naming a LOCATION and letting the location stand in for `this is a receipt`. Rather than narrow the path a fourth time, the predicate moved to an invariant this gate owns: a path is excluded only when its committed blob parses as a JSON object carrying a 64-hex `candidate_sha256`, so a script, a fixture, or an unbound JSON file committed there is bound like anything else. Backward compatibility was measured, not assumed: recomputed at the previous round's own fork point, the old and new path sets produce the identical candidate hash its committed ledger records. Supersedes by pointer the merge-queue half of the row above, which recorded that the CI step now also runs for `merge_group`: the workflow's `on:` never subscribed to that event, so in a merge-queue run the workflow would not start at all and the condition read as coverage while providing none. Restoring the trigger was rejected rather than done, because a merge_group HEAD combines several queued pull requests while each committed ledger binds one individual candidate, so no ledger binds the aggregate and every otherwise-valid queued request would be refused -- the trigger would make the sentence true and the system worse. The unreachable branch is removed and the real coverage boundary is stated where the step lives. RED-baseline (applied, differential, two mutants each attributed to its own cases and nothing else): restoring the whole-subtree exclusion reds exactly the three committed-history cases -- rewriting an earlier plan, rewriting an earlier receipt, deleting an earlier receipt; removing the committed-tree refusal reds exactly the three dirty-tree cases; degenerating the receipt predicate to always-true reds exactly the three smuggled-file cases; three unbound executable paths (a root Makefile, a README, a release script) each red against the original whitelist and are refused after; the added-evidence case stays green throughout, which is what proves the self-reference exclusion survived. Control is 34 passing with no case disturbed. A fourth round then broke the content predicate too -- a JSON file carrying any 64-hex `candidate_sha256` is accepted as a receipt -- and that one is NOT fixed, deliberately. Four shapes of this exclusion have now been broken in four rounds, and every one of them was a proxy for `this is a controller-generated receipt` over a file the candidate itself supplies, which is the already-recorded boundary that a gate living inside the candidate cannot authenticate what it reads. This repository's own standard is that a class recurring across rounds is a question about the design rather than a fifth patch, so the residual is recorded for a person: accept it, or replace the exclusion mechanism outright -- binding the tree as of the commit before the evidence lands would need no exclusion predicate at all. What did close is real: history can no longer be rewritten unnoticed, and a script or binary can no longer ride in under an evidence directory. |
|
|
522
522
|
| A succession may not carry the chain id of the chain it succeeds, and the controller refuses it at mint rather than leaving the refusal to the closeout validator: the validator only sees a lane it reads whole, while the controller mints one receipt at a time, so a caller that never closes a ledger never reaches that check | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md` (entrypoint unchanged this round; lands in scripts/review_gate.py and its suite). Observed failure: the previous round recorded this as a non-blocking deferral with its reason -- fixing it would have moved the controller digest and forfeited that round's ability to close its own ledger with the succession round it introduced. The severity recorded then is the one that holds now, and it is narrower than it first reads: this is not an open bypass, because `validate_extraction_review_state.py` already refuses a succession whose chain id equals the wrapper chain's. What lands is the same refusal at the point the receipt is made, which is the only place it applies to a controller run that never reaches a closeout. The equality direction is not inferred: the existing validator refusal uses the same predicate and the same words, so the intended semantics is that the two ids must differ. RED-baseline (applied, differential): a succession minted with its predecessor's own chain id reds against the pre-fix controller and is refused with its own diagnostic after, with the suite moving from 261 to 262 passing and no pre-existing case disturbed. |
|
|
523
|
+
| The merge gate binds a landing candidate larger than one review packet through a committed landing partition manifest: path partitions whose changed files together equal the candidate's exactly once, each recomputed with the controller's own freeze and bound by its own validated ledger, so the reviewer's byte ceiling is no longer the pull request's ceiling | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py, its suite, references/dual-track-review-gate.md, references/extraction-quickstart.md, and a .github/workflows/ci.yml comment). Observed failure: the gate defined the landing candidate as one packet hash, and the controller caps a packet at what one reviewer can read whole, so a candidate larger than that could not be frozen and no ledger could ever bind it -- a release whose whole diff was three times the ceiling had to land as eight separate merges, each splitting the pull request where the review side already permitted splitting the packet. The two identities are different sizes: base..HEAD has no natural byte limit, a reviewer's input does. A manifest committed under a round's evidence directory names the partitions and, for each, the hash that `--print-candidate --paths` already answers; the gate refuses any manifest whose parts do not add up to the whole -- a changed file in no partition, a changed file in two, a partition whose recorded hash no longer reproduces, a base other than the fork point, an aggregate hash that does not reproduce its partitions, or a partition path shaped like a pathspec. The manifest carries a top-level 64-hex `candidate_sha256` (the aggregate identity) so it satisfies the existing receipt predicate and committing it moves no partition; the exclusion predicate is unchanged and the accepted caller-controlled-evidence residual is not widened. `--print-manifest --partition ...` renders the manifest with every hash computed by the gate, so the canonical form lives in one place. Merge-queue aggregation of several pull requests into one HEAD is a different aggregate and stays unsolved, as the workflow comment now states. RED-baseline (applied, differential): the partition cases red 18 against the pre-fix gate with the existing 35 undisturbed; five in-place mutants -- coverage equality, disjointness, aggregate recomputation, base equality, partition-hash recomputation -- each red exactly their own cases (2, 2, 1, 1, 1) and nothing else, restored and verified after each. The candidate that triggered the round, measured at 623,458 bytes, renders as six freezable partitions. |
|
|
524
|
+
| The partition manifest refuses a wildcard partition path and requires the partition union to EQUAL the reviewed changed set, not merely contain it: git reads `*`, `?`, `[` and `\\` as glob syntax even in a non-magic pathspec, so a wildcard partition chooses its own coverage, and under a narrowed `--paths` scope a partition can reach changed files outside the reviewed set with no uncovered file and no overlap to refuse | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md` (entrypoint unchanged; lands in scripts/review_ledger_binding.py and its suite). Observed failure: the round's own dual-track lane found both -- the independent review reported that `validate_partition_path` rejected only leading pathspec magic while `lane-*` passed through to git as a glob, and that `partition_coverage` checked only `changed_all - owner`, so with `--paths skills/a` a partition naming `.` covered changed files outside the scope and passed; the adversarial challenge independently hit the wildcard class on the same frozen candidate. Both are the same shape: the manifest was allowed to influence what git enumerated on its behalf. Fix: refuse the metacharacters as a path (never handed to git), and refuse a union larger than the reviewed set with the surplus named. Held until the challenge ran, then applied as one batch that moved the candidate, so the lane owes and runs one succession challenge bound to what lands. RED-baseline (applied, differential): the four wildcard shapes plus the rendering case red against the pre-fix gate and are refused after; the out-of-scope case reds against the pre-fix gate with a freeze error on the oversized `.` partition and is refused before freezing after; disabling the wildcard check reds exactly the five wildcard cases and disabling the equality check reds exactly the out-of-scope case, suite otherwise at 60 passing. |
|
|
525
|
+
| A multi-component failure is localized to one boundary in a single instrumented run — entry/exit data and the received env/config logged at every component boundary, run once, the first wrong boundary owns the search — before any hypothesis fans out across the chain, because locating is the expensive phase | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#must be localized to one boundary in a single instrumented run | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (SRE troubleshooting chapter: simplify and reduce, inject known data at component boundaries, bisection over the component chain) plus an independent installed process pack's multi-component evidence-gathering step, mechanism verified stack-agnostic. RED (measured on 54e0f36): `NO_HITS: component boundar\|entry and exit\|each boundary\|enters and\|exits` across the package while the Instrument step named only "targeted logs"; head carries exactly one list line with the anchor. Baseline is instruction presence, not a behavioral run. Entrypoint grew 4416→4943 body words: every addition is a decision point at its firing step; method detail and verified sources went to the diagnosis playbook reference (Localization Playbook, Probe Ordering, Sources) and the Phase B sanitization re-list was consolidated into a pointer to fund the headroom. |
|
|
526
|
+
| Probe order is decided by discriminating power, then cost, then risk — the cheapest, safest probe whose outcome rules out the most alternatives runs first, likely-and-cheap before exotic — and every system-changing active probe is recorded and reverted before the next observation | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#Probe order must be decided by discriminating power, then cost, then risk | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (SRE troubleshooting chapter, test-and-treat design: mutually exclusive alternatives, decreasing likelihood weighed against risk, side effects of active tests). RED (measured on 54e0f36): `NO_HITS: likelihood\|cheap\|order of\|discriminat\|revert\|pre-test\|restore` across the package — the three-strike rule governed stopping and the falsification rule governed probe validity, nothing governed probe ORDER; head carries exactly one anchored list line. |
|
|
527
|
+
| The hypothesis log is kept inside the diagnosis loop — hypothesis, falsifier, probe cost and side effects, result — so a new hypothesis is checked against recorded observations before it costs a probe and the three-strike count reads from the log instead of memory | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#The hypothesis log must be kept inside the loop | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (Debugging Book introduction, keep-a-log; SRE chapter, take clear notes of ideas, tests and results). RED (measured on 54e0f36): `NO_HITS: audit trail\|running log\|log of`; the only log surfaces were the closeout evidence template and the escalation handoff packet; head carries exactly one anchored list line and the template gained a hypothesis-log field. |
|
|
528
|
+
| A diagnosis licenses a fix only when it explains both causality (how the defect produced this failure on the failing path) and incorrectness (why the code, data, or config is wrong against its contract); a change that removes the failure without the second half is a symptom patch, and an unlinked genuine defect is a different bug that is recorded, never shipped as this cause | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#record it, never ship it as this cause | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (Debugging Book introduction, Checking Diagnoses: fix if and only if the diagnosis shows causality and incorrectness). RED (measured on 54e0f36): `NO_HITS: incorrectness\|why the code is wrong\|why it is wrong\|why the code was wrong`; the reachability-is-not-causation rule covered only the unlinked-defect half; head carries exactly one anchored list line. |
|
|
529
|
+
| A test that passes alone and fails in the suite is bisected over the tests that run before it (halve the preceding set until the polluting test or shared state remains), and the same halving isolates a failing input, config, or dataset when no orderable commit range exists | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#must be bisected over the tests that run before it | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (Debugging Book, reducing failure-inducing inputs — delta debugging) plus an independent installed process pack's polluter-finding script, mechanism verified stack-agnostic. RED (measured on 54e0f36): `NO_HITS: polluter\|order-dependen\|test order`; the package called passes-alone-fails-in-suite a symptom and named no localization move; head carries exactly one anchored list line and the playbook's Localization table carries the recipe. |
|
|
530
|
+
| A hypothesis about a runtime value or state is settled by observing it (breakpoint, print, assertion, trace attribute at the exact point), never by inferring from source what the value must be | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#must be settled by observing it | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: independent agentic-debugging research (agents rewrite conditioned on the error message; interactive tool access improves repair and is under-used) plus the SRE chapter's what/where/why observation discipline. RED (measured on 54e0f36): the 2 hits for `breakpoint` listed the debugger only as an instrumentation option and the code-reading warning fired only on the production telemetry path; head carries exactly one anchored list line covering the local-runtime path. |
|
|
531
|
+
| A wrong value is traced upstream to the first point where a correct input produced a wrong output, and the fix lands at that transition; validation added where the symptom surfaced is defence in depth, never the fix | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#must be traced upstream to the first point where it became wrong | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (Debugging Book introduction, fault propagation from defect to failure) plus an independent installed process pack's backward-tracing reference, ≥2 independent sources. RED (measured on 54e0f36): `NO_HITS: correct.*faulty\|transition`; the Non-Negotiable rule forbade stopping at the wrong line but named no direction of travel; head carries exactly one anchored list line. |
|
|
532
|
+
| A production symptom that cannot be re-triggered in place is not blocked on reproduction: the failing run's own telemetry is the reproduction substitute, suggestive race/deadlock evidence is admitted at its grade, and the cause still owes a falsifying probe before any fix | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#is not blocked on reproduction | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (SRE chapter: some tests are only suggestive; telemetry-first examination). RED (measured on 54e0f36): `NO_HITS: irreproducible\|not reproducible\|cannot be reproduced`; the observability-driven path existed under Instrument but the Reproduce step never routed to it, so an agent could stall at remediation-for-reproduction on a symptom that is diagnosable from telemetry; head carries exactly one anchored list line. |
|
|
533
|
+
| A commit bisection narrows its search by pathspec and by every known-good commit before the first checkout, and a half-finished search is handed off through `git bisect log` / `replay` rather than restarted | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#Do not bisect the whole history | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external primary (`git-scm.com/docs/git-bisect`: cutting down bisection with pathspec and multiple good commits; bisect log and replay). RED (measured on 54e0f36): `NO_HITS: pathspec\|bisect log\|replay`; the two `bisect` hits were the exit-code and `--first-parent` passages; head carries exactly one anchored list line. |
|
|
534
|
+
| In the AI-assisted diagnosis discipline the final cause verdict and the regression test belong to whoever ran the verification commands and read their output — the agent when the agent verified — and a cause proposed by any model, including the diagnosing agent's own analysis, stays a hypothesis until then | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#a cause proposed by any model, including your own analysis, stays a hypothesis until then | `updated` | Owner key `defect-diagnosis/SKILL.md`. W-sweep finding: the base text assigned the verdict to "the human", which for this skill's primary reader (an agent) licensed punting the verdict to the user and contradicted the same block's "YOU verify each candidate" sentence and the repository's autonomy goal. RED (measured on 54e0f36): `grep -c "the human still owns"` = 1 in the package; head = 0, and the replacement line is the anchor. No recorded incident; benchmark-derived. |
|
|
535
|
+
| The persisted-evidence sanitization rule carries its category list by pointer to the external-send rule (Phase A item (d)) plus the one category only it named (config values) instead of re-listing the seven categories inline | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#never pasted into shared diagnosis evidence | `updated` | Owner key `defect-diagnosis/SKILL.md`. Consolidation with a zero-loss obligation map: secrets/tokens, customer data / PII, credential-bearing values, internal hostnames / IPs / URLs / paths, raw SQL and query bodies, request/response bodies, env values, proprietary identifiers all survive verbatim in Phase A item (d); config values survive inline in the consolidated sentence; the incident-store, retention, and link-not-paste obligations of the same bullet are byte-identical. Reviewer to confirm no trigger/scope/routing/validation/acceptance change. |
|
|
536
|
+
| Boundary-walk instrumentation logs only allowlisted, redacted metadata (ids, sizes, status codes, field presence, config keys received) and never raw bodies, headers, secrets, PII, or env/config values; the single instrumented run applies only when the chain can be re-run safely with every boundary reachable, otherwise the telemetry path, partial boundary evidence, or layer narrowing is used with the visibility gap recorded; pathspec bisect narrowing applies only when evidence confines the cause to those paths and falls back to the full range when no reproducing commit is found; the hypothesis log separates the prediction only a cause produces from the falsifier that cannot occur if it is true | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#never raw bodies, headers, secrets, PII, or env/config values | `updated` | Owner key `defect-diagnosis/SKILL.md`. Review-round tightening of this round's own additions: independent review (round 1, codex) found that the first draft licensed raw entry/exit logging before copy-time sanitization, made the pathspec restriction unconditional, made the single instrumented run block the telemetry route, and labelled a discriminating prediction as the falsifier. RED (measured on the round-1 candidate 2317757d…): the boundary bullet contained no redaction predicate and the bisect bullet no fallback predicate; head carries exactly one anchored list line and the playbook table carries separate Prediction and Falsifier columns. Dispositions recorded as `fixed` in the round evidence directory. |
|
|
537
|
+
| The diagnosis playbook's localization table condenses the entrypoint and must never loosen a condition SKILL.md states — the boundary-walk recipe carries the safe-rerun condition, the redacted-metadata restriction, and the telemetry/partial-evidence fallback, the commit-bisection recipe carries the evidence-confined pathspec condition and the full-range rerun, and SKILL.md wins when the two disagree | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/references/diagnosis-playbook.md#a recipe here must never loosen a condition SKILL.md states | `updated` | Owner key `defect-diagnosis/SKILL.md`. Succession challenge (round 3, codex) found the reference table still prescribing raw entry/exit logging and unconditional pathspec narrowing after the entrypoint had been tightened — a mirror drift between entrypoint and reference within one round. RED (measured on candidate 753ad96d…): the two table rows carried none of the entrypoint's conditions; head carries the mirrored conditions in both rows plus one anchored drift-guard list line. Disposition recorded as `fixed` in the round evidence directory. |
|
|
538
|
+
| Probe order is one rule — safety is a filter (reject any probe outside the safety boundary first), then rank by alternatives ruled out per unit of cost, with likelihood and residual risk as tie-breakers; suite bisection keeps the original order and, when neither half fails alone, keeps both halves and reduces by smaller chunks toward a minimal polluting sequence; the upstream trace fixes the correct→faulty transition only when it is owned and changeable and otherwise records the upstream cause and enforces the contract at the nearest owned boundary | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#enforce the contract at the nearest owned boundary | `updated` | Owner key `defect-diagnosis/SKILL.md`. Second wrapper chain (review + challenge, codex) on the fixed candidate: three competing probe orderings in one bullet, a halving loop with no failing half for two-test pollution, and an upstream-trace rule that demanded a fix at an unowned producer. RED (measured on candidate f9da56b1…): the three defects were present verbatim; head carries the single ordering rule, the order-preserving reduction, and the owned-boundary qualification, each mirrored in the playbook table. Dispositions recorded in the round evidence directory; two review findings about the round's own evidence packaging are `accepted_tradeoff` against the repository's recorded evidence-is-caller-controlled and enforcement-inside-the-candidate boundaries. |
|
|
539
|
+
| Suite reduction is order-preserving delta debugging: halve the preceding tests and keep a failing half; when neither half fails alone, remove one chunk at a time and keep the reduced set whenever the failure persists without that chunk, then halve the chunk size and repeat until every remaining chunk is needed (a minimal ordered polluting subsequence) or the shared fixture/state is found — this supersedes the halving-only summary in the preceding round's tightening row, which did not guarantee reduction for a jointly-caused two-test pollution | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#remove one chunk at a time and keep the reduced set | `updated` | Owner key `defect-diagnosis/SKILL.md`. Human-authorized continuation lane (the maintainer authorized further external rounds in this session after the landing lane's final succession returned this finding). RED (measured on candidate cdd133da…): the recipe named halving and finer chunks but no chunk-removal step, so an ordered A+D pollution out of A–D could not reduce; head carries the complement step in the entrypoint and the playbook row. Word budget funded by three gloss trims (bisect script determinism, `exit 125` idiom, `--first-parent`, CI trigger-variant gloss, LLM hallucination sentence) with the conditions, consequences, and actions of each kept. |
|
|
540
|
+
| Suite reduction applies only to a failure that reproduces on every run under a fixed serial order; parallel or intermittent failures keep the failing schedule and validate each kept or dropped subset over repeated runs per the flaky rule, or route to concurrency diagnosis; the playbook's probe-ordering paragraph reproduces the entrypoint's single ordering rule; the consolidated persisted-evidence sanitization rule names credential-bearing values explicitly; the blameless-postmortem sentence cites the SRE chapter's own reason instead of an unsourced empirical claim | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#only for a failure that reproduces on every run under a fixed serial order | `updated` | Owner key `defect-diagnosis/SKILL.md`. Human-authorized continuation lane, wrapper chain (review + challenge, codex) on candidate aa34665b…: the reduction recipe assumed a serial deterministic suite, the playbook probe paragraph still carried the cheapest-and-safest ordering, the pointer consolidation narrowed credential-bearing values to variable values, and one empirical sentence had no primary source. RED (measured on aa34665b…): all four present verbatim; head carries the precondition (entrypoint and playbook row), the mirrored ordering, the explicit category, and the SRE-sourced sentence with its excerpt in the round's attribution evidence. |
|
|
541
|
+
| Fan-out is a walked five-item gate, not a cost note: a genuine constraint must exist, slices are cut by context boundary (never role splits over one feature, never shared state/files/contracts/sequencing), shared implicit decisions are pre-made and carried in every brief, every brief carries an effort budget with width starting small, and the task must be worth the 3–10× single-agent premium | `multi-agent-delegation` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/multi-agent-delegation/SKILL.md#never parallelize work that shares state, files, migrations, contracts, or sequencing | `updated` | Owner key `multi-agent-delegation/SKILL.md`. Benchmark round against public primaries (Anthropic multi-agent research 2025-06 and when-to-multi-agent 2026-01, Google agent-scaling study arXiv 2512.08296, Cognition 2025-06, Claude Code agent-teams docs). RED baseline at origin/dev f156079: `grep -rEil 'scale effort\|effort budget\|effort scal' skills/multi-agent-delegation` → NO_HITS and `grep -rEil 'implicit decision' skills/multi-agent-delegation` → NO_HITS, so a controller following the skill had no rule budgeting worker effort or pre-deciding shared choices, and the four cost sub-bullets duplicated the playbook verbatim (same-facet double write). With change: the gate replaces the duplicated sub-bullets; rationale, sources, and the 15×→3–10× baseline correction live once in the playbook; zero-loss map of the four retired obligations recorded in the round charter. |
|
|
542
|
+
| Delegation execution gains three verification-side rules: large or load-bearing worker outputs go to a durable artifact and the controller verifies from the artifact rather than the relayed summary; a failed return is classified by the multi-agent failure taxonomy (MAST mapping in harness-patterns §2) to pick the fix layer before re-dispatch; integration checks slices for divergent implicit decisions, not only merge conflicts; peer-messaging topology is reserved for workers that must exchange findings, default hub-and-spoke with the controller as validation point | `multi-agent-delegation` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/multi-agent-delegation/SKILL.md#verifies from the artifact, never from the relayed summary | `updated` | Owner key `multi-agent-delegation/SKILL.md`. Sources: Anthropic research-system appendix (subagent output to filesystem to avoid the telephone game; end-state evaluation), MAST arXiv 2503.13657 (14 modes, most failures from system design), Claude Code agent-teams doc (lead must not start implementing while teammates run). RED baseline at f156079: `grep -rEil 'game of telephone\|MAST' skills/multi-agent-delegation` → NO_HITS; the review step verified diffs but had no artifact-not-relay rule and no failure-classification step. observed-failure: no — benchmark-derived, no incident this round. |
|
|
543
|
+
| The sub-agent isolation checklist keeps its internal names but maps them to the literature taxonomy (MAST: three categories, fourteen modes) with a fix-layer column, and the external-source list records the four new primaries (Anthropic when-to-multi-agent 2026-01, Google agent-scaling 2025-12, MAST NeurIPS 2025, Cognition 2025-06) so a later round re-verifies against named sources instead of re-borrowing | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/harness-patterns-and-eval.md#必须先按 MAST 类别定位该修哪一层 | `updated` | Owner key `skill-extraction-workflow/references/harness-patterns-and-eval.md`. Reference-only borrow record; the executable landing is in `multi-agent-delegation` (review step classifies by MAST before re-dispatch). RED baseline at f156079: `grep -rEil 'MAST' skills/skill-extraction-workflow` → NO_HITS. Functional-equivalent check recorded per mode in the round's verdict table (every mode had a scattered counterpart; the missing piece was the classification step and fix-layer routing). |
|
|
544
|
+
| Prompt-cache design precedes miss attribution: static-first ordering with the provider's prefix hierarchy, byte-stable append-only prefix (no volatile tokens, deterministic serialization, no in-place rewrites, mode toggles counted as prefix changes), tool set fixed within a loop with masking over redefinition, provider cache contract respected (minimum length verified from usage fields, breakpoint cap, TTL ordering), cache-read share tracked per route, batch caching treated as best-effort; the tool-dispatch reference gains the turn-boundary rule for surfacing/evicting dynamic tools and the context-freshness reference gains the placement rule (objective and constraints near the end, stable block at the start, distractor-bearing long-context evals) | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/llm-client-gateway.md#never rewrite earlier turns in place | `updated` | Owner key `llm-inference-integration/SKILL.md`. Sources: Claude prompt-caching docs (prefix order tools→system→messages, invalidation table, minimum lengths, four breakpoints, TTL ordering, usage fields), Manus context-engineering (100:1 prefill ratio, append-only, mask-don't-remove), Chroma context-rot and arXiv 2307.03172 (placement). RED baseline at f156079: `grep -rEil 'prefix cach' skills/llm-inference-integration` → NO_HITS and `cache_control\|cache breakpoint` hit only the attribution section — the skill could diagnose a miss but had no rule preventing one. SKILL.md step 3 gains the pointer so the rule fires at gateway design time. |
|
|
545
|
+
| Capacity work uses the standard per-phase vocabulary (TTFT, TPOT/ITL, E2EL) with the averaging caveat, gates rollout and batch tuning on goodput (requests meeting every SLO) rather than raw throughput, carries a serving-lever table that names which metric each hosted-inference lever moves and its caveat (continuous batching, paged KV cache, prefix caching, speculative decoding, prefill/decode disaggregation, quantization, prefix-aware routing), routes latency-insensitive volume to provider batch endpoints under their distinct contract, ramps traffic to avoid acceleration limits, and pins in-flight agent runs to their started version during deploys | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/inference-capacity-operations.md#report tails (p95/p99) per phase, never one blended latency | `updated` | Owner key `llm-inference-integration/SKILL.md`. Sources: BentoML inference handbook (metric definitions, request- vs token-weighted averages), DistServe arXiv 2401.09670 and vLLM disaggregated-prefill doc (TTFT/ITL tuned separately, no throughput gain), vLLM speculative-decoding and prefix-caching docs, NVIDIA inference-optimization blog, K8s Gateway API Inference Extension and AIBrix (prefix/KV-aware routing), Claude batch and rate-limit docs, Anthropic research-system post (rainbow deploys). RED baseline at f156079: `grep -rEil 'goodput\|continuous batch\|speculative decod\|batch API\|Batches API' skills/llm-inference-integration` → NO_HITS; the load checklist used the non-standard phrase 'first-token latency and final-token latency' (W-type vocabulary drift, replaced). |
|
|
546
|
+
| Eval reliability names the judge-bias controls (position swap or randomization with order-consistent verdicts, length-penalizing rubric or normalization, cross-family judge or agreement when the candidate shares the judge's family, per-version human-agreement reporting, judge swap treated as suite migration), the statistical minimum (standard error or confidence interval beside every score, clustered errors for grouped questions, paired differences, power-sized eval sets), and the agent-eval choices (pass@k versus pass^k declared before measuring, end-state or checkpoint grading, saturation graduation, transcript reading before trusting a score) | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#treat a judge model or prompt swap as an eval-suite migration | `updated` | Owner key `llm-inference-integration/SKILL.md`. Sources: MT-Bench arXiv 2306.05685 (position, verbosity, self-enhancement biases), self-preference bias arXiv 2410.21819, Anthropic statistical-approach-to-evals 2024-11 (SEM, clustered SE, paired differences, power), Anthropic demystifying-evals 2026-01 (graders, pass@k/pass^k, capability vs regression, saturation, transcripts), Anthropic research-system appendix (end-state evaluation). RED baseline at f156079: `grep -rEil 'position bias\|self.preference\|pass\^k\|pass@k\|clustered standard\|power analysis' skills/llm-inference-integration` → NO_HITS; the judge rule said only 'calibrate and watch for drift'. SKILL.md step 5 gains the pointer so the controls fire at eval design time. |
|
|
547
|
+
| Every agent design runs the lethal-trifecta test (private data + untrusted content + external channel) and, when it holds, must remove a capability or impose a named structural injection-defense pattern; each design review walks the OWASP LLM Top 10 (2025) against its owning rule, adding the system-prompt-leakage rule (no secrets, credentials, or authorization logic in the system prompt); SDK building blocks note the guardrail execution-mode choice (parallel guardrails can trip after tools ran) and the error-amplification reason to keep a validating orchestrator on the path to the user | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/retrieval-agent-safety.md#you must remove one capability or impose a structural pattern | `updated` | Owner key `llm-inference-integration/SKILL.md`. Sources: Willison lethal trifecta 2025-06, arXiv 2506.08837 design patterns, OWASP LLM Top 10 2025 list, OpenAI Agents SDK guardrails doc (parallel vs blocking execution), Google agent-scaling study (error amplification independent vs centralized). RED baseline at f156079: `grep -rEil 'trifecta\|OWASP\|excessive agency\|unbounded consumption' skills/llm-inference-integration` → NO_HITS; nine of the ten OWASP entries had owning rules but no enumeration walk reached them and system-prompt leakage had no rule. SKILL.md step 3 gains the trifecta pointer. |
|
|
548
|
+
| Review-round tightening of the benchmark landing: the append-only prefix rule yields to mandatory invalidation (compaction, privacy deletion, revoked authorization, safety or policy updates rewrite the prefix as a new cache generation with a baseline reset); a tool discovered mid-loop takes effect at the next model invocation of the same loop with its schema appended after the last cache breakpoint; the lethal-trifecta pattern must provably cut one edge with a negative test and context minimization counts only when the private data is absent at tool-selection and action time; order-inconsistent pairwise judgments stay in the denominator as ties or abstentions with the inconsistency rate reported; the OWASP supply-chain and poisoning mappings name enforceable checks (inventory, pin, verify, approve, roll back every model, adapter, prompt, tool, skill, and dependency; integrity validation and change monitoring of every authorized source) | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/llm-client-gateway.md#a stable prefix is never a reason to keep revoked | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-1 wrapper chain on candidate d5919e7: independent review (codex) raised four P1 and one P2 on the landed text — append-only conflicting with compaction and privacy deletion, an under-specified turn boundary for discovered tools, a trifecta rule satisfiable by an ineffective pattern, and silent exclusion of order-inconsistent judge pairs — plus one packet-evidence finding dispositioned accepted_tradeoff (gate outputs cannot bind the tree containing them); the same-candidate adversarial challenge raised one P1 on the OWASP mapping. RED baseline is the reviewed candidate itself (the pre-fix wording is in the round-1 and round-2 receipts under the round's evidence directory). All five applied in this commit; the succession challenge binds the post-fix candidate. |
|
|
549
|
+
| Succession-round tightening: the tool-set mutation boundary is stated once and identically in the gateway prompt-cache design and the tool-dispatch dynamic-tool rules — mutation is forbidden only during an in-flight invocation, a tool discovered mid-loop becomes callable at the next model invocation, and because tool definitions head the cached prefix that change is a new cache generation whose miss is accepted only when the tool is genuinely needed, with pre-declared schemas and masked availability as the prefix-stable alternative; the earlier suffix-only re-prefill claim is withdrawn | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/llm-client-gateway.md#Do not mutate the tool set during an in-flight invocation | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-1 succession challenge (codex) on candidate b21e253 found the round-1 tool-boundary fix contradicting the gateway rule (tools head the prefix, so a schema appended mid-loop cannot re-prefill only a suffix); recorded needs_human_decision in the lane-1 ledger, then fixed under the maintainer's continuation authorization (continuation-authorization-1.md). RED baseline is the contradicting wording in the lane-1 succession receipt. |
|
|
550
|
+
| Provider batch-endpoint routing is a per-provider checklist with cited answers (completion window and expiry, result ordering, cancel semantics, billed terminal states, own rate limits, spend-limit overshoot), never a universal contract copied from one provider; and this row supersedes the tool-boundary sentence of the lane-1 tightening row above (the sentence 'a tool discovered mid-loop takes effect at the next model invocation of the same loop with its schema appended after the last cache breakpoint' is withdrawn — tool definitions head the cached prefix, so that change is a new cache generation at the next invocation, as the succession-round tightening row states; ledger rows are append-only, so the withdrawal is recorded here by pointer rather than by editing the earlier row) | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/inference-capacity-operations.md#never assumed from another provider | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-2 (human-authorized continuation) review (codex) on candidate 564824a: the batch-endpoint bullet read as one generic contract (P2) and the lane-1 tightening row still carried the withdrawn suffix-only sentence (P1); the packet-evidence finding is accepted_tradeoff as before; the same-candidate challenge found the plan-then-execute wording fixing only tool choice (P1), so the plan now binds operation, destination, argument fields, and data flow, and the negative test injects destination and payload changes. RED baseline is the reviewed wording in the lane-2 round-1 receipt. |
|
|
551
|
+
| The entrypoint's prompt-cache pointer states the tool-set boundary exactly as the references do (mutated never during an in-flight invocation, between invocations only as a new cache generation) — this row supersedes the phrase 'tool set fixed within a loop' in the prompt-cache design row above, which is withdrawn by pointer because ledger rows are append-only; and the per-provider batch checklist adds create idempotency (client request key or server-side deduplication), lookup-based reconciliation of an ambiguous submission before any retry, and terminal usage reconciliation by item and batch id | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/inference-capacity-operations.md#so a retry never bills a duplicate batch | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-2 succession challenge (codex) on candidate f4ad9df found the entrypoint and the register row lagging the reconciled boundary (P1) and the batch checklist silent on create idempotency and post-ambiguity reconciliation (P1); fixed in lane 3 under the maintainer's continuation authorization (continuation-authorization-2.md). RED baseline is the reviewed wording in the lane-2 succession receipt. |
|
|
552
|
+
| Cache-usage arithmetic runs only on provider-adapter-normalized disjoint counters (cache-read, cache-creation, uncached remainder derived by subtraction where a provider's total is inclusive), with an absent cache field recorded as unknown rather than zero; and the batch checklist requires an explicit capability decision when a provider offers neither idempotent creation nor an authoritative lookup key (forbid automatic retry of an ambiguous submission with a persisted submission_unknown state, or decline the endpoint), stable item and submission identifiers persisted before sending | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/llm-client-gateway.md#recorded as unknown, never as zero or as not cached | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-3 (human-authorized continuation) review and same-candidate challenge (codex) on candidate 04d8916: the usage equation double-counted cache reads for providers whose input total is inclusive and misread absent fields as not cached (P1, both rounds); the batch checklist had no path when lookup-before-retry cannot run (P1); the packet-evidence finding is accepted_tradeoff as before. RED baseline is the reviewed wording in the lane-3 round-1 receipt. |
|
|
553
|
+
| History-preserving eviction yields to mandatory invalidation: the dynamic-tool rule never to evict a tool the loop's history cites holds only on capacity or recency grounds, while a revoked authorization or a privacy, safety, or policy update removes the tool's definition and dependent prompt material, resets the cache generation, and restarts the loop or fails closed if the retained history cannot stay valid without it | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/agent-tool-dispatch.md#reset the cache generation, and restart the loop or fail closed | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-3 succession challenge (codex) on candidate 3a44680 found the unqualified never-evict rule retaining a revoked tool against the gateway override (P1); fixed in lane 4 under the maintainer's standing instruction (continuation-authorization-3.md). RED baseline is the reviewed wording in the lane-3 succession receipt. |
|
|
554
|
+
| Every model invocation and tool call is bound to the authorization/tool generation it was issued under, and a completion's tool calls are re-authorized against the current generation before any side effect (stale-generation calls rejected) so a revocation during an in-flight invocation cannot execute through it | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/agent-tool-dispatch.md#a call issued under a stale generation is rejected, never executed | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-4 same-candidate challenge (codex) on candidate f0f3b71: an in-flight invocation could return a call to a tool revoked mid-invocation (P1); the review plan's register-row count was stale (P1, fixed in the caller-held plan's evidence rows; the acceptance sentence is scope-bound and superseded by the evidence row). RED baseline is the reviewed wording in the lane-4 round-1 receipt. |
|
|
555
|
+
| The delegation fan-out gate forbids parallel work that shares state, files, migrations, contracts, or sequencing through writes, while parallel read-only use of one artifact (independent investigations; review plus challenge over one diff) stays allowed when the outputs are independent | `multi-agent-delegation` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/multi-agent-delegation/SKILL.md#shared-write work runs sequentially or stays local | `updated` | Owner key `multi-agent-delegation/SKILL.md`. Lane-4 review (codex) on candidate f0f3b71: gate item 2's shared-files prohibition contradicted the read-only fan-out the playbook and execution flow permit (P1); fixed with zero-loss trims elsewhere so the entrypoint stays at 4988 body words. RED baseline is the reviewed wording in the lane-4 round-1 receipt. |
|
|
556
|
+
| A revocation, narrowing, or policy invalidation advances the authorization/tool generation atomically before any in-flight completion is accepted, and a completion's tool calls are re-authorized against current policy on the exact operation, destination, arguments, and data scope before any side effect, so the stale-generation rejection can always fire | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/agent-tool-dispatch.md#advances the authorization/tool generation atomically | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-4 succession challenge (codex) found the generation-binding clause never requiring the generation to advance on revocation, so an issued call's generation could equal the current one and the rejection never fire (P1); fixed in lane 5 under the maintainer's standing instruction (continuation-authorization-4.md). RED baseline is the reviewed wording in the lane-4 succession receipt. |
|
|
557
|
+
| Side-effect admission shares the revocation's serialization boundary: advancing the authorization/tool generation fences or cancels calls accepted but not yet started, and only a call admitted under an unchanged generation enters the irreversible handler; and pairwise judging runs both orders for every pair before an inconsistency rate is claimed, a single randomized order per pair permitting only an aggregate position-effect analysis | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/agent-tool-dispatch.md#only a call admitted under an unchanged generation enters the irreversible handler | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-5 review and same-candidate challenge (codex) on candidate 56342b9: the judge rule allowed a single randomized order yet demanded a per-pair inconsistency rate (P1), and re-authorization was not serialized with a concurrent revocation before the irreversible handler (P1); the packet-evidence finding is accepted_tradeoff as before. RED baseline is the reviewed wording in the lane-5 receipts. |
|
|
558
|
+
| Fan-out width is counted in workers that each own one bounded slice — tightly related items may sit inside that one slice — and every brief must carry an effort budget, so a width rule can never produce multi-task workers whose ownership, deadline, and failed-return classification cannot be attributed to one bounded unit | `multi-agent-delegation` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/multi-agent-delegation/SKILL.md#each owning one bounded slice | `updated` | Owner key `multi-agent-delegation/SKILL.md`. Lane-5 review (codex) on candidate 56342b9: gate item 4's 'several tasks each' contradicted execution step 3's one bounded task per agent (P1); fixed with zero-loss trims so the entrypoint stays within the 5000-word gate. RED baseline is the reviewed wording in the lane-5 round-1 receipt. |
|
|
559
|
+
| Revocation inside a tool-bearing loop is stated as one fencing invariant rather than accumulated ordering patches: the generation advance and the call's final generation check at the irreversible handler's commit boundary are serialized by the same lock or fence, a call rechecks immediately before crossing that boundary, and the lease and fencing-token mechanics already required for stale agents are reused rather than re-derived | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/agent-tool-dispatch.md#revocation is a fencing problem, not a prompt problem | `updated` | Owner key `llm-inference-integration/SKILL.md`. Four consecutive same-class findings (lanes 1, 3, 4, 5: append-only vs invalidation, eviction vs invalidation, in-flight revocation, admission ordering) showed the class was a concurrency protocol being specified one patch at a time; the lane-5 succession's 'started is ambiguous' finding (P1) is fixed by the invariant form and by routing the mechanics to the existing fencing-token rule, per the same-class-recurrence design rule. RED baseline is the reviewed wording in the lane-5 succession receipt. |
|
|
560
|
+
| The code-then-execute pattern cuts the trifecta edge only when the privileged code, its allowed sinks, and its permitted data flows are generated and frozen before any untrusted content is read, untrusted input entering afterwards only as non-instruction typed data with the negative test restoring that ordering; and the serving-lever table states prefill/decode disaggregation's throughput effect as engine- and workload-dependent to be measured on the target engine, not as a categorical claim | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/retrieval-agent-safety.md#generated and frozen before any untrusted content is read | `updated` | Owner key `llm-inference-integration/SKILL.md`. Lane-6 review (codex) on candidate 09a040c: the code-then-execute option did not require code and sinks to be frozen before exposure (P1) and the disaggregation caveat was categorical (P2); the packet-evidence finding is accepted_tradeoff as before; the missing lane-5 authorization record and the authorization-chain wording were corrected in the evidence directory. RED baseline is the reviewed wording in the lane-6 round-1 receipt. |
|