@ccoalm/ccl-skills 0.3.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +34 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_abort_leak_state_helpers.sh +148 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +271 -34
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-data-acquisition.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-disclosure-channels.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +12 -12
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +7 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +7 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +47 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +28 -29
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +165 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +23 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +303 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +222 -91
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh +85 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +6 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +12 -12
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +455 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +20 -20
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_liveness_predicate_gate.sh +288 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +4 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/deliverable-doc-genre-skeletons.md +133 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/doc-charter-first.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +318 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/AGENTS.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/doc-lint.py +246 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/figure-lint.py +1092 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/mutation_probe.sh +100 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/test_figure_and_doc_lint.sh +375 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/control.md +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/empty-header.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fake-header.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fenced-noise.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig-dangling.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig-orphan-captioned.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig-orphan.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig.png +0 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/imbalance.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/no-unit.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/should-be-chart.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/tables-only-clean.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/unfilled.md +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/wide-table.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/bad-viewbox.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/blackmarker.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/control.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/crossings.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/cvd-confusable.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/decorative-line.svg +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/edge-no-arrow.svg +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/edge-vague.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/figure-contract.json +21 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/figure-is-a-list.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/flow-mixed.svg +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/low-contrast.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/malformed.svg +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-aria.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-group.svg +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-legend.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-title.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-viewbox.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/offcontract-shape.svg +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/overflow.svg +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/transformed.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/ungrouped-card.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/unlabeled-edge.svg +13 -0
- package/dist/assets/release.json +276 -31
- package/package.json +1 -1
|
@@ -113,3 +113,15 @@ Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds
|
|
|
113
113
|
**Why recency is a hazard, not a rule.** Vendor guidance reports that models *tend to follow* whichever instruction sits later — an observation about behaviour, not a licence to resolve conflicts by position. Two ways position becomes dangerous if read as a rule: a later permissive line beats an earlier stricter one (directly contradicting `Conflict Resolution`, which keeps the stricter data-loss/security/contract guard); and text embedded in **untrusted data** — a diff under review, a retrieved document, tool output — sits later within the same authority level and would win by placement alone, which is prompt injection with extra steps. Treat recency as a bias to design against: put the load-bearing rule where the decision happens, and never let placement confer authority.
|
|
114
114
|
|
|
115
115
|
**How to use them.** Vendor guidance and benchmarks are **hypotheses with good provenance** — they tell you what to test on your own corpus, they do not substitute for testing it. Citing them as settled is the same error as landing an unverified claim; so is overruling one with an underpowered probe.
|
|
116
|
+
|
|
117
|
+
## 继承的约定读法:失败形态
|
|
118
|
+
|
|
119
|
+
一道建立在**继承来的错误读法**之上的闸——把 `AGENTS.md` 读成存在「根索引」要求(实际并不存在)——在多轮 dual-track 里以不断打补丁的方式存活,直到用户问「参考网上优秀实践了么」,一手源(nearest-file-wins)才证伪了整个前提。
|
|
120
|
+
|
|
121
|
+
教训:前一个 agent 或前一次提交对某个具名约定的读法是 **hypothesis-grade**;在其之上 fix-forward 会把原始错误一并传播。
|
|
122
|
+
|
|
123
|
+
## 只据内部语料的「深度提炼」:失败形态
|
|
124
|
+
|
|
125
|
+
一次仅以内部 launch-SOP 为源的「深度提炼」会落出**看起来完整**的规则集,却从未核过这套技能声称代表的**公开实践**;下一个用户于是问「参考网上优秀实践了么 / did you check industry practice」。
|
|
126
|
+
|
|
127
|
+
判据:凡规则带 state-of-the-art 主张,charter 的 evidence plan 必须含权威外部源类;只编码内部运行约束、且明标内部范围不作行业主张的规则,不受此限。
|
|
@@ -6,7 +6,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
6
6
|
|
|
7
7
|
```
|
|
8
8
|
0. Charter → ~/.<host>/skills/.extraction-work/<project>-charter.md
|
|
9
|
-
Purpose / Scope / Depth /
|
|
9
|
+
Purpose / Scope / Depth / Result baseline / Open questions
|
|
10
10
|
│
|
|
11
11
|
1. Source register → ~/.<host>/skills/.extraction-work/<project>-source-register.md
|
|
12
12
|
Each row: source id / class / status / target skill / extracted mechanisms
|
|
@@ -43,7 +43,8 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
43
43
|
|
|
44
44
|
- Owner: maintainer
|
|
45
45
|
- Location: `~/.<host>/skills/.extraction-work/<project>-charter.md`
|
|
46
|
-
- Required fields: Purpose, Scope, Depth,
|
|
46
|
+
- Required fields: Purpose, Scope, Depth, Result classification, matching analysis, Failure modes or success-reuse conditions, Lifecycle impact, Evidence plan, Completion standard.
|
|
47
|
+
- Result classification: failure/correction → Deep RCA (widen → counterfactual-test → control); stable success → mechanism + non-luck evidence + reuse conditions + firing point + owner; unstable/insufficient evidence → observation only.
|
|
47
48
|
- Full structure: `SKILL.md` Core Workflow Step 0.
|
|
48
49
|
- Output: a file the maintainer can re-read in 3 months and understand what they were trying to do.
|
|
49
50
|
|
|
@@ -73,3 +73,11 @@ Honest scope, from the eval that landed this (2 arms x 5 fixtures x 6 samples in
|
|
|
73
73
|
Observed shape: a session produces a reader-facing deliverable end-to-end through a platform/tool skill pack (a collaborative-doc platform, spreadsheet, or browser pack), and the loaded tool skill supplies enough "already being guided" feeling that nobody ever asks which skill owns the DELIVERABLE's quality bar — so the finalization owner's gates (doc charter, completeness audit, closeout sweep, structural-reference re-resolution) stay dormant for the whole effort, surfacing only when the user asks "did you run the doc-optimization skill". This is the behavioral twin of the digest-masks-corpus trap: the tool layer's presence masks the owner layer's absence.
|
|
74
74
|
|
|
75
75
|
Firing point: **producing or first-publishing a reader-facing deliverable is an owner-check transition** — before the first publish to a collaborative or reader-visible surface, answer "which skill owns this deliverable's quality bar" as a separate question from "which tool writes it"; a tool-skill invocation never discharges that check, and a deliverable with no matching owner routes to the finalization skill's Draft mode rather than proceeding ownerless. Mechanical anchors: the no-owner-deliverable triggers on the finalization skill's routing surface (`tighten-doc` description), and this workflow's closeout gate when the session later lands skill changes. Honesty bound as elsewhere in this section: between those anchors the check is recognition-dependent — do not overclaim it as a mechanical gate.
|
|
76
|
+
|
|
77
|
+
## gap-list 形态为什么最易滑过
|
|
78
|
+
|
|
79
|
+
自检触发点列举的是「沉淀 / 提炼 / 复盘 / distil」这类**自述措辞**。但一轮提炼最常见的第一个产出不是这些词,而是**一张缺口清单**——「外部源有 X、Y、Z,我们没有」。它读起来像在**答一个覆盖问题**,不像在提炼,所以 charter-before-findings 那条规则从不觉得被触发。
|
|
80
|
+
|
|
81
|
+
但按该规则自己的定义,**针对外部源的缺口清单就是它所说的 findings 回合**:一旦产出,charter 就只能事后补写。
|
|
82
|
+
|
|
83
|
+
观测实例:一轮里先产出四条「外部有我们没有」的缺口,之后才 invoke 提炼工作流;改前的触发词表逐字检索该轮实际措辞得零命中。
|
|
@@ -147,13 +147,13 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
|
|
|
147
147
|
|
|
148
148
|
**用时机**:操作层月级别稳定 + 要做一次 system-wide 换层/瘦身时;把它当"换层前的对抗性回归证据",不是频繁迭代期的日常闸。
|
|
149
149
|
|
|
150
|
-
### 3.4
|
|
150
|
+
### 3.4 Health signal dashboard(描述性 roll-up)
|
|
151
151
|
|
|
152
|
-
3.1–3.3
|
|
152
|
+
3.1–3.3 都保留各自的判定对象;跨时间还需要一个快速入口,显示本次有哪些信号在场、哪些维度变化,方便继续下钻。
|
|
153
153
|
|
|
154
|
-
**外部锚**:OpenSSF Scorecard
|
|
154
|
+
**外部锚**:OpenSSF Scorecard 用每个 check 0–10、风险加权聚合和历史变化展示代码仓信号。这里仅借它的**展示形态**,不继承“一个总分代表整体健康”的解释。
|
|
155
155
|
|
|
156
|
-
**我们的映射(route-not-copy
|
|
156
|
+
**我们的映射(route-not-copy)**:`eval-health.rb` 展示四个信号,并保留一个兼容既有实现的加权 0–10 值与同尺子变化 ——
|
|
157
157
|
|
|
158
158
|
| 维度 | 风险/权重 | 0–10 来源 |
|
|
159
159
|
|---|---|---|
|
|
@@ -164,12 +164,13 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
|
|
|
164
164
|
|
|
165
165
|
`composite = Σ(score·weight) / Σ(weight)`,只算在场维(skip 维权重重分,同 OpenSSF/gstack)。确定性两维自动跑,T2/T3 喂报告进来(否则 skip)。契约见 [eval-routing.md](eval-routing.md) 的 Health roll-up 节。
|
|
166
166
|
|
|
167
|
-
|
|
167
|
+
**三条不可省的护栏**:
|
|
168
168
|
|
|
169
169
|
1. **advisory,不当 gate**(Goodhart):度量变成 target 就被博弈。综合分**永不接门禁**,二元门禁(结构 + T1 blocking)仍独立挡 merge;历史文件 git-ignore,免得"committed 的数"招人调数不修仓。
|
|
170
170
|
2. **corpus/version 守卫**:T2/T3 的 task-bank、golden-traces **本身会变**。加 10 条简单 task,pass-rate 涨了但仓没变好 —— **尺子换了**。每条历史记 `corpus` 指纹(输入内容 hash)+ 在场 `dims`;**趋势只跟 `(corpus, dims)` 全同的历史比**,否则 baseline reset 不偷偷比。等价于 gstack-health "尺子变了就从新基线重新追"。
|
|
171
|
+
3. **不平均掉质量失败**:结构、路由、真实回放和任务结果的语义不同;任何质量或安全阻断仍由自己的门禁决定,不能被其它维度的高值抵消。
|
|
171
172
|
|
|
172
|
-
|
|
173
|
+
**用**:周期性看一眼信号变化并下钻。**不用**:别据此单独声称仓库整体变好/变差,别把它当通过标准、别 committed、别跨 corpus 硬比。
|
|
173
174
|
|
|
174
175
|
---
|
|
175
176
|
|
|
@@ -297,6 +297,24 @@ A carve-out is a contract, not a category badge. Without the evidence row, fall
|
|
|
297
297
|
|
|
298
298
|
**Occurrences that promoted it**: `scripts/owner-dispatch/owner-dispatch.sh` (state keyed per-worktree → boundary record written in a worktree unreadable from the primary checkout) and `hooks/guard-edit-isolation.sh` (path-name match → the repo's only edit-time hard-deny gate silently disabled for checkouts under a `worktrees/`-named path).
|
|
299
299
|
|
|
300
|
+
## Anti-pattern 28 — Process liveness decided by an EXISTENCE test that a corpse answers
|
|
301
|
+
|
|
302
|
+
**Symptom**: a probe or fixture concludes "this process is still alive" — or "it is a live orphan" — from a test that only proves the pid is still in the process table: `kill -0 "$pid"`, or reading `ps -o ppid=` and comparing it against init's pid. The same suite's verdict scans then exclude zombies (`$stat !~ /Z/`) when deciding the process is *gone*.
|
|
303
|
+
|
|
304
|
+
**Why bad**: an exited-but-unreaped process is still in the table. It answers `kill -0`, and its ppid still reads — as `1` once it is reparented. So one process state gets two answers: the precondition check says "alive, scenario built", the verdict scan says "gone, nothing leaked", and an assertion between them that samples liveness at one instant says "already dead" and fails. The red then names the code under test for something it did not do, and the failure is intermittent because it depends on when the OS reaps. Worse, the precondition passing means the probe goes on to assert about a scenario it never actually built.
|
|
305
|
+
|
|
306
|
+
**Fix**: consult process **state**, not existence.
|
|
307
|
+
- One vocabulary for the whole probe — a helper returning `live` / `zombie` / `absent` (`ps -o stat=`; empty → absent, `*Z*` → zombie, else live) — used by *every* liveness question, so the checks cannot answer the same pid differently.
|
|
308
|
+
- Preconditions require `live`. A corpse is not an orphan; accepting one is how a probe goes on to assert about a scenario that was never constructed.
|
|
309
|
+
- Better still, stop asking the process at all: assert on an artifact the run leaves behind — a work dir the cleanup path would have deleted, a marker the fixture writes only on the code path under test. That answers *why* a process ended, which no liveness sample can.
|
|
310
|
+
- Never assert liveness by an instantaneous sample; reap lag on a loaded runner needs a bounded grace period. (`testing-strategy/references/ci-fixtures-and-flake-control.md` owns that rule — this row is its firing path.)
|
|
311
|
+
|
|
312
|
+
**Grep**: `find . -name 'test_*.sh' -o -name 'test.sh' | xargs grep -nE 'ps[[:space:]]+-o[[:space:]]+ppid=.*=[[:space:]]*"?1"?[[:space:]]*\]'` (non-comment hits with no `stat=` / `*_state` consult within two lines). Machine-enforced by `check-ccl-skills.sh` (`liveness_predicate_scan`). Scope is test scripts only: outside them a `kill -0` before signalling asks about existence, which is the right question there. Five things are deliberately **not** caught and stay human/challenge checks — `while kill -0 "$pid"` watchdog loops (two exist here; both wait on a direct child and `wait` for it immediately after, so the shell reaps it and the loop ends), a bare `kill -0` liveness branch inside a loop body, any dynamic spelling, a state-helper mention in a **trailing** comment (whole-line comments are dropped, but telling a trailing `#` from `${var#prefix}` needs a shell parser — the same call Anti-pattern 27 makes), and a process-state read — direct `ps -o stat=` or via a helper — that inspects a *different* pid than the oracle tests, or whose result never reaches the verdict. Both are the same irreducible gap: proving the read actually governs the decision needs dataflow over a parsed shell, not a text window, so this is the limit the mechanical gate stops at by design. Both are **exercised** by `test_liveness_predicate_gate.sh` (P13/P14) rather than only described here — they assert the gate does not fire, so tightening it turns them red and forces this paragraph to be updated instead of quietly going stale.
|
|
313
|
+
|
|
314
|
+
The waiver is a **proxy** for "this site consults process state", not an invariant, and five successive review rounds each found it too loose in a different way — a comment naming the helper, an unrelated `*_state` token, a helper that never reads state, a bare `stat=` assignment, and a hollow helper borrowing an unrelated read elsewhere in the file. Tying a call to its definition needs a shell parser, so the honest disposition is a narrow predicate with these limits stated rather than another round of widening. The mechanical gate is the deterministic catch for the exact recurring shape; this checklist and the adversarial challenge remain the comprehensive net.
|
|
315
|
+
|
|
316
|
+
**Occurrences that promoted it**: the code-review abort-leak probe red CI three times, each on a different leg-2 assertion, each asserting the probe's environment rather than the suite — the reparent check accepted a zombie, the "still alive right after the kill" check rejected that same zombie, and the verdict scan reported it gone. The rule forbidding this already existed in `testing-strategy`, and the suite's own hang cases followed it while the probe did not: the gap was enforcement, not content.
|
|
317
|
+
|
|
300
318
|
## How to use this checklist
|
|
301
319
|
|
|
302
320
|
Before committing any skill or reference change touching operational/architectural rules:
|
|
@@ -7,10 +7,10 @@ Keep source-specific provenance outside the distributed repository. Public skill
|
|
|
7
7
|
| Field | Required answer |
|
|
8
8
|
| --- | --- |
|
|
9
9
|
| Task or extraction name | Name the reusable workflow or source family with a source-neutral label. |
|
|
10
|
-
| Purpose | State the future failure or
|
|
10
|
+
| Purpose | State the future failure or drift this extraction prevents, or the evidenced success mechanism it preserves. |
|
|
11
11
|
| Scope | List included source classes, target skills, sibling boundaries, and exclusions. |
|
|
12
12
|
| Depth | Record wording cleanup, targeted check, file refresh, artifact inventory, full workflow extraction, or tooling change. |
|
|
13
|
-
|
|
|
13
|
+
| Result analysis | Classify failure/correction, stable success, or unstable/insufficient evidence, then record the matching RCA, success attribution, or observation-only boundary. |
|
|
14
14
|
| Lifecycle impact | Name the affected intake, design, implementation, testing, launch, iteration, onboarding, and documentation stages. |
|
|
15
15
|
| Evidence plan | List source categories and how each is inspected, routed, excluded, or marked unavailable. |
|
|
16
16
|
| Completion standard | Name the scenario, command, review, and source-map evidence required to finish. |
|
|
@@ -22,8 +22,28 @@ Target-output map:
|
|
|
22
22
|
|
|
23
23
|
Required for upstream-owner skill changes:
|
|
24
24
|
|
|
25
|
+
**Declaration fragments (045).** A row's behavior cell is semicolon-delimited. Alongside
|
|
26
|
+
`behavioral-evidence:` / `observed-failure:` / `firing-path:`, a row for a changed upstream owner
|
|
27
|
+
carries `result-class: failure|stable-success|insufficient-evidence`, and — when the round changes
|
|
28
|
+
a routing surface (a SKILL.md frontmatter `description` entry, or `eval/routing-tasks.jsonl`) —
|
|
29
|
+
`bank-evidence: <locator>` or `bank-evidence: downscoped:<token>` with that token recorded in the
|
|
30
|
+
round's spec. `impact-chain-gate.rb` enforces both. The value stays the author's call; its
|
|
31
|
+
**absence** does not, which is the whole point: an omitted claim is one a reviewer cannot refuse.
|
|
32
|
+
A `bank-evidence` locator must not point back into the owner's own package — the change is not
|
|
33
|
+
evidence about itself.
|
|
34
|
+
|
|
35
|
+
**A round is judged by the grammar its own head declares.** The gate looks for this paragraph's
|
|
36
|
+
`result-class:` definition in the ledger at the round's head; a round that predates it is not held
|
|
37
|
+
to it. That is deliberate and is not a grandfather clause keyed on dates or commit ids: adding a
|
|
38
|
+
required field would otherwise retroactively refuse every historical round on replay, which the
|
|
39
|
+
verdict-differential suite correctly reports as a regression. Rounds landing after this paragraph
|
|
40
|
+
carry the obligation; rounds before it were never told.
|
|
41
|
+
|
|
25
42
|
| Upstream rule | Downstream owner | Expected executable behavior | Status (updated, unchanged, routed, or not-applicable) | Evidence |
|
|
26
43
|
| --- | --- | --- | --- | --- |
|
|
44
|
+
| For model-controlled result text, `line start` and `whole-string start` are different trust boundaries: multiline matching lets a later payload line select a transport-auth decision, so the textual fallback permits only leading whitespace from offset zero and any textual preamble fails closed unless a structured transport status independently classifies it | `code-review` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/code-review/scripts/test_classify_envelope.sh | updated | A review proposed `re.MULTILINE` for leading-preamble tolerance, but `result_text()` returns only `env["result"]` and does not concatenate transport fields. The candidate therefore keeps whole-string anchoring: leading spaces and line breaks are accepted, while `review preamble` followed by the exact 401 phrase remains a generic error. Scratch mutations separately prove removing the whitespace allowance reds its regression and enabling `re.M` reds the preamble control. Structured `api_error_status=401` remains the vocabulary-independent auth arm. Disposition: `specs/044-review-auth-fallback/plan.md`. |
|
|
45
|
+
| A transport-error text fallback is safe only when it pins the observed transport shape and carries benign near-miss cases; searching generic authentication phrases inside model-controlled result text can turn review content into an auth decision and unnecessarily widen packet egress | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_classify_envelope.sh | updated | Independent review found the broad `authentication failed` / expired-token search could classify an errored review whose payload merely discussed authentication as `auth`. The expression is now anchored to the observed line-start `Failed to authenticate. API Error: 401` transport shape, while structured `api_error_status=401` stays the primary vocabulary-independent arm. RED/GREEN adds two benign near-miss fixtures that remain generic errors; a scratch mutation restoring the broad phrase search turns the named precision fixture RED. Plan and disposition: `specs/044-review-auth-fallback/plan.md`. |
|
|
46
|
+
| A Claude result envelope that explicitly reports authentication failure must enter the existing bounded auth-path remediation contract even when its subtype is `success`; otherwise a logged-in-but-unreadable or expired OAuth path is mislabeled as a generic local tool failure, creating contradictory fallback metadata that the controller correctly refuses | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_classify_envelope.sh | updated | Artifact classification and frozen decision table: `specs/044-review-auth-fallback/plan.md`. Observed baseline: an exact-candidate review returned an explicit 401/OAuth-expired result while local auth status still reported logged in; the classifier emitted generic `error:success`, the wrapper paired `reason_code=local_tool_failure` with fallback metadata, and the controller stopped before another client. Minimal correction: `skills/code-review/scripts/classify_envelope.py` maps structured `api_error_status=401` and the exact errored-envelope 401 transport phrase, with optional leading whitespace, to the existing `auth` token; it does not read credentials or alter the gate's terminal non-auth, concern, packet, binding, or tool-boundary paths. RED/GREEN: exact, whitespace-prefixed, and structured-status fixtures fail on the relevant unmodified branch and pass on the candidate; separately removing each classifier arm in scratch makes its named fixture RED. The structured comparison normalizes through a local `api_status_int`, mirroring the sibling normalizer in `parse_probe_result.py`, because a bare `== 401` matches only a Python int and a `"401"` serialization would miss both arms and re-create the generic `error:success` token this row exists to remove; string `401`/`429` and a non-numeric status are pinned, and reverting the normalization reds the string fixture. An earlier draft of this row claimed wrapper parse precedence protects a completed verdict from status 401 — that description was wrong and is corrected: `parse_review_json.py` refuses ANY envelope carrying a non-null `api_error_status`, so such an envelope is never a clean verdict and this change only renames the reason it is refused under; a `run_fail` twin of the existing 429 row now pins 401 plus a valid `structured_output` as refused instead of leaving it to prose. Integration evidence: `test_claude_review_probe.sh`, `test_review_gate.sh`, and the complete `make -j3 test-code-review` family pass, including auth host-retry/fallback and abort-leak contracts. |
|
|
27
47
|
| Requirement records stay with the product workflow while generic Wiki/Base mechanics remain resource-specific | `lark-wiki` / `lark-base` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#Create or reuse only the requested resource | routed | `product-rd-workflow/SKILL.md`; `bootstrap.md` |
|
|
28
48
|
| Structured testcase delivery includes its testcase Base lifecycle | `test-artifact-management` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/test-artifact-management/SKILL.md#but it is optional and must not trigger creation | updated | `test-artifact-management/SKILL.md`; `test-artifact-management/references/bitable-setup.md` |
|
|
29
49
|
| Review-client compatibility is capability-based and leaves model selection to host configuration | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/kimi_review.sh | updated | `code-review/SKILL.md`; `code-review/scripts/kimi_review.sh`; `code-review/scripts/opencode_review.sh`; `code-review/scripts/test_review_client_compat.py` |
|
|
@@ -295,3 +315,28 @@ and inverted the sense (production, not product), and the coordinator now shares
|
|
|
295
315
|
| When every successive bar built to make a mechanical exemption safe is broken by independent review, the signal is that the exempting capability should not exist rather than that another bar is owed; an evidence class may still force an honest label and a real obligation map while leaving the unverifiable judgement to a named risk owner | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; implementation and regressions in `skill-extraction-workflow/scripts/impact-chain-gate.rb` and `skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh`. RED baseline: restoring the automatic floor lift turns the suite red on the compliant-withdrawal case, which must now stay blocked; the other fifteen cases and the repo's four existing impact-chain tests are unchanged either way. Seven review rounds and ten findings are recorded in `specs/043-evidence-class-source-refuted/frozen-acceptance.md` |
|
|
296
316
|
| A rule may fire perfectly and still be false: an entrypoint claim asserting that only the single most-salient prose rule applies while co-resident ones stay dormant is unsupported by either major vendor's published guidance, and the corollary it generated — that appending a clear rule does not add compliance — is contradicted outright, so it must be withdrawn rather than re-explained. Withdrawing an unsupported rationale is a `semantic-control` change, not a behaviour delta: every obligation the bullet carried is preserved verbatim and the entrypoint shrinks. **The `observed-failure: no` classification is a judgement worth challenging**: the withdrawn claim did mislead a reader into building a redirection on it, but this field asks whether the owning rule failed to *fire*, and it fired — it was simply wrong. The repo's evidence taxonomy has no class for 「一条正确触发但内容为假的规则」; recording it here rather than mislabelling it as a RED-baseline | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/SKILL.md#Descriptive, not permissive | updated | `skill-extraction-workflow/SKILL.md` 的机制 bullet(1128 → 825 字节,入口净 −304);出处、证据分级与被撤回的原文保留在 `skill-extraction-workflow/references/external-practice-controls.md#instruction-following-mechanisms`;一手源为 Anthropic 的 context-engineering 文与 OpenAI 的 GPT-4.1 / GPT-5.1 prompting guide;本轮对该撤回做过的三次行为测量全部作废并记账于 `specs/042-skill-corpus-optimization/batch-1-result.md` 与 `evidence/AGENTS.md` |
|
|
297
317
|
| A suite that adds a sibling test without registering it in a lane ships a false green: the file exists, reads as covered, and never runs, so the registration self-audit must be treated as a merge-blocking check rather than a lint nicety — and a clone-based multi-case suite belongs in the heavy lane, not the pre-commit one | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the lane registry is in `skill-extraction-workflow/scripts/test_check_ccl_regressions.sh` and the audit in `skill-extraction-workflow/scripts/test_regression_runner_registration.sh`. RED baseline: CI demonstrated it — `regression-fast` and `regression-heavy` both failed with `test_*.sh not registered in fast_tests/heavy_tests` naming `test_impact_chain_source_refuted.sh`, which the previous round had added and never registered; both lanes pass once the suite is registered in the heavy lane, and reverting the registration turns them red again |
|
|
318
|
+
| A learning workflow that requires a failure RCA for every result forces stable success through an invented bad-outcome story, so the workflow first classifies failure/correction, stable success, or insufficient evidence; only failures run RCA, while stable success lands only with mechanism, non-luck evidence, reuse conditions, firing point, and owner | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#is a claim the **independent review must accept** | updated | The owner rule and canonical templates are synchronized in `skill-extraction-workflow/SKILL.md`, `references/source-to-skill-extraction.md`, `references/extraction-quickstart.md`, and this register template. The entrypoint still names task/session summaries and lessons-learned requests as triggers and keeps missing classification or matching analysis at `interim`; the change broadens the analysis, not the trigger. Controlled A/B with Claude Code 2.1.235, hooks disabled and tools empty, used the same three-times-successful documentation scenario and varied only the quoted governing rule: the base failure-only rule produced `analysis_type=pre_extraction_rca_on_success`, invented a future bad outcome and proposed a new merge-blocking prevention gate; the candidate classified mechanism attribution, left `future_bad_outcome=null`, and returned the evidenced mechanism, non-luck evidence, reuse boundary, firing point, and owner. The existing LARGE-retrospective SUSTAIN rule is the semantic control: it already required mechanism + non-luck evidence + owner, so this round generalizes that owned invariant rather than creating a second success-learning method. Implementer self-review frozen before independent review — acceptance: (1) failure, stable success, and insufficient evidence select different analyses without a parallel 0→1 flow; (2) F4 keeps T2/T3 advisory and lets only an existing owner, risk, or review gate adopt their evidence as a per-change acceptance condition; the roll-up remains navigation, never a universal score or warn→block path; (3) delegation docs preserve pre-dispatch confidentiality, executable leaf containment, and scoped reviewer verdicts; (4) testing/self-review additions remain owner-linked. Candidate scope: `Makefile`, `README.md`, `docs/ARCHITECTURE.md`, `docs/f4-skill-effectiveness-harness.md`, `docs/feature-delivery-handbook.md`, `docs/multi-agent-delegation-handbook.md`, `docs/skill-extraction-handbook.md`, `docs/skills-theory-foundations.md`, `docs/technical-review-handbook.md`, `docs/testing-handbook.md`, `skills/skill-extraction-workflow/SKILL.md`, `references/eval-routing.md`, `references/extraction-quickstart.md`, `references/harness-patterns-and-eval.md`, `references/source-register.md`, `references/source-to-skill-extraction.md`, `scripts/eval-health.rb`. The same frozen candidate also carries three review-classifier and test files belonging to the co-resident review-auth-fallback slice; that slice is adjudicated by its own rows above plus `specs/044-review-auth-fallback/plan.md`, not by this one. Saying so is deliberate — an earlier draft of this row said "no excluded candidate file", which was false for the frozen diff and would have told a reader those files were outside the candidate and separately landed. They are named without package paths on purpose: a path here would read as this row declaring a second changed upstream owner, which is the ambiguity the impact-chain gate refuses. Edge/failure paths checked: success with no non-luck evidence remains observation; a T2/T3 miss never changes runner exit and becomes a landing condition only when an existing gate adopts it; skipped dashboard dimensions stay visible; prompt redaction leaves affected review obligations controller-side. Known residuals: vendor behavior remains time-bound to the linked official pages; the Wiki industry page already carried the current distinction and needed read-back rather than another rewrite; no dedicated F4 Wiki page exists, so 06 carries the reader-facing boundary. |
|
|
319
|
+
| A probe's baseline assertions must state facts the run LEAVES BEHIND, not what was true at the instant it looked: an existence test answers yes about an exited-but-unreaped process, so a precondition spelled that way accepts a corpse as a live orphan and the probe goes on to assert about a scenario it never built, while the verdict scan in the same file excludes zombies and reports the same pid gone — one process state, three answers, and a red that names the code under test for something it did not do. Give the whole probe ONE live/zombie/absent vocabulary, require `live` where the scenario is constructed, and prove WHY a process ended from an artifact of the code path under test (a work dir the cleanup would have deleted, a marker the fixture writes only past its bound) rather than from a liveness sample. A scenario the probe fails to BUILD is retried, never reported as a failed assertion | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate_abort_leak.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `skills/code-review/scripts/test_review_gate_abort_leak.sh` (a `wrapper_state()` helper used by every liveness question, a `live`-requiring reparent check, a work-dir-survives-SIGKILL assertion replacing "the suite exited within 30s", a bound-marker assertion replacing "still alive right after the kill", an arming step that freezes the wrapper's process group across the abort and READS the marker rather than deleting it — while the group is stopped nothing can write that file, so the ordering is enforced instead of sampled between racing events, and no evidence is destroyed — bounded setup retry, and its own deadline for the reparent wait, which previously reused the controller wait's countdown) and `skills/code-review/scripts/test_review_gate.sh` (both hang stubs record `<client>_hang_bound_reached` past their countdown). Round plan: `specs/046-abort-leak-baseline-invariants/plan.md`. Observed failure: the CI run identified in that plan red on `leg2: no reaper ran, so the wrapper is still alive right after the kill` while the leg's own behavioural assertion passed 13ms later on its first loop iteration; three reds total, each on a different leg-2 assertion, each asserting the probe's environment. Verified mechanism, reproduced rather than inferred: a real zombie answers `kill -0`, reports a ppid, and is excluded by both `$stat !~ /Z/` scans — so the reparent check passes, the instantaneous liveness check fails, and the verdict check reports gone, with no timing coincidence needed to explain the 19ms between them. Three candidate causes for the wrapper's early exit were rejected on evidence and recorded as such: the fixture's 25s bound (measured `since_detect=0s` across five runs, including under load average 54 on 18 cores), the suite-group SIGKILL (`review_gate.py:545` is the only spawn path and passes `start_new_session=True`, so the wrapper holds its own pgid), and the controller's own 5-12s wrapper timeout (a forced 15s delay left the wrapper `Ss` at `etime=00:15`). The trigger on that CI run remains unidentified; the bound marker makes the next occurrence name its own cause. RED-baseline (applied, differential): reverting the candidate stub's bound to unbounded reds leg2/fallback on both verdicts with `state=live bound_marker=absent`; reverting the claude stub's bound reds leg2/claude; removing only the marker write reds exactly one assertion — the bound verdict — while the residue verdict stays green, which is the discrimination the deleted timing proxy could not make. Controls green after each revert: `make test-code-review-abort-leak-1`, `-2`, and the full `test_review_gate.sh` |
|
|
320
|
+
| A rule that is already written and already followed elsewhere in the same repository can still not fire on new code, and "the owner skill states it" is enforcement absent, not enforcement present: the fix is the trigger, not another restatement. Promote the class to a mechanical gate scoped to where the predicate is a TEST VERDICT, pick the one spelling whose meaning is unambiguous so precision stays high, and state the recall limits rather than broadening the match | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_liveness_predicate_gate.sh | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the change lands in `skills/skill-extraction-workflow/scripts/check-ccl-skills.sh` (`liveness_predicate_scan`, Anti-pattern 28, same shape as the Anti-pattern 27 gate it sits beside), `references/recurring-anti-patterns-checklist.md` (the Anti-pattern 28 row the gate cites), the new `scripts/test_liveness_predicate_gate.sh`, and its registration in `scripts/test_check_ccl_regressions.sh`. Observed failure: the rule forbidding an instantaneous liveness sample already existed in `testing-strategy` (its CI-fixtures and flake-control reference, named without a package path because that owner is unchanged this round and a path there reads as a second upstream claim) and the review-gate suite's own hang cases followed it, yet the probe written in the SAME round violated it in two of its four liveness sites and red CI three times. The predicate is the orphan-oracle shape only — a `ps -o ppid=` read compared against init's pid on the same line — because that spelling has exactly one meaning, while a ppid read used to identify a parent is the common correct use; a `stat=` consult within the window, or a call to a state helper the same file defines, clears the line, so the documented fix is also the way out. Precision on the current corpus is 100% (one hit, the real defect); five recall limits are named in the gate comment and the checklist row rather than papered over, per the checklist's own promotion guidance that a false positive costs more than a recall gap on a BLOCKING gate. RED-baseline (applied, differential): restoring the pre-fix probe takes the checker to `liveness_predicate_scan_failed` naming line 285 — the verified defect site — while the fixed tree reports `liveness_predicate_scan_ok` and `ccl_skill_check_clean_ok`; attribution is differential in that the failure is raised by this gate and no other. The behaviour suite's fifteen probes pin both directions, and building it under seven review rounds caught five real defects in the gate itself: the sorted hit list lost its trailing newline so the last hit and the diagnosis ran together, and the `trap 'rm -f "${var:-/dev/null}"' EXIT` idiom would have run `rm -f /dev/null` when the variable was unset — harmless unprivileged, a deleted device node in a root container — now a guarded cleanup function that names no fallback path; a whole-line comment naming the helper cleared a real hit; any `*_state` token counted as remediation; and an unrestricted `.*=` swallowed the `!` of `!=`, flagging the negated assertion that does the right thing |
|
|
321
|
+
| Stopping a process GROUP is not an atomic state transition, so reading the LEADER's state after the signal is a proxy for the condition rather than the condition: the leader can report stopped while a sibling has not yet been scheduled to handle it, and when that sibling is the countdown itself it can still reach the write the freeze exists to prevent. Wait for every member that is neither stopped nor already gone, bounded, and treat "still running" as a lost scenario rather than a verdict | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate_abort_leak.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `skills/code-review/scripts/test_review_gate_abort_leak.sh` (`wrapper_group_unstopped`, and an arming step that waits for the whole group). Supersedes the group-freeze half of the earlier row in this round, which checked only the leader. Observed failure: raised as P1 by the independent review lane against the pushed candidate, on the leg whose stub runs its countdown in a backgrounded child sharing the wrapper's group — the half of the mechanism the leader check cannot see. Evidence: with the fixture bound cut to 3s against the CLAUDE stub specifically (the one with the child), all five leg-2 assertions still pass, which they could not if the child were still counting during the abort window; both abort-leak targets and leg 1 stay green |
|
|
322
|
+
| A test harness owns a pid only while that pid is outstanding: once it has been reaped the number is the OS's to reissue, so a cleanup list that still carries it aims its signals at a stranger — and a suite that runs eight-way parallel makes that stranger a sibling lane. Drop ownership at the moment of reaping, which is the same prove-it-now rule the code under test applies before it signals anything | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_abort_leak_state_helpers.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `skills/code-review/scripts/test_abort_leak_state_helpers.sh` (`drop_kid`/`reap_kid`). Observed failure: raised as P1 by the independent review lane against this round's own new test — the harness written to check the probe's ownership discipline violated it, accumulating reaped pids in its cleanup list. RED-baseline: reverting the drop-at-reap change leaves reaped pids in the trap's signal list, which the suite's own accounting shows as entries no longer owned; the hazard is structural rather than timing-reproducible, so the recorded evidence is that accounting rather than a raced kill |
|
|
323
|
+
| A mechanical gate's DOCUMENTED limit belongs in its behaviour suite as a probe asserting the gate does NOT fire, not only in prose: pinned that way, any later tightening turns the probe red and forces the documentation to be corrected, whereas a limit described only in text goes stale silently and is then re-discovered as a finding | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_liveness_predicate_gate.sh | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the change lands in `skills/skill-extraction-workflow/scripts/test_liveness_predicate_gate.sh` (P13/P14) with the limit paragraph in `skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md` pointing at them. Observed failure: the adversarial challenge re-raised the different-pid / discarded-result waiver hole and correctly noted the suite never exercised it, so the limit existed only as prose. RED-baseline: P13/P14 pass against the current predicate and go red against a stricter one — that inversion is the signal they exist to raise |
|
|
324
|
+
| An obligation whose skip leaves NO artifact is not enforced, however normative its wording: the party it constrains states the entry condition, and a reviewer cannot refuse a claim that was never made. Demoting an overclaim must not demote the obligation riding on it — separate the two, and give the surviving obligation a trigger keyed on a fact of the DIFF (which file changed, which key a row carries) rather than on prose. A round is held only to the grammar its own head declares, or adding a required field retroactively refuses every historical round on replay | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; result-class: failure; firing-path: command:skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh | `updated` | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the change lands in `scripts/impact-chain-gate.rb` (two refusals: `impact_chain_result_class_missing`, `impact_chain_bank_evidence_missing`, both scoped to owners the round changed and both gated on the head-declared grammar), `references/source-register.md` (the declaration-fragment paragraph that dates them), and the new `scripts/test_impact_chain_self_adjudication.sh` registered in the heavy lane. Observed failure: round 044 withdrew several overclaims and demoted their obligations in the same move; five consecutive challenge rounds returned one shape, each naming an executable bypass — change a description and never run the bank (no absence to detect), omit the result-class cell (static checks pass, nothing emits interim), or write the class the author prefers (the reviewer has no field to refuse). RED-baseline (applied, differential, re-measured on the final suite — an earlier draft of this row froze a twelve-leg table and an `exactly A2/A3/A5 / B2/B3` partition and went stale as legs were added; independent review caught the drift, and the numbers below are the measured ones): the 28-leg decision table is red on the legs each refusal exists for before the gate change. Disabling `impact_chain_bank_evidence_missing` reds exactly A2 A3 A5 A7 A8 A9 A10 A11 A13 A14 A15 A16 A17 A18 A19; disabling `impact_chain_result_class_missing` reds exactly B2 B3 B5; disabling `impact_chain_grammar_withdrawn` reds exactly G2 G3 G4. Clean partition, no overlap, control green either side of every mutation. Bootstrap: removing `result-class` from this row itself, committed, reds the gate naming this row. Scope limit recorded rather than overclaimed: these close OMISSION, not MISCLASSIFICATION — the value stays the author's, and a locator is not proof the measurement ran. Retroactivity was measured, not assumed: naked, the two triggers newly refuse 34 of the 64 replayed historical integration points (5 for the bank trigger alone); gated on the head-declared grammar the differential returns 64/64 with zero new refusals, and leg G1 pins that property. Round plan: `specs/045-self-adjudicated-obligation-trigger/plan.md` |
|
|
325
|
+
| 交付型文档不是一个体裁:起草期缺「非目标 / 备选与落选理由 / 本方案自身的代价 / 外部先例 / 开放决策带闸 / 证据与正文分离」这些节的首稿,会把缺项以打磨轮的形式付回去,且命题定性写错的返工代价与篇幅成正比。因此把族判定与每族起草期必答项作为形态契约落在定稿技能的 charter 之后,实质完整性的裁决权仍留在各族 owner | `tighten-doc` | behavioral-evidence: RED-baseline; observed-failure: yes; result-class: failure; firing-path: file:skills/tighten-doc/references/doc-charter-first.md#也是 charter 与 Genre 两格必须先锁的原因 | updated | `tighten-doc/SKILL.md` 是 owner key,本轮净减 565 字节:原 Diataxis 段的细则下沉到新增的 `skills/tighten-doc/references/deliverable-doc-genre-skeletons.md`,入口只留两轴判据与指针; `skills/tighten-doc/references/doc-charter-first.md` 的 charter 表新增 Genre 一格并补入「全局修订先于局部润色」。 零损失:原 Diataxis 段的六项义务(单一主导模式、muddled purpose 判据、四模式只用于发现混杂而非强制拆四份、 一页纸默认形态属 how-to/reference、只标记拆分候选不自行改文档集、文档生成技能只是执行器)逐条搬入新文件 §0,无删减。 RED-baseline(applied, differential):同一起草任务(把一个后端服务扩成两区长期并行),仅变更所引治理规则文本一个变量, 在仓外中立目录以 `--safe-mode --tools "" --setting-sources ""` 且禁 CLAUDE.md / auto-memory 运行—— base 产出一份 14 节的合理骨架而九个必答项标记(非目标/备选/落选/取证/开放/退出/降级/禁用词)全部为 0, 且其实施路线一节自命名为「迁移」,正是本类返工最贵的那种命题错误;candidate 九项全部出现。 判分器可失败性由 base 产出真实文档而非拒答证明。首次测量在仓内跑、两侧都读到了新文件,判定污染并作废重测。 外部依据:Google design doc 的 non-goals 与 alternatives、公开 RFC 模板的 drawbacks/prior art/unresolved questions、 arc42 的质量目标与术语表、ISO/IEC/IEEE 42010 的关切—视图覆盖、C4 的一图一层、Minto 金字塔的结论先行、 ISA 230 的收录判据与归卷保管、PRISMA 2020/-S、ICH E9(R1)。 十六轮 dual-track 中一类反复误判(安全事件响应策略)判为设计缺陷并整体删除、改为路由给既有安全 owner,删后该类不再复发 |
|
|
326
|
+
| 由差集得出的缺席结论有两条独立的假阳性来源:发布方收录范围变化,以及匹配口径在标签长出新词时漏配。 而包含式复检只能推翻缺席、不能确认缺席——改名后的新名可能完全不含旧名,任何字面匹配都照不到。 另一面,把某条取数入口写进「下一轮最高性价比」之前要先抽样验它的产出:入口的价值由它实际给出什么决定,不由它的强制力或可达性决定 | `multi-perspective-research` | behavioral-evidence: RED-baseline; observed-failure: yes; result-class: failure; firing-path: file:skills/multi-perspective-research/references/public-data-acquisition.md#证实不了就进待核验清单,不得因为「复检过了」而升格为结论 | updated | `multi-perspective-research/SKILL.md` 是 owner key 且本轮未改;改动落在 `skills/multi-perspective-research/references/public-data-acquisition.md`(并入既有的名单差分条,补第二条假阳性来源与「复检不能确认缺席」的方向限制)与 `skills/multi-perspective-research/references/public-disclosure-channels.md`(在「走过而无产出的渠道要留下原因」一节后补其反面:入口按产出选)。 既有的收录范围变化处置、独立来源证实要求与渠道状态词均原样保留,无义务被删。 RED-baseline(applied, differential,幅度如实记):同一提问(逐期名单中某实体本期未出现,能否得出已退出)在仓外中立目录、`--safe-mode --tools "" --setting-sources ""`、禁 CLAUDE.md 与 auto-memory 下只变更所引取数纪律一个变量。 base 本身已拒绝下结论并想到更名与交叉验证——**差分不在一般判断力上**;base 缺的是可执行机制与处置: 包含式复检、复检只能推翻不能确认这一方向限制、以及证实不了即进待核验清单不进承重结论,四项标记(包含/复检/独立/待核验)base 全 0、candidate 全部出现。 判分器可失败性由 base 产出一份合理且谨慎的答复而非拒答证明。 观察到的失败取自一段多轮公开材料调研的逐轮修订表与自评:某实体被误判为消失(实为归一化在新名下漏配), 以及同一条高权威、零产出的入口被连续两轮写进下一步建议、走完才发现不成立 |
|
|
327
|
+
| 一条规则可以触发正常而内容为假:把「反复润色」几乎全部归因于修订伪装成润色、并断言那是最常见的机制,这个最高级没有证据支撑,且会把一个正当的「再润色一遍」误读成实质未定。撤回该断言,并补上被它掩盖的另一支——润色本身多轮收敛:一致性与口径漂移、跨节重复、密块、元语自证是逐位置缺陷,每一遍改动都可能重新引入前一遍已清掉的类。判据必须钉在**实质的状态**而不是本轮请求的措辞,否则实质未定但只收到措辞请求时会放行润色,抵消同段前半句的前置约束 | `tighten-doc` | behavioral-evidence: RED-baseline; observed-failure: yes; result-class: failure; firing-path: file:skills/tighten-doc/references/doc-charter-first.md#反过来不成立:润色本身就是多轮收敛的 | updated | `tighten-doc/SKILL.md` 是 owner key 且本轮未改;改动只在 `skills/tighten-doc/references/doc-charter-first.md` 的同一条 bullet 内。 前半句的前置约束(实质未定不进润色轮、顺序不可倒)逐字保留,新增的只是第二支与其判据;类目、判法与多轮节拍仍由 `tighten-doc/SKILL.md` 单一持有(DELETE 的元语自证类、closeout 的一坨/跨节重复/族内术语漂移、「用户还能单 paste 挑出同类缺陷 = 清单未真跑」与 Full-pass ≠ token-pass 的注意力摊薄诊断),本行不复述也不新增类。 **observed-failure: yes 的依据**:失败由用户的一手实践报告提出(润色确实要多轮,且类目本仓早已持有),并在探针上复现——该规则触发正常,错的是断言内容。 RED-baseline(applied, differential):仓外中立目录、`--safe-mode --tools "" --setting-sources ""`、禁 CLAUDE.md 与 auto-memory,只变更所引纪律文本一个变量,提问「已润色到第 6 轮、每轮仍能挑出一致性/跨节重复/密块,是不是实质没定」——base 答「大概率是……说明命题/结构还没收敛」并建议停下退回 charter/Genre、别再要求再润色一遍;candidate 答「不必回 charter/Genre……这是润色轮的正常特征」并给出按类分趟走查的做法。**两侧给出方向相反的操作建议**,按用户的一手实践 base 的建议是错的。 判分器可失败性:base 产出的是一段自信连贯的建议而非拒答。 两次 A/B 均**未测出行为差分**并如实记账:(1) 「实质已定但一致性差/跨节重复/密块,要求再润色一遍」——base 自己就判定为真正的润色轮且未回 charter,我假设的误判**未复现**; (2) 把 closeout 表头改成「逐位置判」的候选形态——base 与 candidate 对同一份植入四处缺陷的短文档都全数命中,差分可忽略。 因此该形态提议**未落地**(探针文档过小、注意力摊薄这一真实机制未复现;继续构造更长文档直到显效即是偏向性测量,按测量纪律停手)。 本轮落地的只有真值修正,其正当性来自断言无据本身,不来自行为差分 |
|
|
328
|
+
| 一道公共泄漏闸若把**另一个领域的普通词汇**当作专有标识的代理,它会随那个专有对象失效而变成纯误报来源:该项随初始提交带入、无理由记录,维护者确认它当初是为早期提炼项目的一个专有产品模块名而设,模块已不适用,而词本身是通用技术用语。退休它而非继续维护:真实标识符的权威闸是私有 alias 审计,公共词表只是无该命令环境下的弱代理兜底;删除前先量化今天有多少内容依赖该信号,并用保留项做对照证明闸仍能报错。同类对照就在同一目录:`generic-r0-leak-scan.sh` 的谓词是它自己拥有的凭据形状,有套件;这条借别人的词表,此前无任何测试 | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; result-class: failure; firing-path: command:skills/skill-extraction-workflow/scripts/check-ccl-skills.sh | updated | `skill-extraction-workflow/SKILL.md` 是 owner key 且本轮未改。改动三处:`scripts/check-ccl-skills.sh`(正则去掉一个分支,并就地写下该词表的定性——弱代理兜底、权威闸是私有审计、按退休维护而非增长,因为真正的不变量「公共文本是否指向某个具体组织」本就不是公共正则可判的)、新增 `scripts/test_entrypoint_domain_scan_terms.sh`、以及把它注册进 heavy lane(克隆型套件按既有判例入 heavy;未注册即假绿)。 observed-failure:本轮撰写台账时被该项挡下一次,属正常技术用语被误挡。 四腿(applied, differential):腿1 只读跑真仓证明该闸仍在跑且干净;腿2 从脚本**现读**那条正则(用 awk 取,不用 lookahead——要求 PCRE2 会让没编该特性的 rg 整条 lane 挂掉,而被测闸本身并不需要它),验保留项与非词表分支仍匹配(源绑定,不会与出厂内容漂移);腿3 验退休项在正常行文里不再匹配; 腿4 端到端——在 `--shared` 克隆上**先跑一次不带 fixture 的对照**(须 exit 0 且该闸报干净),再植入命中 fixture 跑,要求非零退出、报出失败 token 且点名该 fixture;没有那次对照,删掉 `exit 1` 再叠加克隆上任何无关的后置失败,同样会给出非零与该 token,腿4 会在坏闸上变绿。**该腿的归因主张经过一次撤回**:最初想证明「它停在这一步」(先用『该闸之后的首个 token 缺席』,后改用『终态成功 token 缺席』),连续四轮被两条 lane 指出同一类洞——黑盒测试看不见控制流,任何 token 形态都只是代理:`exit 1` 被删、再叠加一个由 probe 自身触发的后置失败,非零、失败 token、probe 名、终态 token 缺席会同时成立。按同类跨轮复发即设计信号处理,**撤回该主张而不是再加机制**。腿4 现在只主张两件真实成立的事:植入命中后 (a) 该闸报出失败并点名 probe,(b) 整跑拿不到认证。**不主张这次未签发认证由该闸独占造成**——这是明写的残留,不是已验证性质。 **腿4 是评审逼出来的**:只有腿1-3 时,失败分支被绕过或反转仍会整套全绿(腿1 只证明干净时打印了 ok,腿2-3 只验正则本身)。 四个 applied 变异证明非空洞且归因可分:把退休项加回去只红腿3;删掉一个保留项只红腿2;中和失败分支只红腿4、腿1-3 仍绿;只删 `exit 1`(保留 token 输出)红在腿4 的退出码断言;复合变异(删 `exit 1` + 注入后置失败)被干净克隆对照接住。对照均绿。 **残留风险(accepted,风险 owner = 维护者,非 agent 自受)**:无 `ALIAS_AUDIT_CMD` 的环境(CI、他人机器)今后对「唯一公共信号是该项」的内容失去覆盖。 量化:该闸覆盖面(`skills/*/SKILL.md` 与 `skills/*/references/*.md`)内当前 0 处,故今天无任何内容失去覆盖;风险为前瞻性,且该通用词本就是那个模块名的弱代理 |
|
|
329
|
+
| 交付文档的图与表有起草期可判定形态契约:图种由主张形态推出、画了必须满足记法硬约束(标题/图例/连线单向且标签具体)、版式契约冻结后按契约符合性判一致性(不用魔数阈值)、量级对比与结构关系不得全压进表、图文引用一致 | `tighten-doc` | `tighten-doc` 的图/表形态判据从散文升级为可机械判定:`tighten-doc/references/figure-and-table-craft.md` 定判据与其依据档位(外部权威 / 工程约定 / 查证后禁用),`tighten-doc/scripts/figure-lint.py` 与 `tighten-doc/scripts/doc-lint.py` 实现可判定部分,`tighten-doc/scripts/test_figure_and_doc_lint.sh` 做逐谓词差分自测并注册进 `make test-repo-gates`; result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/tighten-doc/scripts/test_figure_and_doc_lint.sh | `updated` | owner key `tighten-doc/SKILL.md`(「表达形式匹配内容」条就地 merge 判据 + 指针,入口净减过体积闸);`tighten-doc/references/deliverable-doc-genre-skeletons.md` §4/§5 就地 merge。**观察到的失败**(本轮实测):18 张真实架构图里 17 张无图例、12 张连线无标签、7 张文本对比度 2.58:1 低于 WCAG AA 4.5、35 处超长文本零 tspan 必溢出、6 张零分组、5 张画布比例偏离;一份 1059 行调研报告 45 表 0 图、23 处数值列无单位、17 处加粗行冒充小节标题。**RED-baseline 差分是可重跑脚本 `tighten-doc/scripts/mutation_probe.sh` 而非声明**:逐谓词把 code 字面量替换使其消失,24/24 均使套件转红,任一仍绿即退出非 0。该实验本身先失败两轮、两轮都是无效突变:改的是 oracle 不观测的维度(severity 而非 code),以及把脚本崩溃读成绿(判据只看 FAIL 行不看退出码)。**四轮独立 review(34 条 / 29 P1)+ 一轮对抗 challenge(6 条 / 5 P1)全部处置**。对抗轮的结论改变了投放形态:C4-LEGEND 与 C4-EDGE-LABEL 实测**两个方向都会错**(两条目的合法图例被拒 / 注释里出现 legend 即放行;中点附近的分区标题掩盖未标注连线 / 标在别处的合法标签被拒),已**降为非阻断 WARN**——C4 规则本身是 [外],用邻近与色块计数去认它是 [工] 代理,代理不该挡他人提交;要恢复阻断需结构化关联(契约声明的组 / aria-labelledby / textPath),不是把阈值调准。另修:比例先舍入再做「精确」比较等于引入未声明容差(16:9 的 1.7777… 会通过 [1.778])→ 改用未舍入值比较;px() 把 1em 当 1px、12pt 当 12px → 按单位换算,不认识的单位报出;doc-lint 对坏文件返回空发现集 + skipped=true 使损坏文档「干净通过」→ 与 figure-lint 对齐报 READ 并继续整批。**独立评审第一轮 10 条发现全部处置**,7 条 P1 均为真缺陷,最重一条推翻了本轮自己的证据声明——run_expect 加参数漏改调用点致控制组断言被吞、弄脏控制组仍全绿;已改为参数不足判红 + 期望表全量对应,两条回归实测通过。其余修法见 tighten-doc/scripts/AGENTS.md 的 Validation 节 |
|
|
330
|
+
|
|
331
|
+
<!-- 053 轮独立评审第一轮处置(候选 4293dbe7…,reviewer 经原生绑定读取 5 个 owner 技能):10 条发现全部处置,7 条 P1 均为真缺陷。最重一条推翻了本轮自己的证据声明——run_expect 加了参数却漏改调用点,控制组断言被当成参数吃掉,弄脏 control.svg 仍全绿;已改为参数不足直接判红 + 期望表全量对应,两条回归实测通过。另六条 P1:path_mid 对单段路径返回终点致终点节点名被当作连线标签(改按弧长取中点并排除节点内文本);GROUPING 只查全图有无任意 <g>(改为逐卡片判定形状与文字是否同组);契约配色只收文本色(改收全部渲染元素的 fill 与 stroke);契约损坏只记不报致退出码为 0(新增 CONTRACT-INVALID 并做类型校验);比例容差内置 0.02 是魔数(改由契约 ratio_tolerance 声明,未声明即精确匹配);只用首个文件的目录找契约(改为逐文件解析最近契约)。P2 的 CARRIER-IMBALANCE 计数阈值被包装成 ICD 203 要求,已降级为 WARN 且仅在存在量级对比表时提示。 -->
|
|
332
|
+
|
|
333
|
+
<!-- 053 轮独立评审第二轮处置(候选 c696c516…):9 条发现全部处置,8 条 P1。最尖锐一条打在本轮论点正中——ICD203-9 检测被标 [外],但其触发阈值(≥6 行 / 数值占比 >0.8 / ≤3 列)是检查器自定的;一条声称按依据分档的规则自己混了档。修法:载体选择原则标 [外]、识别阈值标 [工],且输出里直接带档位,读者当场可见哪半有依据。两处真绕过口:空的 figure-contract.json({})既不报缺失也不报损坏、三项检查全跳过(现三项字段必填非空);分组判据只要求祖先 g 集合有交集,把整图包进一个顶层 g 即全部通过(现要求卡片专属的最近公共组)。最有价值的修法是把突变清单从**源码派生**而非手写:评审指出手写 24 条漏了 PARSE / C4-VIEWBOX / C4-EDGE-VAGUE,补三条治标、让清单不可能漏才治本;派生后立刻暴露 6 条无覆盖并全部补了 fixture,现 28/28。其余:连线改为先全量识别再校验方向(原只收已有箭头的,无箭头连线根本不进检查,而『每条线单向』正要求它有方向);底色解析不了时报 CONTRAST-UNSUPPORTED 而非回退白底(回退白底会产生假绿或假红);悬空图号按实际编号比对而非数量;viewBox 非法不再抛异常而是报 GEOMETRY。 -->
|
|
334
|
+
|
|
335
|
+
<!-- 053 轮全量试跑与删除收束:两轮独立评审后又做了一件评审做不到的事——把检查器放到 392 份本仓正常文档上跑。结果 718 条发现、318 条 ERROR、命中 150/392 份文件,其中四条阈值型谓词(标题过长 / 单元格塞整段 / 段落内并列枚举 / 列表项多句)独占 96.7% 的发现量与 99.7% 的 ERROR,抽样中无一条是其声称的缺陷。判据错误而非过严:它们都拿宽度或计数当「表达好不好」的代理,而这类阈值经本仓既有 bench 查证无可靠来源——等于把 [禁] 档的东西改名重立。四条全删。删后 23 条、ERROR 0、涉及 13/392。全量试跑另暴露两条前提性错误:围栏代码块内容被当成文档结构;「引用悬空」在零图文档里根本不成立(本轮自己写的判据文档就是第一个受害者)。**方法论落点**:fixture 只证明谓词能报,不证明它报得对——fixture 是作者造的、天然符合作者假设;命中率必须在没有为它准备的真实语料上量。该纪律已写入 tighten-doc/scripts/AGENTS.md 与 craft §9。两轮 diff 评审均未抓到这一类,因为评审看不到全仓命中分布——这是评审的结构性盲区,不是评审失职。 -->
|
|
336
|
+
|
|
337
|
+
<!-- 053 轮独立评审第三轮处置(候选 9d8b44a4):8 条发现全部处置,7 P1。又一次打在证据上——突变探针的源码派生只认 add(...),漏掉走 findings.append({'code':...}) 的两条契约谓词,分母偏小使『全部谓词均有覆盖』再次成为未验证声明。修法不是补两条:派生覆盖全部发射路径,并加 DERIVATION-GAP 等价断言(实际输出过的 code 必须都在派生清单里),该断言经反证有效——把一个 code 改成正则识别不了的形态即当场报缺口。现 26/26。两条『修假阳性时制造假阴性』:整体跳过 defs 使箭头 marker 的契约外颜色查不出(改为只纳入被 marker/use 实际引用的定义,反证通过);strip_fences 把 ```mermaid 起始行也清掉致 mermaid 图恒为 0(改为剥离前计数)。一条新引入的误报源:把『所有有描边的 path』当连线,分隔线与网格线会被要求箭头和标签(改为必须显式声明 class 或自带方向 marker,类名可由契约覆盖,且样式沿祖先继承解析)——与本轮删掉那四条同类,**该修法第一次落地时被后续 patch 覆盖丢失,而台账已先写上「已改」,第四轮评审实测出装饰线仍被判红才发现——台账写了没做的事比缺陷本身严重,现已重新落地并双向验证(装饰线不判、声明为 flow 的线仍判)**,说明『判据是否会误报』要在每次新增判据时重问,不是删过一次就免疫。其余:含 transform 时跳过全部依赖几何的 ERROR 而非硬报;有编号图注却无真实图实例判悬空;退出码契约(0/2/1,含 --json)此前全程被 || true 忽略,现已加回归。 -->
|
|
338
|
+
| 图的布局属性有一支真正做过实验的文献,可进 [外] 档:削减边交叉的收益远大于减弯折与提对称,而正交网格对齐与出边夹角统计上不显著(故不值得为它们做取舍);不可避免的交叉角度越大越好但不必直角 | `tighten-doc` | `tighten-doc` 的图判据补上此前完全缺失的布局面:`tighten-doc/references/figure-and-table-craft.md` §4b 记入两项实证及其三条边界(实验对象是抽象点线图非带标签架构图、给的是排序非阈值、只有边交叉与弯折可机械算),`tighten-doc/scripts/figure-lint.py` 新增 GRAPH-CROSSINGS 谓词只报数不设阈; result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/tighten-doc/scripts/test_figure_and_doc_lint.sh; bank-evidence: file:eval/routing-tasks.jsonl#route-figure-craft-unease | `updated` | owner key `tighten-doc/SKILL.md`(本轮未改其正文,改的是其 reference 与 script)。**观察到的失败**:上一轮(053)跑过一次「核行业最佳实践」并就图表密度、行宽等给出「无可靠来源」结论——该结论对**问过的问题**成立,但整支图布局实证文献从未进入视野,是从外部技能仓的一行引用里发现的。**完备性检查只在自己的问题集上跑,报不出没问过的问题**;便宜的补法是看邻近同行引什么。**RED-baseline**:`mutation_probe.sh` 29/29,新谓词由探针的 DERIVATION-GAP 自动纳入清单并当场验出覆盖;双向 fixture 实测——交叉图报「交叉 1 处、最小交叉角 25°」,不交叉的图不报。**只报数不设阈**是刻意的:实证给的是排序不是门槛,设阈即变成 053 轮删掉那四条谓词的同类。另落人读入口 `docs/figure-and-table-handbook.md`(不进 agent 加载面,仅供队友)与三条图表类路由 eval。 |
|
|
339
|
+
| 能力落进技能正文与脚本,不等于入口开了——routing surface(SKILL.md 的 description)没提到的能力,请求根本到不了这个技能 | `tighten-doc` | `tighten-doc` 的 description 补图表触发词并写死边界:图与表的**表示形式**(图种由主张形态推出、记法硬约束、版式契约、对比度、可跑检查器)归本技能,**系统边界与架构决策本身仍归架构技能**;443→594 字符,未超 800 预算; result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/SKILL.md#description; bank-evidence: file:eval/routing-tasks.jsonl#route-figure-type-choice | `updated` | owner key `tighten-doc/SKILL.md`。**观察到的失败是实测不是推断**:053/054 两轮把判据、reference、两个检查器、29 条谓词、人读手册全部落地,唯独没动 routing surface。**先测后改的差分**:改 description 前跑 Tier-2 routing bank(claude-haiku-4-5 grader)得 **1/3**——「该画时序图还是状态机」被路由到 product-rd-workflow(conf 0.65),「架构图画完总觉得不对劲说不上哪不对」被路由到 grill-me(conf 0.92,措辞完美匹配拷问技能,而它恰是本判据最该接的场景);只有「全是表格要不要补图」命中。改 description 后同一 bank 重测 **3/3**,题未变、顺序为「先测→改→重测」。`make eval-routing` 静态分析零阻断零 advisory(与架构技能无触发词碰撞)。**残余限制如实记**:① co-change 警告结构性存在(本轮确实同时动了 bank 与 description),顺序只能靠本行记录供人核,工具看不到;② grader 是 cheap-LLM、advisory 不阻断、可能错,不是真值裁决;③ 三条任务的 frozen_at_sha 初次填成非 SHA 的 \"054\",被 frozen-drift 排除在回归判定外,已改挂到真实祖先 2a5bc24 后重测仍 3/3。 |
|
|
340
|
+
| 图表判据里唯一有生理机制的一支是色觉障碍:判据不是「避开某些颜色」而是「别让语义只落在同一条混淆线上的色对」,且实算优于模式匹配 | `tighten-doc` | `tighten-doc` 补此前完全缺失的色觉与主题面:`tighten-doc/references/figure-and-table-craft.md` §4c 记混淆线机制与三条边界、§4d 记图值不值得画与退化形态;`tighten-doc/scripts/figure-lint.py` 新增 CVD-DISTANCE / THEME-CONTRAST / THEME-PURE-BLACK / FIGURE-IS-A-LIST / FLOW-DIRECTION-MIXED / VALUE-SHOULD-BE-TOKEN / FIGURE-A11Y-STRUCTURE 七条谓词,契约新增 themes 维度,并按独立评审的七条发现修正后落地; result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/tighten-doc/scripts/test_figure_and_doc_lint.sh | `updated` | owner key `tighten-doc/SKILL.md`(本轮未改其正文,改的是 reference 与 script)。**观察到的失败**:上一轮只读外部仓 README 表格即下结论,正文里有四类 README 看不见的判据——digest-masks-corpus 实例。**实算 vs 模式匹配的差分是实测的**:按外部技能给的「红橙/蓝橙/低饱和」色对模式,在本仓真实 token 集上只命中 warn/critical 一对;改为 Viénot 1999 投影实算后,primary/muted(protan 距 14)与 primary/ok(tritan 距 19)两对也出——模式匹配漏三分之二。参照点:Okabe-Ito 八色最小距离 38、随机八色 10、本仓原 token 集 14。**主题维度**实测:`#111827` 压纯黑仅 1.18:1,说明写死十六进制的 token 表在深色主题下整体失效。**RED-baseline**:`mutation_probe.sh` 37/37,新谓词均由探针的 DERIVATION-GAP 自动纳入并验出覆盖;fixture 全数改用 Okabe-Ito 配色——原配色在 tritan 下仅 19,会让每个 fixture 连带触发 CVD-DISTANCE、谓词隔离不开。**过程瑕疵如实记**:charter 写于前三个源已读之后;对照清单时我对「深色模式」打过一次假勾(实际只落了 CVD、契约主题维度未做),自查时撤回补齐。**独立评审七条全部成立、全部已修**(5×P1):①`themes.*.surface` 只判真值 → `"black"` 过校验后被静默跳过=声明了主题却一次没测;②主题对比度拿全局底色判压在局部卡片上的文字 → 误报,改为只判压在整幅底板上的文字(覆盖≥95% 画布算页面底色);③CVD 比的是全部渲染色(含装饰、边框)并拿 25 当触发条件,而同文件写着「只报数不设阈」——**声明与实现相反**,改为只测契约 `semantic_colors` 显式声明的语义色且无条件报出距离与参照点;④`VALUE-SHOULD-BE-TOKEN` 规则写 >2 次即报、实现要求≥3 个颜色 → 单色重复永不报,已对齐并补一色/多色两个用例;⑤流向判据在几何不可信时仍跑、且忽略 marker-start 与双向边 → 加前提并按 marker 语义定向;⑥主题与 token 两组独立断言用 `\|\| true` 吞掉退出码 → 补退出码断言;⑦发布档要求候选绑定的机器可验证证据 → 产出 `eval/evidence/figure-lint-predicates-2026-08-25/REPLAY.md`(**tracked,在冻结包内**): 列出三条命令与期望退出码/关键行,前两条完全由包内文件决定、评审方可自行复算;第三条依赖维护者私有 alias, 如实标为**不可从包内独立复核**而非当成已验证。此前那版证据放在 gitignore 的 `.work/` 下、根本不在包里, 且第一版把套件退出码取成空(PIPESTATUS 落在子 shell)、又把 gates 写成跑在提交后的 tree 上——正是本条所指的缺陷形态本身。**第二轮 dual-track(review+challenge 同绑候选 1f1e5f06)再报 9 条,两条 lane 独立收敛到同一批**,全部已修: ⑧`CVD-DISTANCE` 走 `WARN` 档 → 退出码非零,「只报数不设阈」这句声明被退出码当场推翻;新增 `INFO` 档,不影响退出码; ⑨`semantic_colors` 未受契约校验 → 数组/字符串在 `.values()` 处直接崩,非法颜色被静默丢成空集,「没报 CVD」分不清是「距离没问题」还是「压根没测」;现校验类型、`#RRGGBB` 语法与 `color_tokens` 子集三项; ⑩`fill="none"` 的描边框被当成卡片底板 → 框住低对比文字的边框把它豁免掉;卡片须为实际填充色; ⑪底色解析不可信时主题检查仍执行 → 圆形/路径/渐变卡片不进 `rects`,「查不到卡片」被当成「文字压在主题底上」,判出假红;改为整条跳过并显式报新谓词 `THEME-UNASSESSED`(未判定必须说成未判定,不能沉默成绿); ⑫两端都没箭头的边按 `d` 的书写顺序被计入流向 → 三条视觉上无向的线仅因端点顺序不同即可触发混合流向;无向边跳过。 **修 ⑧ 时当场炸出探针自身的第二个分母缺陷**:`mutation_probe.sh` 的派生正则只枚举 `ERROR\|WARN` 档, 把一条谓词改成 `INFO` 就让它**静默退出分母**——总数看着没变、覆盖少了一条。 档位改为 `[A-Z]+` 不枚举后为 37/37。**覆盖率的分母若由被测代码的某个属性决定,改那个属性就能无声地让分数变好看**。 **修正后复测**:套件全绿、突变探针 differential_sensitivity=37/37 且 DERIVATION-GAP 无缺口、`make test-repo-gates`=0(含 alias_audit_ok / r0_status=private-ok)。**真实语料复测 55 张 SVG**(非本仓,无契约):`FIGURE-IS-A-LIST` 17、`VALUE-SHOULD-BE-TOKEN` 34、`CVD-*` 0(无契约声明语义色=不测,不是「测过没问题」)、`FLOW-DIRECTION-MIXED` 0——较修正前的 5 条下降,因为 55 张里 45 张含 gradient/非矩形底,几何不可信,那 5 条原本就是拿不可信坐标判出来的。 |
|
|
341
|
+
| 自检触发点漏掉了 gap-list 形态:「列举外部源有而我们没有的」就是 findings 回合,但它读起来像在答覆盖问题,charter-before-findings 从不觉得被触发 | `skill-extraction-workflow` | 自检触发点扩到 gap-list 形态并写明它为何最易滑过;自审纪律补「验证锚点本身可能是错的」——锚点的预期方向若来自你对一手源的解读,该解读是 hypothesis-grade,锚点不过应先质疑锚点而非判实现有 bug; result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#enumerates what an external source has that we lack | `updated` | owner key `skill-extraction-workflow/SKILL.md`。两条均为 **merge 非 append**(并进既有的自检触发点与自审纪律)。**观察到的失败**:①本轮产出四条「外部有我们没有」的缺口清单在先、invoke 提炼工作流在后,而规则明写 charter 要在第一个 findings 回合之前——规则在、没触发,故补触发点而非补规则。②用「红色模拟后应变暗」验 CVD 实现,实测亮度上升;追一手后确认是模拟把红投影到 575nm 黄色不变轴、亮度上升为正确行为,而「红色看起来暗」说的是红与黑难分、是另一个量。**危险在于其余两个锚点通过**——错的尺子量对的实现,比实现出错更难发现。**RED-baseline 差分**:改前的触发点只列举「沉淀 / 提炼 / 复盘 / distil」四种自述措辞,本轮实际产出的是「缺口清单」措辞,逐字不匹配任何一项,故规则在而未触发(本轮即为该差分的观测实例);改后触发点显式含 gap-list 形态。可核对方式:在改前文本上检索本轮的触发措辞得零命中,改后命中。 |
|
|
342
|
+
| 确定性闸对历史形态有前提,而**递给它哪个 ref 是 harness 的选择**:闸按分支自己的一级父链切轮,CI 默认检出的 `refs/pull/N/merge` 的一级父是**目标分支**,于是分支上所有轮塌成一轮,凡「验证依赖本轮很窄」的台账行都会在自己一个字没改的情况下翻红——闸没坏,喂它的历史形态不对 | `skill-extraction-workflow` | 五个跑闸 lane 的 job 全部改为检出分支 head(`ref: pull_request.head.sha \|\| github.sha`,非 PR 事件回落 `github.sha`);新增 `test_ci_checkout_ref_binding.sh` 逐 job 断言该绑定并注册进 fast lane,绑定丢失即闭式失败; result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh | `updated` | owner key `skill-extraction-workflow/SKILL.md`(本轮未改其正文,改的是 scripts 与 `.github/workflows/ci.yml`)。**观察到的失败是 CI 实测**:PR #53(dev→main,053–056 四轮 + npm 0.5.0)`repository-gates` 报 `impact_chain_firing_path_missing`(incomplete 点名的是文档技能那个 owner 的入口),`regression-heavy` 两个套件连带红;同一棵树在线性 dev 上闸全绿。055 那轮只改 tighten-doc 的 description、用 `#description` 锚点,该锚点合法的前提是「该 owner 本轮全部 diff 就是 description」;四轮塌成一轮后 tighten-doc 在同一轮里还改了 `references/` 与 `scripts/`,前提失效。**先试过改闸、证否了**:把 round 派生改成「一级父已被 base 包含时走第二个父」,在 merge-ref 树上确实转绿,但打破了 `test_check_ccl_impact_chain_refscripts.sh` 的 round scoping 8——「台账行写在工作提交之前的 worktree 轮必须塌成一个边界」。两种形态在拓扑上**完全相同**(都是 merge X into Y 且 Y==base),git 里没有可区分它们的不变量,而该 fixture 明写 `--first-parent` 是承重的。按仓里对 form 3 的既定裁决(同类复现以删除收束、不再补代理),改的不是闸而是喂给它的 ref。**RED-baseline 是双向差分**:新测试在修好的 workflow 上绿,把 `.github/workflows/ci.yml` 换回 `origin/dev` 版即报三个 job 未绑定。该测试第一次跑就抓到我只改了三个 job、漏了两条 regression lane——**这正是它存在的理由**。**本轮暴露的验证面缺陷**:053–056 四轮我只跑 `make test-repo-gates`、基线取 `origin/dev`;而 CI 的基线是 `origin/main`、检出是 merge ref、还有 fast/heavy 两条 lane——三处都不同,四轮的绿从未覆盖这个形态 |
|
|
@@ -46,12 +46,12 @@ Use this before reading deeply or editing any skill. The charter is the guardrai
|
|
|
46
46
|
|
|
47
47
|
| Field | Required answer |
|
|
48
48
|
| --- | --- |
|
|
49
|
-
| Purpose | What future failure
|
|
49
|
+
| Purpose | What future failure or drift should this extraction prevent, or what evidenced success mechanism should it preserve and reuse? |
|
|
50
50
|
| Scope | Which skill(s), source classes, users/tasks, and sibling boundaries are in scope? What is out of scope? For an iterative-program source (multi-round research/writing/delivery), also answer: does a project-local `covered-through` watermark exist, and what does this round cover above it? (see the watermark rule under Task Retrospective Extraction) |
|
|
51
51
|
| Depth | Is this wording cleanup/no new source read, targeted check, file-level refresh, node/artifact inventory, full workflow extraction, or generator/tooling change? |
|
|
52
|
-
|
|
|
53
|
-
|
|
|
54
|
-
| Failure mode
|
|
52
|
+
| Result classification | Is the observed result a failure/correction, stable success, or unstable/insufficient evidence? What observation supports that classification? |
|
|
53
|
+
| Matching analysis | Failure/correction: scaled RCA from observed issue to controllable prevention. Stable success: reusable mechanism, non-luck evidence, reuse conditions, firing point, and owner. Unstable/insufficient evidence: what remains unknown and why no executable rule lands yet. |
|
|
54
|
+
| Failure mode or success boundary | What bad output would weak extraction permit, or under which conditions would the success mechanism stop transferring? |
|
|
55
55
|
| Lifecycle impact | Which stages are affected: product intent, design/UX, implementation, debugging, testing, launch acceptance, iteration feedback, team onboarding, and use without source access? |
|
|
56
56
|
| Evidence plan | Which source categories must be inspected, routed, discarded, or marked unavailable? For a task/session retrospective over a session that produced artifacts, the FIRST source class listed MUST be those produced artifacts (deliverables, reports, scripts, datasets — the a0 enumeration, owed at charter time, not only before an exhaustion claim); a session with genuinely no produced artifacts records an explicit `produced artifacts: not-applicable` entry carrying the reason, the minimum checked surfaces (deliverable directories, script/output locations, dataset paths), and a resolvable inventory-check locator — the command or listing that establishes absence — instead of the class. Either way, correction turns and the agent's own summaries are friction-biased digests — they record only what rubbed, so what went RIGHT is structurally invisible in them — and cannot substitute for the artifact class or excuse skipping the check. |
|
|
57
57
|
| Completion standard | What pressure scenario, independent review, command, install check, or source-map evidence proves done? |
|
|
@@ -82,30 +82,30 @@ If the user asks for complete, deep, full, or repeat extraction, or challenges s
|
|
|
82
82
|
|
|
83
83
|
Before a source portfolio can confirm or contradict a reusable rule, classify it as `stable` (in production, not slated for replacement), `evolving` (actively iterating, design not frozen), `legacy-deprecating` (scheduled for retirement), or `mixed`. Only `stable` portfolios can be used as confirmation/contradiction baseline. `Evolving`, `legacy-deprecating`, and `mixed` portfolios are audit/anti-pattern signals unless the extraction is explicitly downscoped to that status; state the long-lived caveat reason instead of implying future verification will upgrade it automatically.
|
|
84
84
|
|
|
85
|
-
## Baseline
|
|
85
|
+
## Result-Learning Baseline For Every Extraction
|
|
86
86
|
|
|
87
|
-
|
|
87
|
+
Classify the result before choosing an analysis method:
|
|
88
88
|
|
|
89
|
-
|
|
|
90
|
-
| --- | --- |
|
|
91
|
-
| Future
|
|
92
|
-
|
|
|
93
|
-
|
|
|
94
|
-
|
|
95
|
-
|
|
89
|
+
| Result class | Required analysis | Landing condition |
|
|
90
|
+
| --- | --- | --- |
|
|
91
|
+
| Failure or correction | Future bad outcome, contributing factors, counterfactual ranking, controllable prevention, firing path, and owner | Evidence shows the control addresses the failure class; known failures use correction RCA |
|
|
92
|
+
| Stable success | Mechanism that produced the result, evidence it recurs and is not luck, reuse conditions and transfer boundary, firing point, and owner | The mechanism is observable and reusable; praise or a single good run is insufficient |
|
|
93
|
+
| Unstable or insufficient evidence | What was observed, competing explanations, and missing evidence | Observation only; no executable rule until the classification becomes supportable |
|
|
94
|
+
|
|
95
|
+
**The classification is not self-elective.** Two of the three classes skip RCA, and the agent choosing the class is the same agent whose work the RCA would examine — so left to the author the cheap classes are always available. Therefore:
|
|
96
96
|
|
|
97
|
-
|
|
97
|
+
- An extraction triggered by a correction, a review finding, a failed run, a regression, or a user pointing out a miss is **`Failure or correction` by default**, whatever the author's own reading of it.
|
|
98
|
+
- Relabelling such an extraction into `Stable success` or `Unstable or insufficient evidence` is a claim the **independent review must accept**; the author recording the relabel is not the adjudication, and neither is a reason written into the round artifact.
|
|
99
|
+
- Every classification carries the observation that would **disconfirm** it — for `Stable success`, what would show the result was luck; for `Unstable`, what evidence would settle it. A class with no disconfirming observation named is unrecorded, not recorded-and-passed.
|
|
100
|
+
- Missing classification, an unaccepted relabel, or missing matching analysis leaves the extraction `interim`, and `interim` is a state the closeout reports rather than a label the author clears.
|
|
98
101
|
|
|
99
|
-
|
|
100
|
-
- Targeted check: future failure, source boundary, and proof.
|
|
101
|
-
- File-level or broader extraction: full table above, plus lifecycle impact, run through the Deep RCA five moves below (the table's single "enabling cause" row becomes the *set* of contributing factors).
|
|
102
|
-
- Broad or multi-skill extraction: full table above, source register, target-output map, independent review, and the Deep RCA five moves below.
|
|
102
|
+
Scale depth to the task. Wording cleanup records one concise classification and owner. A small non-wording failure gets a widen-check plus control; a non-trivial failure uses Deep RCA. A stable-success extraction increases evidence depth with the breadth of the reuse claim. Broad or multi-skill extraction also requires the source register, target-output map, and independent review.
|
|
103
103
|
|
|
104
|
-
If the
|
|
104
|
+
If the analysis reveals a product decision, architecture decision, test strategy, design readiness issue, source-access problem, or sibling-skill update, route it before editing the target skill.
|
|
105
105
|
|
|
106
106
|
### Deep RCA For Extraction
|
|
107
107
|
|
|
108
|
-
|
|
108
|
+
For a failure or correction, a causal account is required; 5 Why is only the **entry technique** to get past a visible symptom. Used alone it has a documented failure mode: it traces ONE linear chain to ONE "root cause", is bounded by the investigator's current knowledge, is non-reproducible (different agents reach different ends), and the word "why" drifts toward "who" (blame) and toward hindsight. Most process/agent failures are not single-cause — overt failure requires several contributing causes to coincide — so for any non-trivial failure extraction run the fuller method below, not just a why-chain. Stable success uses the Result-Learning baseline above instead of inventing a failure. (For pure wording cleanup — the strict wording-only test, no trigger/scope/routing/validation/owner-meaning change — one concise classification is enough.)
|
|
109
109
|
|
|
110
110
|
Do not force exactly five questions, and do not accept a single straight chain. Ask enough "why" to leave the symptom; ask "how/what conditions" to widen; stop a branch once its next action is concrete and owned.
|
|
111
111
|
|
|
@@ -170,18 +170,17 @@ Source-specific prompts (where each branch's RCA should resolve):
|
|
|
170
170
|
|
|
171
171
|
## Task Retrospective Extraction
|
|
172
172
|
|
|
173
|
-
Use this when the user asks to summarize this task, summarize lessons learned, review what went wrong, or turn the current session into reusable team practice.
|
|
173
|
+
Use this when the user asks to summarize this task, summarize lessons learned, review what went wrong or right, or turn the current session into reusable team practice.
|
|
174
174
|
|
|
175
|
-
The current task is a source, but it is not automatically a skill rule. Treat task history as evidence and
|
|
175
|
+
The current task is a source, but it is not automatically a skill rule. Treat task history as evidence and classify each result before choosing RCA, success-mechanism attribution, or observation-only treatment.
|
|
176
176
|
|
|
177
177
|
Required flow:
|
|
178
178
|
|
|
179
179
|
1. Define the task boundary: which user request, implementation slice, review, bug, correction, or validation result is being summarized.
|
|
180
|
-
2.
|
|
181
|
-
-
|
|
182
|
-
-
|
|
183
|
-
-
|
|
184
|
-
- Which skill, validator, shared project doc, memory note, repo doc, or final-response rule owns the prevention?
|
|
180
|
+
2. Classify the result and run the matching analysis:
|
|
181
|
+
- Failure/correction: what bad outcome would repeat, which factors enabled it, which control should have caught it, and which owner must carry the prevention?
|
|
182
|
+
- Stable success: what mechanism produced the result, what proves it was not luck, under which conditions it transfers, where it should fire again, and who owns it?
|
|
183
|
+
- Unstable/insufficient evidence: which explanations remain open and what evidence is missing? Keep it as an observation.
|
|
185
184
|
- For delivery-chain failures, ask why the requirement/contract was not defined correctly, why implementation could proceed by inference, why unit/contract/integration/E2E tests or review/MR readiness did not block it, and why any earlier retrospective missed the deeper cause; land prevention at every failed owning layer, not only one target skill.
|
|
186
185
|
3. Classify each lesson:
|
|
187
186
|
- `skill`: reusable agent behavior that belongs in an existing or new skill.
|
|
@@ -195,9 +194,9 @@ Required flow:
|
|
|
195
194
|
|
|
196
195
|
Minimum retrospective table:
|
|
197
196
|
|
|
198
|
-
| Task
|
|
197
|
+
| Task result | Result analysis | Lesson classification | Durable owner | Verification |
|
|
199
198
|
| --- | --- | --- | --- | --- |
|
|
200
|
-
| What happened
|
|
199
|
+
| What happened; failure, stable success, or insufficient evidence | RCA; or success mechanism + non-luck evidence + reuse boundary; or observation-only reason | skill / validator / project artifact / memory / final response only | File, skill, script, shared artifact, memory note for local preference only, or no-skill reason | Diff, command, review, or explicit non-skill reason |
|
|
201
200
|
|
|
202
201
|
### LARGE-Session Lesson Axes And The Delivery-State Axis
|
|
203
202
|
|