@ccoalm/ccl-skills 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +93 -16
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +10 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +249 -5
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +394 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +2 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +6 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +3 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-capability-composition.md +128 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +35 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +54 -2
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +11 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +7 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +6 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +50 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +1 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +1 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +6 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +1 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +59 -1
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +1 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +3 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +43 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +16 -4
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +391 -14
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +74 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +74 -33
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +11 -4
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +421 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +576 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +176 -0
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_lanes.sh +101 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +6 -2
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +2 -2
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +19 -2
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +1 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +3 -3
  37. package/dist/assets/release.json +63 -33
  38. package/package.json +1 -1
@@ -0,0 +1,176 @@
1
+ #!/usr/bin/env bash
2
+ # 043:`source-refuted` 证据类的冻结验收用例(specs/043-evidence-class-source-refuted/frozen-acceptance.md)。
3
+ #
4
+ # A-F 在闸被触碰之前冻结;G/H 由独立评审与挑战各指出一处滥用路径后补入(**加严**,
5
+ # 不是放宽)。通过条件是全部命中——任何一条不符即本轮
6
+ # 未完成,**不得**以"主要目标达成"结案,也不得为让 A 变绿而放宽 B–F。
7
+ #
8
+ # 全部确定性,不调模型。每用例一条分支,互不影响本地未提交改动。
9
+ set -u
10
+ ROOT="$(cd "$(dirname "$0")/../../.." && pwd -P)"
11
+ GATE="$ROOT/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb"
12
+ TMP="$(mktemp -d)"; trap 'rm -rf "$TMP"' EXIT INT TERM
13
+ fails=0
14
+ fail() { echo "FAIL: $*" >&2; fails=$((fails+1)); }
15
+ ok() { echo " ok $*"; }
16
+
17
+ REPO="$TMP/repo"
18
+ git clone -q "$ROOT" "$REPO"
19
+ git -C "$REPO" config user.email test@example.invalid
20
+ git -C "$REPO" config user.name "Test User"
21
+
22
+ OWNER_REF="skills/product-rd-workflow/SKILL.md"
23
+ REGISTER="skills/skill-extraction-workflow/references/source-register.md"
24
+ ZLOSS="specs/043-evidence-class-source-refuted/zero-loss-fixture.md"
25
+
26
+ # 种子:给 owner 的 reference 加一段可被后续用例整段删掉的文本(保证净字节 < 0),
27
+ # 以及一份非空的零损失义务对照 fixture。
28
+ WITHDRAWN_TEXT='A fixture claim that a later case withdraws wholesale; it is long enough that removing it makes the owner package shrink by a clear margin.'
29
+ mkdir -p "$REPO/$(dirname "$ZLOSS")"
30
+ printf '# 零损失义务对照\n\n被撤回的原文(逐字):\n\n> %s\n\n| before 义务 | 去向 |\n| --- | --- |\n| 该段承载的义务 | survives-verbatim: 同文件上一段 |\n' "$WITHDRAWN_TEXT" > "$REPO/$ZLOSS"
31
+ printf '\n%s\n' "$WITHDRAWN_TEXT" >> "$REPO/$OWNER_REF"
32
+ printf '\n<!-- fixture-withdrawable-start -->\nA fixture claim that a later case withdraws wholesale; it is long enough that removing it makes the owner package shrink by a clear margin, which is what the net-bytes floor checks.\n<!-- fixture-withdrawable-end -->\n' >> "$REPO/$OWNER_REF"
33
+ git -C "$REPO" add -A >/dev/null
34
+ git -C "$REPO" commit -qm "seed fixtures for source-refuted cases"
35
+ git -C "$REPO" branch fixture-base HEAD
36
+
37
+ new_case() { git -C "$REPO" switch -q -C "$1" fixture-base; git -C "$REPO" branch --set-upstream-to=fixture-base "$1" >/dev/null 2>&1; }
38
+ commit_case() { git -C "$REPO" add -A >/dev/null; git -C "$REPO" commit -qm "$1"; }
39
+ run_gate() { set +e; OUT="$(env -u CCL_SKILL_BASE_REF ruby "$GATE" "$REPO" 2>&1)"; RC=$?; set -e 2>/dev/null || true; }
40
+ add_row() { printf '%s\n' "$1" >> "$REPO/$REGISTER"; }
41
+ withdraw_fixture() { # **纯删除**:只删那一行,不加任何内容。说明文字进 reference。
42
+ python3 - "$REPO/$OWNER_REF" "$WITHDRAWN_TEXT" <<'PYX'
43
+ import io,sys
44
+ p, txt = sys.argv[1], sys.argv[2]
45
+ s = io.open(p, encoding='utf-8').read()
46
+ io.open(p, 'w', encoding='utf-8').write(s.replace("\n" + txt + "\n", "\n", 1))
47
+ PYX
48
+ }
49
+
50
+ SRC="https://example.invalid/primary-source-that-refutes-the-claim"
51
+ FIRING="firing-path: file:skills/product-rd-workflow/SKILL.md#must survive verbatim above"
52
+ FIRING_F="firing-path: file:skills/product-rd-workflow/SKILL.md#must record a RED-baseline row before it lands"
53
+
54
+ # ---- A:合规撤回 —— 有类时必须绿,无类时必须红 ----
55
+ new_case case-a; withdraw_fixture
56
+ add_row "| A fixture claim is refuted by its primary source and is withdrawn wholesale | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
57
+ commit_case "case A: compliant source-refuted withdrawal"; run_gate
58
+ # A 的期望值在裁决后由绿改红:该类不再顶起 RED 底线(见 frozen-acceptance.md 的裁决记录)。
59
+ # 合规撤回现在仍然被拦——它要么配一条 RED 行,要么由具名风险 owner 人工放行。
60
+ [ "$RC" != "0" ] && ok "A 合规撤回 -> 仍红(该类不顶底线,撤回需人工放行)" || fail "A 不得自动放行"
61
+
62
+ # ---- B:真行为变更披该标签 —— 任何时候都必须红 ----
63
+ new_case case-b
64
+ printf '\n- Every fixture change must obtain approval before it lands.\n' >> "$REPO/$OWNER_REF"
65
+ add_row "| A newly added normative rule wearing the withdrawal label | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
66
+ commit_case "case B: real behaviour change wearing the label"; run_gate
67
+ [ "$RC" != "0" ] && ok "B 真行为变更披标签 -> 仍红" || fail "B 不得通过(净增却用撤回类)"
68
+
69
+ # ---- C:缺一手源 URL —— 必须红 ----
70
+ new_case case-c; withdraw_fixture
71
+ add_row "| A withdrawal whose row cites no refuting source | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; \`product-rd-workflow/SKILL.md\` |"
72
+ commit_case "case C: missing refuting source"; run_gate
73
+ [ "$RC" != "0" ] && ok "C 缺一手源 -> 仍红" || fail "C 不得通过(无源即无据)"
74
+
75
+ # ---- D:缺零损失指针 —— 必须红 ----
76
+ new_case case-d; withdraw_fixture
77
+ add_row "| A withdrawal whose row carries no zero-loss map | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
78
+ commit_case "case D: missing zero-loss pointer"; run_gate
79
+ [ "$RC" != "0" ] && ok "D 缺零损失指针 -> 仍红" || fail "D 不得通过(义务保全无对照)"
80
+
81
+ # ---- E:observed-failure: yes —— 必须红 ----
82
+ new_case case-e; withdraw_fixture
83
+ add_row "| A withdrawal self-declaring an observed failure | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: yes | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
84
+ commit_case "case E: observed-failure yes"; run_gate
85
+ [ "$RC" != "0" ] && ok "E observed-failure: yes -> 仍红" || fail "E 不得通过(分类语义不得自选)"
86
+
87
+ # ---- F:既有类不得回归 ----
88
+ new_case case-f
89
+ printf '\n- Every fixture delivery must record a RED-baseline row before it lands.\n' >> "$REPO/$OWNER_REF"
90
+ add_row "| An ordinary behaviour-changing rule with a RED-baseline row | \`downstream-executor\` | behavioral-evidence: RED-baseline; observed-failure: yes; $FIRING_F | \`updated\` | \`product-rd-workflow/SKILL.md\` |"
91
+ commit_case "case F: existing RED-baseline path"; run_gate
92
+ if [ "$RC" = "0" ]; then ok "F 既有 RED-baseline 路径 -> 绿(无回归)"; else echo "--- F 完整输出 ---"; printf '%s\n' "$OUT"; fail "F 回归"; fi
93
+
94
+ # ---- G:混合行 —— 同 owner 一条合规撤回 + 一条真行为变更的 semantic-control。必须红 ----
95
+ # 独立评审 REVIEW-1:any? 顶起底线时,真行为变更会搭便车。
96
+ new_case case-g; withdraw_fixture
97
+ printf '\n- Every unrelated fixture delivery must obtain approval before it lands. zzmixedrule\n' >> "$REPO/$OWNER_REF"
98
+ add_row "| A compliant withdrawal alongside an unrelated change | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
99
+ add_row "| An unrelated behaviour change riding along on a stable-control label | \`downstream-executor\` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/product-rd-workflow/SKILL.md#must obtain approval before it lands | \`updated\` | \`product-rd-workflow/SKILL.md\` |"
100
+ commit_case "case G: mixed rows"; run_gate
101
+ [ "$RC" != "0" ] && ok "G 混合行 -> 仍红" || fail "G 不得通过(真行为变更搭撤回的便车)"
102
+
103
+ # ---- H:抵消式新增 + 假锚点 —— 必须红 ----
104
+ # 独立挑战 CHALLENGE-1:删无关文字凑净负、加一条新规则、指针写不存在的锚。
105
+ new_case case-h; withdraw_fixture
106
+ printf '\n- Every smuggled fixture rule must be rejected by the withdrawal class. zzsmuggled\n' >> "$REPO/$OWNER_REF"
107
+ add_row "| A smuggled rule hidden behind an offsetting deletion | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`README.md#no-such-anchor-exists-here\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
108
+ commit_case "case H: offset addition with bogus anchor"; run_gate
109
+ [ "$RC" != "0" ] && ok "H 抵消式新增 + 假锚点 -> 仍红" || fail "H 不得通过(新增规范行 / 锚点不解析)"
110
+
111
+ # ---- I:脚本改动、零删除 —— 必须红(旧代理谓词看不见脚本)----
112
+ new_case case-i
113
+ printf '\n# fixture: an added shell line that changes behaviour without matching any prose predicate\n' >> "$REPO/skills/product-rd-workflow/scripts/check-agent-contract-coverage.sh"
114
+ add_row "| A script change wearing the withdrawal label with nothing deleted | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
115
+ commit_case "case I: script change, no deletion"; run_gate
116
+ [ "$RC" != "0" ] && ok "I 脚本改动零删除 -> 仍红" || fail "I 不得通过(非 .md、且无删除)"
117
+
118
+ # ---- J:`Always` 规则、零删除 —— 必须红(旧谓词故意排除 always)----
119
+ new_case case-j
120
+ printf '\n- Always authenticate fixture requests before dispatch.\n' >> "$REPO/$OWNER_REF"
121
+ add_row "| An Always-phrased rule wearing the withdrawal label with nothing deleted | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
122
+ commit_case "case J: Always rule, no deletion"; run_gate
123
+ [ "$RC" != "0" ] && ok "J Always 规则零删除 -> 仍红" || fail "J 不得通过(有新增行)"
124
+
125
+ # ---- K:改权限 + 一次真删除 —— 必须红 ----
126
+ # 独立评审:chmod 在 --name-status 里显示为 M、在 --numstat 里记 0/0。
127
+ new_case case-k; withdraw_fixture
128
+ chmod 755 "$REPO/skills/product-rd-workflow/references/adr-convention.md"
129
+ add_row "| A mode change riding along with a genuine deletion | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
130
+ commit_case "case K: chmod plus deletion"; run_gate
131
+ [ "$RC" != "0" ] && ok "K 改权限 + 删除 -> 仍红" || fail "K 不得通过(mode 变更未被 numstat 反映)"
132
+
133
+ # ---- L:指针路径穿越出仓 —— 必须红 ----
134
+ new_case case-l; withdraw_fixture
135
+ add_row "| A withdrawal whose zero-loss pointer escapes the repository | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`../../../etc/hosts#localhost\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
136
+ commit_case "case L: traversal pointer"; run_gate
137
+ [ "$RC" != "0" ] && ok "L 指针穿越出仓 -> 仍红" || fail "L 不得通过(路径逃出仓库)"
138
+
139
+ # ---- M:锚点是存在但无意义的单字符子串 —— 必须红 ----
140
+ new_case case-m; withdraw_fixture
141
+ add_row "| A withdrawal whose anchor is a bare existing substring | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#a\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
142
+ commit_case "case M: one-character substring anchor"; run_gate
143
+ [ "$RC" != "0" ] && ok "M 单字符子串锚 -> 仍红" || fail "M 不得通过(锚点未落在标题上)"
144
+
145
+ # ---- N:锚点是存在但不相干的整词 —— 必须红 ----
146
+ new_case case-n; withdraw_fixture
147
+ add_row "| A withdrawal whose anchor is an unrelated existing word | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#survives-verbatim\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
148
+ commit_case "case N: unrelated existing substring anchor"; run_gate
149
+ [ "$RC" != "0" ] && ok "N 不相干整词锚 -> 仍红" || fail "N 不得通过(锚点未落在标题上)"
150
+
151
+ # ---- O:既有 semantic-control 路径不得回归(配对 RED 行)----
152
+ new_case case-o
153
+ printf '\n- Every fixture control rule must be recorded before it lands. zzsemctl\n' >> "$REPO/$OWNER_REF"
154
+ add_row "| An unchanged control alongside a RED-baseline row | \`downstream-executor\` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: file:skills/product-rd-workflow/SKILL.md#must be recorded before it lands | \`updated\` | \`product-rd-workflow/SKILL.md\` |"
155
+ add_row "| The behaviour change the control accompanies | \`downstream-executor\` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#must be recorded before it lands | \`updated\` | \`product-rd-workflow/SKILL.md\` |"
156
+ commit_case "case O: semantic-control beside RED-baseline"; run_gate
157
+ [ "$RC" = "0" ] && ok "O semantic-control + RED-baseline -> 绿(无回归)" || { echo "--- O ---"; printf '%s\n' "$OUT" | tail -3; fail "O 既有路径回归"; }
158
+
159
+ # ---- P:对照表没有逐字交代被删内容 —— 必须红 ----
160
+ new_case case-p; withdraw_fixture
161
+ python3 - "$REPO/$ZLOSS" <<'PYP'
162
+ import io,sys
163
+ p=sys.argv[1]
164
+ io.open(p,'w',encoding='utf-8').write("# 零损失义务对照\n\n| before 义务 | 去向 |\n| --- | --- |\n| 某条义务 | survives-verbatim: 别处 |\n")
165
+ PYP
166
+ add_row "| A withdrawal whose map does not quote what was removed | \`downstream-executor\` | behavioral-evidence: source-refuted; observed-failure: no | \`updated\` | zero-loss: \`$ZLOSS#零损失义务对照\`; refuting source: $SRC; \`product-rd-workflow/SKILL.md\` |"
167
+ commit_case "case P: map does not account for deleted text"; run_gate
168
+ [ "$RC" != "0" ] && ok "P 对照表未逐字交代被删内容 -> 仍红" || fail "P 不得通过(删了什么没抄出来)"
169
+
170
+ echo
171
+ if [ "$fails" = "0" ]; then
172
+ echo "all frozen cases behaved as specified"
173
+ else
174
+ echo "frozen cases did not all behave as specified"
175
+ exit 1
176
+ fi
@@ -0,0 +1,101 @@
1
+ #!/usr/bin/env bash
2
+ # Lane-semantics regression for test_check_ccl_regressions.sh (036 challenge
3
+ # P2): --heavy-only must run exactly the heavy_tests entries and no fast entry,
4
+ # propagate a failing heavy suite, and end with its own token; --fast must run
5
+ # exactly the fast_tests entries and no heavy entry. Runs the real runner
6
+ # against a stub fixture tree via REGRESSION_SCRIPTS_DIR, so it proves the mode
7
+ # dispatch in seconds without executing the real multi-minute suites. The stub
8
+ # list is derived from the runner's own arrays, so a lane edit cannot drift the
9
+ # fixture.
10
+ set -euo pipefail
11
+
12
+ SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
13
+ RUNNER="$SCRIPT_DIR/test_check_ccl_regressions.sh"
14
+
15
+ fail() { printf 'FAIL: %s\n' "$1" >&2; exit 1; }
16
+
17
+ extract_lane() {
18
+ # Bare entry tokens of one array block, comments stripped.
19
+ sed -n "/^${1}=(/,/^)/p" "$RUNNER" | sed -e '/^#/d' -e '/=($/d' -e '/^)$/d' \
20
+ -e 's/[[:space:]]*#.*$//' -e 's/^[[:space:]]*//' -e '/^$/d'
21
+ }
22
+
23
+ fast_lane="$(extract_lane fast_tests)"
24
+ heavy_lane="$(extract_lane heavy_tests)"
25
+ [ -n "$fast_lane" ] || fail "could not extract fast_tests from the runner"
26
+ [ -n "$heavy_lane" ] || fail "could not extract heavy_tests from the runner"
27
+
28
+ # Exactly-once contract at the array level (r4 challenge P2): a suite listed
29
+ # twice in one lane, or in both lanes, would run twice in CI while every
30
+ # per-mode comparison here still passed — reject it before building stubs.
31
+ dup_fast="$(sort <<<"$fast_lane" | uniq -d)"
32
+ [ -z "$dup_fast" ] || fail "duplicate entries inside fast_tests: $dup_fast"
33
+ dup_heavy="$(sort <<<"$heavy_lane" | uniq -d)"
34
+ [ -z "$dup_heavy" ] || fail "duplicate entries inside heavy_tests: $dup_heavy"
35
+ overlap="$(comm -12 <(sort <<<"$fast_lane") <(sort <<<"$heavy_lane"))"
36
+ [ -z "$overlap" ] || fail "entries registered in both lanes (would run twice per pipeline): $overlap"
37
+
38
+ TMP="$(mktemp -d)"
39
+ trap 'rm -rf "$TMP"' EXIT
40
+ FIX="$TMP/skills/skill-extraction-workflow/scripts"
41
+ mkdir -p "$FIX"
42
+ LOG="$TMP/ran.log"
43
+
44
+ write_stub() {
45
+ local entry="$1"
46
+ local path="$FIX/$entry"
47
+ mkdir -p "$(dirname "$path")"
48
+ printf '#!/usr/bin/env bash\nprintf "%%s\\n" "%s" >> "%s"\n' "$entry" "$LOG" >"$path"
49
+ }
50
+ while IFS= read -r entry; do write_stub "$entry"; done <<<"$fast_lane"
51
+ while IFS= read -r entry; do write_stub "$entry"; done <<<"$heavy_lane"
52
+
53
+ run_mode() {
54
+ : >"$LOG"
55
+ REGRESSION_SCRIPTS_DIR="$FIX" bash "$RUNNER" "$1" >"$TMP/out.$2" 2>&1
56
+ }
57
+
58
+ # --heavy-only: exactly the heavy lane, its own token, no fast token.
59
+ run_mode --heavy-only heavy || fail "--heavy-only exited nonzero on passing stubs"
60
+ diff <(sort <<<"$heavy_lane") <(sort "$LOG") >/dev/null \
61
+ || fail "--heavy-only did not run exactly the heavy_tests entries"
62
+ grep -q '^test_check_ccl_regressions_heavy_only_ok$' "$TMP/out.heavy" \
63
+ || fail "--heavy-only missing its completion token"
64
+ if grep -q 'test_check_ccl_regressions_fast_ok\|test_check_ccl_regressions_full_ok' "$TMP/out.heavy"; then
65
+ fail "--heavy-only emitted a fast/full token"
66
+ fi
67
+
68
+ # --fast: exactly the fast lane, fast token, no heavy entry.
69
+ run_mode --fast fast || fail "--fast exited nonzero on passing stubs"
70
+ diff <(sort <<<"$fast_lane") <(sort "$LOG") >/dev/null \
71
+ || fail "--fast did not run exactly the fast_tests entries"
72
+ grep -q '^test_check_ccl_regressions_fast_ok$' "$TMP/out.fast" \
73
+ || fail "--fast missing its completion token"
74
+
75
+ # --full: both lanes, full token.
76
+ run_mode --full full || fail "--full exited nonzero on passing stubs"
77
+ diff <(sort <(printf '%s\n%s\n' "$fast_lane" "$heavy_lane")) <(sort "$LOG") >/dev/null \
78
+ || fail "--full did not run exactly fast_tests + heavy_tests"
79
+ grep -q '^test_check_ccl_regressions_full_ok$' "$TMP/out.full" \
80
+ || fail "--full missing its completion token"
81
+
82
+ # Failure propagation: a red heavy stub must fail --heavy-only.
83
+ first_heavy="$(head -n 1 <<<"$heavy_lane")"
84
+ printf '#!/usr/bin/env bash\nexit 1\n' >"$FIX/$first_heavy"
85
+ if REGRESSION_SCRIPTS_DIR="$FIX" bash "$RUNNER" --heavy-only >"$TMP/out.red" 2>&1; then
86
+ fail "--heavy-only exited zero although a heavy stub failed"
87
+ fi
88
+ write_stub "$first_heavy"
89
+
90
+ # Unregistered-sibling enforcement (036 challenge P1): an unregistered
91
+ # test_*.sh must hard-fail every execution mode, independent of the
92
+ # registration guard test staying registered.
93
+ : >"$FIX/test_unregistered_probe.sh"
94
+ if REGRESSION_SCRIPTS_DIR="$FIX" bash "$RUNNER" --fast >"$TMP/out.unreg" 2>&1; then
95
+ fail "--fast exited zero although an unregistered sibling test exists"
96
+ fi
97
+ grep -q 'test_unregistered_probe.sh' "$TMP/out.unreg" \
98
+ || fail "unregistered-sibling failure did not name the offending file"
99
+ rm "$FIX/test_unregistered_probe.sh"
100
+
101
+ echo "regression_runner_lanes_ok"
@@ -39,7 +39,11 @@ printf 'not-frontmatter\n' > "$tmp/repo/skills/sample-skill/fixtures/SKILL.md"
39
39
  mkdir -p "$tmp/repo/vendor/thing"
40
40
  printf 'not-frontmatter\n' > "$tmp/repo/vendor/thing/SKILL.md"
41
41
  out="$(bash "$VALIDATE" "$tmp/repo" 2>&1)" || fail "repo-root run exited non-zero: $out"
42
- printf '%s\n' "$out" | grep -q 'validating 1 skill root(s)' \
42
+ # Pure-bash substring match, not `printf | grep -q`: grep -q exits on its first
43
+ # hit, the producer takes SIGPIPE, and under `set -o pipefail` the SUCCESS path
44
+ # then reports failure. Timing-dependent, so it stayed latent until this lane
45
+ # ran concurrently (specs/037-ci-intra-job-parallel/plan.md).
46
+ [[ "$out" == *"validating 1 skill root(s)"* ]] \
43
47
  || fail "expected exactly 1 root (nested fixture must be excluded): $out"
44
48
 
45
49
  # (2) a tree with no SKILL.md still fails loudly (exit 2, no silent green).
@@ -47,7 +51,7 @@ mkdir -p "$tmp/empty/sub/deeper"
47
51
  if out="$(bash "$VALIDATE" "$tmp/empty" 2>&1)"; then
48
52
  fail "empty tree unexpectedly passed: $out"
49
53
  fi
50
- printf '%s\n' "$out" | grep -q 'no_skill_roots_found' \
54
+ [[ "$out" == *no_skill_roots_found* ]] \
51
55
  || fail "empty tree missing loud no_skill_roots_found: $out"
52
56
 
53
57
  echo "PASS: validate-skill discovers skills/<name>/SKILL.md and stays loud on empty trees"
@@ -43,9 +43,9 @@ Use this skill to decide whether a specialized or non-functional test belongs in
43
43
  - Every critical scenario needs at least one automated assertion at the lowest layer that can prove the risk, plus real-flow smoke only when cross-boundary behavior matters; built/installed deliverables need one smoke via the published entry path (`references/e2e-real-flow-testing.md`, Published Entry Path).
44
44
  - Happy-path tests are insufficient for high-risk workflows. Build a compact risk matrix and cover the triggered failure classes — the canonical failure-class list lives in `references/scenario-testing.md`.
45
45
  - Composite / multi-stage pipelines (a capability built from several modules, services, or model stages chained together) need acceptance at three layers, not just one: **module-level** (the changed stage's own correctness, errors, latency, fallback), **chain-level** (upstream/downstream input-output contracts, version pass-through, end-to-end recovery and rollback), and **product-level** (the user-visible outcome / acceptance baseline holds). A change to one sub-module cannot pass on its local metric alone if the end-to-end product result regresses; require an end-to-end check whenever the pipeline composition or a stage contract changes. Scope and the unchanged-contract refactor exemption: `references/scenario-testing.md`. (For AI/inference pipelines the component-vs-end-to-end split and "launch follows the product baseline, not the best component metric" rule are owned by `llm-inference-integration`; this rule is the stack-agnostic testing form.)
46
- - Maintain a frozen regression set of real past failures (every fixed bug / incident / confirmed bad case becomes a case) and run it on **every release**, not ad hoc — "we tested that once" is not regression coverage. Tier it so the gate stays affordable and complement it with a periodic adversarial pass over code considered "done" (tiering and adversarial-pass procedure: `references/ci-fixtures-and-flake-control.md`). For AI features, the online-bad-case → regression-set re-injection mechanics are owned by `llm-inference-integration`; this rule is the general gate that the regression suite is release-blocking.
46
+ - Maintain a frozen regression set of real past failures (every fixed bug / incident / confirmed bad case becomes a case) and run it on **every release**, not ad hoc. Tier it so the gate stays affordable, plus a periodic adversarial pass over code considered "done" (tiering and adversarial-pass procedure: `references/ci-fixtures-and-flake-control.md`). For AI features, the online-bad-case → regression-set re-injection mechanics are owned by `llm-inference-integration`; the release-blocking gate stays here.
47
47
  - Use the repository's own test wrapper when it exists; it encodes env, codegen, fixture, timeout, and CI parity.
48
- - Use the right test double: stub to supply inputs, fake to emulate dependency behavior, and spy/mock to verify interaction only when the interaction is the contract; double the expensive/nondeterministic/unsafe/privileged/unavailable boundary and keep the rest real on isolated test-owned resources (`references/test-code-authoring-patterns.md` §5).
48
+ - Use the right test double: stub to supply inputs, fake to emulate dependency behavior, and spy/mock to verify interaction only when the interaction is the contract; double the expensive/nondeterministic/unsafe/privileged/unavailable boundary and keep the rest real on isolated test-owned resources (`references/test-code-authoring-patterns.md` §5; external-provider recovery paths: `references/ci-fixtures-and-flake-control.md`, Fault-Injection Layers).
49
49
  - Use integration tests where mocks would hide contract, transaction, serialization, permission, data-shape, or runtime failures.
50
50
  - Use E2E only for critical user/caller workflows and release confidence, not every branch through the browser.
51
51
  - A passing test must assert the outcome. Clicking a button, calling an endpoint, or seeing status 200 is not enough. **Every property a test NAMES — in ANY name a reader or a runner sees: the function name, a docstring or comment, and equally an `it(...)`/`describe(...)` title, a Gherkin scenario name, a pytest parameter id, or a subtest/table-case label (`..._is_bijective`, `..._is_pinned`, "rejects out-of-range input") — is a claim, and each named property owes a killing mutation: the concrete implementation change that would make this test fail.**
@@ -4,7 +4,7 @@
4
4
 
5
5
  Use layered gates:
6
6
 
7
- - Fast PR gate: format, compile/typecheck, lint/static checks, focused unit tests, deterministic codegen clean check.
7
+ - Fast PR gate: format, compile/typecheck, lint/static checks, duplication and dead-code gates where configured, focused unit tests, deterministic codegen clean check.
8
8
  - Integration gate: DB/Redis/MQ/API/dependency adapter tests with stable local containers or provisioned CI infra.
9
9
  - Scenario gate: selected acceptance/risk scenarios mapped to unit, contract, integration, component, or E2E commands. Keep it small enough to diagnose failures quickly.
10
10
  - E2E/release gate: critical browser/API workflows and smoke tests in an isolated environment.
@@ -12,6 +12,8 @@ Use layered gates:
12
12
 
13
13
  Use the same command wrapper locally and in CI when feasible. If CI invokes a wrapper such as `scripts/run_tests.sh`, `scripts/dev.sh test`, `make test`, or service-local package commands, use that wrapper for local verification unless debugging a lower-level runner.
14
14
 
15
+ Duplication and dead-code gates run with explicit configuration, not defaults: the copy-paste detector gets a minimum-token/line floor and exits non-zero, test code is excluded, and justified duplication is fenced by explicit inline ignore markers (an untracked "dedupe later" is how parallel implementations drift); an unused-export/dead-code gate runs beside it, and each finding axis gets exactly one owning tool so two gates do not contest the same class. Derived committed artifacts (generated notices, catalogs, indexes) get caught stale at the pre-commit hook — not discovered later as a test-lane failure the author no longer connects to the edit. Scope that hook's trigger to every input the generator reads, including the generator script itself. The recommended form is reject-and-instruct: the hook fails and prints the exact regeneration command, so the author reruns and stages deliberately. Auto-staging the regenerated output demands genuine atomicity over every state the hook reads or writes (staged-input identity, output worktree file, output index entry — an exclusive repository lock plus crash-safe rollback), which most hook tooling cannot guarantee; without that guarantee it silently commits or overwrites content the author did not select, so keep it out of the default path. Keep a freshness assertion over the committed tree in the test lane as the gate for everything hooks cannot see (deletions, uninstalled hooks, aborted or partial-commit drift): the hook is convenience; the test lane is the gate.
16
+
15
17
  ## Frozen Regression Set And Adversarial Passes
16
18
 
17
19
  Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests.
@@ -22,9 +24,19 @@ The proactive complement — the adversarial pass over code already considered "
22
24
 
23
25
  Use `test-data-and-determinism.md` as the canonical source for fixture shape, anonymization, data builders, golden-file normalization, and deterministic clocks/randomness/ordering.
24
26
 
27
+ ### Fault-Injection Layers For External-Provider Recovery Paths
28
+
29
+ Recovery behavior against an external provider (a model API, payment/storage backend, streaming dependency) needs its fault permutations proven below the live layer. Layer the fixtures; prove each fault class at the most protocol-real layer that can still script it deterministically:
30
+
31
+ - **Protocol-real fault server** (wire boundary): a scriptable in-test server speaking the provider's real HTTP/streaming protocol, where each accepted request consumes one scripted behavior from an explicit fault taxonomy — transport faults (connection reset, mid-stream disconnect, stall), protocol faults (malformed payload or stream event, wrong content type, truncated stream), and semantic faults (empty-but-successful response, rate limit, auth error, context/quota limits). The fixture stays policy-free: it never retries or interprets the policy under test, so the retry/backoff/error-mapping behavior the test observes is entirely the product's. When the product under test may retry or parallelize requests, arrival-order consumption misassigns faults to the wrong attempt — bind each scripted behavior to its intended attempt (correlation key or strict single-flight sequencing) and fail the test on unexpected, out-of-order, or unconsumed behaviors.
32
+ - **Recorded replay** (adapter seam): reconstruct dependency responses from recorded real sessions and short-circuit the adapter, so scenario suites run keyless and deterministic. A recording cannot reconstruct every case — a thrown mid-stream error, a hang awaiting cancellation — so the replay format carries explicit override entries for those, not silent absence. Record from synthetic test accounts and non-production sessions — production or customer traffic is never a fixture source, since redaction cannot reliably scrub sensitive values from legitimate free-text fields; on top of that, capture through a schema allowlist with credentials and sensitive fields redacted before persistence (anonymization canon: `test-data-and-determinism.md`) and gate fixture commits on a secret/PII scan.
33
+ - **Live credentialed e2e**: wiring sanity only, under the env-gated live-lane rules in Verification Report below — keyless environments skip visibly, the credentialed lane treats a missing secret as a preflight failure (never a skip), and it runs on synthetic/test-scoped credentials, tenants, and resources per those rules, never production ones.
34
+
35
+ The live layer proves wiring and credentials, never the fault matrix.
36
+
25
37
  ## Flake Control
26
38
 
27
- - Replace sleeps with explicit conditions. When an assertion still depends on time, make the ORACLE a state whose existence proves the invariant — the helper process still alive, the record not yet written, the lock still held — rather than "finished within N seconds"; a wall-clock margin only holds while the runner is fast enough, so it turns into a red on a loaded runner and, worse, into a silent pass when the margin is generous. Where a fixture's own lifetime is what the assertion races (a sleeping helper, a TTL, a retention window), push that lifetime far past any plausible run so it stops being a deadline, and keep any wall-clock bound as a coarse backstop only. Sample process state, not just liveness: an unreaped zombie still answers `kill -0`, so "still running" and "already gone" both need the state check, and reap lag on a loaded runner needs a bounded grace period instead of an instantaneous sample.
39
+ - Replace sleeps with explicit conditions. When an assertion still depends on time, make the ORACLE a state whose existence proves the invariant — the helper process still alive, the record not yet written, the lock still held — rather than "finished within N seconds"; a wall-clock margin only holds while the runner is fast enough, so it turns into a red on a loaded runner and, worse, into a silent pass when the margin is generous. Where a fixture's own lifetime is what the assertion races (a sleeping helper, a TTL, a retention window), push that lifetime far past any plausible run so it stops being a deadline, and keep any wall-clock bound as a coarse backstop only — far past, but never UNBOUNDED: the thing that normally ends such a helper is whatever the suite is testing, and when that reaper dies first (an aborted run, a cancelled job, a suspended host) the fixture's own lifetime becomes the only thing that will ever end it, so an infinite loop is how a test helper outlives the machine rather than the run. The suite also owns reaping what it started, and cannot delegate that to the component under test: a helper deliberately made signal-immune, or started in its own session/process group where a signal aimed at the suite cannot reach it, has no other reaper — trap the abort signals as well as EXIT, identify what to reap by something the run owns (its own private temp path, its recorded process groups) rather than by a shared name that would hit a concurrent lane, and prove ownership again at the moment of signalling because a pid or group id read earlier is a number the OS recycles. Assert at exit that nothing the run started is still alive: every case can pass while the process table is not clean, so a green suite is not the acceptance object here. Sample process state, not just liveness: an unreaped zombie still answers `kill -0`, so "still running" and "already gone" both need the state check, and reap lag on a loaded runner needs a bounded grace period instead of an instantaneous sample.
28
40
  - Isolate tests by namespace, temp directory, tenant, user, DB schema, or unique id.
29
41
  - Clean up with test cleanup hooks.
30
42
  - Avoid shared global state or reset it explicitly.
@@ -43,6 +55,11 @@ Use `test-data-and-determinism.md` as the canonical source for fixture shape, an
43
55
  - Use coverage to find untested risk, not as a target to hit.
44
56
  - High line coverage with weak assertions is not safety.
45
57
  - For high-risk pure logic, consider mutation testing or equivalent assertion-strength checks.
58
+ - A failing coverage gate names exact locations: print every uncovered statement, branch path, and function as a clickable `path:line:col` (a custom reporter when the runner's built-in failure names only the file), and print nothing when green — a gate that names only the file forces the author to re-run coverage locally just to find the gap.
59
+ - Enforce thresholds per file rather than repo-aggregate when the goal is stopping a well-covered big file from subsidizing a bare one; the threshold value stays the team's choice (floor-not-goal and cleanup exceptions: `test-code-authoring-patterns.md` §6).
60
+ - A platform- or capability-conditional coverage exemption derives from the same probe that makes the corresponding suites skip, so the exemption is active exactly when those tests cannot run; a hand-maintained exemption list drifts into exempting files whose tests actually run. Guard the shared probe's common-mode failure: distinguish "capability genuinely absent" from "probe errored", and a CI lane that declares the capability treats an unavailable result as a failure, never as skip-plus-exemption.
61
+ - When instrumented runs are slow, partition the suite across parallel coverage processes and merge reports; thresholds run once against the merged report, never inside a partition (a partition sees only its slice and would false-fail). The merge validates that every expected partition reported exactly once for the current run: use a run-scoped artifact directory (cleaning only that private directory) and bind every report and the expected-partition manifest to the same run identity (run id / commit / config), so a missing, stale, foreign, or duplicate partition artifact fails the run instead of silently shrinking or backfilling the report.
62
+ - Perf/stress/high-cardinality diagnostic suites stay outside the coverage and PR gates by config-level include inventory — a separate runner config that enumerates them — not by runtime skip marks: an inventory is auditable; skip marks rot silently (skip semantics: Verification Report below).
46
63
 
47
64
  ## Local Evidence Selection
48
65
 
@@ -60,6 +60,7 @@ When the deliverable ships as a built or installed artifact — a package `bin`,
60
60
  - Prefer one happy path plus high-risk negative paths over many shallow click-throughs.
61
61
  - A click-through without assertions is not E2E evidence.
62
62
  - If a scenario can be proven with a stable API/contract/integration test and only needs one browser smoke for confidence, do not duplicate all permutations in the browser.
63
+ - External-provider fault/recovery permutations (disconnects, malformed streams, rate limits) belong at the protocol-real fault-server and recorded-replay layers (`ci-fixtures-and-flake-control.md`, Fault-Injection Layers); the live credentialed e2e keeps one wiring sanity path, not the fault matrix.
63
64
 
64
65
  ## Failure Handling
65
66
 
@@ -90,7 +90,7 @@ def test_export_request_returns_signed_url_when_user_has_quota():
90
90
  | **Conditional Test Logic**(Meszaros) | 测试体内有 if/else/loop | 实际测的是什么不明 |
91
91
  | **Mystery Guest**(Meszaros 子项) | 依赖外部文件/数据但未声明 | 不可复现 |
92
92
 
93
- **用**:code review / 重构 / 排查 flaky test 时按这清单查。Slow Test 阈值只对**默认快速单测目标**(unit 层)严卡;integration / E2E / host-smoke / benchmark 测必有独立的更宽 budget,按 marker 分离(如 `@pytest.mark.integration` / Go `-short` 区分),不混入 unit 套时间预算。
93
+ **用**:code review / 重构 / 排查 flaky test 时按这清单查。Slow Test 阈值只对**默认快速单测目标**(unit 层)严卡;integration / E2E / host-smoke / benchmark 测必有独立的更宽 budget,按 marker 分离(如 `@pytest.mark.integration` / Go `-short` 区分)或按 runner 配置级 include 清单隔成独立套(perf/stress 类车道优先用清单——清单可审计,运行时 skip 标记会静默腐烂,见 `ci-fixtures-and-flake-control.md` Coverage As Signal),不混入 unit 套时间预算。
94
94
 
95
95
  **不用**:写新测试时不必预先记住名字 — 用 §1 §2 §7 等正面规则反过来就避开了大半。
96
96
 
@@ -235,11 +235,11 @@ internal_helper_mock.parse.assert_called_once() # 重构改 parse 就挂
235
235
  - **不把整库平均覆盖率作 PR gate**(个别 PR 不应承担整体覆盖率波动)
236
236
 
237
237
  **落地**:
238
- - CI floor:仓库整体 line ≥ X%
238
+ - CI floor:仓库整体 line ≥ X%;要防"大文件覆盖率补贴裸文件"时改用 per-file 门(阈值团队定,见 `ci-fixtures-and-flake-control.md` Coverage As Signal)
239
239
  - **PR 覆盖率不可下降原则有例外**:删冗余测试 / 移除已废弃代码 / 合并重复 fixture 等清理类 PR 覆盖率下降允许,但 PR 描述必须给 **preservation proof row**:`deleted: <test path/name>; preserved scenario: <TC-ID or behavior statement>; replacement: <test path or "still covered by <existing test>">; oracle parity: <断言形状等价说明>; evidence: <command output / report ref>`。没 row = 当作场景失守 = 阻挡。Reviewer 用 row 验证:跑 replacement test 应能 catch 删掉的 mutant;equivalent assertion 不是"两个都跑过",是"等价业务断言"
240
240
  - critical-path 模块单独定 ratchet(如 `<critical-module> line ≥ 90% 且 branch ≥ 80%`)— 模块名按 repo 实际填
241
241
  - 用 mutation testing 周期性检查(见 source-to-case-workflows §C.1)— 比追 100% line 更省力且更真实
242
- - coverage 报告 + uncovered lines 进 PR comment(不要靠开发者主动看)
242
+ - coverage 报告 + uncovered lines 进 PR comment(不要靠开发者主动看);覆盖门失败输出指名到可点击的 `path:line:col`(runner 内建只报文件名时加自定义 reporter,绿时静默)——只报文件名等于让作者本地重跑一遍才能找到缺口
243
243
  - 覆盖门下的 uncovered line **先当删除候选、再当补测候选**:门在正确地标记死代码/不可达分支时,补一个测试只是把死代码钉死;行覆盖是必要不充分——证明行跑过了,不证明功能按交付形态工作
244
244
 
245
245
  ---