@ccoalm/ccl-skills 0.3.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +34 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_abort_leak_state_helpers.sh +148 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +271 -34
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-data-acquisition.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-disclosure-channels.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +12 -12
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +7 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +7 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +47 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +28 -29
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +165 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +23 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +303 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +222 -91
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh +85 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +6 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +12 -12
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +455 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +20 -20
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_liveness_predicate_gate.sh +288 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +4 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/deliverable-doc-genre-skeletons.md +133 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/doc-charter-first.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +318 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/AGENTS.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/doc-lint.py +246 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/figure-lint.py +1092 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/mutation_probe.sh +100 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/test_figure_and_doc_lint.sh +375 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/control.md +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/empty-header.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fake-header.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fenced-noise.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig-dangling.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig-orphan-captioned.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig-orphan.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/fig.png +0 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/imbalance.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/no-unit.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/should-be-chart.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/tables-only-clean.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/unfilled.md +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/doc/wide-table.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/bad-viewbox.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/blackmarker.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/control.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/crossings.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/cvd-confusable.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/decorative-line.svg +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/edge-no-arrow.svg +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/edge-vague.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/figure-contract.json +21 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/figure-is-a-list.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/flow-mixed.svg +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/low-contrast.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/malformed.svg +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-aria.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-group.svg +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-legend.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-title.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/no-viewbox.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/offcontract-shape.svg +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/overflow.svg +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/transformed.svg +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/ungrouped-card.svg +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/scripts/tests/svg/unlabeled-edge.svg +13 -0
- package/dist/assets/release.json +276 -31
- package/package.json +1 -1
|
@@ -13,6 +13,14 @@
|
|
|
13
13
|
# leg 2 SIGKILL the suite: nothing can trap that, so no reaper runs at all and the
|
|
14
14
|
# wrapper must die on the fixture's own lifetime bound instead.
|
|
15
15
|
#
|
|
16
|
+
# Every verdict below is a fact the run leaves behind — the work dir the EXIT trap would
|
|
17
|
+
# have deleted, the marker the fixture writes only past its countdown — never an
|
|
18
|
+
# observation of what was true at the instant the probe looked. Liveness questions all go
|
|
19
|
+
# through wrapper_state() so they cannot answer the same process differently, and a
|
|
20
|
+
# scenario the probe fails to BUILD is retried rather than reported as a failed assertion.
|
|
21
|
+
# All three of those rules are here because the earlier spellings red this gate on CI
|
|
22
|
+
# while the suite was behaving correctly.
|
|
23
|
+
#
|
|
16
24
|
# Leg 2 exists because leg 1 alone stays green if the fixture's bound were restored to
|
|
17
25
|
# an unbounded loop: the trap would still reap it. It was an unbounded loop that turned
|
|
18
26
|
# this defect into a wrapper observed alive for 39 hours with ppid=1, ignoring SIGTERM.
|
|
@@ -59,10 +67,48 @@ esac
|
|
|
59
67
|
# with the probe still green, so CI points its two jobs at different stubs.
|
|
60
68
|
ABORT_LEAK_PROBE_CLIENT="${ABORT_LEAK_PROBE_CLIENT:-claude}"
|
|
61
69
|
case "$ABORT_LEAK_PROBE_CLIENT" in
|
|
62
|
-
claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh
|
|
63
|
-
|
|
70
|
+
claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh
|
|
71
|
+
PROBE_BOUND_MARKER=claude_hang_bound_reached ;;
|
|
72
|
+
fallback) PROBE_BEHAVIOR_FILE=kimi_behavior; PROBE_WRAPPER=kimi_review.sh
|
|
73
|
+
PROBE_BOUND_MARKER=kimi_hang_bound_reached ;;
|
|
64
74
|
*) echo "ABORT_LEAK_PROBE_CLIENT must be claude or fallback" >&2; exit 2 ;;
|
|
65
75
|
esac
|
|
76
|
+
# How many times a leg may rebuild its scenario before giving up. Only the SETUP is
|
|
77
|
+
# retried: constructing "an orphaned wrapper with no reaper left" depends on the probe
|
|
78
|
+
# winning races against the controller and the host, and losing one says nothing about
|
|
79
|
+
# the suite. An assertion ABOUT the suite is never retried, so a fixture whose bound was
|
|
80
|
+
# reverted fails on every attempt and cannot be retried into green.
|
|
81
|
+
SETUP_ATTEMPTS="${ABORT_LEAK_PROBE_SETUP_ATTEMPTS:-3}"
|
|
82
|
+
|
|
83
|
+
# ONE process-state vocabulary for the whole probe. Three separate spellings of "is it
|
|
84
|
+
# there?" used to disagree about the same pid: `kill -0` succeeds for a zombie, reading
|
|
85
|
+
# ppid succeeds for a zombie, and the two verdict scans exclude zombies. An already-dead
|
|
86
|
+
# wrapper awaiting reaping therefore satisfied "reparented to init", failed "still alive",
|
|
87
|
+
# and satisfied "gone" — one process state, three answers, and a red that named the suite
|
|
88
|
+
# for something the suite had not done. Every liveness question below goes through here.
|
|
89
|
+
wrapper_state() {
|
|
90
|
+
local st rc
|
|
91
|
+
case "${1:-}" in ''|*[!0-9]*) printf 'absent\n'; return 0 ;; esac
|
|
92
|
+
st="$(ps -o stat= -p "$1" 2>/dev/null)"; rc=$?
|
|
93
|
+
# `ps` exits 1 for "no such process", which is a real answer. Anything above that is
|
|
94
|
+
# the TOOL failing, not the process being gone — and swallowing that into `absent`
|
|
95
|
+
# would report a suite exited because `ps` could not run. Unknown is reported as its
|
|
96
|
+
# own state and every consumer treats it as still present, because the expensive
|
|
97
|
+
# direction of this error is declaring something gone that is not.
|
|
98
|
+
if [ "$rc" -gt 1 ]; then printf 'unknown\n'; return 0; fi
|
|
99
|
+
case "$st" in
|
|
100
|
+
'') printf 'absent\n' ;;
|
|
101
|
+
*Z*) printf 'zombie\n' ;;
|
|
102
|
+
# A stopped process is neither running nor gone, and leg 2 deliberately puts the
|
|
103
|
+
# wrapper in this state while it arms. Folding it into `live` would let the arming
|
|
104
|
+
# step report success without SIGSTOP having actually landed.
|
|
105
|
+
# `T` (job-control stop) or `t` (tracing stop), anywhere in the field: neither letter
|
|
106
|
+
# appears among the flag characters either ps appends, so this cannot catch a
|
|
107
|
+
# running process by accident.
|
|
108
|
+
*[Tt]*) printf 'stopped\n' ;;
|
|
109
|
+
*) printf 'live\n' ;;
|
|
110
|
+
esac
|
|
111
|
+
}
|
|
66
112
|
|
|
67
113
|
fails=0
|
|
68
114
|
check() {
|
|
@@ -178,6 +224,21 @@ suite_work_dir() {
|
|
|
178
224
|
done
|
|
179
225
|
return 1
|
|
180
226
|
}
|
|
227
|
+
# The fixture writes this file only after its bounded countdown runs to completion, so
|
|
228
|
+
# its presence says the fixture's OWN lifetime bound ended the wrapper and its absence
|
|
229
|
+
# says something else did. That is the fact leg 2 needs, and it is a fact about which
|
|
230
|
+
# code path ran — not about what was true at the instant the probe happened to look.
|
|
231
|
+
bound_marker_path() {
|
|
232
|
+
local work
|
|
233
|
+
work="$(suite_work_dir)" || return 1
|
|
234
|
+
printf '%s\n' "$work/state/$PROBE_BOUND_MARKER"
|
|
235
|
+
}
|
|
236
|
+
bound_marker_state() {
|
|
237
|
+
local path
|
|
238
|
+
path="$(bound_marker_path)" || { printf 'no-work-dir\n'; return 0; }
|
|
239
|
+
if [ -e "$path" ]; then printf 'present\n'; else printf 'absent\n'; fi
|
|
240
|
+
}
|
|
241
|
+
|
|
181
242
|
live_wrapper() {
|
|
182
243
|
local work behavior pid parent_cmd
|
|
183
244
|
work="$(suite_work_dir)" || return 0
|
|
@@ -203,7 +264,10 @@ orphan_target_wrapper() {
|
|
|
203
264
|
local hang_bound="$1" candidate controller_pid controller_ok deadline suite_died
|
|
204
265
|
PROBE_TMP="$(mktemp -d "${TMPDIR:-/tmp}/review-gate-abort-probe.XXXXXX")" || return 1
|
|
205
266
|
PROBE_TMP_REAL="$(cd "$PROBE_TMP" && pwd -P)" || return 1
|
|
267
|
+
# Both, not just the pid: a retry that left the previous attempt's group id in place
|
|
268
|
+
# would have the verdict scan looking for members of a group this attempt never built.
|
|
206
269
|
wrapper_pid=""
|
|
270
|
+
wrapper_pgid=""
|
|
207
271
|
|
|
208
272
|
# Job control so the suite gets its own process group and the probe can signal that
|
|
209
273
|
# group without signalling itself.
|
|
@@ -219,7 +283,11 @@ orphan_target_wrapper() {
|
|
|
219
283
|
while [ "$(date +%s)" -lt "$deadline" ]; do
|
|
220
284
|
candidate="$(live_wrapper)"
|
|
221
285
|
if [ -n "$candidate" ]; then wrapper_pid="$candidate"; break; fi
|
|
222
|
-
|
|
286
|
+
# Third site of the same class: a suite that has exited but whose corpse this shell
|
|
287
|
+
# has not reaped still answers `kill -0`, so the bare existence test would keep this
|
|
288
|
+
# loop polling for a dead suite until the whole reach budget elapsed and then report
|
|
289
|
+
# the wrong one of the two failures below.
|
|
290
|
+
[ "$(wrapper_state "$suite_pid")" = live ] || { suite_died=1; break; }
|
|
223
291
|
sleep 0.1
|
|
224
292
|
done
|
|
225
293
|
if [ -z "$wrapper_pid" ]; then
|
|
@@ -281,29 +349,53 @@ orphan_target_wrapper() {
|
|
|
281
349
|
fi
|
|
282
350
|
sleep 0.5
|
|
283
351
|
done
|
|
352
|
+
# Its own deadline: this loop used to reuse the one the controller wait had already
|
|
353
|
+
# been counting down, so a controller that took nine of those ten seconds to die left
|
|
354
|
+
# the reparent check one second to succeed in.
|
|
355
|
+
deadline=$(( $(date +%s) + 10 ))
|
|
284
356
|
while :; do
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
357
|
+
# A corpse is not an orphan. The old spelling of this check read ppid, which a zombie
|
|
358
|
+
# answers, so a wrapper some reaper had already killed passed for a live orphan and
|
|
359
|
+
# the leg went on to make assertions about a scenario it had never built.
|
|
360
|
+
if [ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "1" ] &&
|
|
361
|
+
[ "$(wrapper_state "$wrapper_pid")" = live ]; then
|
|
362
|
+
return 0
|
|
363
|
+
fi
|
|
364
|
+
case "$(wrapper_state "$wrapper_pid")" in
|
|
365
|
+
live) : ;;
|
|
366
|
+
*)
|
|
367
|
+
printf 'abort_leak_setup_lost: wrapper %s was already %s before the abort — a reaper reached it first (bound marker: %s)\n' \
|
|
368
|
+
"$wrapper_pid" "$(wrapper_state "$wrapper_pid")" "$(bound_marker_state)" >&2
|
|
369
|
+
return 1
|
|
370
|
+
;;
|
|
371
|
+
esac
|
|
290
372
|
if [ "$(date +%s)" -ge "$deadline" ]; then
|
|
291
|
-
echo "
|
|
373
|
+
echo "abort_leak_setup_lost: wrapper $wrapper_pid was never reparented" >&2
|
|
292
374
|
return 1
|
|
293
375
|
fi
|
|
294
376
|
sleep 0.5
|
|
295
377
|
done
|
|
296
378
|
}
|
|
297
379
|
|
|
380
|
+
# Same one vocabulary as the wrapper: a killed suite whose corpse the probe's own shell
|
|
381
|
+
# has not reaped yet still answers `kill -0`, so spelling this as `kill -0` made a bookkeeping
|
|
382
|
+
# lag inside the probe look like a suite that refused to die.
|
|
298
383
|
await_suite_exit() {
|
|
299
384
|
local deadline
|
|
300
385
|
deadline=$(( $(date +%s) + GRACE ))
|
|
301
|
-
|
|
386
|
+
# GONE is `absent` or `zombie` — a corpse cannot run cleanup — and everything else,
|
|
387
|
+
# `stopped` included, is still present. Testing `!= live` instead would call a STOPPED
|
|
388
|
+
# suite exited: the work dir would still be there, the wrapper would still reach its
|
|
389
|
+
# bound, and every leg-2 assertion could pass while the suite was never killed at all.
|
|
390
|
+
# Introducing a third process state without revisiting the checks that had two is the
|
|
391
|
+
# same defect this round is about, one file over.
|
|
392
|
+
while :; do
|
|
393
|
+
case "$(wrapper_state "$suite_pid")" in absent|zombie) return 0 ;; esac
|
|
302
394
|
[ "$(date +%s)" -ge "$deadline" ] && break
|
|
303
395
|
sleep 0.5
|
|
304
396
|
done
|
|
305
|
-
|
|
306
|
-
return
|
|
397
|
+
case "$(wrapper_state "$suite_pid")" in absent|zombie) return 0 ;; esac
|
|
398
|
+
return 1
|
|
307
399
|
}
|
|
308
400
|
|
|
309
401
|
# The verdict is about the GROUP, not the leader. The wrapper backgrounds a TERM-immune
|
|
@@ -316,6 +408,15 @@ wrapper_group_members_alive() {
|
|
|
316
408
|
ps -eo pid=,pgid=,stat= 2>/dev/null |
|
|
317
409
|
awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ { print $1 }'
|
|
318
410
|
}
|
|
411
|
+
# Group members still RUNNING after a group stop. A group signal is not an atomic
|
|
412
|
+
# transition: the leader can report stopped while a sibling has not been scheduled to
|
|
413
|
+
# handle it yet, and for the claude stub that sibling is the countdown itself. Zombies
|
|
414
|
+
# and stopped members are excluded; anything left is still able to reach the marker write.
|
|
415
|
+
wrapper_group_unstopped() {
|
|
416
|
+
[ -n "${wrapper_pgid:-}" ] || return 0
|
|
417
|
+
ps -eo pid=,pgid=,stat= 2>/dev/null |
|
|
418
|
+
awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ && $3 !~ /[Tt]/ { print $1 }'
|
|
419
|
+
}
|
|
319
420
|
wrapper_gone_within() {
|
|
320
421
|
local deadline members residue
|
|
321
422
|
deadline=$(( $(date +%s) + $1 ))
|
|
@@ -335,15 +436,39 @@ wrapper_gone_within() {
|
|
|
335
436
|
}
|
|
336
437
|
|
|
337
438
|
diagnose_wrapper() {
|
|
338
|
-
printf 'abort leak diagnostic (%s): wrapper=%s %s\n'
|
|
439
|
+
printf 'abort leak diagnostic (%s): wrapper=%s state=%s bound_marker=%s ps=[%s]\n' \
|
|
440
|
+
"$1" "$wrapper_pid" "$(wrapper_state "$wrapper_pid")" "$(bound_marker_state)" \
|
|
339
441
|
"$(ps -o pid=,ppid=,etime=,stat= -p "$wrapper_pid" 2>/dev/null | tr -s ' ')" >&2
|
|
340
442
|
}
|
|
341
443
|
|
|
444
|
+
# Build the scenario, retrying only the CONSTRUCTION of it. Losing a race to some reaper
|
|
445
|
+
# leaves the probe with nothing to assert about, and reporting that as a failed assertion
|
|
446
|
+
# is what made this gate red at random on three separate assertions — each of them a
|
|
447
|
+
# statement about the probe's environment wearing the wording of a statement about the
|
|
448
|
+
# suite. A lost setup is retried from a fresh suite run; only running out of attempts is
|
|
449
|
+
# a failure, and it says so in those words.
|
|
450
|
+
# $3, when given, is a function run after a successful orphan that must also succeed for
|
|
451
|
+
# the attempt to count — the place for any arming step whose failure means the scenario
|
|
452
|
+
# was lost rather than the suite misbehaved.
|
|
453
|
+
orphan_with_retry() {
|
|
454
|
+
local hang_bound="$1" leg="$2" arm="${3:-}" attempt=1
|
|
455
|
+
while [ "$attempt" -le "$SETUP_ATTEMPTS" ]; do
|
|
456
|
+
if orphan_target_wrapper "$hang_bound" && { [ -z "$arm" ] || "$arm"; }; then return 0; fi
|
|
457
|
+
printf '%s: setup attempt %s/%s did not build the scenario; rebuilding\n' \
|
|
458
|
+
"$leg" "$attempt" "$SETUP_ATTEMPTS" >&2
|
|
459
|
+
cleanup; PROBE_TMP=""; PROBE_TMP_REAL=""; suite_pgid=""; suite_pid=""; suite_start=""
|
|
460
|
+
attempt=$((attempt+1))
|
|
461
|
+
done
|
|
462
|
+
printf '%s: could not build the scenario in %s attempts\n' "$leg" "$SETUP_ATTEMPTS" >&2
|
|
463
|
+
return 1
|
|
464
|
+
}
|
|
465
|
+
|
|
342
466
|
# ---- leg 1: a trappable abort — the suite's own cleanup must reap the wrapper --------
|
|
343
467
|
if [ "$ABORT_LEAK_PROBE_LEG" = 1 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
|
|
344
468
|
leg1_orphaned=0
|
|
345
|
-
|
|
346
|
-
check "leg1: the
|
|
469
|
+
orphan_with_retry "" leg1 && leg1_orphaned=1
|
|
470
|
+
check "leg1: the scenario was built — controller removed, live wrapper reparented to init" \
|
|
471
|
+
'[ "$leg1_orphaned" = 1 ]'
|
|
347
472
|
if [ "$leg1_orphaned" = 1 ]; then
|
|
348
473
|
signal_suite TERM
|
|
349
474
|
leg1_suite_gone=0; await_suite_exit && leg1_suite_gone=1
|
|
@@ -357,31 +482,143 @@ fi
|
|
|
357
482
|
|
|
358
483
|
# ---- leg 2: an untrappable abort — only the fixture's own bound can end the wrapper --
|
|
359
484
|
if [ "$ABORT_LEAK_PROBE_LEG" = 2 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
|
|
485
|
+
leg2_work=""
|
|
486
|
+
# Arm the leg, in this order: prove the wrapper is live, freeze its group, then read the
|
|
487
|
+
# marker. Freezing first is what makes the read meaningful — while the group is stopped
|
|
488
|
+
# nothing can write that file, so "absent" is a stable fact rather than a sample taken
|
|
489
|
+
# between two racing events. Absent means the bound has not fired, which is exactly the
|
|
490
|
+
# precondition leg 2 needs; present means it already fired and the scenario is rebuilt.
|
|
491
|
+
# The ordering this establishes — controller gone, bound not yet fired, abort next — is
|
|
492
|
+
# ENFORCED rather than observed, and needs no clock and no timestamp comparison, which
|
|
493
|
+
# second-granularity stamps could not have given honestly anyway.
|
|
494
|
+
leg2_arm() {
|
|
495
|
+
leg2_work="$(suite_work_dir || true)"
|
|
496
|
+
[ -n "$leg2_work" ] && [ -d "$leg2_work/state" ] || {
|
|
497
|
+
echo "abort_leak_setup_lost: leg2 found no suite work dir to arm against" >&2
|
|
498
|
+
return 1
|
|
499
|
+
}
|
|
500
|
+
[ "$(wrapper_state "$wrapper_pid")" = live ] || {
|
|
501
|
+
printf 'abort_leak_setup_lost: wrapper %s was %s before arming\n' \
|
|
502
|
+
"$wrapper_pid" "$(wrapper_state "$wrapper_pid")" >&2
|
|
503
|
+
return 1
|
|
504
|
+
}
|
|
505
|
+
# STOP the wrapper for the length of the abort. Checking "is it still live" and then
|
|
506
|
+
# killing the suite is a race no shell can close: between the two the countdown can
|
|
507
|
+
# finish, write the marker, and exit, after which every assertion passes while the
|
|
508
|
+
# untrappable-abort path never ran. A stopped process cannot reach the write at all,
|
|
509
|
+
# so the ordering stops being something to observe and becomes something enforced —
|
|
510
|
+
# which is the whole point of this round. The group, not the leader: the claude stub
|
|
511
|
+
# runs its countdown in a backgrounded child that shares the wrapper's group, and
|
|
512
|
+
# stopping only the leader would leave that child counting.
|
|
513
|
+
# Ownership before signalling, proved the way `probe_processes` proves it: re-match the
|
|
514
|
+
# pid against the live table by this run's private path, then require it to LEAD the
|
|
515
|
+
# group being signalled. Those two together mean the group is the validated wrapper's
|
|
516
|
+
# own, which is what makes a group signal safe here.
|
|
517
|
+
# (`probe_group_is_ours` is not used for this: measured on macOS it answers no for a
|
|
518
|
+
# group whose leader `probe_processes` matches by the same private path in the same
|
|
519
|
+
# instant. That is pre-existing and only makes `cleanup` fall back to per-pid kills —
|
|
520
|
+
# which still reap, since every member carries the path — so it is reported rather
|
|
521
|
+
# than changed under this round's scope.)
|
|
522
|
+
probe_processes | grep -qx "$wrapper_pid" || {
|
|
523
|
+
echo "abort_leak_setup_lost: wrapper $wrapper_pid no longer matches this run's private path" >&2
|
|
524
|
+
return 1
|
|
525
|
+
}
|
|
526
|
+
[ "$wrapper_pgid" = "$wrapper_pid" ] || {
|
|
527
|
+
printf 'abort_leak_setup_lost: wrapper %s does not lead group %s; refusing to stop the group\n' \
|
|
528
|
+
"$wrapper_pid" "$wrapper_pgid" >&2
|
|
529
|
+
return 1
|
|
530
|
+
}
|
|
531
|
+
kill -STOP -"$wrapper_pgid" 2>/dev/null || true
|
|
532
|
+
# Wait for the WHOLE group, not just the leader. Reading only the leader is the same
|
|
533
|
+
# proxy-for-the-condition mistake this round exists to remove: the leader can be stopped
|
|
534
|
+
# while the countdown sibling is still running and still able to write the marker.
|
|
535
|
+
leg2_stop_deadline=$(( $(date +%s) + 5 ))
|
|
536
|
+
while :; do
|
|
537
|
+
leg2_unstopped="$(wrapper_group_unstopped | tr '\n' ' ')"
|
|
538
|
+
case "$leg2_unstopped" in *[0-9]*) : ;; *) break ;; esac
|
|
539
|
+
if [ "$(date +%s)" -ge "$leg2_stop_deadline" ]; then
|
|
540
|
+
printf 'abort_leak_setup_lost: group %s still running after SIGSTOP: %s\n' \
|
|
541
|
+
"$wrapper_pgid" "$leg2_unstopped" >&2
|
|
542
|
+
kill -CONT -"$wrapper_pgid" 2>/dev/null || true
|
|
543
|
+
return 1
|
|
544
|
+
fi
|
|
545
|
+
sleep 0.2
|
|
546
|
+
done
|
|
547
|
+
[ "$(wrapper_state "$wrapper_pid")" = stopped ] || {
|
|
548
|
+
printf 'abort_leak_setup_lost: wrapper %s did not stop (state %s)\n' \
|
|
549
|
+
"$wrapper_pid" "$(wrapper_state "$wrapper_pid")" >&2
|
|
550
|
+
kill -CONT -"$wrapper_pgid" 2>/dev/null || true
|
|
551
|
+
return 1
|
|
552
|
+
}
|
|
553
|
+
# Observe the barrier; never manufacture it. Deleting the marker here destroyed the one
|
|
554
|
+
# piece of evidence that distinguishes the two cases: for the claude stub the countdown
|
|
555
|
+
# runs in a child, so the child can finish and write the marker while the wrapper is
|
|
556
|
+
# still unwinding and therefore still reads live. Removing it then made a scenario that
|
|
557
|
+
# was already lost — the bound fired BEFORE the abort — look like a clean run whose
|
|
558
|
+
# marker never came back, i.e. a false RED against a suite that behaved correctly.
|
|
559
|
+
# Read after the freeze, when nothing can be writing that file: absent means the bound
|
|
560
|
+
# has not fired yet, which is the precondition. Present means it already has, so the
|
|
561
|
+
# scenario is lost and gets rebuilt. The suite wipes state files at each case start, so
|
|
562
|
+
# a marker here belongs to this case, and every retry gets a fresh private tmp anyway.
|
|
563
|
+
[ "$(bound_marker_state)" = absent ] || {
|
|
564
|
+
printf 'abort_leak_setup_lost: %s already present — the bound fired before the abort\n' \
|
|
565
|
+
"$PROBE_BOUND_MARKER" >&2
|
|
566
|
+
kill -CONT -"$wrapper_pgid" 2>/dev/null || true
|
|
567
|
+
return 1
|
|
568
|
+
}
|
|
569
|
+
return 0
|
|
570
|
+
}
|
|
571
|
+
# Resume the wrapper once the suite is gone. From here its countdown runs with no
|
|
572
|
+
# controller, no suite, and no trap anywhere — exactly the state leg 2 is about.
|
|
573
|
+
leg2_resume_wrapper() {
|
|
574
|
+
[ -n "${wrapper_pgid:-}" ] && [ "$wrapper_pgid" = "${wrapper_pid:-}" ] &&
|
|
575
|
+
{ kill -CONT -"$wrapper_pgid" 2>/dev/null || true; }
|
|
576
|
+
return 0
|
|
577
|
+
}
|
|
360
578
|
leg2_orphaned=0
|
|
361
|
-
|
|
362
|
-
check "leg2: the
|
|
579
|
+
orphan_with_retry "$LEG2_HANG_BOUND" leg2 leg2_arm && leg2_orphaned=1
|
|
580
|
+
check "leg2: the scenario was built — controller removed, live wrapper reparented to init" \
|
|
581
|
+
'[ "$leg2_orphaned" = 1 ]'
|
|
363
582
|
if [ "$leg2_orphaned" = 1 ]; then
|
|
364
583
|
signal_suite KILL
|
|
584
|
+
# The suite must really be gone, and this stays an ASSERTION rather than a warning:
|
|
585
|
+
# the work-dir check below passes vacuously against a suite that is still running,
|
|
586
|
+
# because a live suite has not reached its EXIT trap either. Demoting this to a
|
|
587
|
+
# diagnostic would leave that check reading "no cleanup ran" whenever the kill missed.
|
|
588
|
+
# What made the old version of this flaky was the zombie ambiguity inside
|
|
589
|
+
# await_suite_exit, not the assertion itself — that is fixed at the helper, so the
|
|
590
|
+
# obligation can be kept instead of traded away.
|
|
365
591
|
leg2_suite_gone=0; await_suite_exit && leg2_suite_gone=1
|
|
366
|
-
|
|
367
|
-
#
|
|
368
|
-
#
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
592
|
+
# Resume only now: the wrapper was held stopped across the abort so its countdown
|
|
593
|
+
# could not have completed before it, and everything observed from here happens with
|
|
594
|
+
# no reaper of any kind left alive.
|
|
595
|
+
leg2_resume_wrapper
|
|
596
|
+
check "leg2: the killed suite is gone" '[ "$leg2_suite_gone" = 1 ]'
|
|
597
|
+
|
|
598
|
+
# No trap ran — structurally, not by the clock. The suite's EXIT trap ends in
|
|
599
|
+
# `rm -rf "$WORK"`, so the work dir outliving a SIGKILLed suite is the trap's absence
|
|
600
|
+
# made visible. The old spelling asked whether the suite exited inside a 30s window,
|
|
601
|
+
# which is a fact about the host's scheduler rather than about whether cleanup ran.
|
|
602
|
+
# Read together with the assertion above, the pair says: the suite is gone AND it left
|
|
603
|
+
# its work dir behind — which only an untrapped death produces.
|
|
604
|
+
leg2_no_cleanup=0
|
|
605
|
+
[ "$leg2_suite_gone" = 1 ] && [ -n "$leg2_work" ] && [ -d "$leg2_work" ] && leg2_no_cleanup=1
|
|
606
|
+
check "leg2: no cleanup ran — the killed suite's work dir survives it" \
|
|
607
|
+
'[ "$leg2_no_cleanup" = 1 ]'
|
|
608
|
+
|
|
381
609
|
leg2_self_exit=0; wrapper_gone_within "$(( LEG2_HANG_BOUND + GRACE ))" && leg2_self_exit=1
|
|
382
610
|
[ "$leg2_self_exit" = 1 ] || diagnose_wrapper leg2
|
|
383
|
-
check "leg2: an unreaped wrapper
|
|
384
|
-
|
|
611
|
+
check "leg2: an unreaped wrapper leaves nothing behind" '[ "$leg2_self_exit" = 1 ]'
|
|
612
|
+
|
|
613
|
+
# WHY it ended, not WHEN it was last seen. The fixture writes this only on the far side
|
|
614
|
+
# of its bounded countdown, so present means the bound ended it and absent means a
|
|
615
|
+
# reaper did — the distinction the deleted "still alive right after the kill" check was
|
|
616
|
+
# trying to draw by looking at the process at one instant, which a corpse answers wrong.
|
|
617
|
+
# An unbounded fixture never reaches the write, so this still reds the reverted bound.
|
|
618
|
+
leg2_bound_marker="$(bound_marker_state)"
|
|
619
|
+
[ "$leg2_bound_marker" = present ] || diagnose_wrapper leg2-bound
|
|
620
|
+
check "leg2: the fixture's own lifetime bound is what ended the wrapper" \
|
|
621
|
+
'[ "$leg2_bound_marker" = present ]'
|
|
385
622
|
fi
|
|
386
623
|
|
|
387
624
|
fi
|
|
@@ -478,7 +478,9 @@ totB = sum(v[1] for v in rows.values())
|
|
|
478
478
|
- **同人异名**:同一个人在不同期用了不同写法(全名/常用名/中间名有无)→ 差分会同时报"一个人离开、一个人加入"。必须先做归一,且归一规则要写出来
|
|
479
479
|
- **收录范围本身会变**:一份名单页面收谁、不收谁,发布方随时可能调整。**"从名单里消失"不等于"离开了"**——它可能只是不再被收录。差分结果只能当**线索**:每一条要用**独立于这份名单的来源**证实(同一上游数据的另一个展示面不算独立,那是循环确认),并写明该来源的截至日期。证实不了的按 `SKILL.md` 的既有处置走——进待核验清单、不得进入承重结论;**别新造一个"未确认"标签**
|
|
480
480
|
|
|
481
|
-
|
|
481
|
+
- **匹配方法自己会造出"消失"**:按 token 集合 / 精确串做归一时,标签**长出新词**(原名被包进一个更长的名字里)会漏配,差集于是报出一个其实还在的实体"消失了"。**凡由差集得出的缺席结论,落笔前先过一遍包含式复检**(把旧名当子串在新全集里查一遍)。**复检只能推翻缺席,不能确认缺席。** 命中且确认与原实体同一(新名把旧名包进去、指向同一对象)→ 判漏配,缺席结论作废;命中但确认是另一个恰好含该字串的实体,或根本没命中 → **只说明这条路没找到它**,不等于它真的不在:改名后的新名可能完全不含旧名,任何字面匹配都照不到。所以缺席结论的成立仍走上一条的既有处置——用独立于该名单的来源证实,证实不了就进待核验清单,不得因为「复检过了」而升格为结论。这与上一条方向相反:上一条是发布方不再收录,这一条是**我方的匹配口径**造的假缺席——两条都要过,只处理其一仍会出错。缺席结论同时要写明检索边界(查了哪些面、到什么截至日期),"未出现在任何一手材料"这类全称否定不写边界即为过强。
|
|
482
|
+
|
|
483
|
+
这三层不处理,一张"逐年进出表"会看起来非常有说服力而实际大半是噪声。
|
|
482
484
|
|
|
483
485
|
### 3.5 取到的是不是全量
|
|
484
486
|
|
|
@@ -95,3 +95,5 @@
|
|
|
95
95
|
## 三、走过而无产出的渠道要留下原因
|
|
96
96
|
|
|
97
97
|
写清是哪一种:**在声明的检索边界内未命中**、**取到但不足以关闭**(只有标题摘要、正文删节、字段范围不足、版本冲突、真实性无法核验)、**被挡住**(需要密钥或账号、接口下线、被反爬拦截)、还是**没走**。只写"无产出"会把后三种伪装成第一种。这四种都不是闭合判断——按 `SKILL.md` 的覆盖矩阵去裁决哪条主张还差什么。
|
|
98
|
+
|
|
99
|
+
**反过来,把某条渠道写进「下一轮最高性价比」之前,先小成本抽样验它到底产出什么。** 入口的价值由**它实际给出什么**决定,不由它的强制力、权威性或可达性决定——「有强制披露义务」「链接打得开」都不是产出证据。判据:能不能举出该渠道上**已取到的、与本轮载重主张同类**的一条内容;举不出就先抽样几份看内容,再决定要不要为它开一轮。这条防的是把同一个高权威、零产出的入口连续两轮写进建议清单,走完才发现不成立。
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md
CHANGED
|
@@ -40,8 +40,8 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
40
40
|
|
|
41
41
|
### Evidence, RCA, charter & attribution(证据 / RCA / charter / 出处核实)
|
|
42
42
|
|
|
43
|
-
- **Charter-before-editing red-line**:
|
|
44
|
-
-
|
|
43
|
+
- **Charter-before-editing red-line**: do not read sources or edit skills until the complete charter in `references/source-to-skill-extraction.md#extraction-charter` is filled cell-by-cell. A remembered field subset is not a charter. Step 0 owns the procedure; this rule owns the hard stop. Core-Rule canonicality governs same-facet drift only and never narrows that field set to this bullet.
|
|
44
|
+
- Classify every result — including task/session summaries and lessons-learned requests, not only failures — as **failure/correction**, **stable success**, or **unstable/insufficient evidence**. Failure runs RCA; stable success requires mechanism, non-luck evidence, reuse conditions, firing point, and owner; insufficient evidence stays an observation. Never invent a failure story to justify learning from success. The author classifies, so the class is not self-elective: correction/finding/failure-triggered work defaults to `failure/correction`, and only independent review may accept a relabel into an RCA-skipping class. Missing classification, unaccepted relabel, or missing analysis leaves it `interim`. Full method: `references/source-to-skill-extraction.md#result-learning-baseline-for-every-extraction`.
|
|
45
45
|
- **RCA must go wider than one 5-Why chain.** 5 Why is the entry technique to get past a symptom, but a single linear chain to one "root cause" is its documented failure mode: real process/agent failures need multiple concurrent causes, the stop point is arbitrary, and "why" drifts toward "who"/blame.
|
|
46
46
|
- For any non-trivial extraction, RCA must (a) **widen** — enumerate the multiple contributing factors across categories (trigger/routing, stale process-model, missing mechanical control, missing feedback, latent authored-earlier condition, detection gap) before deepening one — a straight chain with no branches means you stopped early; (b) **counterfactually test** each candidate to rank causal weight — *if removed or changed, would the failure still happen?* — keeping necessary/sufficient factors **and** failed redundant safeguards as secondary controls / defence-in-depth, and dropping only genuine coincidence (a one-trace factor is a hypothesis — mark it probabilistic, don't hard-drop); (c) frame prevention as a **mechanical control on the failure CLASS** — an enforced constraint plus the feedback that confirms it fired on a surface the next agent actually reaches in time (not a clause buried in a deep reference) — not agent diligence, preferring the highest-leverage **practical** control (a named artifact, an owner who can change it today, an observable check — never deleting a useful narrow gate or inflating one miss into an over-broad hook) over the first patchable point.
|
|
47
47
|
- Reject hindsight causes ("agent careless" / "need more attention" / "should have known"): ask why the action made sense given what the agent could see, not what it should have done. Full method, category prompts, stopping points, and sources: `references/source-to-skill-extraction.md` (Deep RCA For Extraction).
|
|
@@ -56,8 +56,8 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
56
56
|
- Think across the full delivery lifecycle before editing: product intent, design/UX, implementation, debugging, test strategy, launch acceptance, iteration feedback, team onboarding, and normal users without source access. A rule that improves only one slice while leaving another slice ambiguous is incomplete or belongs in a narrower skill.
|
|
57
57
|
- Evidence must come before new rules. Do not add a new conceptual layer, workflow gate, or strong claim first and then backfill supporting sources. If a useful rule appears before source review, keep it as a working hypothesis and do not land it until evidence confirms it, narrows it, or routes it elsewhere. For subjective design, UX, frontend/client, product, architecture, or review rules, unverified external expertise is not enough to land executable guidance.
|
|
58
58
|
- **Product-agnostic / industry-practice skills require an external authoritative source class in the evidence plan, not internal corpus alone.** An extraction sourced only from one internal corpus (an SOP, one repo, one project doc) shows what *this org* does, not whether the skill matches the public state of the art.
|
|
59
|
-
- The trigger is a *public-best-practice / state-of-the-art claim* (architecture, testing-strategy, LLM/inference, observability, release, security, design), not every rule: a rule that encodes an internal-only operating constraint or a postmortem-derived guard, stated with explicit internal scope and no state-of-art claim, does not need external grounding. For rules that do claim to represent industry practice, the charter's evidence plan MUST include authoritative external sources (standards, canonical vendor/tool docs, widely-cited literature, or ≥2 independent practitioner sources) used to confirm, refine, or contradict each such rule — or record per-rule why external grounding is not applicable. Verify any named attribution per the attribution rule; generic established terms (e.g. a well-known named problem or method confirmed by ≥2 independent sources) may be used without person-attribution. Failure shape
|
|
60
|
-
- **Implementing or depending on a named external convention/spec/format verifies it against the primary source FIRST — before building, not after.** This fires on *implementing the named thing itself* — a convention/spec/standard/file-format/protocol such as `AGENTS.md`/CODEOWNERS placement, an RFC, a wire or file format, or a tool's config contract — even with no best-practice claim (distinct from the industry-practice trigger above, which fires on a state-of-the-art *claim*). A prior agent's or a prior commit's reading of that convention is **hypothesis-grade**: re-verify against the primary source before extending it, because a fix-forward built on an inherited interpretation propagates the original error. When the convention uses a term with stack-specific meanings (e.g. "package" = a manifest-bearing directory in npm but *every directory* in Go; likewise module/project/workspace), map it to each concrete target stack before encoding scope. Failure shape
|
|
59
|
+
- The trigger is a *public-best-practice / state-of-the-art claim* (architecture, testing-strategy, LLM/inference, observability, release, security, design), not every rule: a rule that encodes an internal-only operating constraint or a postmortem-derived guard, stated with explicit internal scope and no state-of-art claim, does not need external grounding. For rules that do claim to represent industry practice, the charter's evidence plan MUST include authoritative external sources (standards, canonical vendor/tool docs, widely-cited literature, or ≥2 independent practitioner sources) used to confirm, refine, or contradict each such rule — or record per-rule why external grounding is not applicable. Verify any named attribution per the attribution rule; generic established terms (e.g. a well-known named problem or method confirmed by ≥2 independent sources) may be used without person-attribution. Failure shape 见 `references/external-practice-controls.md`。
|
|
60
|
+
- **Implementing or depending on a named external convention/spec/format verifies it against the primary source FIRST — before building, not after.** This fires on *implementing the named thing itself* — a convention/spec/standard/file-format/protocol such as `AGENTS.md`/CODEOWNERS placement, an RFC, a wire or file format, or a tool's config contract — even with no best-practice claim (distinct from the industry-practice trigger above, which fires on a state-of-the-art *claim*). A prior agent's or a prior commit's reading of that convention is **hypothesis-grade**: re-verify against the primary source before extending it, because a fix-forward built on an inherited interpretation propagates the original error. When the convention uses a term with stack-specific meanings (e.g. "package" = a manifest-bearing directory in npm but *every directory* in Go; likewise module/project/workspace), map it to each concrete target stack before encoding scope. Failure shape 见 `references/external-practice-controls.md`。
|
|
61
61
|
- Match evidence claims to evidence depth. "Full", "complete", "all", "re-read", and "source inventory" claims require named source categories, inspected artifacts, and concrete observations. Use "targeted check" or "no new source read" when that is the real coverage.
|
|
62
62
|
- **A curated digest is one source class, not the repo — its exhaustion is not the repo's exhaustion (digest-masks-corpus trap).** A high-quality maintainer digest (`AGENTS.md`/`CLAUDE.md`, README, CONTRIBUTING, architecture/design doc) that *summarizes* a larger code corpus is a **distinct source class** from the code; a strong digest masks how much went unread. Gates: **(1)** an "exhausted / complete / no-gap / fully-extracted" claim requires the **code corpus as its own register row with a terminal status** (deep-read, inventory+owner-mapping, or a downscope citing an actual user instruction — not self-declared); until then scope the claim ("digest-layer covered; code corpus `pending`") — the sweep is usually inventory+owner-mapping+novelty-spot depth and commonly low-yield, so record that outcome, don't skip the row. **(2) Enumerate the source's OWN top-level structure** before any exhausted claim — a doc's `##`/`###` sections, a repo's top-level dirs (or the next unit: TOC/pages/anchors/line-chunks for a doc; package/module/test/script/config for a repo) — and mark each `read`/`skipped`; un-enumerated structure = unsupported claim (the trap recurs even within one artifact).
|
|
63
63
|
The invariant under both shapes is **an exhaustion claim must be scoped to a unit you actually enumerated** — for the digest/corpus shape that unit is the artifact's own structure; the rule keeps its `digest-masks-corpus` name for continuity, so do not skip it just because no digest is present. **When the claim is over an ACQUISITION CHANNEL SET rather than one artifact** ("the public sources are mined out", "there is no more data"), walking a seed list of channel classes is the cheap way to catch a class you never considered — but a seed list is not a universe, so **the honest output is which classes you walked and with what search boundary, never an exhaustion claim**; the tell that this is the live shape is that each pushback surfaces a class you had not considered rather than another artifact in a known class. Closure and downgrade stay with whatever coverage gate the owning skill already has — do not introduce a parallel status vocabulary here (variant (c) in `references/coverage-exhaustion-traps.md`).
|
|
@@ -147,13 +147,13 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
147
147
|
- Either way the landing must be a **shared** artifact — CCL skill, shared reference, validator, checklist, or project template — naming the exact trigger/gate teammates will hit; for a reusable routing/process/team failure, classify a memory-only landing as insufficient — local-only and "I'll remember next time" count the same (local memory supplements user/workspace context only). A candidate that turns out NOT genuinely reusable may be `discarded` with evidence. When neither path fires, ordinary `routed`/`unchanged` disposition to a different owning skill stays available per bullet A step 3.
|
|
148
148
|
- For any analysis-parse-fix-test-challenge loop, separate five stages explicitly: analysis, parse/decompose, fix, test/verification, and challenge. Add a replay step when validating reusable lessons: rerun the same task shape or a close analog through the proposed workflow and check whether the required outputs and gates still appear in order. Keep four outputs explicit: the project-level fix, the test/verification evidence, the challenge findings, and the reusable workflow lesson. If the same pattern can recur across different domains, lift only the workflow lesson into `skill-extraction-workflow`; keep domain-specific implementation details in the owning project or target skill. See `references/analysis-parse-fix-test-challenge-replay.md` for the replay validation runbook.
|
|
149
149
|
- **The agent failing to self-invoke this workflow (the user had to point out that `skill-extraction-workflow` should have been used) is a tracked failure class that recurs across the session / different tasks, not only within one extraction thread** (so the "twice in one extraction thread" scope above does not catch it). **Honesty:** an in-the-moment self-trigger is recognition-dependent — the always-on bootstrap layer raises its salience but is NOT a mechanical gate; do not overclaim a passive rule "fixes" the recurrence.
|
|
150
|
-
- **One self-detectable firing point does exist and must be used: the moment YOUR OWN output names 沉淀 / 提炼 / 复盘 / "distil this into a skill"
|
|
150
|
+
- **One self-detectable firing point does exist and must be used: the moment YOUR OWN output names 沉淀 / 提炼 / 复盘 / "distil this into a skill", OR **enumerates what an external source has that we lack** (a gap list vs another pack; see `references/firing-point-placement.md`), that naming is a trigger to RECOGNISE the owner and load it** — not a licence to widen scope: shared-skill edits still need the authority you already have, so when the user's request covered only a status review or a narrow fix, record the extraction as `pending` with the owner named and ask rather than self-authorising a shared-skill change off your own suggestion.
|
|
151
151
|
- The mechanical backstops are (a) the closeout gate — a committed skill-change with neither a visible in-session `skill-extraction-workflow` invocation nor the round's durable charter/target-output record is `interim` (per the closeout gate's evidence forms) — and (b) **user-signal escalation**: you generally cannot self-count misses you did not notice, so a user-pointed-out under-trigger is a recurrence check (was there a similar miss earlier this session, even on another task?) and, if so, escalates to tightening the always-on discipline rather than landing another narrow per-case trigger.
|
|
152
152
|
- **Firing-point-placement corollary:** when the SAME meta-class (a precise gate walked past at the routing → pre-code/design transition) recurs at a *new* lifecycle sub-point despite prior bootstrap-salience + the closeout gate, the durable lever is **moving the owning gate's firing point ONTO the transition itself** (pre-substance-draft AND pre-first-impl-edit) and sharpening *name→invoke* — naming/knowing an owner is NOT invoking/loading it, and a named-but-unloaded owner's mechanical rules never fire — at the SAME transition, NOT another bootstrap/per-case bullet or more prose.
|
|
153
153
|
- **Record-field corollary (the forgery surface):** when you land an owner gate as a *field in a record* — a checklist row, a boundary-record line, a CLI flag taking owner names, a "decision:" slot — that field is fillable without invoking the owner, and filling it is what *feels* like discharging the gate; any field naming an owner therefore carries an explicit invoke bar on its triggered values.
|
|
154
154
|
- The self-detect firing point's authority boundary and observed shape, the record-field corollary expansion (the invoke-bar coverage set-diff mechanics, the delegation-dispatch worked case), the worked recurrence-chain, and the landed owner-dispatch implementation: `references/firing-point-placement.md`.
|
|
155
155
|
- **Run your own adversary to convergence BEFORE any "done / fixed / passing / covered / converged / complete" claim — your own such claim is the least-trustworthy thing you emit.** For any non-trivial completion/coverage/convergence claim, you must have already run — **yourself, not deferred to the user** — the verification or adversarial pass that would catch its failure, to a **clean fresh result** (a first clean pass on the current candidate, never a "confirm my fix" pass), OR **downgrade the claim to `interim` and name what you ran vs. didn't**. "Covered / converged / already handled" is a claim, not a status — back it with firing-path or clean-pass evidence or do not emit it; this self-adversary duty never narrows the mandatory dual-track challenge (it is the always-on generalization of self-audit-to-convergence, not a replacement for the gate).
|
|
156
|
-
That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified, and re-reading your own prose can only ever confirm that the prose is self-consistent with itself. **A mutation you did not APPLY is a hypothesis, not evidence** — bound its blast radius (never disable an authorization, idempotency, or deletion guard and exercise it against a shared or live dependency; mutate against isolated dependencies or at the lowest layer that avoids them, and where neither is possible record the property `unverified`). **Prove the oracle can fail before trusting its clean verdict** — point the check at something you know is broken and watch it report that; a check that can only ever say clean is no evidence, and whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration. **A validated oracle is still clean only over the DIMENSIONS it crossed** — proving it can fail says nothing about the axis you never varied, so a clean run is reported with the dimensions it covers, and the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering — `testing-strategy` owns that list) before values. If no contradicting observation exists, the property is `unverified` and must be labelled that way rather than counted as audited.
|
|
156
|
+
That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified, and re-reading your own prose can only ever confirm that the prose is self-consistent with itself. **A mutation you did not APPLY is a hypothesis, not evidence** — bound its blast radius (never disable an authorization, idempotency, or deletion guard and exercise it against a shared or live dependency; mutate against isolated dependencies or at the lowest layer that avoids them, and where neither is possible record the property `unverified`). **Prove the oracle can fail before trusting its clean verdict** — point the check at something you know is broken and watch it report that; a check that can only ever say clean is no evidence, and whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration. **A failing anchor is first a question about the ANCHOR, not a verdict on the implementation** (§Self-audit). **A validated oracle is still clean only over the DIMENSIONS it crossed** — proving it can fail says nothing about the axis you never varied, so a clean run is reported with the dimensions it covers, and the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering — `testing-strategy` owns that list) before values. If no contradicting observation exists, the property is `unverified` and must be labelled that way rather than counted as audited.
|
|
157
157
|
A scoped "X verified; Y not run" is an interim checkpoint, **not** `done`/`complete`/`landed`: `Y not run` blocks a done/complete/landed claim unless a **risk owner — the user/maintainer, never the agent self-accepting — explicitly accepts the gap AND it is tracked to that owner** (agent self-labeling "risk accepted" or "deferred" does not qualify; scoping is a downgrade, never a license to call the narrowed slice done). **Recurrence signal:** a user prompting you to keep digging / verify / disputing a "covered/converged/done" is a premature-completion signal — on the **2nd** such correction in a session (even across different tasks) escalate to tightening this discipline, not just fixing the one case (per the repeated-correction escalation above).
|
|
158
158
|
The full self-adversary method — the mutation enumeration, the applied-mutation discipline, the independent-oracle validation, the re-owe-after-fixes rule, the graded-verdict calibration, and the recognition-dependent honesty caveat: `references/dual-track-review-gate.md` §Self-audit.
|
|
159
159
|
- Automatically trigger durable learning when extraction work exposes a reusable failure — **and when ordinary delivery work does, capture it here too, but without extraction taking over the delivery**: let the active owner (`product-rd-workflow` / `defect-diagnosis` / `testing-strategy` / …) handle the immediate work first, then route the durable lesson here. **For a premature-stop correction after affirmative continuation**, immediate recovery means first rerun the active owner's current continuation/blocking gate in full (for product R&D, Pre-Final Continuation Gate steps 1–6) against current state, then follow its observable outcome — proceeding only when a literal binding exists (the original proposed-next action/scope plus literal assent, preserved in the visible conversation or quoted exactly in trusted host-owned session/compaction state — never reconstructed, broadened, or substituted — or the user's correction literally naming the paused action and scope) — a semantic compaction paraphrase or a bare "why did you stop" complaint is not path-(b) authority, and a `blocked:` recovery without the step-1 evidence and a specific missing authority/ambiguity is invalid — asking again when neither binds, the user intervened, or scope/gates changed, and never copying real conversation text into a shared repository record. Do not let correction RCA or extraction extend a still-authorized delivery, and do not let stale assent bypass a newly pending or inconclusive gate. After delivery recovery, correction RCA plus the durable prevention landing and verification are still due before the turn can be reported complete; otherwise report `interim`. The full binding rules, the `continuing:`-line form, and the invalid-`blocked:`-recovery rule: `references/resume-paused-delivery.md`.
|
|
@@ -166,7 +166,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
166
166
|
|
|
167
167
|
### Validation & the dual-track gate(验证 / dual-track 门)
|
|
168
168
|
|
|
169
|
-
- Static validation is not extraction validation.
|
|
169
|
+
- Static validation is not extraction validation. Nontrivial closeout shows the matching result analysis, target-output and sibling decisions, landed diff, commands, and independent review/challenge; otherwise report `interim` even if static checks pass.
|
|
170
170
|
- For a whole-session/task-retrospective extraction over operational delivery that changed repositories, branches, MRs, pipelines, releases, or deployable artifacts, closeout validation must show one of: `delivery-state rows` with changed artifact, branch/worktree, remote/MR, CI/local verification, cancelled/retried, residual-risk, and next-action state; or `artifact/status axis: not-applicable` with the reason. Missing delivery-state evidence downgrades the extraction to `interim`; static validation and clean independent review do not close it.
|
|
171
171
|
- A rule that exists but did not trigger is a validation-gate defect, not proof that the workflow is adequate.
|
|
172
172
|
- For any correction where the missed step was covered by any rule in this workflow that a reasonable reader would apply to the scenario, add or tighten a closeout gate that would have blocked the exact premature final answer.
|
|
@@ -215,9 +215,9 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
215
215
|
|
|
216
216
|
## Extraction Workflow
|
|
217
217
|
|
|
218
|
-
0. Set the extraction charter and
|
|
218
|
+
0. Set the extraction charter and result-learning baseline.
|
|
219
219
|
- Hard stop: before source reads / before edits, record the charter from `references/source-to-skill-extraction.md#extraction-charter` — **open that table and fill it cell-by-cell; a charter written from memory of the field names is not a charter and must not be recorded as one.** Each field's real constraints live only in its own cell (Evidence plan's produced-artifact-first rule, Scope's watermark validation, RCA depth scaling), so a from-memory charter reproduces the field list and none of the gates, while looking complete. Trivial wording cleanup still records Depth explicitly.
|
|
220
|
-
-
|
|
220
|
+
- Result baseline: classify first, then use `#result-learning-baseline-for-every-extraction`; failure RCA depth comes from `#deep-rca-for-extraction`.
|
|
221
221
|
- Task/session incidents: use `#task-retrospective-extraction` and its delivery-chain RCA prompts before deciding whether the lesson lands in this workflow, a sibling skill, validator, memory, project artifact, or final-response only.
|
|
222
222
|
- Full/complete/deep asks require the source-register shape before reads: source groups, inclusion/exclusion, minimum artifact depth, owner skill, completion evidence, and batch-progress status tracking (`pending`/`read`/`deep-read`/`excluded`/`unavailable`/`routed`) from `#full-coverage-source-register-protocol`; close or explicitly downscope every required batch before saying "complete".
|
|
223
223
|
- Target-output map derives from Lifecycle impact: every affected stage gets an owner target or explicit no-update reason before source-derived editing starts (`#target-output-map`).
|
|
@@ -239,7 +239,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
239
239
|
- The map must include every plausible owner per affected lifecycle stage, including sibling skills in the same stage. If no skill owns a stage, write the no-output reason; do not silently omit the stage.
|
|
240
240
|
- For any extraction beyond wording-only cleanup, include the provenance-to-target diff shape before editing: source mechanism, provenance row, target file, executable landing, test or acceptance owner, and status.
|
|
241
241
|
- Trigger situations and users/tasks it should serve.
|
|
242
|
-
- What future failure
|
|
242
|
+
- What future failure or drift it should prevent, or which evidenced success mechanism it should preserve and reuse.
|
|
243
243
|
- For subjective or high-impact skills such as design, UX, frontend/client, product workflow, architecture, or review, define pressure scenarios and acceptance criteria before editing the skill.
|
|
244
244
|
- For UI/UX or client-facing skills, the pressure scenario must ask whether a person without source access can produce a good-looking and behaviorally sound screen: clear visual hierarchy, fitting density, risk-matched feedback, recoverable state transitions, responsive/device adaptation, and rendered acceptance evidence.
|
|
245
245
|
|
|
@@ -282,8 +282,8 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
282
282
|
|
|
283
283
|
6. Validate before landing.
|
|
284
284
|
- YAML frontmatter parses and description is trigger-focused.
|
|
285
|
-
- Extraction charter is satisfied: purpose, scope, depth,
|
|
286
|
-
-
|
|
285
|
+
- Extraction charter is satisfied: purpose, scope, depth, result classification and matching analysis, evidence plan, and completion standard are either met or explicitly downscoped.
|
|
286
|
+
- Result-learning gate: classification and matching analysis satisfy `#result-learning-baseline-for-every-extraction`; missing, mismatched, or insufficient-evidence-as-rule fails validation.
|
|
287
287
|
- Delivery-chain RCA gate: for incident or task-retrospective extraction, validation must show definition, implementation, verification, review/MR or release readiness, and retrospective-quality causes were checked or explicitly ruled out. If the extraction workflow itself missed the deeper cause, the workflow fix must be landed and validated before finalizing.
|
|
288
288
|
- Blocked-verification gate: any `unavailable`, `skipped`, or `blocked` verification claim must include remediation commands already attempted, observed result, residual risk, and next unblock action. If remediation was feasible but not attempted, validation fails and the test remains pending rather than unavailable.
|
|
289
289
|
- Blocked-source gate: any `unavailable`, `skipped`, `blocked`, timed-out, partial, or failed source-read claim must include the smaller/different read strategy already attempted, observed result, recovered evidence, residual gap, and next unblock action. If remediation was feasible but not attempted, validation fails and the source row remains pending.
|
|
@@ -505,3 +505,11 @@ The challenge pass is **structurally different** from review — it must be invo
|
|
|
505
505
|
## Cost note
|
|
506
506
|
|
|
507
507
|
Challenge pass at high reasoning typically costs 2-5× review pass in tokens. For a ~3 kLOC reference diff, expect ~250-500k tokens on challenge vs ~50-100k on review. The value of one P0 finding caught before landing dwarfs the cost difference; do not skip on cost.
|
|
508
|
+
|
|
509
|
+
## 错误锚点:用错的尺子量对的实现
|
|
510
|
+
|
|
511
|
+
验证 oracle 用的**锚点本身可能是错的**,而这比 oracle 出错更难发现——因为**其余锚点会继续通过**。
|
|
512
|
+
|
|
513
|
+
当锚点的预期方向来自**你对一手源的解读**时,该解读是 hypothesis-grade。锚点不过,第一步应是质疑锚点、回到源的机制陈述重新推导预期,而不是判实现有 bug。否则两种后果:要么去「修」一个正确的实现,要么——更危险——因为其余锚点通过而接受一个错的。
|
|
514
|
+
|
|
515
|
+
观测实例:验证色觉障碍模拟时,用「红色模拟后应变暗」作锚点,实测亮度上升。回到一手后确认:模拟把颜色投影到**单侧二色视者双眼一致的不变轴**(protan/deutan 取 475nm 与 575nm),红投到黄轴、亮度上升是算法的**正确行为**;而「红色看起来暗」说的是**红与黑难以区分**,是另一个量。另外两个锚点(灰阶不变、已知色对靠拢)当时都通过。
|
|
@@ -50,13 +50,13 @@
|
|
|
50
50
|
- **防作弊**:runner 校验每 task 的 `frozen_at_sha` 是 HEAD 祖先(非祖先 = drift,排除出回归判定);同一改动若同时动 task-bank 和 SKILL.md description 会显式告警(防"改 skill 顺手改测试让它过")。
|
|
51
51
|
- `--baseline <json>` 给 diff 式报告(newly_failed / newly_passed)。
|
|
52
52
|
|
|
53
|
-
### Bank
|
|
53
|
+
### Bank 用例修复的测量协议(测量必做;落地裁决归本轮实际门禁)
|
|
54
54
|
|
|
55
|
-
|
|
55
|
+
以下是生成可比较证据的默认协议。**降级的是「F4 自己充当统一合并门禁」这个声称,不是「必须测、且必须有人裁决」这个义务**——这两件事分开:落地判断交给本轮实际的 owner/风险/评审门禁,但**测量本身不可选**。任何动 routing 面(SKILL.md description、task-bank 判定面)的改动都必须按下列协议产出证据;没跑就是没收敛,不得进入独立评审、也不得声称本轮无回归。采用不同样本量时,须随工件记录理由,且该理由与本轮证据一同进入独立评审——「记了理由」本身不是豁免,自审通过的理由不构成已裁决:
|
|
56
56
|
|
|
57
57
|
1. 动任何 description 之前必须先跑 **≥10 轮有效观测**的稳定性基线,把稳定失败与抖动分开;抖动不得作为修改依据(grader 超时/不可解析轮不算有效观测,须补跑)。
|
|
58
58
|
2. 改后通过数必须在**最终措辞**上重测:中间稿的通过数在措辞再变的那一刻作废,不得挪用到最终候选的证据里。
|
|
59
|
-
3.
|
|
59
|
+
3. 受影响邻居用例集默认改前/改后各 **≥3 轮**,集合须含期望 owner 自己的兄弟用例与高词面重叠的他 owner 用例;邻居回归作为独立 finding 交由本轮实际门禁处置——**该 finding 须以 blocking 记入本轮 dual-track 评审记录,且只能由独立评审方豁免,不能由实现者自行判定「本轮没有门禁采用这组证据」而放行**。降级的是「F4 自己充当合并门禁」这一声称,不是「回归必须被人裁决」这一义务;后者若也随之消失,这一条就只剩被裁决方自审。
|
|
60
60
|
4. 每轮判决必须连同 **runner 调用、grader 模型身份、候选身份**(commit 或描述内容指纹)与**原始逐轮工件的持久定位符**一并记入轮记录;没有定位符的通过数只能标注为 operator-reported,不得据以宣称修复轮已 concluded。
|
|
61
61
|
|
|
62
62
|
## Tier-3:hub golden trace 真 agent 回放(已落地,advisory,人工判定)
|
|
@@ -71,16 +71,15 @@
|
|
|
71
71
|
- **随机性**:agent 非确定;判定先人工、nightly 起步,有稳定史前不自动 gate。防作弊同 T2(`frozen_at_sha` 祖先校验)。
|
|
72
72
|
- **双用途**:除回归外,Tier-3 还可当**改技能前的 RED-baseline**(改前手动跑触发场景看真 agent 是否真路由错,改后看 compliance)——可选;只有真观察到 miss 才算 RED(PASS/INCONCLUSIVE 不算),小 N + 非确定有噪声,手动跑两次自己留两份报告。落地 + 防作弊注意见 [validation-and-landing.md](validation-and-landing.md) "Optional real-agent RED-baseline"。
|
|
73
73
|
|
|
74
|
-
## Health roll-up
|
|
74
|
+
## Health roll-up:描述性仪表盘(已落地,advisory)
|
|
75
75
|
|
|
76
|
-
把上面各信号卷成**一个加权 0–10
|
|
76
|
+
把上面各信号卷成**一个加权 0–10 显示值 + 同尺子变化**,用于定位值得继续检查的维度。它借用 OpenSSF Scorecard 的呈现形态,但不把不同性质的 F4 信号变成“仓库整体变好/变差”的总判决。映射与限制见 [harness-patterns-and-eval.md](harness-patterns-and-eval.md) §3.4。
|
|
77
77
|
|
|
78
78
|
`scripts/eval-health.rb <repo-root> [--trace-json p] [--bank-json p] [--history p] [--no-write] [--json p] [--quiet]`。`make eval-health`。
|
|
79
79
|
|
|
80
80
|
- **维度(各 0–10,按风险加权,OpenSSF 风格)**:`structural` 权重 10(Critical,`validate-skill.sh` pass/fail)· `routing_static` 权重 10(Critical,T1 blocking=0 满分、有 blocking 砸到 3、advisory 轻罚)· `trace` 权重 7.5(High,T3 pass/considered)· `bank` 权重 5(Medium,T2 pass/tasks)。
|
|
81
81
|
- **只跑确定性两维**(structural + routing_static,无 LLM、快);`trace`/`bank` 需 `claude` 且随机,**不自动跑** —— 用 `--trace-json` / `--bank-json` 把 T3/T2 报告喂进来,否则该维 **skip,权重按比例重分**给在场维(诚实标注 `dims=…`)。
|
|
82
|
-
- `composite = Σ(score_i·w_i) / Σ(w_i)`,只对在场维求和;skip 维自动从分母剔除。band
|
|
83
|
-
- **advisory
|
|
82
|
+
- `composite = Σ(score_i·w_i) / Σ(w_i)`,只对在场维求和;skip 维自动从分母剔除。band 只是浏览提示,不得当作验收等级。
|
|
83
|
+
- **advisory,永不因显示值阻断**:退出码 `0` = 跑完 **或** 没有可算的维(都不是失败);`2` = 用法/setup 错(`<repo-root>` 不对或无 `skills/` 目录)。低值、坏报告、空历史都不会非 0。畸形的 T2/T3 报告(非对象 / 非整数 / `pass>total`)直接 **skip,不静默打分**。**绝不接进 `check-ccl-skills.sh`**。结构校验、T1 blocking、任务验收和行为证据分别独立判定;质量失败不能被其它维度的高值平均掉。
|
|
84
84
|
- **corpus/version 守卫(防跨变更 task-bank 当稳定指标比)**:每条历史记 `corpus`(= task-bank + golden-traces 输入内容的指纹)+ `repo_sha` + 在场 `dims`。**趋势 delta(IMPROVING/DECLINING)只跟最近一条 `(corpus, dims)` 都相同的历史比**;否则打印"baseline reset, not compared",不偷偷比。这样"加了 10 条简单 task → 分涨了"不会被读成真进步(尺子换了)。
|
|
85
85
|
- **历史文件**:默认 `eval/health-history.jsonl`,**git-ignore**(同 Goodhart 理由:committed 的数会招"调数不修仓";也免 append 把树搞脏)。`--history` 可改路径,`--no-write` 不落盘。
|
|
86
|
-
- **基线**(本机、本 corpus):确定性两维 = `9.5/10`(structural 10 + routing_static 9,advisory=1)。
|