@ccoalm/ccl-skills 0.2.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +34 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_abort_leak_state_helpers.sh +148 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +271 -34
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +8 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +7 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +7 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +29 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +28 -29
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +153 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +23 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +303 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +222 -91
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +6 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +12 -12
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +455 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +20 -20
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_liveness_predicate_gate.sh +288 -0
- package/dist/assets/release.json +60 -25
- package/dist/cli.js +23 -2
- package/dist/update-notice.d.ts +71 -0
- package/dist/update-notice.js +173 -0
- package/dist/version-check.d.ts +2 -0
- package/dist/version-check.js +7 -0
- package/package.json +1 -1
|
@@ -13,6 +13,14 @@
|
|
|
13
13
|
# leg 2 SIGKILL the suite: nothing can trap that, so no reaper runs at all and the
|
|
14
14
|
# wrapper must die on the fixture's own lifetime bound instead.
|
|
15
15
|
#
|
|
16
|
+
# Every verdict below is a fact the run leaves behind — the work dir the EXIT trap would
|
|
17
|
+
# have deleted, the marker the fixture writes only past its countdown — never an
|
|
18
|
+
# observation of what was true at the instant the probe looked. Liveness questions all go
|
|
19
|
+
# through wrapper_state() so they cannot answer the same process differently, and a
|
|
20
|
+
# scenario the probe fails to BUILD is retried rather than reported as a failed assertion.
|
|
21
|
+
# All three of those rules are here because the earlier spellings red this gate on CI
|
|
22
|
+
# while the suite was behaving correctly.
|
|
23
|
+
#
|
|
16
24
|
# Leg 2 exists because leg 1 alone stays green if the fixture's bound were restored to
|
|
17
25
|
# an unbounded loop: the trap would still reap it. It was an unbounded loop that turned
|
|
18
26
|
# this defect into a wrapper observed alive for 39 hours with ppid=1, ignoring SIGTERM.
|
|
@@ -59,10 +67,48 @@ esac
|
|
|
59
67
|
# with the probe still green, so CI points its two jobs at different stubs.
|
|
60
68
|
ABORT_LEAK_PROBE_CLIENT="${ABORT_LEAK_PROBE_CLIENT:-claude}"
|
|
61
69
|
case "$ABORT_LEAK_PROBE_CLIENT" in
|
|
62
|
-
claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh
|
|
63
|
-
|
|
70
|
+
claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh
|
|
71
|
+
PROBE_BOUND_MARKER=claude_hang_bound_reached ;;
|
|
72
|
+
fallback) PROBE_BEHAVIOR_FILE=kimi_behavior; PROBE_WRAPPER=kimi_review.sh
|
|
73
|
+
PROBE_BOUND_MARKER=kimi_hang_bound_reached ;;
|
|
64
74
|
*) echo "ABORT_LEAK_PROBE_CLIENT must be claude or fallback" >&2; exit 2 ;;
|
|
65
75
|
esac
|
|
76
|
+
# How many times a leg may rebuild its scenario before giving up. Only the SETUP is
|
|
77
|
+
# retried: constructing "an orphaned wrapper with no reaper left" depends on the probe
|
|
78
|
+
# winning races against the controller and the host, and losing one says nothing about
|
|
79
|
+
# the suite. An assertion ABOUT the suite is never retried, so a fixture whose bound was
|
|
80
|
+
# reverted fails on every attempt and cannot be retried into green.
|
|
81
|
+
SETUP_ATTEMPTS="${ABORT_LEAK_PROBE_SETUP_ATTEMPTS:-3}"
|
|
82
|
+
|
|
83
|
+
# ONE process-state vocabulary for the whole probe. Three separate spellings of "is it
|
|
84
|
+
# there?" used to disagree about the same pid: `kill -0` succeeds for a zombie, reading
|
|
85
|
+
# ppid succeeds for a zombie, and the two verdict scans exclude zombies. An already-dead
|
|
86
|
+
# wrapper awaiting reaping therefore satisfied "reparented to init", failed "still alive",
|
|
87
|
+
# and satisfied "gone" — one process state, three answers, and a red that named the suite
|
|
88
|
+
# for something the suite had not done. Every liveness question below goes through here.
|
|
89
|
+
wrapper_state() {
|
|
90
|
+
local st rc
|
|
91
|
+
case "${1:-}" in ''|*[!0-9]*) printf 'absent\n'; return 0 ;; esac
|
|
92
|
+
st="$(ps -o stat= -p "$1" 2>/dev/null)"; rc=$?
|
|
93
|
+
# `ps` exits 1 for "no such process", which is a real answer. Anything above that is
|
|
94
|
+
# the TOOL failing, not the process being gone — and swallowing that into `absent`
|
|
95
|
+
# would report a suite exited because `ps` could not run. Unknown is reported as its
|
|
96
|
+
# own state and every consumer treats it as still present, because the expensive
|
|
97
|
+
# direction of this error is declaring something gone that is not.
|
|
98
|
+
if [ "$rc" -gt 1 ]; then printf 'unknown\n'; return 0; fi
|
|
99
|
+
case "$st" in
|
|
100
|
+
'') printf 'absent\n' ;;
|
|
101
|
+
*Z*) printf 'zombie\n' ;;
|
|
102
|
+
# A stopped process is neither running nor gone, and leg 2 deliberately puts the
|
|
103
|
+
# wrapper in this state while it arms. Folding it into `live` would let the arming
|
|
104
|
+
# step report success without SIGSTOP having actually landed.
|
|
105
|
+
# `T` (job-control stop) or `t` (tracing stop), anywhere in the field: neither letter
|
|
106
|
+
# appears among the flag characters either ps appends, so this cannot catch a
|
|
107
|
+
# running process by accident.
|
|
108
|
+
*[Tt]*) printf 'stopped\n' ;;
|
|
109
|
+
*) printf 'live\n' ;;
|
|
110
|
+
esac
|
|
111
|
+
}
|
|
66
112
|
|
|
67
113
|
fails=0
|
|
68
114
|
check() {
|
|
@@ -178,6 +224,21 @@ suite_work_dir() {
|
|
|
178
224
|
done
|
|
179
225
|
return 1
|
|
180
226
|
}
|
|
227
|
+
# The fixture writes this file only after its bounded countdown runs to completion, so
|
|
228
|
+
# its presence says the fixture's OWN lifetime bound ended the wrapper and its absence
|
|
229
|
+
# says something else did. That is the fact leg 2 needs, and it is a fact about which
|
|
230
|
+
# code path ran — not about what was true at the instant the probe happened to look.
|
|
231
|
+
bound_marker_path() {
|
|
232
|
+
local work
|
|
233
|
+
work="$(suite_work_dir)" || return 1
|
|
234
|
+
printf '%s\n' "$work/state/$PROBE_BOUND_MARKER"
|
|
235
|
+
}
|
|
236
|
+
bound_marker_state() {
|
|
237
|
+
local path
|
|
238
|
+
path="$(bound_marker_path)" || { printf 'no-work-dir\n'; return 0; }
|
|
239
|
+
if [ -e "$path" ]; then printf 'present\n'; else printf 'absent\n'; fi
|
|
240
|
+
}
|
|
241
|
+
|
|
181
242
|
live_wrapper() {
|
|
182
243
|
local work behavior pid parent_cmd
|
|
183
244
|
work="$(suite_work_dir)" || return 0
|
|
@@ -203,7 +264,10 @@ orphan_target_wrapper() {
|
|
|
203
264
|
local hang_bound="$1" candidate controller_pid controller_ok deadline suite_died
|
|
204
265
|
PROBE_TMP="$(mktemp -d "${TMPDIR:-/tmp}/review-gate-abort-probe.XXXXXX")" || return 1
|
|
205
266
|
PROBE_TMP_REAL="$(cd "$PROBE_TMP" && pwd -P)" || return 1
|
|
267
|
+
# Both, not just the pid: a retry that left the previous attempt's group id in place
|
|
268
|
+
# would have the verdict scan looking for members of a group this attempt never built.
|
|
206
269
|
wrapper_pid=""
|
|
270
|
+
wrapper_pgid=""
|
|
207
271
|
|
|
208
272
|
# Job control so the suite gets its own process group and the probe can signal that
|
|
209
273
|
# group without signalling itself.
|
|
@@ -219,7 +283,11 @@ orphan_target_wrapper() {
|
|
|
219
283
|
while [ "$(date +%s)" -lt "$deadline" ]; do
|
|
220
284
|
candidate="$(live_wrapper)"
|
|
221
285
|
if [ -n "$candidate" ]; then wrapper_pid="$candidate"; break; fi
|
|
222
|
-
|
|
286
|
+
# Third site of the same class: a suite that has exited but whose corpse this shell
|
|
287
|
+
# has not reaped still answers `kill -0`, so the bare existence test would keep this
|
|
288
|
+
# loop polling for a dead suite until the whole reach budget elapsed and then report
|
|
289
|
+
# the wrong one of the two failures below.
|
|
290
|
+
[ "$(wrapper_state "$suite_pid")" = live ] || { suite_died=1; break; }
|
|
223
291
|
sleep 0.1
|
|
224
292
|
done
|
|
225
293
|
if [ -z "$wrapper_pid" ]; then
|
|
@@ -281,29 +349,53 @@ orphan_target_wrapper() {
|
|
|
281
349
|
fi
|
|
282
350
|
sleep 0.5
|
|
283
351
|
done
|
|
352
|
+
# Its own deadline: this loop used to reuse the one the controller wait had already
|
|
353
|
+
# been counting down, so a controller that took nine of those ten seconds to die left
|
|
354
|
+
# the reparent check one second to succeed in.
|
|
355
|
+
deadline=$(( $(date +%s) + 10 ))
|
|
284
356
|
while :; do
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
357
|
+
# A corpse is not an orphan. The old spelling of this check read ppid, which a zombie
|
|
358
|
+
# answers, so a wrapper some reaper had already killed passed for a live orphan and
|
|
359
|
+
# the leg went on to make assertions about a scenario it had never built.
|
|
360
|
+
if [ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "1" ] &&
|
|
361
|
+
[ "$(wrapper_state "$wrapper_pid")" = live ]; then
|
|
362
|
+
return 0
|
|
363
|
+
fi
|
|
364
|
+
case "$(wrapper_state "$wrapper_pid")" in
|
|
365
|
+
live) : ;;
|
|
366
|
+
*)
|
|
367
|
+
printf 'abort_leak_setup_lost: wrapper %s was already %s before the abort — a reaper reached it first (bound marker: %s)\n' \
|
|
368
|
+
"$wrapper_pid" "$(wrapper_state "$wrapper_pid")" "$(bound_marker_state)" >&2
|
|
369
|
+
return 1
|
|
370
|
+
;;
|
|
371
|
+
esac
|
|
290
372
|
if [ "$(date +%s)" -ge "$deadline" ]; then
|
|
291
|
-
echo "
|
|
373
|
+
echo "abort_leak_setup_lost: wrapper $wrapper_pid was never reparented" >&2
|
|
292
374
|
return 1
|
|
293
375
|
fi
|
|
294
376
|
sleep 0.5
|
|
295
377
|
done
|
|
296
378
|
}
|
|
297
379
|
|
|
380
|
+
# Same one vocabulary as the wrapper: a killed suite whose corpse the probe's own shell
|
|
381
|
+
# has not reaped yet still answers `kill -0`, so spelling this as `kill -0` made a bookkeeping
|
|
382
|
+
# lag inside the probe look like a suite that refused to die.
|
|
298
383
|
await_suite_exit() {
|
|
299
384
|
local deadline
|
|
300
385
|
deadline=$(( $(date +%s) + GRACE ))
|
|
301
|
-
|
|
386
|
+
# GONE is `absent` or `zombie` — a corpse cannot run cleanup — and everything else,
|
|
387
|
+
# `stopped` included, is still present. Testing `!= live` instead would call a STOPPED
|
|
388
|
+
# suite exited: the work dir would still be there, the wrapper would still reach its
|
|
389
|
+
# bound, and every leg-2 assertion could pass while the suite was never killed at all.
|
|
390
|
+
# Introducing a third process state without revisiting the checks that had two is the
|
|
391
|
+
# same defect this round is about, one file over.
|
|
392
|
+
while :; do
|
|
393
|
+
case "$(wrapper_state "$suite_pid")" in absent|zombie) return 0 ;; esac
|
|
302
394
|
[ "$(date +%s)" -ge "$deadline" ] && break
|
|
303
395
|
sleep 0.5
|
|
304
396
|
done
|
|
305
|
-
|
|
306
|
-
return
|
|
397
|
+
case "$(wrapper_state "$suite_pid")" in absent|zombie) return 0 ;; esac
|
|
398
|
+
return 1
|
|
307
399
|
}
|
|
308
400
|
|
|
309
401
|
# The verdict is about the GROUP, not the leader. The wrapper backgrounds a TERM-immune
|
|
@@ -316,6 +408,15 @@ wrapper_group_members_alive() {
|
|
|
316
408
|
ps -eo pid=,pgid=,stat= 2>/dev/null |
|
|
317
409
|
awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ { print $1 }'
|
|
318
410
|
}
|
|
411
|
+
# Group members still RUNNING after a group stop. A group signal is not an atomic
|
|
412
|
+
# transition: the leader can report stopped while a sibling has not been scheduled to
|
|
413
|
+
# handle it yet, and for the claude stub that sibling is the countdown itself. Zombies
|
|
414
|
+
# and stopped members are excluded; anything left is still able to reach the marker write.
|
|
415
|
+
wrapper_group_unstopped() {
|
|
416
|
+
[ -n "${wrapper_pgid:-}" ] || return 0
|
|
417
|
+
ps -eo pid=,pgid=,stat= 2>/dev/null |
|
|
418
|
+
awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ && $3 !~ /[Tt]/ { print $1 }'
|
|
419
|
+
}
|
|
319
420
|
wrapper_gone_within() {
|
|
320
421
|
local deadline members residue
|
|
321
422
|
deadline=$(( $(date +%s) + $1 ))
|
|
@@ -335,15 +436,39 @@ wrapper_gone_within() {
|
|
|
335
436
|
}
|
|
336
437
|
|
|
337
438
|
diagnose_wrapper() {
|
|
338
|
-
printf 'abort leak diagnostic (%s): wrapper=%s %s\n'
|
|
439
|
+
printf 'abort leak diagnostic (%s): wrapper=%s state=%s bound_marker=%s ps=[%s]\n' \
|
|
440
|
+
"$1" "$wrapper_pid" "$(wrapper_state "$wrapper_pid")" "$(bound_marker_state)" \
|
|
339
441
|
"$(ps -o pid=,ppid=,etime=,stat= -p "$wrapper_pid" 2>/dev/null | tr -s ' ')" >&2
|
|
340
442
|
}
|
|
341
443
|
|
|
444
|
+
# Build the scenario, retrying only the CONSTRUCTION of it. Losing a race to some reaper
|
|
445
|
+
# leaves the probe with nothing to assert about, and reporting that as a failed assertion
|
|
446
|
+
# is what made this gate red at random on three separate assertions — each of them a
|
|
447
|
+
# statement about the probe's environment wearing the wording of a statement about the
|
|
448
|
+
# suite. A lost setup is retried from a fresh suite run; only running out of attempts is
|
|
449
|
+
# a failure, and it says so in those words.
|
|
450
|
+
# $3, when given, is a function run after a successful orphan that must also succeed for
|
|
451
|
+
# the attempt to count — the place for any arming step whose failure means the scenario
|
|
452
|
+
# was lost rather than the suite misbehaved.
|
|
453
|
+
orphan_with_retry() {
|
|
454
|
+
local hang_bound="$1" leg="$2" arm="${3:-}" attempt=1
|
|
455
|
+
while [ "$attempt" -le "$SETUP_ATTEMPTS" ]; do
|
|
456
|
+
if orphan_target_wrapper "$hang_bound" && { [ -z "$arm" ] || "$arm"; }; then return 0; fi
|
|
457
|
+
printf '%s: setup attempt %s/%s did not build the scenario; rebuilding\n' \
|
|
458
|
+
"$leg" "$attempt" "$SETUP_ATTEMPTS" >&2
|
|
459
|
+
cleanup; PROBE_TMP=""; PROBE_TMP_REAL=""; suite_pgid=""; suite_pid=""; suite_start=""
|
|
460
|
+
attempt=$((attempt+1))
|
|
461
|
+
done
|
|
462
|
+
printf '%s: could not build the scenario in %s attempts\n' "$leg" "$SETUP_ATTEMPTS" >&2
|
|
463
|
+
return 1
|
|
464
|
+
}
|
|
465
|
+
|
|
342
466
|
# ---- leg 1: a trappable abort — the suite's own cleanup must reap the wrapper --------
|
|
343
467
|
if [ "$ABORT_LEAK_PROBE_LEG" = 1 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
|
|
344
468
|
leg1_orphaned=0
|
|
345
|
-
|
|
346
|
-
check "leg1: the
|
|
469
|
+
orphan_with_retry "" leg1 && leg1_orphaned=1
|
|
470
|
+
check "leg1: the scenario was built — controller removed, live wrapper reparented to init" \
|
|
471
|
+
'[ "$leg1_orphaned" = 1 ]'
|
|
347
472
|
if [ "$leg1_orphaned" = 1 ]; then
|
|
348
473
|
signal_suite TERM
|
|
349
474
|
leg1_suite_gone=0; await_suite_exit && leg1_suite_gone=1
|
|
@@ -357,31 +482,143 @@ fi
|
|
|
357
482
|
|
|
358
483
|
# ---- leg 2: an untrappable abort — only the fixture's own bound can end the wrapper --
|
|
359
484
|
if [ "$ABORT_LEAK_PROBE_LEG" = 2 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
|
|
485
|
+
leg2_work=""
|
|
486
|
+
# Arm the leg, in this order: prove the wrapper is live, freeze its group, then read the
|
|
487
|
+
# marker. Freezing first is what makes the read meaningful — while the group is stopped
|
|
488
|
+
# nothing can write that file, so "absent" is a stable fact rather than a sample taken
|
|
489
|
+
# between two racing events. Absent means the bound has not fired, which is exactly the
|
|
490
|
+
# precondition leg 2 needs; present means it already fired and the scenario is rebuilt.
|
|
491
|
+
# The ordering this establishes — controller gone, bound not yet fired, abort next — is
|
|
492
|
+
# ENFORCED rather than observed, and needs no clock and no timestamp comparison, which
|
|
493
|
+
# second-granularity stamps could not have given honestly anyway.
|
|
494
|
+
leg2_arm() {
|
|
495
|
+
leg2_work="$(suite_work_dir || true)"
|
|
496
|
+
[ -n "$leg2_work" ] && [ -d "$leg2_work/state" ] || {
|
|
497
|
+
echo "abort_leak_setup_lost: leg2 found no suite work dir to arm against" >&2
|
|
498
|
+
return 1
|
|
499
|
+
}
|
|
500
|
+
[ "$(wrapper_state "$wrapper_pid")" = live ] || {
|
|
501
|
+
printf 'abort_leak_setup_lost: wrapper %s was %s before arming\n' \
|
|
502
|
+
"$wrapper_pid" "$(wrapper_state "$wrapper_pid")" >&2
|
|
503
|
+
return 1
|
|
504
|
+
}
|
|
505
|
+
# STOP the wrapper for the length of the abort. Checking "is it still live" and then
|
|
506
|
+
# killing the suite is a race no shell can close: between the two the countdown can
|
|
507
|
+
# finish, write the marker, and exit, after which every assertion passes while the
|
|
508
|
+
# untrappable-abort path never ran. A stopped process cannot reach the write at all,
|
|
509
|
+
# so the ordering stops being something to observe and becomes something enforced —
|
|
510
|
+
# which is the whole point of this round. The group, not the leader: the claude stub
|
|
511
|
+
# runs its countdown in a backgrounded child that shares the wrapper's group, and
|
|
512
|
+
# stopping only the leader would leave that child counting.
|
|
513
|
+
# Ownership before signalling, proved the way `probe_processes` proves it: re-match the
|
|
514
|
+
# pid against the live table by this run's private path, then require it to LEAD the
|
|
515
|
+
# group being signalled. Those two together mean the group is the validated wrapper's
|
|
516
|
+
# own, which is what makes a group signal safe here.
|
|
517
|
+
# (`probe_group_is_ours` is not used for this: measured on macOS it answers no for a
|
|
518
|
+
# group whose leader `probe_processes` matches by the same private path in the same
|
|
519
|
+
# instant. That is pre-existing and only makes `cleanup` fall back to per-pid kills —
|
|
520
|
+
# which still reap, since every member carries the path — so it is reported rather
|
|
521
|
+
# than changed under this round's scope.)
|
|
522
|
+
probe_processes | grep -qx "$wrapper_pid" || {
|
|
523
|
+
echo "abort_leak_setup_lost: wrapper $wrapper_pid no longer matches this run's private path" >&2
|
|
524
|
+
return 1
|
|
525
|
+
}
|
|
526
|
+
[ "$wrapper_pgid" = "$wrapper_pid" ] || {
|
|
527
|
+
printf 'abort_leak_setup_lost: wrapper %s does not lead group %s; refusing to stop the group\n' \
|
|
528
|
+
"$wrapper_pid" "$wrapper_pgid" >&2
|
|
529
|
+
return 1
|
|
530
|
+
}
|
|
531
|
+
kill -STOP -"$wrapper_pgid" 2>/dev/null || true
|
|
532
|
+
# Wait for the WHOLE group, not just the leader. Reading only the leader is the same
|
|
533
|
+
# proxy-for-the-condition mistake this round exists to remove: the leader can be stopped
|
|
534
|
+
# while the countdown sibling is still running and still able to write the marker.
|
|
535
|
+
leg2_stop_deadline=$(( $(date +%s) + 5 ))
|
|
536
|
+
while :; do
|
|
537
|
+
leg2_unstopped="$(wrapper_group_unstopped | tr '\n' ' ')"
|
|
538
|
+
case "$leg2_unstopped" in *[0-9]*) : ;; *) break ;; esac
|
|
539
|
+
if [ "$(date +%s)" -ge "$leg2_stop_deadline" ]; then
|
|
540
|
+
printf 'abort_leak_setup_lost: group %s still running after SIGSTOP: %s\n' \
|
|
541
|
+
"$wrapper_pgid" "$leg2_unstopped" >&2
|
|
542
|
+
kill -CONT -"$wrapper_pgid" 2>/dev/null || true
|
|
543
|
+
return 1
|
|
544
|
+
fi
|
|
545
|
+
sleep 0.2
|
|
546
|
+
done
|
|
547
|
+
[ "$(wrapper_state "$wrapper_pid")" = stopped ] || {
|
|
548
|
+
printf 'abort_leak_setup_lost: wrapper %s did not stop (state %s)\n' \
|
|
549
|
+
"$wrapper_pid" "$(wrapper_state "$wrapper_pid")" >&2
|
|
550
|
+
kill -CONT -"$wrapper_pgid" 2>/dev/null || true
|
|
551
|
+
return 1
|
|
552
|
+
}
|
|
553
|
+
# Observe the barrier; never manufacture it. Deleting the marker here destroyed the one
|
|
554
|
+
# piece of evidence that distinguishes the two cases: for the claude stub the countdown
|
|
555
|
+
# runs in a child, so the child can finish and write the marker while the wrapper is
|
|
556
|
+
# still unwinding and therefore still reads live. Removing it then made a scenario that
|
|
557
|
+
# was already lost — the bound fired BEFORE the abort — look like a clean run whose
|
|
558
|
+
# marker never came back, i.e. a false RED against a suite that behaved correctly.
|
|
559
|
+
# Read after the freeze, when nothing can be writing that file: absent means the bound
|
|
560
|
+
# has not fired yet, which is the precondition. Present means it already has, so the
|
|
561
|
+
# scenario is lost and gets rebuilt. The suite wipes state files at each case start, so
|
|
562
|
+
# a marker here belongs to this case, and every retry gets a fresh private tmp anyway.
|
|
563
|
+
[ "$(bound_marker_state)" = absent ] || {
|
|
564
|
+
printf 'abort_leak_setup_lost: %s already present — the bound fired before the abort\n' \
|
|
565
|
+
"$PROBE_BOUND_MARKER" >&2
|
|
566
|
+
kill -CONT -"$wrapper_pgid" 2>/dev/null || true
|
|
567
|
+
return 1
|
|
568
|
+
}
|
|
569
|
+
return 0
|
|
570
|
+
}
|
|
571
|
+
# Resume the wrapper once the suite is gone. From here its countdown runs with no
|
|
572
|
+
# controller, no suite, and no trap anywhere — exactly the state leg 2 is about.
|
|
573
|
+
leg2_resume_wrapper() {
|
|
574
|
+
[ -n "${wrapper_pgid:-}" ] && [ "$wrapper_pgid" = "${wrapper_pid:-}" ] &&
|
|
575
|
+
{ kill -CONT -"$wrapper_pgid" 2>/dev/null || true; }
|
|
576
|
+
return 0
|
|
577
|
+
}
|
|
360
578
|
leg2_orphaned=0
|
|
361
|
-
|
|
362
|
-
check "leg2: the
|
|
579
|
+
orphan_with_retry "$LEG2_HANG_BOUND" leg2 leg2_arm && leg2_orphaned=1
|
|
580
|
+
check "leg2: the scenario was built — controller removed, live wrapper reparented to init" \
|
|
581
|
+
'[ "$leg2_orphaned" = 1 ]'
|
|
363
582
|
if [ "$leg2_orphaned" = 1 ]; then
|
|
364
583
|
signal_suite KILL
|
|
584
|
+
# The suite must really be gone, and this stays an ASSERTION rather than a warning:
|
|
585
|
+
# the work-dir check below passes vacuously against a suite that is still running,
|
|
586
|
+
# because a live suite has not reached its EXIT trap either. Demoting this to a
|
|
587
|
+
# diagnostic would leave that check reading "no cleanup ran" whenever the kill missed.
|
|
588
|
+
# What made the old version of this flaky was the zombie ambiguity inside
|
|
589
|
+
# await_suite_exit, not the assertion itself — that is fixed at the helper, so the
|
|
590
|
+
# obligation can be kept instead of traded away.
|
|
365
591
|
leg2_suite_gone=0; await_suite_exit && leg2_suite_gone=1
|
|
366
|
-
|
|
367
|
-
#
|
|
368
|
-
#
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
592
|
+
# Resume only now: the wrapper was held stopped across the abort so its countdown
|
|
593
|
+
# could not have completed before it, and everything observed from here happens with
|
|
594
|
+
# no reaper of any kind left alive.
|
|
595
|
+
leg2_resume_wrapper
|
|
596
|
+
check "leg2: the killed suite is gone" '[ "$leg2_suite_gone" = 1 ]'
|
|
597
|
+
|
|
598
|
+
# No trap ran — structurally, not by the clock. The suite's EXIT trap ends in
|
|
599
|
+
# `rm -rf "$WORK"`, so the work dir outliving a SIGKILLed suite is the trap's absence
|
|
600
|
+
# made visible. The old spelling asked whether the suite exited inside a 30s window,
|
|
601
|
+
# which is a fact about the host's scheduler rather than about whether cleanup ran.
|
|
602
|
+
# Read together with the assertion above, the pair says: the suite is gone AND it left
|
|
603
|
+
# its work dir behind — which only an untrapped death produces.
|
|
604
|
+
leg2_no_cleanup=0
|
|
605
|
+
[ "$leg2_suite_gone" = 1 ] && [ -n "$leg2_work" ] && [ -d "$leg2_work" ] && leg2_no_cleanup=1
|
|
606
|
+
check "leg2: no cleanup ran — the killed suite's work dir survives it" \
|
|
607
|
+
'[ "$leg2_no_cleanup" = 1 ]'
|
|
608
|
+
|
|
381
609
|
leg2_self_exit=0; wrapper_gone_within "$(( LEG2_HANG_BOUND + GRACE ))" && leg2_self_exit=1
|
|
382
610
|
[ "$leg2_self_exit" = 1 ] || diagnose_wrapper leg2
|
|
383
|
-
check "leg2: an unreaped wrapper
|
|
384
|
-
|
|
611
|
+
check "leg2: an unreaped wrapper leaves nothing behind" '[ "$leg2_self_exit" = 1 ]'
|
|
612
|
+
|
|
613
|
+
# WHY it ended, not WHEN it was last seen. The fixture writes this only on the far side
|
|
614
|
+
# of its bounded countdown, so present means the bound ended it and absent means a
|
|
615
|
+
# reaper did — the distinction the deleted "still alive right after the kill" check was
|
|
616
|
+
# trying to draw by looking at the process at one instant, which a corpse answers wrong.
|
|
617
|
+
# An unbounded fixture never reaches the write, so this still reds the reverted bound.
|
|
618
|
+
leg2_bound_marker="$(bound_marker_state)"
|
|
619
|
+
[ "$leg2_bound_marker" = present ] || diagnose_wrapper leg2-bound
|
|
620
|
+
check "leg2: the fixture's own lifetime bound is what ended the wrapper" \
|
|
621
|
+
'[ "$leg2_bound_marker" = present ]'
|
|
385
622
|
fi
|
|
386
623
|
|
|
387
624
|
fi
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md
CHANGED
|
@@ -40,8 +40,8 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
40
40
|
|
|
41
41
|
### Evidence, RCA, charter & attribution(证据 / RCA / charter / 出处核实)
|
|
42
42
|
|
|
43
|
-
- **Charter-before-editing red-line**:
|
|
44
|
-
-
|
|
43
|
+
- **Charter-before-editing red-line**: do not read sources or edit skills until the complete charter in `references/source-to-skill-extraction.md#extraction-charter` is filled cell-by-cell. A remembered field subset is not a charter. Step 0 owns the procedure; this rule owns the hard stop. Core-Rule canonicality governs same-facet drift only and never narrows that field set to this bullet.
|
|
44
|
+
- Classify every result — including task/session summaries and lessons-learned requests, not only failures — as **failure/correction**, **stable success**, or **unstable/insufficient evidence**. Failure runs RCA; stable success requires mechanism, non-luck evidence, reuse conditions, firing point, and owner; insufficient evidence stays an observation. Never invent a failure story to justify learning from success. The author classifies, so the class is not self-elective: correction/finding/failure-triggered work defaults to `failure/correction`, and only independent review may accept a relabel into an RCA-skipping class. Missing classification, unaccepted relabel, or missing analysis leaves it `interim`. Full method: `references/source-to-skill-extraction.md#result-learning-baseline-for-every-extraction`.
|
|
45
45
|
- **RCA must go wider than one 5-Why chain.** 5 Why is the entry technique to get past a symptom, but a single linear chain to one "root cause" is its documented failure mode: real process/agent failures need multiple concurrent causes, the stop point is arbitrary, and "why" drifts toward "who"/blame.
|
|
46
46
|
- For any non-trivial extraction, RCA must (a) **widen** — enumerate the multiple contributing factors across categories (trigger/routing, stale process-model, missing mechanical control, missing feedback, latent authored-earlier condition, detection gap) before deepening one — a straight chain with no branches means you stopped early; (b) **counterfactually test** each candidate to rank causal weight — *if removed or changed, would the failure still happen?* — keeping necessary/sufficient factors **and** failed redundant safeguards as secondary controls / defence-in-depth, and dropping only genuine coincidence (a one-trace factor is a hypothesis — mark it probabilistic, don't hard-drop); (c) frame prevention as a **mechanical control on the failure CLASS** — an enforced constraint plus the feedback that confirms it fired on a surface the next agent actually reaches in time (not a clause buried in a deep reference) — not agent diligence, preferring the highest-leverage **practical** control (a named artifact, an owner who can change it today, an observable check — never deleting a useful narrow gate or inflating one miss into an over-broad hook) over the first patchable point.
|
|
47
47
|
- Reject hindsight causes ("agent careless" / "need more attention" / "should have known"): ask why the action made sense given what the agent could see, not what it should have done. Full method, category prompts, stopping points, and sources: `references/source-to-skill-extraction.md` (Deep RCA For Extraction).
|
|
@@ -166,7 +166,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
166
166
|
|
|
167
167
|
### Validation & the dual-track gate(验证 / dual-track 门)
|
|
168
168
|
|
|
169
|
-
- Static validation is not extraction validation.
|
|
169
|
+
- Static validation is not extraction validation. Nontrivial closeout shows the matching result analysis, target-output and sibling decisions, landed diff, commands, and independent review/challenge; otherwise report `interim` even if static checks pass.
|
|
170
170
|
- For a whole-session/task-retrospective extraction over operational delivery that changed repositories, branches, MRs, pipelines, releases, or deployable artifacts, closeout validation must show one of: `delivery-state rows` with changed artifact, branch/worktree, remote/MR, CI/local verification, cancelled/retried, residual-risk, and next-action state; or `artifact/status axis: not-applicable` with the reason. Missing delivery-state evidence downgrades the extraction to `interim`; static validation and clean independent review do not close it.
|
|
171
171
|
- A rule that exists but did not trigger is a validation-gate defect, not proof that the workflow is adequate.
|
|
172
172
|
- For any correction where the missed step was covered by any rule in this workflow that a reasonable reader would apply to the scenario, add or tighten a closeout gate that would have blocked the exact premature final answer.
|
|
@@ -215,9 +215,9 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
215
215
|
|
|
216
216
|
## Extraction Workflow
|
|
217
217
|
|
|
218
|
-
0. Set the extraction charter and
|
|
218
|
+
0. Set the extraction charter and result-learning baseline.
|
|
219
219
|
- Hard stop: before source reads / before edits, record the charter from `references/source-to-skill-extraction.md#extraction-charter` — **open that table and fill it cell-by-cell; a charter written from memory of the field names is not a charter and must not be recorded as one.** Each field's real constraints live only in its own cell (Evidence plan's produced-artifact-first rule, Scope's watermark validation, RCA depth scaling), so a from-memory charter reproduces the field list and none of the gates, while looking complete. Trivial wording cleanup still records Depth explicitly.
|
|
220
|
-
-
|
|
220
|
+
- Result baseline: classify first, then use `#result-learning-baseline-for-every-extraction`; failure RCA depth comes from `#deep-rca-for-extraction`.
|
|
221
221
|
- Task/session incidents: use `#task-retrospective-extraction` and its delivery-chain RCA prompts before deciding whether the lesson lands in this workflow, a sibling skill, validator, memory, project artifact, or final-response only.
|
|
222
222
|
- Full/complete/deep asks require the source-register shape before reads: source groups, inclusion/exclusion, minimum artifact depth, owner skill, completion evidence, and batch-progress status tracking (`pending`/`read`/`deep-read`/`excluded`/`unavailable`/`routed`) from `#full-coverage-source-register-protocol`; close or explicitly downscope every required batch before saying "complete".
|
|
223
223
|
- Target-output map derives from Lifecycle impact: every affected stage gets an owner target or explicit no-update reason before source-derived editing starts (`#target-output-map`).
|
|
@@ -239,7 +239,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
239
239
|
- The map must include every plausible owner per affected lifecycle stage, including sibling skills in the same stage. If no skill owns a stage, write the no-output reason; do not silently omit the stage.
|
|
240
240
|
- For any extraction beyond wording-only cleanup, include the provenance-to-target diff shape before editing: source mechanism, provenance row, target file, executable landing, test or acceptance owner, and status.
|
|
241
241
|
- Trigger situations and users/tasks it should serve.
|
|
242
|
-
- What future failure
|
|
242
|
+
- What future failure or drift it should prevent, or which evidenced success mechanism it should preserve and reuse.
|
|
243
243
|
- For subjective or high-impact skills such as design, UX, frontend/client, product workflow, architecture, or review, define pressure scenarios and acceptance criteria before editing the skill.
|
|
244
244
|
- For UI/UX or client-facing skills, the pressure scenario must ask whether a person without source access can produce a good-looking and behaviorally sound screen: clear visual hierarchy, fitting density, risk-matched feedback, recoverable state transitions, responsive/device adaptation, and rendered acceptance evidence.
|
|
245
245
|
|
|
@@ -282,8 +282,8 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
282
282
|
|
|
283
283
|
6. Validate before landing.
|
|
284
284
|
- YAML frontmatter parses and description is trigger-focused.
|
|
285
|
-
- Extraction charter is satisfied: purpose, scope, depth,
|
|
286
|
-
-
|
|
285
|
+
- Extraction charter is satisfied: purpose, scope, depth, result classification and matching analysis, evidence plan, and completion standard are either met or explicitly downscoped.
|
|
286
|
+
- Result-learning gate: classification and matching analysis satisfy `#result-learning-baseline-for-every-extraction`; missing, mismatched, or insufficient-evidence-as-rule fails validation.
|
|
287
287
|
- Delivery-chain RCA gate: for incident or task-retrospective extraction, validation must show definition, implementation, verification, review/MR or release readiness, and retrospective-quality causes were checked or explicitly ruled out. If the extraction workflow itself missed the deeper cause, the workflow fix must be landed and validated before finalizing.
|
|
288
288
|
- Blocked-verification gate: any `unavailable`, `skipped`, or `blocked` verification claim must include remediation commands already attempted, observed result, residual risk, and next unblock action. If remediation was feasible but not attempted, validation fails and the test remains pending rather than unavailable.
|
|
289
289
|
- Blocked-source gate: any `unavailable`, `skipped`, `blocked`, timed-out, partial, or failed source-read claim must include the smaller/different read strategy already attempted, observed result, recovered evidence, residual gap, and next unblock action. If remediation was feasible but not attempted, validation fails and the source row remains pending.
|
|
@@ -50,13 +50,13 @@
|
|
|
50
50
|
- **防作弊**:runner 校验每 task 的 `frozen_at_sha` 是 HEAD 祖先(非祖先 = drift,排除出回归判定);同一改动若同时动 task-bank 和 SKILL.md description 会显式告警(防"改 skill 顺手改测试让它过")。
|
|
51
51
|
- `--baseline <json>` 给 diff 式报告(newly_failed / newly_passed)。
|
|
52
52
|
|
|
53
|
-
### Bank
|
|
53
|
+
### Bank 用例修复的测量协议(测量必做;落地裁决归本轮实际门禁)
|
|
54
54
|
|
|
55
|
-
|
|
55
|
+
以下是生成可比较证据的默认协议。**降级的是「F4 自己充当统一合并门禁」这个声称,不是「必须测、且必须有人裁决」这个义务**——这两件事分开:落地判断交给本轮实际的 owner/风险/评审门禁,但**测量本身不可选**。任何动 routing 面(SKILL.md description、task-bank 判定面)的改动都必须按下列协议产出证据;没跑就是没收敛,不得进入独立评审、也不得声称本轮无回归。采用不同样本量时,须随工件记录理由,且该理由与本轮证据一同进入独立评审——「记了理由」本身不是豁免,自审通过的理由不构成已裁决:
|
|
56
56
|
|
|
57
57
|
1. 动任何 description 之前必须先跑 **≥10 轮有效观测**的稳定性基线,把稳定失败与抖动分开;抖动不得作为修改依据(grader 超时/不可解析轮不算有效观测,须补跑)。
|
|
58
58
|
2. 改后通过数必须在**最终措辞**上重测:中间稿的通过数在措辞再变的那一刻作废,不得挪用到最终候选的证据里。
|
|
59
|
-
3.
|
|
59
|
+
3. 受影响邻居用例集默认改前/改后各 **≥3 轮**,集合须含期望 owner 自己的兄弟用例与高词面重叠的他 owner 用例;邻居回归作为独立 finding 交由本轮实际门禁处置——**该 finding 须以 blocking 记入本轮 dual-track 评审记录,且只能由独立评审方豁免,不能由实现者自行判定「本轮没有门禁采用这组证据」而放行**。降级的是「F4 自己充当合并门禁」这一声称,不是「回归必须被人裁决」这一义务;后者若也随之消失,这一条就只剩被裁决方自审。
|
|
60
60
|
4. 每轮判决必须连同 **runner 调用、grader 模型身份、候选身份**(commit 或描述内容指纹)与**原始逐轮工件的持久定位符**一并记入轮记录;没有定位符的通过数只能标注为 operator-reported,不得据以宣称修复轮已 concluded。
|
|
61
61
|
|
|
62
62
|
## Tier-3:hub golden trace 真 agent 回放(已落地,advisory,人工判定)
|
|
@@ -71,16 +71,15 @@
|
|
|
71
71
|
- **随机性**:agent 非确定;判定先人工、nightly 起步,有稳定史前不自动 gate。防作弊同 T2(`frozen_at_sha` 祖先校验)。
|
|
72
72
|
- **双用途**:除回归外,Tier-3 还可当**改技能前的 RED-baseline**(改前手动跑触发场景看真 agent 是否真路由错,改后看 compliance)——可选;只有真观察到 miss 才算 RED(PASS/INCONCLUSIVE 不算),小 N + 非确定有噪声,手动跑两次自己留两份报告。落地 + 防作弊注意见 [validation-and-landing.md](validation-and-landing.md) "Optional real-agent RED-baseline"。
|
|
73
73
|
|
|
74
|
-
## Health roll-up
|
|
74
|
+
## Health roll-up:描述性仪表盘(已落地,advisory)
|
|
75
75
|
|
|
76
|
-
把上面各信号卷成**一个加权 0–10
|
|
76
|
+
把上面各信号卷成**一个加权 0–10 显示值 + 同尺子变化**,用于定位值得继续检查的维度。它借用 OpenSSF Scorecard 的呈现形态,但不把不同性质的 F4 信号变成“仓库整体变好/变差”的总判决。映射与限制见 [harness-patterns-and-eval.md](harness-patterns-and-eval.md) §3.4。
|
|
77
77
|
|
|
78
78
|
`scripts/eval-health.rb <repo-root> [--trace-json p] [--bank-json p] [--history p] [--no-write] [--json p] [--quiet]`。`make eval-health`。
|
|
79
79
|
|
|
80
80
|
- **维度(各 0–10,按风险加权,OpenSSF 风格)**:`structural` 权重 10(Critical,`validate-skill.sh` pass/fail)· `routing_static` 权重 10(Critical,T1 blocking=0 满分、有 blocking 砸到 3、advisory 轻罚)· `trace` 权重 7.5(High,T3 pass/considered)· `bank` 权重 5(Medium,T2 pass/tasks)。
|
|
81
81
|
- **只跑确定性两维**(structural + routing_static,无 LLM、快);`trace`/`bank` 需 `claude` 且随机,**不自动跑** —— 用 `--trace-json` / `--bank-json` 把 T3/T2 报告喂进来,否则该维 **skip,权重按比例重分**给在场维(诚实标注 `dims=…`)。
|
|
82
|
-
- `composite = Σ(score_i·w_i) / Σ(w_i)`,只对在场维求和;skip 维自动从分母剔除。band
|
|
83
|
-
- **advisory
|
|
82
|
+
- `composite = Σ(score_i·w_i) / Σ(w_i)`,只对在场维求和;skip 维自动从分母剔除。band 只是浏览提示,不得当作验收等级。
|
|
83
|
+
- **advisory,永不因显示值阻断**:退出码 `0` = 跑完 **或** 没有可算的维(都不是失败);`2` = 用法/setup 错(`<repo-root>` 不对或无 `skills/` 目录)。低值、坏报告、空历史都不会非 0。畸形的 T2/T3 报告(非对象 / 非整数 / `pass>total`)直接 **skip,不静默打分**。**绝不接进 `check-ccl-skills.sh`**。结构校验、T1 blocking、任务验收和行为证据分别独立判定;质量失败不能被其它维度的高值平均掉。
|
|
84
84
|
- **corpus/version 守卫(防跨变更 task-bank 当稳定指标比)**:每条历史记 `corpus`(= task-bank + golden-traces 输入内容的指纹)+ `repo_sha` + 在场 `dims`。**趋势 delta(IMPROVING/DECLINING)只跟最近一条 `(corpus, dims)` 都相同的历史比**;否则打印"baseline reset, not compared",不偷偷比。这样"加了 10 条简单 task → 分涨了"不会被读成真进步(尺子换了)。
|
|
85
85
|
- **历史文件**:默认 `eval/health-history.jsonl`,**git-ignore**(同 Goodhart 理由:committed 的数会招"调数不修仓";也免 append 把树搞脏)。`--history` 可改路径,`--no-write` 不落盘。
|
|
86
|
-
- **基线**(本机、本 corpus):确定性两维 = `9.5/10`(structural 10 + routing_static 9,advisory=1)。
|
|
@@ -6,7 +6,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
6
6
|
|
|
7
7
|
```
|
|
8
8
|
0. Charter → ~/.<host>/skills/.extraction-work/<project>-charter.md
|
|
9
|
-
Purpose / Scope / Depth /
|
|
9
|
+
Purpose / Scope / Depth / Result baseline / Open questions
|
|
10
10
|
│
|
|
11
11
|
1. Source register → ~/.<host>/skills/.extraction-work/<project>-source-register.md
|
|
12
12
|
Each row: source id / class / status / target skill / extracted mechanisms
|
|
@@ -43,7 +43,8 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
|
|
|
43
43
|
|
|
44
44
|
- Owner: maintainer
|
|
45
45
|
- Location: `~/.<host>/skills/.extraction-work/<project>-charter.md`
|
|
46
|
-
- Required fields: Purpose, Scope, Depth,
|
|
46
|
+
- Required fields: Purpose, Scope, Depth, Result classification, matching analysis, Failure modes or success-reuse conditions, Lifecycle impact, Evidence plan, Completion standard.
|
|
47
|
+
- Result classification: failure/correction → Deep RCA (widen → counterfactual-test → control); stable success → mechanism + non-luck evidence + reuse conditions + firing point + owner; unstable/insufficient evidence → observation only.
|
|
47
48
|
- Full structure: `SKILL.md` Core Workflow Step 0.
|
|
48
49
|
- Output: a file the maintainer can re-read in 3 months and understand what they were trying to do.
|
|
49
50
|
|
|
@@ -147,13 +147,13 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
|
|
|
147
147
|
|
|
148
148
|
**用时机**:操作层月级别稳定 + 要做一次 system-wide 换层/瘦身时;把它当"换层前的对抗性回归证据",不是频繁迭代期的日常闸。
|
|
149
149
|
|
|
150
|
-
### 3.4
|
|
150
|
+
### 3.4 Health signal dashboard(描述性 roll-up)
|
|
151
151
|
|
|
152
|
-
3.1–3.3
|
|
152
|
+
3.1–3.3 都保留各自的判定对象;跨时间还需要一个快速入口,显示本次有哪些信号在场、哪些维度变化,方便继续下钻。
|
|
153
153
|
|
|
154
|
-
**外部锚**:OpenSSF Scorecard
|
|
154
|
+
**外部锚**:OpenSSF Scorecard 用每个 check 0–10、风险加权聚合和历史变化展示代码仓信号。这里仅借它的**展示形态**,不继承“一个总分代表整体健康”的解释。
|
|
155
155
|
|
|
156
|
-
**我们的映射(route-not-copy
|
|
156
|
+
**我们的映射(route-not-copy)**:`eval-health.rb` 展示四个信号,并保留一个兼容既有实现的加权 0–10 值与同尺子变化 ——
|
|
157
157
|
|
|
158
158
|
| 维度 | 风险/权重 | 0–10 来源 |
|
|
159
159
|
|---|---|---|
|
|
@@ -164,12 +164,13 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
|
|
|
164
164
|
|
|
165
165
|
`composite = Σ(score·weight) / Σ(weight)`,只算在场维(skip 维权重重分,同 OpenSSF/gstack)。确定性两维自动跑,T2/T3 喂报告进来(否则 skip)。契约见 [eval-routing.md](eval-routing.md) 的 Health roll-up 节。
|
|
166
166
|
|
|
167
|
-
|
|
167
|
+
**三条不可省的护栏**:
|
|
168
168
|
|
|
169
169
|
1. **advisory,不当 gate**(Goodhart):度量变成 target 就被博弈。综合分**永不接门禁**,二元门禁(结构 + T1 blocking)仍独立挡 merge;历史文件 git-ignore,免得"committed 的数"招人调数不修仓。
|
|
170
170
|
2. **corpus/version 守卫**:T2/T3 的 task-bank、golden-traces **本身会变**。加 10 条简单 task,pass-rate 涨了但仓没变好 —— **尺子换了**。每条历史记 `corpus` 指纹(输入内容 hash)+ 在场 `dims`;**趋势只跟 `(corpus, dims)` 全同的历史比**,否则 baseline reset 不偷偷比。等价于 gstack-health "尺子变了就从新基线重新追"。
|
|
171
|
+
3. **不平均掉质量失败**:结构、路由、真实回放和任务结果的语义不同;任何质量或安全阻断仍由自己的门禁决定,不能被其它维度的高值抵消。
|
|
171
172
|
|
|
172
|
-
|
|
173
|
+
**用**:周期性看一眼信号变化并下钻。**不用**:别据此单独声称仓库整体变好/变差,别把它当通过标准、别 committed、别跨 corpus 硬比。
|
|
173
174
|
|
|
174
175
|
---
|
|
175
176
|
|
|
@@ -297,6 +297,24 @@ A carve-out is a contract, not a category badge. Without the evidence row, fall
|
|
|
297
297
|
|
|
298
298
|
**Occurrences that promoted it**: `scripts/owner-dispatch/owner-dispatch.sh` (state keyed per-worktree → boundary record written in a worktree unreadable from the primary checkout) and `hooks/guard-edit-isolation.sh` (path-name match → the repo's only edit-time hard-deny gate silently disabled for checkouts under a `worktrees/`-named path).
|
|
299
299
|
|
|
300
|
+
## Anti-pattern 28 — Process liveness decided by an EXISTENCE test that a corpse answers
|
|
301
|
+
|
|
302
|
+
**Symptom**: a probe or fixture concludes "this process is still alive" — or "it is a live orphan" — from a test that only proves the pid is still in the process table: `kill -0 "$pid"`, or reading `ps -o ppid=` and comparing it against init's pid. The same suite's verdict scans then exclude zombies (`$stat !~ /Z/`) when deciding the process is *gone*.
|
|
303
|
+
|
|
304
|
+
**Why bad**: an exited-but-unreaped process is still in the table. It answers `kill -0`, and its ppid still reads — as `1` once it is reparented. So one process state gets two answers: the precondition check says "alive, scenario built", the verdict scan says "gone, nothing leaked", and an assertion between them that samples liveness at one instant says "already dead" and fails. The red then names the code under test for something it did not do, and the failure is intermittent because it depends on when the OS reaps. Worse, the precondition passing means the probe goes on to assert about a scenario it never actually built.
|
|
305
|
+
|
|
306
|
+
**Fix**: consult process **state**, not existence.
|
|
307
|
+
- One vocabulary for the whole probe — a helper returning `live` / `zombie` / `absent` (`ps -o stat=`; empty → absent, `*Z*` → zombie, else live) — used by *every* liveness question, so the checks cannot answer the same pid differently.
|
|
308
|
+
- Preconditions require `live`. A corpse is not an orphan; accepting one is how a probe goes on to assert about a scenario that was never constructed.
|
|
309
|
+
- Better still, stop asking the process at all: assert on an artifact the run leaves behind — a work dir the cleanup path would have deleted, a marker the fixture writes only on the code path under test. That answers *why* a process ended, which no liveness sample can.
|
|
310
|
+
- Never assert liveness by an instantaneous sample; reap lag on a loaded runner needs a bounded grace period. (`testing-strategy/references/ci-fixtures-and-flake-control.md` owns that rule — this row is its firing path.)
|
|
311
|
+
|
|
312
|
+
**Grep**: `find . -name 'test_*.sh' -o -name 'test.sh' | xargs grep -nE 'ps[[:space:]]+-o[[:space:]]+ppid=.*=[[:space:]]*"?1"?[[:space:]]*\]'` (non-comment hits with no `stat=` / `*_state` consult within two lines). Machine-enforced by `check-ccl-skills.sh` (`liveness_predicate_scan`). Scope is test scripts only: outside them a `kill -0` before signalling asks about existence, which is the right question there. Five things are deliberately **not** caught and stay human/challenge checks — `while kill -0 "$pid"` watchdog loops (two exist here; both wait on a direct child and `wait` for it immediately after, so the shell reaps it and the loop ends), a bare `kill -0` liveness branch inside a loop body, any dynamic spelling, a state-helper mention in a **trailing** comment (whole-line comments are dropped, but telling a trailing `#` from `${var#prefix}` needs a shell parser — the same call Anti-pattern 27 makes), and a process-state read — direct `ps -o stat=` or via a helper — that inspects a *different* pid than the oracle tests, or whose result never reaches the verdict. Both are the same irreducible gap: proving the read actually governs the decision needs dataflow over a parsed shell, not a text window, so this is the limit the mechanical gate stops at by design. Both are **exercised** by `test_liveness_predicate_gate.sh` (P13/P14) rather than only described here — they assert the gate does not fire, so tightening it turns them red and forces this paragraph to be updated instead of quietly going stale.
|
|
313
|
+
|
|
314
|
+
The waiver is a **proxy** for "this site consults process state", not an invariant, and five successive review rounds each found it too loose in a different way — a comment naming the helper, an unrelated `*_state` token, a helper that never reads state, a bare `stat=` assignment, and a hollow helper borrowing an unrelated read elsewhere in the file. Tying a call to its definition needs a shell parser, so the honest disposition is a narrow predicate with these limits stated rather than another round of widening. The mechanical gate is the deterministic catch for the exact recurring shape; this checklist and the adversarial challenge remain the comprehensive net.
|
|
315
|
+
|
|
316
|
+
**Occurrences that promoted it**: the code-review abort-leak probe red CI three times, each on a different leg-2 assertion, each asserting the probe's environment rather than the suite — the reparent check accepted a zombie, the "still alive right after the kill" check rejected that same zombie, and the verdict scan reported it gone. The rule forbidding this already existed in `testing-strategy`, and the suite's own hang cases followed it while the probe did not: the gap was enforcement, not content.
|
|
317
|
+
|
|
300
318
|
## How to use this checklist
|
|
301
319
|
|
|
302
320
|
Before committing any skill or reference change touching operational/architectural rules:
|