@ccoalm/ccl-skills 0.3.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +34 -3
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_abort_leak_state_helpers.sh +148 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +28 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +7 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +13 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +271 -34
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +8 -8
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +7 -8
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +3 -2
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +7 -6
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +18 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +29 -2
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +28 -29
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +153 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +23 -10
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +303 -1
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +222 -91
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +6 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +6 -1
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +1 -1
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +12 -12
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +455 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +20 -20
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_liveness_predicate_gate.sh +288 -0
  25. package/dist/assets/release.json +39 -24
  26. package/package.json +1 -1
@@ -13,6 +13,14 @@
13
13
  # leg 2 SIGKILL the suite: nothing can trap that, so no reaper runs at all and the
14
14
  # wrapper must die on the fixture's own lifetime bound instead.
15
15
  #
16
+ # Every verdict below is a fact the run leaves behind — the work dir the EXIT trap would
17
+ # have deleted, the marker the fixture writes only past its countdown — never an
18
+ # observation of what was true at the instant the probe looked. Liveness questions all go
19
+ # through wrapper_state() so they cannot answer the same process differently, and a
20
+ # scenario the probe fails to BUILD is retried rather than reported as a failed assertion.
21
+ # All three of those rules are here because the earlier spellings red this gate on CI
22
+ # while the suite was behaving correctly.
23
+ #
16
24
  # Leg 2 exists because leg 1 alone stays green if the fixture's bound were restored to
17
25
  # an unbounded loop: the trap would still reap it. It was an unbounded loop that turned
18
26
  # this defect into a wrapper observed alive for 39 hours with ppid=1, ignoring SIGTERM.
@@ -59,10 +67,48 @@ esac
59
67
  # with the probe still green, so CI points its two jobs at different stubs.
60
68
  ABORT_LEAK_PROBE_CLIENT="${ABORT_LEAK_PROBE_CLIENT:-claude}"
61
69
  case "$ABORT_LEAK_PROBE_CLIENT" in
62
- claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh ;;
63
- fallback) PROBE_BEHAVIOR_FILE=kimi_behavior; PROBE_WRAPPER=kimi_review.sh ;;
70
+ claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh
71
+ PROBE_BOUND_MARKER=claude_hang_bound_reached ;;
72
+ fallback) PROBE_BEHAVIOR_FILE=kimi_behavior; PROBE_WRAPPER=kimi_review.sh
73
+ PROBE_BOUND_MARKER=kimi_hang_bound_reached ;;
64
74
  *) echo "ABORT_LEAK_PROBE_CLIENT must be claude or fallback" >&2; exit 2 ;;
65
75
  esac
76
+ # How many times a leg may rebuild its scenario before giving up. Only the SETUP is
77
+ # retried: constructing "an orphaned wrapper with no reaper left" depends on the probe
78
+ # winning races against the controller and the host, and losing one says nothing about
79
+ # the suite. An assertion ABOUT the suite is never retried, so a fixture whose bound was
80
+ # reverted fails on every attempt and cannot be retried into green.
81
+ SETUP_ATTEMPTS="${ABORT_LEAK_PROBE_SETUP_ATTEMPTS:-3}"
82
+
83
+ # ONE process-state vocabulary for the whole probe. Three separate spellings of "is it
84
+ # there?" used to disagree about the same pid: `kill -0` succeeds for a zombie, reading
85
+ # ppid succeeds for a zombie, and the two verdict scans exclude zombies. An already-dead
86
+ # wrapper awaiting reaping therefore satisfied "reparented to init", failed "still alive",
87
+ # and satisfied "gone" — one process state, three answers, and a red that named the suite
88
+ # for something the suite had not done. Every liveness question below goes through here.
89
+ wrapper_state() {
90
+ local st rc
91
+ case "${1:-}" in ''|*[!0-9]*) printf 'absent\n'; return 0 ;; esac
92
+ st="$(ps -o stat= -p "$1" 2>/dev/null)"; rc=$?
93
+ # `ps` exits 1 for "no such process", which is a real answer. Anything above that is
94
+ # the TOOL failing, not the process being gone — and swallowing that into `absent`
95
+ # would report a suite exited because `ps` could not run. Unknown is reported as its
96
+ # own state and every consumer treats it as still present, because the expensive
97
+ # direction of this error is declaring something gone that is not.
98
+ if [ "$rc" -gt 1 ]; then printf 'unknown\n'; return 0; fi
99
+ case "$st" in
100
+ '') printf 'absent\n' ;;
101
+ *Z*) printf 'zombie\n' ;;
102
+ # A stopped process is neither running nor gone, and leg 2 deliberately puts the
103
+ # wrapper in this state while it arms. Folding it into `live` would let the arming
104
+ # step report success without SIGSTOP having actually landed.
105
+ # `T` (job-control stop) or `t` (tracing stop), anywhere in the field: neither letter
106
+ # appears among the flag characters either ps appends, so this cannot catch a
107
+ # running process by accident.
108
+ *[Tt]*) printf 'stopped\n' ;;
109
+ *) printf 'live\n' ;;
110
+ esac
111
+ }
66
112
 
67
113
  fails=0
68
114
  check() {
@@ -178,6 +224,21 @@ suite_work_dir() {
178
224
  done
179
225
  return 1
180
226
  }
227
+ # The fixture writes this file only after its bounded countdown runs to completion, so
228
+ # its presence says the fixture's OWN lifetime bound ended the wrapper and its absence
229
+ # says something else did. That is the fact leg 2 needs, and it is a fact about which
230
+ # code path ran — not about what was true at the instant the probe happened to look.
231
+ bound_marker_path() {
232
+ local work
233
+ work="$(suite_work_dir)" || return 1
234
+ printf '%s\n' "$work/state/$PROBE_BOUND_MARKER"
235
+ }
236
+ bound_marker_state() {
237
+ local path
238
+ path="$(bound_marker_path)" || { printf 'no-work-dir\n'; return 0; }
239
+ if [ -e "$path" ]; then printf 'present\n'; else printf 'absent\n'; fi
240
+ }
241
+
181
242
  live_wrapper() {
182
243
  local work behavior pid parent_cmd
183
244
  work="$(suite_work_dir)" || return 0
@@ -203,7 +264,10 @@ orphan_target_wrapper() {
203
264
  local hang_bound="$1" candidate controller_pid controller_ok deadline suite_died
204
265
  PROBE_TMP="$(mktemp -d "${TMPDIR:-/tmp}/review-gate-abort-probe.XXXXXX")" || return 1
205
266
  PROBE_TMP_REAL="$(cd "$PROBE_TMP" && pwd -P)" || return 1
267
+ # Both, not just the pid: a retry that left the previous attempt's group id in place
268
+ # would have the verdict scan looking for members of a group this attempt never built.
206
269
  wrapper_pid=""
270
+ wrapper_pgid=""
207
271
 
208
272
  # Job control so the suite gets its own process group and the probe can signal that
209
273
  # group without signalling itself.
@@ -219,7 +283,11 @@ orphan_target_wrapper() {
219
283
  while [ "$(date +%s)" -lt "$deadline" ]; do
220
284
  candidate="$(live_wrapper)"
221
285
  if [ -n "$candidate" ]; then wrapper_pid="$candidate"; break; fi
222
- kill -0 "$suite_pid" 2>/dev/null || { suite_died=1; break; }
286
+ # Third site of the same class: a suite that has exited but whose corpse this shell
287
+ # has not reaped still answers `kill -0`, so the bare existence test would keep this
288
+ # loop polling for a dead suite until the whole reach budget elapsed and then report
289
+ # the wrong one of the two failures below.
290
+ [ "$(wrapper_state "$suite_pid")" = live ] || { suite_died=1; break; }
223
291
  sleep 0.1
224
292
  done
225
293
  if [ -z "$wrapper_pid" ]; then
@@ -281,29 +349,53 @@ orphan_target_wrapper() {
281
349
  fi
282
350
  sleep 0.5
283
351
  done
352
+ # Its own deadline: this loop used to reuse the one the controller wait had already
353
+ # been counting down, so a controller that took nine of those ten seconds to die left
354
+ # the reparent check one second to succeed in.
355
+ deadline=$(( $(date +%s) + 10 ))
284
356
  while :; do
285
- [ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "1" ] && return 0
286
- kill -0 "$wrapper_pid" 2>/dev/null || {
287
- echo "abort leak setup: wrapper $wrapper_pid died with its controller" >&2
288
- return 1
289
- }
357
+ # A corpse is not an orphan. The old spelling of this check read ppid, which a zombie
358
+ # answers, so a wrapper some reaper had already killed passed for a live orphan and
359
+ # the leg went on to make assertions about a scenario it had never built.
360
+ if [ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "1" ] &&
361
+ [ "$(wrapper_state "$wrapper_pid")" = live ]; then
362
+ return 0
363
+ fi
364
+ case "$(wrapper_state "$wrapper_pid")" in
365
+ live) : ;;
366
+ *)
367
+ printf 'abort_leak_setup_lost: wrapper %s was already %s before the abort — a reaper reached it first (bound marker: %s)\n' \
368
+ "$wrapper_pid" "$(wrapper_state "$wrapper_pid")" "$(bound_marker_state)" >&2
369
+ return 1
370
+ ;;
371
+ esac
290
372
  if [ "$(date +%s)" -ge "$deadline" ]; then
291
- echo "abort leak setup: wrapper $wrapper_pid was never reparented" >&2
373
+ echo "abort_leak_setup_lost: wrapper $wrapper_pid was never reparented" >&2
292
374
  return 1
293
375
  fi
294
376
  sleep 0.5
295
377
  done
296
378
  }
297
379
 
380
+ # Same one vocabulary as the wrapper: a killed suite whose corpse the probe's own shell
381
+ # has not reaped yet still answers `kill -0`, so spelling this as `kill -0` made a bookkeeping
382
+ # lag inside the probe look like a suite that refused to die.
298
383
  await_suite_exit() {
299
384
  local deadline
300
385
  deadline=$(( $(date +%s) + GRACE ))
301
- while kill -0 "$suite_pid" 2>/dev/null; do
386
+ # GONE is `absent` or `zombie` — a corpse cannot run cleanup — and everything else,
387
+ # `stopped` included, is still present. Testing `!= live` instead would call a STOPPED
388
+ # suite exited: the work dir would still be there, the wrapper would still reach its
389
+ # bound, and every leg-2 assertion could pass while the suite was never killed at all.
390
+ # Introducing a third process state without revisiting the checks that had two is the
391
+ # same defect this round is about, one file over.
392
+ while :; do
393
+ case "$(wrapper_state "$suite_pid")" in absent|zombie) return 0 ;; esac
302
394
  [ "$(date +%s)" -ge "$deadline" ] && break
303
395
  sleep 0.5
304
396
  done
305
- kill -0 "$suite_pid" 2>/dev/null && return 1
306
- return 0
397
+ case "$(wrapper_state "$suite_pid")" in absent|zombie) return 0 ;; esac
398
+ return 1
307
399
  }
308
400
 
309
401
  # The verdict is about the GROUP, not the leader. The wrapper backgrounds a TERM-immune
@@ -316,6 +408,15 @@ wrapper_group_members_alive() {
316
408
  ps -eo pid=,pgid=,stat= 2>/dev/null |
317
409
  awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ { print $1 }'
318
410
  }
411
+ # Group members still RUNNING after a group stop. A group signal is not an atomic
412
+ # transition: the leader can report stopped while a sibling has not been scheduled to
413
+ # handle it yet, and for the claude stub that sibling is the countdown itself. Zombies
414
+ # and stopped members are excluded; anything left is still able to reach the marker write.
415
+ wrapper_group_unstopped() {
416
+ [ -n "${wrapper_pgid:-}" ] || return 0
417
+ ps -eo pid=,pgid=,stat= 2>/dev/null |
418
+ awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ && $3 !~ /[Tt]/ { print $1 }'
419
+ }
319
420
  wrapper_gone_within() {
320
421
  local deadline members residue
321
422
  deadline=$(( $(date +%s) + $1 ))
@@ -335,15 +436,39 @@ wrapper_gone_within() {
335
436
  }
336
437
 
337
438
  diagnose_wrapper() {
338
- printf 'abort leak diagnostic (%s): wrapper=%s %s\n' "$1" "$wrapper_pid" \
439
+ printf 'abort leak diagnostic (%s): wrapper=%s state=%s bound_marker=%s ps=[%s]\n' \
440
+ "$1" "$wrapper_pid" "$(wrapper_state "$wrapper_pid")" "$(bound_marker_state)" \
339
441
  "$(ps -o pid=,ppid=,etime=,stat= -p "$wrapper_pid" 2>/dev/null | tr -s ' ')" >&2
340
442
  }
341
443
 
444
+ # Build the scenario, retrying only the CONSTRUCTION of it. Losing a race to some reaper
445
+ # leaves the probe with nothing to assert about, and reporting that as a failed assertion
446
+ # is what made this gate red at random on three separate assertions — each of them a
447
+ # statement about the probe's environment wearing the wording of a statement about the
448
+ # suite. A lost setup is retried from a fresh suite run; only running out of attempts is
449
+ # a failure, and it says so in those words.
450
+ # $3, when given, is a function run after a successful orphan that must also succeed for
451
+ # the attempt to count — the place for any arming step whose failure means the scenario
452
+ # was lost rather than the suite misbehaved.
453
+ orphan_with_retry() {
454
+ local hang_bound="$1" leg="$2" arm="${3:-}" attempt=1
455
+ while [ "$attempt" -le "$SETUP_ATTEMPTS" ]; do
456
+ if orphan_target_wrapper "$hang_bound" && { [ -z "$arm" ] || "$arm"; }; then return 0; fi
457
+ printf '%s: setup attempt %s/%s did not build the scenario; rebuilding\n' \
458
+ "$leg" "$attempt" "$SETUP_ATTEMPTS" >&2
459
+ cleanup; PROBE_TMP=""; PROBE_TMP_REAL=""; suite_pgid=""; suite_pid=""; suite_start=""
460
+ attempt=$((attempt+1))
461
+ done
462
+ printf '%s: could not build the scenario in %s attempts\n' "$leg" "$SETUP_ATTEMPTS" >&2
463
+ return 1
464
+ }
465
+
342
466
  # ---- leg 1: a trappable abort — the suite's own cleanup must reap the wrapper --------
343
467
  if [ "$ABORT_LEAK_PROBE_LEG" = 1 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
344
468
  leg1_orphaned=0
345
- orphan_target_wrapper "" && leg1_orphaned=1
346
- check "leg1: the controller was removed and the wrapper reparented to init" '[ "$leg1_orphaned" = 1 ]'
469
+ orphan_with_retry "" leg1 && leg1_orphaned=1
470
+ check "leg1: the scenario was built — controller removed, live wrapper reparented to init" \
471
+ '[ "$leg1_orphaned" = 1 ]'
347
472
  if [ "$leg1_orphaned" = 1 ]; then
348
473
  signal_suite TERM
349
474
  leg1_suite_gone=0; await_suite_exit && leg1_suite_gone=1
@@ -357,31 +482,143 @@ fi
357
482
 
358
483
  # ---- leg 2: an untrappable abort — only the fixture's own bound can end the wrapper --
359
484
  if [ "$ABORT_LEAK_PROBE_LEG" = 2 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
485
+ leg2_work=""
486
+ # Arm the leg, in this order: prove the wrapper is live, freeze its group, then read the
487
+ # marker. Freezing first is what makes the read meaningful — while the group is stopped
488
+ # nothing can write that file, so "absent" is a stable fact rather than a sample taken
489
+ # between two racing events. Absent means the bound has not fired, which is exactly the
490
+ # precondition leg 2 needs; present means it already fired and the scenario is rebuilt.
491
+ # The ordering this establishes — controller gone, bound not yet fired, abort next — is
492
+ # ENFORCED rather than observed, and needs no clock and no timestamp comparison, which
493
+ # second-granularity stamps could not have given honestly anyway.
494
+ leg2_arm() {
495
+ leg2_work="$(suite_work_dir || true)"
496
+ [ -n "$leg2_work" ] && [ -d "$leg2_work/state" ] || {
497
+ echo "abort_leak_setup_lost: leg2 found no suite work dir to arm against" >&2
498
+ return 1
499
+ }
500
+ [ "$(wrapper_state "$wrapper_pid")" = live ] || {
501
+ printf 'abort_leak_setup_lost: wrapper %s was %s before arming\n' \
502
+ "$wrapper_pid" "$(wrapper_state "$wrapper_pid")" >&2
503
+ return 1
504
+ }
505
+ # STOP the wrapper for the length of the abort. Checking "is it still live" and then
506
+ # killing the suite is a race no shell can close: between the two the countdown can
507
+ # finish, write the marker, and exit, after which every assertion passes while the
508
+ # untrappable-abort path never ran. A stopped process cannot reach the write at all,
509
+ # so the ordering stops being something to observe and becomes something enforced —
510
+ # which is the whole point of this round. The group, not the leader: the claude stub
511
+ # runs its countdown in a backgrounded child that shares the wrapper's group, and
512
+ # stopping only the leader would leave that child counting.
513
+ # Ownership before signalling, proved the way `probe_processes` proves it: re-match the
514
+ # pid against the live table by this run's private path, then require it to LEAD the
515
+ # group being signalled. Those two together mean the group is the validated wrapper's
516
+ # own, which is what makes a group signal safe here.
517
+ # (`probe_group_is_ours` is not used for this: measured on macOS it answers no for a
518
+ # group whose leader `probe_processes` matches by the same private path in the same
519
+ # instant. That is pre-existing and only makes `cleanup` fall back to per-pid kills —
520
+ # which still reap, since every member carries the path — so it is reported rather
521
+ # than changed under this round's scope.)
522
+ probe_processes | grep -qx "$wrapper_pid" || {
523
+ echo "abort_leak_setup_lost: wrapper $wrapper_pid no longer matches this run's private path" >&2
524
+ return 1
525
+ }
526
+ [ "$wrapper_pgid" = "$wrapper_pid" ] || {
527
+ printf 'abort_leak_setup_lost: wrapper %s does not lead group %s; refusing to stop the group\n' \
528
+ "$wrapper_pid" "$wrapper_pgid" >&2
529
+ return 1
530
+ }
531
+ kill -STOP -"$wrapper_pgid" 2>/dev/null || true
532
+ # Wait for the WHOLE group, not just the leader. Reading only the leader is the same
533
+ # proxy-for-the-condition mistake this round exists to remove: the leader can be stopped
534
+ # while the countdown sibling is still running and still able to write the marker.
535
+ leg2_stop_deadline=$(( $(date +%s) + 5 ))
536
+ while :; do
537
+ leg2_unstopped="$(wrapper_group_unstopped | tr '\n' ' ')"
538
+ case "$leg2_unstopped" in *[0-9]*) : ;; *) break ;; esac
539
+ if [ "$(date +%s)" -ge "$leg2_stop_deadline" ]; then
540
+ printf 'abort_leak_setup_lost: group %s still running after SIGSTOP: %s\n' \
541
+ "$wrapper_pgid" "$leg2_unstopped" >&2
542
+ kill -CONT -"$wrapper_pgid" 2>/dev/null || true
543
+ return 1
544
+ fi
545
+ sleep 0.2
546
+ done
547
+ [ "$(wrapper_state "$wrapper_pid")" = stopped ] || {
548
+ printf 'abort_leak_setup_lost: wrapper %s did not stop (state %s)\n' \
549
+ "$wrapper_pid" "$(wrapper_state "$wrapper_pid")" >&2
550
+ kill -CONT -"$wrapper_pgid" 2>/dev/null || true
551
+ return 1
552
+ }
553
+ # Observe the barrier; never manufacture it. Deleting the marker here destroyed the one
554
+ # piece of evidence that distinguishes the two cases: for the claude stub the countdown
555
+ # runs in a child, so the child can finish and write the marker while the wrapper is
556
+ # still unwinding and therefore still reads live. Removing it then made a scenario that
557
+ # was already lost — the bound fired BEFORE the abort — look like a clean run whose
558
+ # marker never came back, i.e. a false RED against a suite that behaved correctly.
559
+ # Read after the freeze, when nothing can be writing that file: absent means the bound
560
+ # has not fired yet, which is the precondition. Present means it already has, so the
561
+ # scenario is lost and gets rebuilt. The suite wipes state files at each case start, so
562
+ # a marker here belongs to this case, and every retry gets a fresh private tmp anyway.
563
+ [ "$(bound_marker_state)" = absent ] || {
564
+ printf 'abort_leak_setup_lost: %s already present — the bound fired before the abort\n' \
565
+ "$PROBE_BOUND_MARKER" >&2
566
+ kill -CONT -"$wrapper_pgid" 2>/dev/null || true
567
+ return 1
568
+ }
569
+ return 0
570
+ }
571
+ # Resume the wrapper once the suite is gone. From here its countdown runs with no
572
+ # controller, no suite, and no trap anywhere — exactly the state leg 2 is about.
573
+ leg2_resume_wrapper() {
574
+ [ -n "${wrapper_pgid:-}" ] && [ "$wrapper_pgid" = "${wrapper_pid:-}" ] &&
575
+ { kill -CONT -"$wrapper_pgid" 2>/dev/null || true; }
576
+ return 0
577
+ }
360
578
  leg2_orphaned=0
361
- orphan_target_wrapper "$LEG2_HANG_BOUND" && leg2_orphaned=1
362
- check "leg2: the controller was removed and the wrapper reparented to init" '[ "$leg2_orphaned" = 1 ]'
579
+ orphan_with_retry "$LEG2_HANG_BOUND" leg2 leg2_arm && leg2_orphaned=1
580
+ check "leg2: the scenario was built — controller removed, live wrapper reparented to init" \
581
+ '[ "$leg2_orphaned" = 1 ]'
363
582
  if [ "$leg2_orphaned" = 1 ]; then
364
583
  signal_suite KILL
584
+ # The suite must really be gone, and this stays an ASSERTION rather than a warning:
585
+ # the work-dir check below passes vacuously against a suite that is still running,
586
+ # because a live suite has not reached its EXIT trap either. Demoting this to a
587
+ # diagnostic would leave that check reading "no cleanup ran" whenever the kill missed.
588
+ # What made the old version of this flaky was the zombie ambiguity inside
589
+ # await_suite_exit, not the assertion itself — that is fixed at the helper, so the
590
+ # obligation can be kept instead of traded away.
365
591
  leg2_suite_gone=0; await_suite_exit && leg2_suite_gone=1
366
- check "leg2: the killed suite exits without running any cleanup" '[ "$leg2_suite_gone" = 1 ]'
367
- # Precondition: with no trap able to run, the wrapper must still be here. If it is
368
- # already gone, something else reaped it and the bound was never exercised.
369
- # `kill -0` succeeds for a zombie, and the verdict below excludes zombies — so a wrapper
370
- # that had already died would satisfy "still alive" here and "gone" there, passing the
371
- # leg without the bound ever being exercised. Require a live, non-zombie process.
372
- leg2_survived_abort=0
373
- if kill -0 "$wrapper_pid" 2>/dev/null; then
374
- case "$(ps -o stat= -p "$wrapper_pid" 2>/dev/null || true)" in
375
- ''|*Z*) : ;;
376
- *) leg2_survived_abort=1 ;;
377
- esac
378
- fi
379
- check "leg2: no reaper ran, so the wrapper is still alive right after the kill" \
380
- '[ "$leg2_survived_abort" = 1 ]'
592
+ # Resume only now: the wrapper was held stopped across the abort so its countdown
593
+ # could not have completed before it, and everything observed from here happens with
594
+ # no reaper of any kind left alive.
595
+ leg2_resume_wrapper
596
+ check "leg2: the killed suite is gone" '[ "$leg2_suite_gone" = 1 ]'
597
+
598
+ # No trap ran — structurally, not by the clock. The suite's EXIT trap ends in
599
+ # `rm -rf "$WORK"`, so the work dir outliving a SIGKILLed suite is the trap's absence
600
+ # made visible. The old spelling asked whether the suite exited inside a 30s window,
601
+ # which is a fact about the host's scheduler rather than about whether cleanup ran.
602
+ # Read together with the assertion above, the pair says: the suite is gone AND it left
603
+ # its work dir behind — which only an untrapped death produces.
604
+ leg2_no_cleanup=0
605
+ [ "$leg2_suite_gone" = 1 ] && [ -n "$leg2_work" ] && [ -d "$leg2_work" ] && leg2_no_cleanup=1
606
+ check "leg2: no cleanup ran — the killed suite's work dir survives it" \
607
+ '[ "$leg2_no_cleanup" = 1 ]'
608
+
381
609
  leg2_self_exit=0; wrapper_gone_within "$(( LEG2_HANG_BOUND + GRACE ))" && leg2_self_exit=1
382
610
  [ "$leg2_self_exit" = 1 ] || diagnose_wrapper leg2
383
- check "leg2: an unreaped wrapper still exits on the fixture's own lifetime bound" \
384
- '[ "$leg2_self_exit" = 1 ]'
611
+ check "leg2: an unreaped wrapper leaves nothing behind" '[ "$leg2_self_exit" = 1 ]'
612
+
613
+ # WHY it ended, not WHEN it was last seen. The fixture writes this only on the far side
614
+ # of its bounded countdown, so present means the bound ended it and absent means a
615
+ # reaper did — the distinction the deleted "still alive right after the kill" check was
616
+ # trying to draw by looking at the process at one instant, which a corpse answers wrong.
617
+ # An unbounded fixture never reaches the write, so this still reds the reverted bound.
618
+ leg2_bound_marker="$(bound_marker_state)"
619
+ [ "$leg2_bound_marker" = present ] || diagnose_wrapper leg2-bound
620
+ check "leg2: the fixture's own lifetime bound is what ended the wrapper" \
621
+ '[ "$leg2_bound_marker" = present ]'
385
622
  fi
386
623
 
387
624
  fi
@@ -40,8 +40,8 @@ Use this skill to turn observed experience into durable agent skills without cop
40
40
 
41
41
  ### Evidence, RCA, charter & attribution(证据 / RCA / charter / 出处核实)
42
42
 
43
- - **Charter-before-editing red-line**: shallow extraction is not useful — do not read sources or edit skills until the extraction charter is set (minimum elements: purpose / scope / depth / root cause / evidence plan / completion standard). The **complete** charter field set and per-standard detail are the *procedure* facet owned by Step 0 + `references/source-to-skill-extraction.md#extraction-charter` — the "Core Rule is canonical on conflict" contract does NOT narrow that field set (it governs same-facet drift, not which surface owns the template). This rule is the always-on gate, not a second copy of that template.
44
- - Run RCA before every extraction, including "summarize this task" and "summarize lessons learned" requests, not only after a failure. Define the future bad outcome, why the current skills/process would allow it, what source coverage or validation would prevent it, and which skill layer owns the prevention rule. Scale the RCA depth to the task: pure wording cleanup can use one sentence; source-derived, task-retrospective, or multi-skill work needs the full baseline RCA. If no RCA is recorded, the extraction is not complete.
43
+ - **Charter-before-editing red-line**: do not read sources or edit skills until the complete charter in `references/source-to-skill-extraction.md#extraction-charter` is filled cell-by-cell. A remembered field subset is not a charter. Step 0 owns the procedure; this rule owns the hard stop. Core-Rule canonicality governs same-facet drift only and never narrows that field set to this bullet.
44
+ - Classify every result — including task/session summaries and lessons-learned requests, not only failures — as **failure/correction**, **stable success**, or **unstable/insufficient evidence**. Failure runs RCA; stable success requires mechanism, non-luck evidence, reuse conditions, firing point, and owner; insufficient evidence stays an observation. Never invent a failure story to justify learning from success. The author classifies, so the class is not self-elective: correction/finding/failure-triggered work defaults to `failure/correction`, and only independent review may accept a relabel into an RCA-skipping class. Missing classification, unaccepted relabel, or missing analysis leaves it `interim`. Full method: `references/source-to-skill-extraction.md#result-learning-baseline-for-every-extraction`.
45
45
  - **RCA must go wider than one 5-Why chain.** 5 Why is the entry technique to get past a symptom, but a single linear chain to one "root cause" is its documented failure mode: real process/agent failures need multiple concurrent causes, the stop point is arbitrary, and "why" drifts toward "who"/blame.
46
46
  - For any non-trivial extraction, RCA must (a) **widen** — enumerate the multiple contributing factors across categories (trigger/routing, stale process-model, missing mechanical control, missing feedback, latent authored-earlier condition, detection gap) before deepening one — a straight chain with no branches means you stopped early; (b) **counterfactually test** each candidate to rank causal weight — *if removed or changed, would the failure still happen?* — keeping necessary/sufficient factors **and** failed redundant safeguards as secondary controls / defence-in-depth, and dropping only genuine coincidence (a one-trace factor is a hypothesis — mark it probabilistic, don't hard-drop); (c) frame prevention as a **mechanical control on the failure CLASS** — an enforced constraint plus the feedback that confirms it fired on a surface the next agent actually reaches in time (not a clause buried in a deep reference) — not agent diligence, preferring the highest-leverage **practical** control (a named artifact, an owner who can change it today, an observable check — never deleting a useful narrow gate or inflating one miss into an over-broad hook) over the first patchable point.
47
47
  - Reject hindsight causes ("agent careless" / "need more attention" / "should have known"): ask why the action made sense given what the agent could see, not what it should have done. Full method, category prompts, stopping points, and sources: `references/source-to-skill-extraction.md` (Deep RCA For Extraction).
@@ -166,7 +166,7 @@ Use this skill to turn observed experience into durable agent skills without cop
166
166
 
167
167
  ### Validation & the dual-track gate(验证 / dual-track 门)
168
168
 
169
- - Static validation is not extraction validation. For any nontrivial or correction-driven extraction, the closeout must explicitly show the correction/baseline RCA, target-output map, sibling-generalization or no-sibling decision, landed-rule diff, validation commands, and independent review/challenge result. If those evidence rows are absent, report the work as `interim` even when `validate-skill.sh` or `check-ccl-skills.sh` passes.
169
+ - Static validation is not extraction validation. Nontrivial closeout shows the matching result analysis, target-output and sibling decisions, landed diff, commands, and independent review/challenge; otherwise report `interim` even if static checks pass.
170
170
  - For a whole-session/task-retrospective extraction over operational delivery that changed repositories, branches, MRs, pipelines, releases, or deployable artifacts, closeout validation must show one of: `delivery-state rows` with changed artifact, branch/worktree, remote/MR, CI/local verification, cancelled/retried, residual-risk, and next-action state; or `artifact/status axis: not-applicable` with the reason. Missing delivery-state evidence downgrades the extraction to `interim`; static validation and clean independent review do not close it.
171
171
  - A rule that exists but did not trigger is a validation-gate defect, not proof that the workflow is adequate.
172
172
  - For any correction where the missed step was covered by any rule in this workflow that a reasonable reader would apply to the scenario, add or tighten a closeout gate that would have blocked the exact premature final answer.
@@ -215,9 +215,9 @@ Use this skill to turn observed experience into durable agent skills without cop
215
215
 
216
216
  ## Extraction Workflow
217
217
 
218
- 0. Set the extraction charter and RCA baseline.
218
+ 0. Set the extraction charter and result-learning baseline.
219
219
  - Hard stop: before source reads / before edits, record the charter from `references/source-to-skill-extraction.md#extraction-charter` — **open that table and fill it cell-by-cell; a charter written from memory of the field names is not a charter and must not be recorded as one.** Each field's real constraints live only in its own cell (Evidence plan's produced-artifact-first rule, Scope's watermark validation, RCA depth scaling), so a from-memory charter reproduces the field list and none of the gates, while looking complete. Trivial wording cleanup still records Depth explicitly.
220
- - RCA baseline: use `#baseline-rca-for-every-extraction`; scale by failure shape. Deep RCA anchor for any non-trivial failure: widen to multiple factors -> counterfactual-test each -> land a mechanical control with firing path (`#deep-rca-for-extraction`).
220
+ - Result baseline: classify first, then use `#result-learning-baseline-for-every-extraction`; failure RCA depth comes from `#deep-rca-for-extraction`.
221
221
  - Task/session incidents: use `#task-retrospective-extraction` and its delivery-chain RCA prompts before deciding whether the lesson lands in this workflow, a sibling skill, validator, memory, project artifact, or final-response only.
222
222
  - Full/complete/deep asks require the source-register shape before reads: source groups, inclusion/exclusion, minimum artifact depth, owner skill, completion evidence, and batch-progress status tracking (`pending`/`read`/`deep-read`/`excluded`/`unavailable`/`routed`) from `#full-coverage-source-register-protocol`; close or explicitly downscope every required batch before saying "complete".
223
223
  - Target-output map derives from Lifecycle impact: every affected stage gets an owner target or explicit no-update reason before source-derived editing starts (`#target-output-map`).
@@ -239,7 +239,7 @@ Use this skill to turn observed experience into durable agent skills without cop
239
239
  - The map must include every plausible owner per affected lifecycle stage, including sibling skills in the same stage. If no skill owns a stage, write the no-output reason; do not silently omit the stage.
240
240
  - For any extraction beyond wording-only cleanup, include the provenance-to-target diff shape before editing: source mechanism, provenance row, target file, executable landing, test or acceptance owner, and status.
241
241
  - Trigger situations and users/tasks it should serve.
242
- - What future failure, drift, or repeated explanation it should prevent.
242
+ - What future failure or drift it should prevent, or which evidenced success mechanism it should preserve and reuse.
243
243
  - For subjective or high-impact skills such as design, UX, frontend/client, product workflow, architecture, or review, define pressure scenarios and acceptance criteria before editing the skill.
244
244
  - For UI/UX or client-facing skills, the pressure scenario must ask whether a person without source access can produce a good-looking and behaviorally sound screen: clear visual hierarchy, fitting density, risk-matched feedback, recoverable state transitions, responsive/device adaptation, and rendered acceptance evidence.
245
245
 
@@ -282,8 +282,8 @@ Use this skill to turn observed experience into durable agent skills without cop
282
282
 
283
283
  6. Validate before landing.
284
284
  - YAML frontmatter parses and description is trigger-focused.
285
- - Extraction charter is satisfied: purpose, scope, depth, root cause, evidence plan, and completion standard are either met or explicitly downscoped.
286
- - RCA gate: every extraction, including task-summary or lesson-learned extraction, has a visible baseline RCA or correction RCA. Missing RCA is a validation failure, not a documentation nit.
285
+ - Extraction charter is satisfied: purpose, scope, depth, result classification and matching analysis, evidence plan, and completion standard are either met or explicitly downscoped.
286
+ - Result-learning gate: classification and matching analysis satisfy `#result-learning-baseline-for-every-extraction`; missing, mismatched, or insufficient-evidence-as-rule fails validation.
287
287
  - Delivery-chain RCA gate: for incident or task-retrospective extraction, validation must show definition, implementation, verification, review/MR or release readiness, and retrospective-quality causes were checked or explicitly ruled out. If the extraction workflow itself missed the deeper cause, the workflow fix must be landed and validated before finalizing.
288
288
  - Blocked-verification gate: any `unavailable`, `skipped`, or `blocked` verification claim must include remediation commands already attempted, observed result, residual risk, and next unblock action. If remediation was feasible but not attempted, validation fails and the test remains pending rather than unavailable.
289
289
  - Blocked-source gate: any `unavailable`, `skipped`, `blocked`, timed-out, partial, or failed source-read claim must include the smaller/different read strategy already attempted, observed result, recovered evidence, residual gap, and next unblock action. If remediation was feasible but not attempted, validation fails and the source row remains pending.
@@ -50,13 +50,13 @@
50
50
  - **防作弊**:runner 校验每 task 的 `frozen_at_sha` 是 HEAD 祖先(非祖先 = drift,排除出回归判定);同一改动若同时动 task-bank 和 SKILL.md description 会显式告警(防"改 skill 顺手改测试让它过")。
51
51
  - `--baseline <json>` 给 diff 式报告(newly_failed / newly_passed)。
52
52
 
53
- ### Bank 用例修复的测量纪律(normative)
53
+ ### Bank 用例修复的测量协议(测量必做;落地裁决归本轮实际门禁)
54
54
 
55
- 对 bank 失败用例做路由面修复时,以下每条都是修复轮的过闸条件,不是建议:
55
+ 以下是生成可比较证据的默认协议。**降级的是「F4 自己充当统一合并门禁」这个声称,不是「必须测、且必须有人裁决」这个义务**——这两件事分开:落地判断交给本轮实际的 owner/风险/评审门禁,但**测量本身不可选**。任何动 routing 面(SKILL.md description、task-bank 判定面)的改动都必须按下列协议产出证据;没跑就是没收敛,不得进入独立评审、也不得声称本轮无回归。采用不同样本量时,须随工件记录理由,且该理由与本轮证据一同进入独立评审——「记了理由」本身不是豁免,自审通过的理由不构成已裁决:
56
56
 
57
57
  1. 动任何 description 之前必须先跑 **≥10 轮有效观测**的稳定性基线,把稳定失败与抖动分开;抖动不得作为修改依据(grader 超时/不可解析轮不算有效观测,须补跑)。
58
58
  2. 改后通过数必须在**最终措辞**上重测:中间稿的通过数在措辞再变的那一刻作废,不得挪用到最终候选的证据里。
59
- 3. 受影响邻居用例集必须改前/改后各 **≥3 轮**,集合须含期望 owner 自己的兄弟用例与高词面重叠的他 owner 用例;任何邻居回归都阻断本轮。
59
+ 3. 受影响邻居用例集默认改前/改后各 **≥3 轮**,集合须含期望 owner 自己的兄弟用例与高词面重叠的他 owner 用例;邻居回归作为独立 finding 交由本轮实际门禁处置——**该 finding 须以 blocking 记入本轮 dual-track 评审记录,且只能由独立评审方豁免,不能由实现者自行判定「本轮没有门禁采用这组证据」而放行**。降级的是「F4 自己充当合并门禁」这一声称,不是「回归必须被人裁决」这一义务;后者若也随之消失,这一条就只剩被裁决方自审。
60
60
  4. 每轮判决必须连同 **runner 调用、grader 模型身份、候选身份**(commit 或描述内容指纹)与**原始逐轮工件的持久定位符**一并记入轮记录;没有定位符的通过数只能标注为 operator-reported,不得据以宣称修复轮已 concluded。
61
61
 
62
62
  ## Tier-3:hub golden trace 真 agent 回放(已落地,advisory,人工判定)
@@ -71,16 +71,15 @@
71
71
  - **随机性**:agent 非确定;判定先人工、nightly 起步,有稳定史前不自动 gate。防作弊同 T2(`frozen_at_sha` 祖先校验)。
72
72
  - **双用途**:除回归外,Tier-3 还可当**改技能前的 RED-baseline**(改前手动跑触发场景看真 agent 是否真路由错,改后看 compliance)——可选;只有真观察到 miss 才算 RED(PASS/INCONCLUSIVE 不算),小 N + 非确定有噪声,手动跑两次自己留两份报告。落地 + 防作弊注意见 [validation-and-landing.md](validation-and-landing.md) "Optional real-agent RED-baseline"。
73
73
 
74
- ## Health roll-up:综合分 + 趋势(已落地,advisory)
74
+ ## Health roll-up:描述性仪表盘(已落地,advisory)
75
75
 
76
- 把上面各信号卷成**一个加权 0–10 分 + 趋势**,答"仓库在变好还是变差"。这是 OpenSSF Scorecard 给代码仓用的形态(每 check 0–10 → 按风险加权聚合 0–10 → 跨时间追);我们把它套到 **F4 技能 eval 信号**上,不是去打分代码工具。映射依据见 [harness-patterns-and-eval.md](harness-patterns-and-eval.md) §3.4。
76
+ 把上面各信号卷成**一个加权 0–10 显示值 + 同尺子变化**,用于定位值得继续检查的维度。它借用 OpenSSF Scorecard 的呈现形态,但不把不同性质的 F4 信号变成“仓库整体变好/变差”的总判决。映射与限制见 [harness-patterns-and-eval.md](harness-patterns-and-eval.md) §3.4。
77
77
 
78
78
  `scripts/eval-health.rb <repo-root> [--trace-json p] [--bank-json p] [--history p] [--no-write] [--json p] [--quiet]`。`make eval-health`。
79
79
 
80
80
  - **维度(各 0–10,按风险加权,OpenSSF 风格)**:`structural` 权重 10(Critical,`validate-skill.sh` pass/fail)· `routing_static` 权重 10(Critical,T1 blocking=0 满分、有 blocking 砸到 3、advisory 轻罚)· `trace` 权重 7.5(High,T3 pass/considered)· `bank` 权重 5(Medium,T2 pass/tasks)。
81
81
  - **只跑确定性两维**(structural + routing_static,无 LLM、快);`trace`/`bank` 需 `claude` 且随机,**不自动跑** —— 用 `--trace-json` / `--bank-json` 把 T3/T2 报告喂进来,否则该维 **skip,权重按比例重分**给在场维(诚实标注 `dims=…`)。
82
- - `composite = Σ(score_i·w_i) / Σ(w_i)`,只对在场维求和;skip 维自动从分母剔除。band:≥9 CLEAN / ≥7 WARNING / ≥4 NEEDS WORK / <4 CRITICAL。
83
- - **advisory,永不因分数阻断**:退出码 `0` = 跑完 **或** 没有可算的维(都不是失败);`2` = 用法/setup 错(`<repo-root>` 不对或无 `skills/` 目录)。低分、坏报告、空历史都不会非 0。畸形的 T2/T3 报告(非对象 / 非整数 / `pass>total`)直接 **skip,不静默打分**。**绝不接进 `check-ccl-skills.sh`**。二元门禁(结构校验 + T1 blocking)仍是真值、独立挡 merge,与这个分无关。理由是 Goodhart 律(measure 变成 target 就不再是好 measure):一旦这个综合分成了 merge gate,人会去调分而不是修仓。
82
+ - `composite = Σ(score_i·w_i) / Σ(w_i)`,只对在场维求和;skip 维自动从分母剔除。band 只是浏览提示,不得当作验收等级。
83
+ - **advisory,永不因显示值阻断**:退出码 `0` = 跑完 **或** 没有可算的维(都不是失败);`2` = 用法/setup 错(`<repo-root>` 不对或无 `skills/` 目录)。低值、坏报告、空历史都不会非 0。畸形的 T2/T3 报告(非对象 / 非整数 / `pass>total`)直接 **skip,不静默打分**。**绝不接进 `check-ccl-skills.sh`**。结构校验、T1 blocking、任务验收和行为证据分别独立判定;质量失败不能被其它维度的高值平均掉。
84
84
  - **corpus/version 守卫(防跨变更 task-bank 当稳定指标比)**:每条历史记 `corpus`(= task-bank + golden-traces 输入内容的指纹)+ `repo_sha` + 在场 `dims`。**趋势 delta(IMPROVING/DECLINING)只跟最近一条 `(corpus, dims)` 都相同的历史比**;否则打印"baseline reset, not compared",不偷偷比。这样"加了 10 条简单 task → 分涨了"不会被读成真进步(尺子换了)。
85
85
  - **历史文件**:默认 `eval/health-history.jsonl`,**git-ignore**(同 Goodhart 理由:committed 的数会招"调数不修仓";也免 append 把树搞脏)。`--history` 可改路径,`--no-write` 不落盘。
86
- - **基线**(本机、本 corpus):确定性两维 = `9.5/10`(structural 10 + routing_static 9,advisory=1)。
@@ -6,7 +6,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
6
6
 
7
7
  ```
8
8
  0. Charter → ~/.<host>/skills/.extraction-work/<project>-charter.md
9
- Purpose / Scope / Depth / RCA baseline / Open questions
9
+ Purpose / Scope / Depth / Result baseline / Open questions
10
10
  │
11
11
  1. Source register → ~/.<host>/skills/.extraction-work/<project>-source-register.md
12
12
  Each row: source id / class / status / target skill / extracted mechanisms
@@ -43,7 +43,8 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
43
43
 
44
44
  - Owner: maintainer
45
45
  - Location: `~/.<host>/skills/.extraction-work/<project>-charter.md`
46
- - Required fields: Purpose, Scope, Depth, Root cause (Deep RCA: widen → counterfactual-test → control), Failure modes, Lifecycle impact, Evidence plan, Completion standard.
46
+ - Required fields: Purpose, Scope, Depth, Result classification, matching analysis, Failure modes or success-reuse conditions, Lifecycle impact, Evidence plan, Completion standard.
47
+ - Result classification: failure/correction → Deep RCA (widen → counterfactual-test → control); stable success → mechanism + non-luck evidence + reuse conditions + firing point + owner; unstable/insufficient evidence → observation only.
47
48
  - Full structure: `SKILL.md` Core Workflow Step 0.
48
49
  - Output: a file the maintainer can re-read in 3 months and understand what they were trying to do.
49
50
 
@@ -147,13 +147,13 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
147
147
 
148
148
  **用时机**:操作层月级别稳定 + 要做一次 system-wide 换层/瘦身时;把它当"换层前的对抗性回归证据",不是频繁迭代期的日常闸。
149
149
 
150
- ### 3.4 Composite health score + trend(综合分,advisory roll-up)
150
+ ### 3.4 Health signal dashboard(描述性 roll-up)
151
151
 
152
- 3.1–3.3 都是**单次**判定(这次改动好不好、这条 trace 偏没偏、这次换层纪律掉没掉);即便 3.3 横跨一整个夹具集,判的仍是**某一次改动**的即时行为。缺的是**跨时间的一个数**:仓库整体在变好还是变差。
152
+ 3.1–3.3 都保留各自的判定对象;跨时间还需要一个快速入口,显示本次有哪些信号在场、哪些维度变化,方便继续下钻。
153
153
 
154
- **外部锚**:OpenSSF Scorecard 给代码仓的做法是 —— 每个 check 0–10,再**按风险加权**聚合成一个 0–10 总分(Critical=10 / High=7.5 / Medium=5 / Low=2.5),并跨时间追。gstack-health 给代码仓同形态(加权 composite + 历史趋势 + dashboard)。**这是成熟业内形态,不是某一家发明的**。
154
+ **外部锚**:OpenSSF Scorecard 用每个 check 0–10、风险加权聚合和历史变化展示代码仓信号。这里仅借它的**展示形态**,不继承“一个总分代表整体健康”的解释。
155
155
 
156
- **我们的映射(route-not-copy:不打分代码工具,套到 F4 信号上)**:`eval-health.rb` 把四个信号卷成一个加权 0–10 + 趋势 ——
156
+ **我们的映射(route-not-copy)**:`eval-health.rb` 展示四个信号,并保留一个兼容既有实现的加权 0–10 值与同尺子变化 ——
157
157
 
158
158
  | 维度 | 风险/权重 | 0–10 来源 |
159
159
  |---|---|---|
@@ -164,12 +164,13 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
164
164
 
165
165
  `composite = Σ(score·weight) / Σ(weight)`,只算在场维(skip 维权重重分,同 OpenSSF/gstack)。确定性两维自动跑,T2/T3 喂报告进来(否则 skip)。契约见 [eval-routing.md](eval-routing.md) 的 Health roll-up 节。
166
166
 
167
- **两条不可省的护栏**(否则这个数会骗人):
167
+ **三条不可省的护栏**:
168
168
 
169
169
  1. **advisory,不当 gate**(Goodhart):度量变成 target 就被博弈。综合分**永不接门禁**,二元门禁(结构 + T1 blocking)仍独立挡 merge;历史文件 git-ignore,免得"committed 的数"招人调数不修仓。
170
170
  2. **corpus/version 守卫**:T2/T3 的 task-bank、golden-traces **本身会变**。加 10 条简单 task,pass-rate 涨了但仓没变好 —— **尺子换了**。每条历史记 `corpus` 指纹(输入内容 hash)+ 在场 `dims`;**趋势只跟 `(corpus, dims)` 全同的历史比**,否则 baseline reset 不偷偷比。等价于 gstack-health "尺子变了就从新基线重新追"。
171
+ 3. **不平均掉质量失败**:结构、路由、真实回放和任务结果的语义不同;任何质量或安全阻断仍由自己的门禁决定,不能被其它维度的高值抵消。
171
172
 
172
- **用**:周期性(nightly / 改动后)看一眼仓库 health 趋势,**辅助判断**,不替代任何门禁。**不用**:别把它当通过标准、别committed、别跨 corpus 硬比。
173
+ **用**:周期性看一眼信号变化并下钻。**不用**:别据此单独声称仓库整体变好/变差,别把它当通过标准、别 committed、别跨 corpus 硬比。
173
174
 
174
175
  ---
175
176
 
@@ -297,6 +297,24 @@ A carve-out is a contract, not a category badge. Without the evidence row, fall
297
297
 
298
298
  **Occurrences that promoted it**: `scripts/owner-dispatch/owner-dispatch.sh` (state keyed per-worktree → boundary record written in a worktree unreadable from the primary checkout) and `hooks/guard-edit-isolation.sh` (path-name match → the repo's only edit-time hard-deny gate silently disabled for checkouts under a `worktrees/`-named path).
299
299
 
300
+ ## Anti-pattern 28 — Process liveness decided by an EXISTENCE test that a corpse answers
301
+
302
+ **Symptom**: a probe or fixture concludes "this process is still alive" — or "it is a live orphan" — from a test that only proves the pid is still in the process table: `kill -0 "$pid"`, or reading `ps -o ppid=` and comparing it against init's pid. The same suite's verdict scans then exclude zombies (`$stat !~ /Z/`) when deciding the process is *gone*.
303
+
304
+ **Why bad**: an exited-but-unreaped process is still in the table. It answers `kill -0`, and its ppid still reads — as `1` once it is reparented. So one process state gets two answers: the precondition check says "alive, scenario built", the verdict scan says "gone, nothing leaked", and an assertion between them that samples liveness at one instant says "already dead" and fails. The red then names the code under test for something it did not do, and the failure is intermittent because it depends on when the OS reaps. Worse, the precondition passing means the probe goes on to assert about a scenario it never actually built.
305
+
306
+ **Fix**: consult process **state**, not existence.
307
+ - One vocabulary for the whole probe — a helper returning `live` / `zombie` / `absent` (`ps -o stat=`; empty → absent, `*Z*` → zombie, else live) — used by *every* liveness question, so the checks cannot answer the same pid differently.
308
+ - Preconditions require `live`. A corpse is not an orphan; accepting one is how a probe goes on to assert about a scenario that was never constructed.
309
+ - Better still, stop asking the process at all: assert on an artifact the run leaves behind — a work dir the cleanup path would have deleted, a marker the fixture writes only on the code path under test. That answers *why* a process ended, which no liveness sample can.
310
+ - Never assert liveness by an instantaneous sample; reap lag on a loaded runner needs a bounded grace period. (`testing-strategy/references/ci-fixtures-and-flake-control.md` owns that rule — this row is its firing path.)
311
+
312
+ **Grep**: `find . -name 'test_*.sh' -o -name 'test.sh' | xargs grep -nE 'ps[[:space:]]+-o[[:space:]]+ppid=.*=[[:space:]]*"?1"?[[:space:]]*\]'` (non-comment hits with no `stat=` / `*_state` consult within two lines). Machine-enforced by `check-ccl-skills.sh` (`liveness_predicate_scan`). Scope is test scripts only: outside them a `kill -0` before signalling asks about existence, which is the right question there. Five things are deliberately **not** caught and stay human/challenge checks — `while kill -0 "$pid"` watchdog loops (two exist here; both wait on a direct child and `wait` for it immediately after, so the shell reaps it and the loop ends), a bare `kill -0` liveness branch inside a loop body, any dynamic spelling, a state-helper mention in a **trailing** comment (whole-line comments are dropped, but telling a trailing `#` from `${var#prefix}` needs a shell parser — the same call Anti-pattern 27 makes), and a process-state read — direct `ps -o stat=` or via a helper — that inspects a *different* pid than the oracle tests, or whose result never reaches the verdict. Both are the same irreducible gap: proving the read actually governs the decision needs dataflow over a parsed shell, not a text window, so this is the limit the mechanical gate stops at by design. Both are **exercised** by `test_liveness_predicate_gate.sh` (P13/P14) rather than only described here — they assert the gate does not fire, so tightening it turns them red and forces this paragraph to be updated instead of quietly going stale.
313
+
314
+ The waiver is a **proxy** for "this site consults process state", not an invariant, and five successive review rounds each found it too loose in a different way — a comment naming the helper, an unrelated `*_state` token, a helper that never reads state, a bare `stat=` assignment, and a hollow helper borrowing an unrelated read elsewhere in the file. Tying a call to its definition needs a shell parser, so the honest disposition is a narrow predicate with these limits stated rather than another round of widening. The mechanical gate is the deterministic catch for the exact recurring shape; this checklist and the adversarial challenge remain the comprehensive net.
315
+
316
+ **Occurrences that promoted it**: the code-review abort-leak probe red CI three times, each on a different leg-2 assertion, each asserting the probe's environment rather than the suite — the reparent check accepted a zombie, the "still alive right after the kill" check rejected that same zombie, and the verdict scan reported it gone. The rule forbidding this already existed in `testing-strategy`, and the suite's own hang cases followed it while the probe did not: the gap was enforcement, not content.
317
+
300
318
  ## How to use this checklist
301
319
 
302
320
  Before committing any skill or reference change touching operational/architectural rules: