@ccoalm/ccl-skills 0.18.11 → 0.18.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (30) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/hooks/headless-background-stop.sh +18 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +7 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/hooks/host-input.py +71 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +9 -4
  5. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_headless_background_stop.sh +150 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +47 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +7 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/references/multi-agent-delegation-playbook.md +1 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +4 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/worktree-mechanics.md +10 -3
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +2 -2
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +6 -1
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-lifecycle-handoff.md +1 -1
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +2 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +9 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-sync-pointers.sh +11 -4
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-paired-eval.py +1546 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +140 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +4 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_sync_pointers.sh +13 -4
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +12 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_skill_paired_eval.py +1052 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_teardown_guard_pins.sh +354 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/SKILL.md +8 -108
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/merge-and-teardown.md +47 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/pre-merge-landing-checks.md +73 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/shared-branch-rebase.md +1 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_sweep.sh +1 -1
  29. package/dist/assets/release.json +59 -24
  30. package/package.json +1 -1
@@ -0,0 +1,18 @@
1
+ #!/usr/bin/env bash
2
+ # Stop guard for sessions nothing re-invokes (claude -p, the SDKs): one block per background
3
+ # task still running at the stop, because the host stops those tasks when the session ends and
4
+ # work waiting on them is lost. Interactive sessions exit here untouched. Reads only the hook
5
+ # input, never the transcript or host configuration.
6
+ set -u
7
+ case "${CLAUDE_CODE_ENTRYPOINT:-}" in
8
+ sdk|sdk-*) ;;
9
+ *) exit 0 ;;
10
+ esac
11
+ HELPER="$(cd "$(dirname "$0")" && pwd)/host-input.py"
12
+ if ! command -v python3 >/dev/null 2>&1 || [ ! -r "$HELPER" ]; then
13
+ printf '%s\n' '{"systemMessage":"Background task check unavailable: Python input normalizer missing."}'
14
+ exit 0
15
+ fi
16
+ python3 "$HELPER" headless-background 2>/dev/null ||
17
+ printf '%s\n' '{"systemMessage":"Background task check unavailable: input normalizer failed."}'
18
+ exit 0
@@ -161,6 +161,13 @@
161
161
  "timeout": 5,
162
162
  "async": false,
163
163
  "statusMessage": "Checking delivery handoff..."
164
+ },
165
+ {
166
+ "type": "command",
167
+ "command": "\"${CLAUDE_PLUGIN_ROOT}/hooks/headless-background-stop.sh\"",
168
+ "timeout": 5,
169
+ "async": false,
170
+ "statusMessage": "Checking background tasks..."
164
171
  }
165
172
  ]
166
173
  }
@@ -612,6 +612,65 @@ def stop_notice(payload, lane, message):
612
612
  return {'systemMessage': message}
613
613
 
614
614
 
615
+ # Claude Code reports `sdk-cli` for `claude -p` and an `sdk` entrypoint for the SDKs; interactive
616
+ # sessions report other values and are re-invoked when a background task finishes.
617
+ NON_INTERACTIVE_ENTRYPOINT = re.compile(r'sdk(?:-[a-z]+)?')
618
+
619
+
620
+ def claim_background_notice(payload, task_id):
621
+ """True the first time this session stops with the task still running, False on a repeat,
622
+ None when state is unavailable. Keyed by session, not transcript, which may not exist."""
623
+ try:
624
+ source = Path(__file__).resolve().with_name('skill-loading.py')
625
+ spec = importlib.util.spec_from_file_location('ccl_stop_state', source)
626
+ module = importlib.util.module_from_spec(spec)
627
+ spec.loader.exec_module(module)
628
+ session = payload.get('session_id')
629
+ if not isinstance(session, str) or not session or len(session) > 1024:
630
+ return None
631
+ state = module.State(module.digest([str(module.ROOT), session, 'headless-background']))
632
+ try:
633
+ return bool(state.claim_attempt('headless-background', [task_id]))
634
+ finally:
635
+ state.close()
636
+ except Exception:
637
+ return None
638
+
639
+
640
+ def headless_background(payload):
641
+ """Block a stop once per background task still running in a session nothing re-invokes; the
642
+ caller has already checked that the entrypoint is one (NON_INTERACTIVE_ENTRYPOINT).
643
+
644
+ A headless session ends at the stop and the host stops its background tasks seconds later,
645
+ so work waiting on their results is lost. Reads only the hook input, so it also works when
646
+ session persistence is off and the transcript-reading hooks cannot run."""
647
+ if not isinstance(payload, dict) or payload.get('hook_event_name') != 'Stop':
648
+ return None
649
+ tasks = payload.get('background_tasks')
650
+ running = [task for task in tasks if isinstance(task, dict) and task.get('status') == 'running'
651
+ and isinstance(task.get('id'), str) and 0 < len(task['id']) <= 128] if isinstance(tasks, list) else []
652
+ fresh = []
653
+ for task in running:
654
+ claimed = claim_background_notice(payload, task['id'])
655
+ # Without state a repeat cannot be told apart, so the host's own retry flag bounds it.
656
+ if claimed or (claimed is None and payload.get('stop_hook_active') is False):
657
+ fresh.append(task)
658
+ if not fresh:
659
+ return None
660
+ def label(task):
661
+ text = re.sub(r'\s+', ' ', str(task.get('description') or task.get('type') or 'task')).strip()
662
+ return f'- {text[:100]} ({task["id"]})'
663
+ listed = [label(task) for task in fresh[:5]] + ([f'- and {len(fresh) - 5} more'] if len(fresh) > 5 else [])
664
+ return {'decision': 'block', 'reason': (
665
+ 'Background task check: this is a headless session. When you stop, the session ends and these '
666
+ 'background tasks are stopped within seconds, with no notification afterwards:\n' + '\n'.join(listed) +
667
+ '\nIf the request still depends on one of them, wait for that task in the foreground: poll its output '
668
+ 'file, or the files it writes, in a bounded foreground loop until it has finished, then act on its '
669
+ 'result. Do not start it again; a second run would repeat its effects. If none is needed, stop it with '
670
+ 'TaskStop or say why it can be dropped. '
671
+ 'This check fires once per task.')}
672
+
673
+
615
674
  def extraction_overflow(payload):
616
675
  path = payload.get('transcript_path')
617
676
  cwd = payload.get('cwd')
@@ -884,6 +943,18 @@ def main():
884
943
  'verifiable': False, 'truncated': True, 'prior_handoff': False,
885
944
  'continuation_contract_visible': False}))
886
945
  return 1
946
+ elif sys.argv[1] == 'headless-background':
947
+ if not NON_INTERACTIVE_ENTRYPOINT.fullmatch(os.environ.get('CLAUDE_CODE_ENTRYPOINT', '')):
948
+ return 0 # an interactive session hears nothing from this check, whatever its input
949
+ try:
950
+ raw = sys.stdin.read(2 * 1024 * 1024 + 1)
951
+ if len(raw) > 2 * 1024 * 1024:
952
+ raise ValueError('oversized input')
953
+ result = headless_background(json.loads(raw))
954
+ if result:
955
+ print(json.dumps(result))
956
+ except (OSError, ValueError, TypeError, IndexError, AttributeError):
957
+ print(json.dumps({'systemMessage': 'Background task check unavailable: input could not be verified.'}))
887
958
  elif sys.argv[1] in ('proposed-next', 'extraction-overflow'):
888
959
  payload = {}
889
960
  try:
@@ -70,8 +70,9 @@ masked=$(printf '%s' "$cmd" | sed -E \
70
70
  # `-f body=` values, echoed docs, `git log --grep`), and an advisory that cries
71
71
  # wolf gets ignored. Raw-API / GraphQL / curl merges are an ACCEPTED best-effort
72
72
  # NON-fire — `glab mr merge` / `gh pr merge` is the near-universal agent merge
73
- # path, and the human-readable cleanup rule in worktree-isolation SKILL.md +
74
- # bootstrap covers EVERY merge path regardless of this reminder.
73
+ # path, and the human-readable cleanup rule in worktree-isolation
74
+ # references/merge-and-teardown.md + bootstrap covers EVERY merge path
75
+ # regardless of this reminder.
75
76
  # `gh help pr merge` / `glab help mr merge` print help (the merge guard's help
76
77
  # denial points there). Remove only those literal invocations, never a prefix,
77
78
  # so a real merge before or after them in the same command still matches.
@@ -129,12 +130,16 @@ if command -v git >/dev/null 2>&1 && git -C "$cwd" rev-parse --git-dir >/dev/nul
129
130
  wt=$(git -C "$cwd" worktree list 2>/dev/null)
130
131
  fi
131
132
 
132
- reminder="🧹 worktree-isolation 收尾提醒(自动):检测到 MR/PR 合并命令。先按本节「已集成判据」确认这次合并**已真正完成**(平台 MR/PR 已在当前 head SHA 上 merged;仅授权、仅排队 auto-merge、或合并失败都不算已集成);确认后,若源分支是临时 feature 分支就立即清理三侧,别攒:
133
+ # The text points at the canonical teardown section and carries the guards that
134
+ # must not be lost at this moment; a digest that drops any of them would
135
+ # out-vote the canonical rule the agent loaded earlier.
136
+ reminder="🧹 worktree-isolation 收尾提醒(自动):检测到 MR/PR 合并命令。动手前先读 worktree-isolation/references/merge-and-teardown.md 的「收尾」节,按其「已集成判据」确认这次合并**已真正完成**(平台 MR/PR 已在当前 head SHA 上 merged;仅授权、仅排队 auto-merge、或合并失败都不算已集成);确认后,若源分支是临时 feature 分支就立即清理三侧,别攒:
137
+ git -C <path> status --ignored -s # 删 worktree 前必须先跑且必须 exit 0;非空先按重算代价判定,贵的产物先救回主检出
133
138
  git worktree remove <path> # 不加 --force(脏树/未合并被拒=安全网)
134
139
  git branch -d <branch> # 不加 -D(未合并被拒=安全网)
135
140
  git push origin --delete <branch> # 远端源分支——破坏性,务必先确认已集成再删
136
141
  git worktree prune && git worktree list && git branch # 验证三侧都没了
137
- 唯一例外:源分支本身是永久/集成分支(如 dev→main promotion,源是 dev)——绝不删。squash 合并测不到祖先则保守保留、先确认已集成。"
142
+ 不删的例外:① 源分支本身是永久/集成分支(如 dev→main promotion,源是 dev);② 分支名含 release 的分支(大小写不敏感,如 release/*、hotfix-release),要删由用户显式指名。squash 合并测不到祖先则保守保留、先确认已集成;worktree 里仍有未完成的外部副作用任务(迁移/部署等)时,等它完成再清。"
138
143
  if [ -n "$wt" ]; then
139
144
  reminder="${reminder}
140
145
  当前 worktrees(挑出刚合并的那个源 worktree 清理):
@@ -0,0 +1,150 @@
1
+ #!/usr/bin/env bash
2
+ # Deterministic behavior suite for hooks/headless-background-stop.sh.
3
+ # Registered in the Makefile `test` target; requires jq to build hook inputs.
4
+ set -u
5
+
6
+ SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
7
+ HOOK="${HEADLESS_BG_HOOK:-$SCRIPT_DIR/headless-background-stop.sh}"
8
+ [ -f "$HOOK" ] || { echo "FAIL: hook not found: $HOOK" >&2; exit 1; }
9
+ command -v jq >/dev/null 2>&1 || { echo "FAIL: jq required for this suite" >&2; exit 1; }
10
+ bash -n "$HOOK" || { echo "FAIL: hook is not syntactically valid" >&2; exit 1; }
11
+
12
+ pass=0; fail=0
13
+ WORK="$(mktemp -d "${TMPDIR:-/tmp}/headless-bg-test.XXXXXX")"
14
+ trap 'rm -rf "$WORK"' EXIT
15
+ STATE="$WORK/state"; mkdir -p "$STATE"
16
+ ERR="$WORK/err"
17
+
18
+ # payload <session> <stop_hook_active> <tasks-json>
19
+ payload() {
20
+ jq -nc --arg s "$1" --argjson a "$2" --argjson t "$3" \
21
+ '{session_id:$s,hook_event_name:"Stop",stop_hook_active:$a,background_tasks:$t}'
22
+ }
23
+ task() { # task <id> <status> [description]
24
+ jq -nc --arg i "$1" --arg st "$2" --arg d "${3:-Run the review gate}" \
25
+ '{id:$i,type:"shell",status:$st,description:$d,command:"sleep 60"}'
26
+ }
27
+
28
+ # run <entrypoint> <tmpdir> <stdin> -> sets OUT; fails the case unless exit 0 with silent stderr
29
+ run() {
30
+ local rc errbytes
31
+ OUT=$(printf '%s' "$3" | CLAUDE_CODE_ENTRYPOINT="$1" TMPDIR="$2" bash "$HOOK" 2>"$ERR"); rc=$?
32
+ errbytes=$(wc -c <"$ERR" | tr -d ' ')
33
+ [ "$rc" = 0 ] && [ "$errbytes" = 0 ] && return 0
34
+ echo "FAIL: hook must exit 0 with silent stderr (rc=$rc, stderr=${errbytes}B)"; fail=$((fail + 1)); return 1
35
+ }
36
+
37
+ # expect <block|quiet|notice> <label> <entrypoint> <tmpdir> <stdin> [needle...]
38
+ expect() {
39
+ local want="$1" label="$2" got needle
40
+ run "$3" "$4" "$5" || return
41
+ if [ -z "$OUT" ]; then
42
+ got=quiet
43
+ elif printf '%s' "$OUT" | jq -e '.decision == "block" and (.reason | type == "string")' >/dev/null 2>&1; then
44
+ got=block
45
+ elif printf '%s' "$OUT" | jq -e '.systemMessage | startswith("Background task check unavailable")' >/dev/null 2>&1; then
46
+ got=notice
47
+ else
48
+ echo "FAIL [$label]: unexpected output: $OUT"; fail=$((fail + 1)); return
49
+ fi
50
+ if [ "$got" != "$want" ]; then
51
+ echo "FAIL [$label]: expected $want, got $got: $OUT"; fail=$((fail + 1)); return
52
+ fi
53
+ shift 5
54
+ for needle in "$@"; do
55
+ if ! printf '%s' "$OUT" | jq -e --arg n "$needle" '.reason | contains($n)' >/dev/null 2>&1; then
56
+ echo "FAIL [$label]: reason lacks '$needle': $OUT"; fail=$((fail + 1)); return
57
+ fi
58
+ done
59
+ pass=$((pass + 1))
60
+ }
61
+
62
+ one="[$(task t1 running 'Run the external review on the diff')]"
63
+ expect block "headless stop with a running task" sdk-cli "$STATE" "$(payload s1 false "$one")" \
64
+ "headless session" "Run the external review on the diff (t1)" "foreground" "Do not start it again" "TaskStop" "once per task"
65
+ if printf '%s' "$OUT" | jq -e '.reason | test("(?i)run(ning)? it again|rerun")' >/dev/null 2>&1; then
66
+ echo "FAIL [headless stop with a running task]: the reason advises starting the task again"; fail=$((fail + 1))
67
+ fi
68
+ expect quiet "the same task at the next stop" sdk-cli "$STATE" "$(payload s1 true "$one")"
69
+ expect quiet "the same task without the host retry flag" sdk-cli "$STATE" "$(payload s1 false "$one")"
70
+ both="[$(task t1 running),$(task t2 running 'Wait for the lane')]"
71
+ expect block "a new task in the same session" sdk-cli "$STATE" "$(payload s1 true "$both")" "Wait for the lane (t2)"
72
+ if printf '%s' "$OUT" | jq -e '.reason | contains("(t1)")' >/dev/null 2>&1; then
73
+ echo "FAIL [a new task in the same session]: an already reported task was listed again"; fail=$((fail + 1))
74
+ fi
75
+ expect block "another session with the same task id" sdk-cli "$STATE" "$(payload s2 false "$one")" "(t1)"
76
+ expect block "an SDK entrypoint" sdk-py "$STATE" "$(payload s3 false "$one")" "(t1)"
77
+
78
+ expect quiet "an interactive session" cli "$STATE" "$(payload s4 false "$one")"
79
+ expect quiet "no entrypoint" "" "$STATE" "$(payload s4 false "$one")"
80
+ expect quiet "a lookalike entrypoint" sdkx "$STATE" "$(payload s4 false "$one")"
81
+ done_task="[$(task t9 completed)]"
82
+ expect quiet "no task still running" sdk-cli "$STATE" "$(payload s5 false "$done_task")"
83
+ expect quiet "no background tasks field" sdk-cli "$STATE" '{"session_id":"s5","hook_event_name":"Stop","stop_hook_active":false}'
84
+ expect quiet "another hook event" sdk-cli "$STATE" "$(jq -nc --argjson t "$one" '{session_id:"s5",hook_event_name:"SubagentStop",stop_hook_active:false,background_tasks:$t}')"
85
+ expect notice "input that is not JSON" sdk-cli "$STATE" 'not json'
86
+
87
+ # Without usable state a repeat cannot be recognized, so the host retry flag bounds it.
88
+ expect block "no state, first stop" sdk-cli "$WORK/missing/dir" "$(payload s6 false "$one")" "(t1)"
89
+ expect quiet "no state, host retry" sdk-cli "$WORK/missing/dir" "$(payload s6 true "$one")"
90
+
91
+ # The helper repeats the entrypoint check, so it holds even when called without the wrapper.
92
+ HELPER="$(dirname "$HOOK")/host-input.py"
93
+ helper_quiet() { # helper_quiet <label> <entrypoint>
94
+ local out
95
+ out=$(printf '%s' "$(payload s9 false "$one")" | CLAUDE_CODE_ENTRYPOINT="$2" TMPDIR="$STATE" python3 "$HELPER" headless-background 2>"$ERR")
96
+ if [ -z "$out" ] && [ ! -s "$ERR" ]; then pass=$((pass + 1)); else echo "FAIL [$1]: helper spoke: $out"; fail=$((fail + 1)); fi
97
+ }
98
+ helper_quiet "the helper alone in an interactive session" cli
99
+ helper_quiet "the helper alone with a lookalike entrypoint" sdkx
100
+ out=$(printf 'not json' | CLAUDE_CODE_ENTRYPOINT=cli TMPDIR="$STATE" python3 "$HELPER" headless-background 2>"$ERR")
101
+ if [ -z "$out" ] && [ ! -s "$ERR" ]; then pass=$((pass + 1)); else echo "FAIL [the helper alone, interactive, bad input]: $out"; fail=$((fail + 1)); fi
102
+
103
+ # The wrapper leaves interactive sessions before the helper: no Python start, no notice about it.
104
+ BASH_BIN="$(command -v bash)"; mkdir -p "$WORK/no-python"; ln -s "$(command -v dirname)" "$WORK/no-python/dirname"
105
+ no_python() { # no_python <label> <entrypoint> <expected output or empty>
106
+ local out
107
+ out=$(printf '%s' "$(payload s10 false "$one")" | CLAUDE_CODE_ENTRYPOINT="$2" PATH="$WORK/no-python" "$BASH_BIN" "$HOOK" 2>"$ERR")
108
+ if [ "$out" = "$3" ] && [ ! -s "$ERR" ]; then pass=$((pass + 1)); else echo "FAIL [$1]: got '$out'"; fail=$((fail + 1)); fi
109
+ }
110
+ no_python "an interactive session without Python" cli ""
111
+ no_python "a headless session without Python" sdk-cli \
112
+ '{"systemMessage":"Background task check unavailable: Python input normalizer missing."}'
113
+
114
+ many="[$(for i in 1 2 3 4 5 6 7; do task "m$i" running; printf ','; done | sed 's/,$//')]"
115
+ expect block "a long task list" sdk-cli "$STATE" "$(payload s7 false "$many")" "(m5)" "and 2 more"
116
+ if printf '%s' "$OUT" | jq -e '.reason | test("\\(m[67]\\)")' >/dev/null 2>&1; then
117
+ echo "FAIL [a long task list]: a task beyond the first five was listed"; fail=$((fail + 1))
118
+ fi
119
+ hundred=$(printf 'x%.0s' $(seq 100))
120
+ long="[$(task c1 running "${hundred}yyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyy")]"
121
+ expect block "a long description" sdk-cli "$STATE" "$(payload s11 false "$long")" "${hundred} (c1)"
122
+
123
+ # Identifiers from the input never name a path: hostile ones stay inside the state root.
124
+ hostile="[$(task '../../../evil-task' running)]"
125
+ expect block "hostile identifiers, first stop" sdk-cli "$STATE" "$(payload '../../escape' false "$hostile")" "(../../../evil-task)"
126
+ expect quiet "hostile identifiers, next stop" sdk-cli "$STATE" "$(payload '../../escape' false "$hostile")"
127
+ if find "$WORK" -name '*evil*' -o -name '*escape*' | grep -q .; then
128
+ echo "FAIL [hostile identifiers]: an input identifier became a path"; fail=$((fail + 1))
129
+ else pass=$((pass + 1)); fi
130
+
131
+ # Two stops racing for the same new task: the claim is atomic, so exactly one blocks.
132
+ race="$(payload race false "[$(task r1 running)]")"
133
+ for i in 1 2 3 4; do
134
+ (printf '%s' "$race" | CLAUDE_CODE_ENTRYPOINT=sdk-cli TMPDIR="$STATE" bash "$HOOK" > "$WORK/race.$i" 2>/dev/null) &
135
+ done
136
+ wait
137
+ blocks=$(cat "$WORK"/race.* | grep -c '"decision": "block"')
138
+ if [ "$blocks" = 1 ]; then pass=$((pass + 1)); else echo "FAIL [racing stops]: $blocks blocks for one task"; fail=$((fail + 1)); fi
139
+
140
+ # A state root that is a symlink is refused, so the guard falls back to the host retry flag.
141
+ mkdir -p "$WORK/sym" "$WORK/elsewhere"; ln -s "$WORK/elsewhere" "$WORK/sym/ccl-skill-loading-$(id -u)"
142
+ expect block "symlinked state root, first stop" sdk-cli "$WORK/sym" "$(payload s12 false "$one")" "(t1)"
143
+ expect quiet "symlinked state root, host retry" sdk-cli "$WORK/sym" "$(payload s12 true "$one")"
144
+ if [ -z "$(ls -A "$WORK/elsewhere")" ]; then pass=$((pass + 1)); else echo "FAIL [symlinked state root]: state written through the link"; fail=$((fail + 1)); fi
145
+ multiline="[$(task n1 running "$(printf 'line one\nline two')")]"
146
+ expect block "a description with a line break" sdk-cli "$STATE" "$(payload s8 false "$multiline")" "line one line two (n1)"
147
+
148
+ echo "pass=$pass fail=$fail"
149
+ [ "$fail" = 0 ] || exit 1
150
+ echo "test_headless_background_stop_ok"
@@ -128,6 +128,53 @@ probe remind 'glab mr merge 123 --yes; glab help mr merge' 'Merged !123'
128
128
  # a successful-looking string response still reminds
129
129
  probe remind 'glab mr merge 123 --yes' 'Merged! https://.../merge_requests/123'
130
130
 
131
+ # --- Reminder TEXT contract: the injected text is what the agent acts on right
132
+ # after the merge, so it must route to the canonical teardown section and
133
+ # carry the guards a lossy digest once dropped — the ignored-artifact scan
134
+ # that must exit 0 before any worktree removal, and the release-name branch
135
+ # exception — instead of claiming a single exception. ---
136
+ reminder_text=$(jq -nc --arg c 'gh pr merge 45 --merge' --arg r 'Merged' \
137
+ '{tool_input:{command:$c},tool_response:$r,cwd:"/tmp"}' | bash "$HOOK" \
138
+ | jq -r '.hookSpecificOutput.additionalContext // empty')
139
+ text_has() { # <label> <needle>
140
+ if printf '%s' "$reminder_text" | grep -Fq -- "$2"; then pass=$((pass+1))
141
+ else fail=$((fail+1)); printf 'FAIL [reminder text lacks %s] %s\n' "$1" "$2" >&2; fi
142
+ }
143
+ text_lacks() { # <label> <needle>
144
+ if printf '%s' "$reminder_text" | grep -Fq -- "$2"; then
145
+ fail=$((fail+1)); printf 'FAIL [reminder text still carries %s] %s\n' "$1" "$2" >&2
146
+ else pass=$((pass+1)); fi
147
+ }
148
+ if [ -n "$reminder_text" ]; then pass=$((pass+1))
149
+ else fail=$((fail+1)); echo 'FAIL [no reminder text extracted]' >&2; fi
150
+ TEARDOWN_REF='worktree-isolation/references/merge-and-teardown.md'
151
+ text_has 'canonical teardown pointer' "$TEARDOWN_REF"
152
+ text_has 'ignored-artifact scan before removal' 'status --ignored -s'
153
+ text_has 'scan exit-0 requirement' '必须 exit 0'
154
+ text_lacks 'single-exception claim' '唯一例外'
155
+ # A guard is its meaning, not one keyword: every needle must sit on the same
156
+ # line, so dropping the "keep" sense while a keyword survives still fails.
157
+ line_has_all() { # <label> <needle>...
158
+ local label="$1"; shift
159
+ if printf '%s\n' "$reminder_text" | awk -v n="$#" 'BEGIN{for(i=1;i<=n;i++) want[i]=ARGV[i]; ARGC=1}
160
+ {hit=1; for(i=1;i<=n;i++) if (index($0, want[i])==0) hit=0; if (hit) found=1}
161
+ END{exit found?0:1}' "$@"; then pass=$((pass+1))
162
+ else fail=$((fail+1)); printf 'FAIL [reminder text lacks %s on one line]\n' "$label" >&2; fi
163
+ }
164
+ line_has_all 'release-name branch kept' '不删的例外' '分支名含 release' '要删由用户显式指名'
165
+ line_has_all 'permanent/integration branch kept' '不删的例外' '源分支本身是永久/集成分支'
166
+ line_has_all 'external-side-effect wait' '未完成的外部副作用任务' '等它完成再清'
167
+ # The scan guards the removal, so it must come first — an order the substring
168
+ # checks above cannot see.
169
+ scan_at=$(printf '%s\n' "$reminder_text" | grep -n -F 'status --ignored -s' | head -1 | cut -d: -f1)
170
+ remove_at=$(printf '%s\n' "$reminder_text" | grep -n -F 'git worktree remove' | head -1 | cut -d: -f1)
171
+ if [ -n "$scan_at" ] && [ -n "$remove_at" ] && [ "$scan_at" -lt "$remove_at" ]; then pass=$((pass+1))
172
+ else fail=$((fail+1)); echo "FAIL [reminder text lacks scan ordered before removal] scan=${scan_at:-none} remove=${remove_at:-none}" >&2; fi
173
+ # The pointer must resolve: the canonical file exists in this tree and still
174
+ # carries the teardown section heading the pointer names.
175
+ if grep -Fq '## 收尾:' "$SCRIPT_DIR/../skills/$TEARDOWN_REF" 2>/dev/null; then pass=$((pass+1))
176
+ else fail=$((fail+1)); echo "FAIL [teardown pointer dangles] skills/$TEARDOWN_REF lacks '## 收尾:'" >&2; fi
177
+
131
178
  printf 'remind-post-merge-cleanup tests: pass=%d fail=%d\n' "$pass" "$fail"
132
179
  [ "$fail" -eq 0 ] || exit 1
133
180
  echo "test_remind_post_merge_cleanup_ok"
@@ -68,6 +68,8 @@ const OPENCODE_HOOK_BINDINGS = Object.freeze({
68
68
  "owner-dispatch-stop.sh": "event:session.idle/session.status",
69
69
  "skill-extraction-gate-stop.sh": "event:session.idle/session.status",
70
70
  "proposed-next-stop.sh": "event:session.idle/session.status",
71
+ // Inert here: it acts only for a Claude Code sdk entrypoint, which runHook never passes on.
72
+ "headless-background-stop.sh": "event:session.idle/session.status",
71
73
  })
72
74
 
73
75
  type HookJson = {
@@ -105,10 +107,13 @@ function runHook(root: string | null, script: keyof typeof OPENCODE_HOOK_BINDING
105
107
  if (!existsSync(path) || !lstatSync(path).isFile() || lstatSync(path).isSymbolicLink()) {
106
108
  return { status: "missing", message: `${script} is missing from the OpenCode hook runtime` }
107
109
  }
110
+ // A Claude Code entrypoint inherited from a parent process does not describe this OpenCode
111
+ // session, so no hook here may act on it.
112
+ const { CLAUDE_CODE_ENTRYPOINT: _inheritedEntrypoint, ...inherited } = process.env
108
113
  const result = spawnSync("bash", [path], {
109
114
  cwd,
110
115
  encoding: "utf8",
111
- env: { ...process.env, CLAUDE_PLUGIN_ROOT: root },
116
+ env: { ...inherited, CLAUDE_PLUGIN_ROOT: root },
112
117
  input: JSON.stringify(payload),
113
118
  maxBuffer: 256 * 1024,
114
119
  timeout,
@@ -664,6 +669,7 @@ export const CclSkills = async (context: {
664
669
  runHook(hooksRoot, "owner-dispatch-stop.sh", stopPayload, directory, 10_000),
665
670
  runHook(hooksRoot, "skill-extraction-gate-stop.sh", stopPayload, directory, 15_000),
666
671
  runHook(hooksRoot, "proposed-next-stop.sh", stopPayload, directory, 5_000),
672
+ runHook(hooksRoot, "headless-background-stop.sh", stopPayload, directory, 5_000),
667
673
  ]
668
674
  const reasons = results
669
675
  .filter((result) => result.output?.decision === "block" && typeof result.output.reason === "string")
@@ -95,6 +95,7 @@ Controller-authored, non-shared edits stay in the owning local workflow. Use thi
95
95
 
96
96
  5. Clean up after integration.
97
97
  - After the reviewed ref is non-interim landed, sync the target checkout. If landing used a mutable source-branch merge without an atomic reviewed-ref guard, do not clean up yet; keep the branch as evidence.
98
+ - Every worktree removal below must run the pre-removal scan that the closeout section (`## 收尾`) of `worktree-isolation/references/merge-and-teardown.md` requires: `git -C <worktree-path> status --ignored -s` must exit 0, and a failed scan counts as no scan. Copy costly gitignored outputs the worker produced (long-running results, collected data, trained artifacts) into the target checkout before removal and drop regenerable ones (dependencies, build and test outputs, caches, logs); a worktree that `git status` reports clean still loses its gitignored files on `git worktree remove`. Record the scan exit status and each kept or dropped entry in `cleanup_proof`.
98
99
  - For local-commit-only handoff, prefer fast-forward or merge that preserves the exact reviewed SHA. First sync or fetch the target checkout, verify its current `HEAD` contains the reviewed SHA, and prove the landed content matches the reviewed delta with `git range-diff` or a scoped `git diff` showing no unexpected delta. Only after that content-equivalence proof and `git merge-base --is-ancestor <reviewed-sha> HEAD` both pass in the synced target checkout may the controller remove the isolated worktree and delete the local branch with `git branch -d` (not `-D`) from the same checkout. There is no remote source branch to prune.
99
100
  - For pushed-branch/MR handoff, check the platform merge mode first. If it is squash, rebase, cherry-pick, or another rewrite mode, use the non-ancestor path below. If the reviewed branch tip equals the recorded reviewed SHA and is a target ancestor, sync the target checkout, verify it contains the reviewed SHA, re-fetch and verify the local source branch tip still equals the reviewed SHA, remove the isolated worktree, delete the local branch with `git branch -d` from that synced target checkout, confirm the remote source branch was removed by the platform or delete it explicitly, run `git fetch --prune`, run `git worktree prune`, and verify `git worktree list` no longer shows the path and the remote source branch is absent. If the branch has post-review commits, review those commits before cleanup or keep the branch.
100
101
  - For squash, rebase, cherry-pick, or platform rewrites where ancestry does not prove integration, require a concrete non-ancestor proof before partial cleanup: platform merge state tied to the reviewed SHA, or an explicit patch-equivalence check such as `git range-diff` / scoped `git diff` showing the reviewed hunks are present on target with no unexpected differences. Under the default no-force-delete policy, remove only the clean worktree, keep the local branch, and report `branch retained: non-ancestor integration`; do not claim full cleanup unless `worktree-isolation` explicitly permits that merge-mode cleanup.
@@ -104,6 +104,10 @@ On hosts providing a current final message, `proposed-next-stop.sh` returns one
104
104
 
105
105
  The hook recognizes declarations, not authorization or actual task completion, and cannot force the model to follow through. OpenCode idle does not expose the required final-message evidence; its Stop behavior remains unverified.
106
106
 
107
+ A session nothing re-invokes gets no second chance at the in-flight-work rule below: Claude Code reports an `sdk` entrypoint for `claude -p` and the SDKs, the session ends at the stop, and the host stops its background tasks seconds later, so work waiting on them is lost.
108
+
109
+ - In such a session, `headless-background-stop.sh` blocks such a stop once per background task still running and names them: wait in the foreground for any the request depends on, without starting it again, and stop the rest with TaskStop or say why they can be dropped. It reads only the hook input, so it also fires when session persistence is off and the transcript-reading reminders above cannot run. Interactive sessions, which the host re-invokes when a task finishes, are left alone.
110
+
107
111
  ## Gate triggers and outcome contract
108
112
 
109
113
  An eligible next slice comes from an explicit status/task/acceptance source or active user continuation, is low-risk, local-only/already-authenticated, in accepted scope, clearly owned and verifiable with existing commands. It needs no destructive action, external purchase/financial commitment, production access, legal/compliance/product-strategy decision or high-impact architecture choice. Existing configured internal developer-self-use metered model/tool accounts are not an external purchase. Apply the following conditions to each action.
@@ -38,12 +38,19 @@ A per-line WIP branch stays local/private until the normal shared-branch-push an
38
38
 
39
39
  ## Closeout Cleanup
40
40
 
41
- At closeout, after the work lands or is abandoned, clean up the worktree and private branch:
41
+ At closeout, after the work lands or is abandoned, clean up the worktree and private branch. The canonical procedure is the closeout section (`## 收尾`) of `worktree-isolation/references/merge-and-teardown.md`, which also holds the integration evidence and the remote-branch rules; its guards apply to every removal:
42
+
43
+ - Before removing the worktree directory, you must run `git -C <path> status --ignored -s` from the primary checkout. It must exit 0; a failed scan counts as no scan, so stop and find the cause instead of reading empty output as nothing to keep. `git worktree remove` without `--force` still deletes gitignored files.
44
+ - Judge each listed entry by what it costs to recreate: drop regenerable outputs (dependency directories, build and test outputs, caches, logs) and copy costly ones (long-running results, collected data, trained artifacts) back to the primary checkout before removal. When unsure, treat an entry as costly.
45
+ - If a task with unfinished external side effects, such as a migration or a deployment, still runs from the worktree, wait for it to finish; never kill it to clean up.
46
+ - Never pass `--force` to `git worktree remove` or use `git branch -D`: a refusal means unmerged or uncommitted work. An abandoned line's branch usually fails `-d`; keep it and report it.
47
+ - Never delete a permanent or integration branch, or any branch whose name contains `release`, as part of this cleanup.
42
48
 
43
49
  ```bash
44
- git worktree remove <path>
50
+ git -C <path> status --ignored -s # must exit 0; copy costly ignored outputs out first
51
+ git worktree remove <path> # no --force
45
52
  git worktree prune
46
- git branch -d <line-branch>
53
+ git branch -d <line-branch> # -d, not -D
47
54
  ```
48
55
 
49
56
  ## Owner Routing
@@ -25,7 +25,7 @@ The read side already defends against oversized files (chunked reads under ~200
25
25
  - An existing over-limit reference is frozen per invariant 4: shrink or stay level; growth blocks. Additions to a frozen reference are funded by consolidating existing text in the same file.
26
26
  - Append-only ledgers are structurally excluded: `references/source-register.md` grows by contract (append-only, supersede-by-pointer, rows never edited), so a line cap would block the ledger discipline itself; the gate skips it and prints a visibility token when it is over the figure. Residual risk, accepted under the same trusted-contributor model as the entrypoint gate: a prose file named `source-register.md` would dodge the cap — review owns that shape.
27
27
  - A new reference over 100 lines must be structured with `##` sections so chunked reads and greps can navigate it; a heading-less long file draws an advisory token (never a block). A table-of-contents list is optional — section structure is the invariant, not a TOC block.
28
- - **Funding an addition by trimming prose means editing text that may be pinned — resolve the pins before rewriting, not after.** The ratchet's per-file freeze makes every addition to a legacy surface a rewrite of something else in the same file, and load-bearing sentences are pinned in two places: declaratively in `../../skill-extraction-workflow/scripts/contract-anchors.tsv`, which the fast repo gate checks, and as `grep -Fq` assertions inside owner suites, whose break a full lane run reports half an hour later. Read BOTH for the file you are about to trim — `awk -F'\t' '$2 == "<path>"' skills/skill-extraction-workflow/scripts/contract-anchors.tsv` lists the registry rows pinning it, and `grep -rn 'grep -Fq' skills/*/scripts/*.sh` finds the suite assertions — because a funded trim that silently retired two pinned wait-contract obligations was reported by the slow lane only, long after the edit. A pinned sentence may be reworded only together with whatever pins it, in the same landing.
28
+ - **Funding an addition by trimming prose means editing text that may be pinned — resolve the pins before rewriting, not after.** The ratchet's per-file freeze makes every addition to a legacy surface a rewrite of something else in the same file, and load-bearing sentences are pinned by literal from places that keep multiplying: `../../skill-extraction-workflow/scripts/contract-anchors.tsv`, the always-on pairs registered in `scripts/check-sync-pointers.sh`, assertions in owner and hook suites (some through helpers such as `assert_contains`, so a search for one assertion idiom misses them), and `file:<path>#<anchor>` firing paths of `source-register.md` rows, which `scripts/register-firing-path-resolution.rb` resolves against the tree. Because the classes grow, search by path rather than by class: you must search every non-Markdown file in the repository for the path you are about to trim or move text out of (`git grep -n -F '<path>' -- ':!*.md' ':!specs/'` lists the registries, gates and suites that name it; round evidence under `specs/` records history and pins nothing) and list the ledger anchors with `grep -o 'file:<path>#[^;|]*' skills/skill-extraction-workflow/references/source-register.md`; after the edit, run `bash skills/skill-extraction-workflow/scripts/check-ccl-skills.sh .`, which resolves the contract anchors, the sync registry and every ledger anchor, plus the suites the path search named. A funded trim that silently retired two pinned wait-contract obligations was reported by the slow lane only, long after the edit, and a relocation whose pin read covered only two of these sources moved a ledger-anchored rule out of its entrypoint. A pinned sentence may be reworded or moved only together with whatever pins it, in the same landing.
29
29
  - Authoring anti-patterns (verified against the official skill-authoring checklist, see verdicts below): time-sensitive facts outside an explicit old-patterns section; inconsistent terminology for one concept; abstract examples where a concrete input/output pair fits; Windows-style paths; unexplained constants; scripts that defer error handling to the model instead of solving it.
30
30
 
31
31
  ## Retirement and relocation signal (usage census)
@@ -33,7 +33,7 @@ The read side already defends against oversized files (chunked reads under ~200
33
33
  The ratchet only stops growth; it never says *what* to retire or relocate, and "not pulling its weight" is an author's opinion until something is measured. Three independent lines converge on the same instrument: context-evolution methods keep per-bullet usage counters (helpful/harmful marks) and prune or merge on them rather than on a single monolithic rewrite that collapses detail; trajectory-distillation work places broadly applicable procedure in the root document and *lower-frequency* detail in auxiliary files, and finds joint consolidation over many traces stronger than order-dependent one-lesson-at-a-time edits; and the official skill-authoring guidance tells authors to watch how the agent navigates a skill — a bundled file the agent never accesses is unnecessary or poorly signaled, one it reads on every run belongs in the entrypoint. The repo instantiation:
34
34
 
35
35
  - `scripts/reference-access-census.sh [--skill <name>] [--days <n>]` reads the host's own agent transcripts (Claude Code and Codex session logs; both consume this tree) and prints, per `SKILL.md`/reference file, how many sessions in the window mentioned it and when it was last touched (the last-touched column must be the newest touching transcript's date, never the oldest or an unparsed value). Counts only — no transcript text, prompts, absolute log paths, or session ids — so the output is safe for a private charter; it is still per-host data and never lands in the shared tree.
36
- - **Placement by firing frequency, not only by kind.** The entrypoint's content-placement rule sorts by kind (trigger / routing / core workflow / non-negotiables stay; detail moves). Add the frequency axis: a rule that fires on a narrow source class or correction type (one client surface, one correction shape, one artifact kind) is *low-frequency detail* even when it is non-negotiable, and must live verbatim in the reference the entrypoint already points at, with a one-bullet summary that keeps the load-bearing obligations inline (`rule-consolidation.md` condensing rule). Ledger `file:` anchors and script pins name a path, so a pinned phrase stays where it is — enumerate them before choosing what moves.
36
+ - **Placement by firing frequency and firing time, not only by kind.** The entrypoint's content-placement rule sorts by kind (trigger / routing / core workflow / non-negotiables stay; detail moves). Add the frequency axis: a rule that fires on a narrow source class or correction type (one client surface, one correction shape, one artifact kind) is *low-frequency detail* even when it is non-negotiable, and must live verbatim in the reference the entrypoint already points at, with a one-bullet summary that keeps the load-bearing obligations inline (`rule-consolidation.md` condensing rule). Ledger `file:` anchors and script pins name a path, so a pinned phrase stays where it is — enumerate them before choosing what moves. The time axis works the same way: a section that applies only at a later point of the same task (push and merge, post-merge teardown) is dead weight at activation and sits mid-context by the time it applies, so move it verbatim to a reference and name that reference where the later point is reached — a firing-point table in the entrypoint and the hook that fires there. Moving a section also breaks every pointer that names it by its old location, such as `SKILL.md「X」` or a link to the entrypoint followed by the section title; no gate resolves those, so before landing you must search the Markdown for the name of each moved section, outside `specs/`, and retarget every hit. Exact command conventions and destructive-operation guards that must be applied verbatim stay inline. A hook or always-on text injected at a firing point is what drives the action at that moment, so it must name the canonical section it summarizes, and what it restates must not read as complete when it is not — a digest that calls one exception the only one, or a command list that skips a mandatory pre-step, out-votes the canonical rule loaded earlier; pin the injected guards and the pointer target in the hook's own suite.
37
37
  - **Advisory, never a gate** (Goodhart, same as the health roll-up): a count that becomes a target gets gamed by mentioning files. A zero-session reference is a relocation/merge *candidate* that still owes the zero-loss obligation map; a high-share reference is a promotion candidate, not an automatic move. Pair the census with the closeout cost row in `extraction-quickstart.md` §4 so that a round records what it cost and what it retired, and the next round can tell whether the corpus is shrinking toward the cap or only holding level.
38
38
 
39
39
  - **A mention count is not an open count — classify HOW the file is reached before relocating on a census figure.** The census matches the path anywhere in a transcript line, so a file that a gate names in its own output, that an agent appends to, or that a bounded `sed`/`tail`/`grep` touches for one row scores the same as one an agent loads whole. Before a census figure justifies a split, a promotion, or a retirement, the read shape must be counted in the same window — whole-file reads versus bounded reads versus gate echo versus writes — and let the whole-read count, not the mention share, carry the read-side cost argument. Observed: the append-only ledger scored the package's highest mention share, and the read-shape count showed whole-file loads in a small minority of those mentions, with bounded reads and gate echo making up the rest — the split its share seemed to demand would have bought nothing on the read side.
@@ -113,8 +113,13 @@ Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds
113
113
  | [OpenAI, *GPT-4.1 Prompting Guide*](https://developers.openai.com/cookbook/examples/gpt4-1_prompting_guide) | conflicting instructions tend to resolve to the one nearer the end; instructions at both ends of long context beat either alone; check-conflicts-first; a single clear sentence usually steers | same class; explicitly model-generation-bound ("GPT-4.1 tends to…") |
114
114
  | [OpenAI, *GPT-5.1 Prompting Guide*](https://cookbook.openai.com/examples/gpt-5/gpt-5-1_prompting_guide) | check-conflicts-first; a published metaprompt recipe for finding contradictions in your own system prompt | same class |
115
115
  | [RECAST](https://arxiv.org/html/2505.19030) | joint satisfaction degrades sharply as constraint count grows — and it is already low at small counts. The metric for "all of them at once" is the paper's **OSR** (§4.1: "the HSR of all constraints, both rule-based and model-based, that are successfully satisfied simultaneously"). Across all 29 model rows of Table 1 the **ceiling** on OSR is **25.0** at Level 1, falling to **19.0 / 13.0 / 13.5** at Levels 2–4, whose constraint counts are **5 / 10 / 15 / all** (§B.4). So at five constraints no model held the whole set more than about a quarter of the time, and by fifteen none exceeded ~13% | benchmark paper proposing its own dataset and method — a low baseline flatters the contribution; these are one hard benchmark's order of magnitude, not a usage failure rate, and its constraints are generation-task instruction constraints rather than preconditions of a procedure, so transfer is by analogy. **Citation corrected 2026-08 against Table 1:** the widely quoted 39.75% is the **Average column of the single best-by-average row** (Gemini-2.5-Pro) — the arithmetic mean of that row's twelve MSR/RSR/OSR cells (sum 477; 477/12 = 39.75, confirmed) — **not** an all-constraints-satisfied rate, and it must not be cited as one. The two orderings differ: Gemini leads on Average while Qwen3-235B-A22B holds the highest Level-1 OSR, so do not carry "best model" across from one column to the other. Cite the OSR ceilings above |
116
-
117
116
  | [IFScale](https://arxiv.org/abs/2507.11538) | instruction-following accuracy degrades as instruction density rises (500 keyword-inclusion instructions; best frontier model 68% at max density across 20 models / 7 providers); three degradation shapes correlated with model size and reasoning; a **bias toward earlier instructions** (primacy), and omission as a distinct error category | benchmark paper on a synthetic keyword-inclusion task — density and primacy transfer by analogy only; it measures a list of independent constraints, not a procedure's preconditions; note it reports primacy where the vendor guides report recency, so position is a bias with no single direction |
117
+ | [AgentIF](https://arxiv.org/abs/2505.16944) | real agent system prompts (707 instructions from 50 applications) average 1,723 words and 11.9 constraints; the best model followed fewer than 30% of them perfectly, adherence fell as instructions grew longer, and condition and tool constraints were the hardest | benchmark paper on production-shaped prompts — the closest public analog to a loaded skill body, but it scores one response per instruction, not a multi-step procedure |
118
+ | [Context Length Alone Hurts LLM Performance Despite Perfect Retrieval](https://arxiv.org/abs/2510.05381) | with every relevant token retrievable, accuracy still fell 13.9–85% as input grew within the advertised window, even when the added tokens were whitespace or masked; reciting the relevant evidence before answering recovered part of it | 5 models on math, QA and code — supports "length itself costs attention", not any particular budget figure |
119
+ | [SkillsBench](https://arxiv.org/abs/2602.12670) (§6, App. F, Tables 8–9) | curated Skills raised an 87-task pass rate from 33.9% to 50.5% across 18 model–harness configurations; by SKILL.md size bucket the lift was compact +19.0, standard +21.5, detailed +14.5 and comprehensive +0.7 points, and by Skills per task 1 / 2–3 / ≥4 gave +18.0 / +19.0 / +10.1; where Skills hurt, the paired-trajectory audit found a heavyweight pipeline crowding out a simpler path, a generic recipe displacing a stronger native strategy, or a brittle framework the agent could not debug; in the three configurations that tested self-generated Skills, they fell below the no-Skills baseline | observational on size and count (task and Skill are confounded, the comprehensive bucket has five tasks) on terminal-based tasks whose authors flag long-horizon workflows as possibly not transferring — a direction to test here, not a budget |
120
+ | [SkillJuror](https://arxiv.org/abs/2606.11543) | holding the knowledge fixed, a concise root that points to on-demand resources versus one flat file raised distinct resources touched per trajectory from 1.18 to 3.85 and added 17 verifier-passing trials of 410 (+4.1%); it helped where resources guide implementation, checking or repair and was weaker where success hinges on exact output conventions, numerical thresholds or long artifact pipelines | controlled variants on 82 tasks — organization changes how the agent searches before it changes outcomes, and the outcome gain is small and task-dependent |
121
+ | [Skill availability and presentation granularity](https://arxiv.org/abs/2605.31408) | having the Skill raised task-mean pass rates by 18–36 points, while moving the same guidance between abstraction levels or adding one worked example changed them by −6.7 to +1.3 points (both abstraction contrasts' bootstrap intervals cross zero) | controlled, 30 tasks × 2 models × 5 trials — do not expect a rewording of the same knowledge to move outcomes measurably |
122
+ | [Anthropic, *Prompting best practices*](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices) | newer Claude models follow the system prompt more closely, so prompts written to stop tools or skills under-triggering may now over-trigger; the fix is to dial back aggressive emphasis ("CRITICAL: You MUST use this tool when…" → "Use this tool when…") | vendor guidance bound to named model generations; not yet measured against this repository's process-heavy skills |
118
123
 
119
124
  **Why recency is a hazard, not a rule.** Vendor guidance reports that models *tend to follow* whichever instruction sits later — an observation about behaviour, not a licence to resolve conflicts by position. Two ways position becomes dangerous if read as a rule: a later permissive line beats an earlier stricter one (directly contradicting `Conflict Resolution`, which keeps the stricter data-loss/security/contract guard); and text embedded in **untrusted data** — a diff under review, a retrieved document, tool output — sits later within the same authority level and would win by placement alone, which is prompt injection with extra steps. Treat recency as a bias to design against: put the load-bearing rule where the decision happens, and never let placement confer authority.
120
125
 
@@ -59,7 +59,7 @@ When the workflow is consumed as a plugin (or any read-only install mechanism),
59
59
  - **Discover the source URL from the install.** The plugin/marketplace registration that shipped with the install records it: the marketplace clone's git remote (`git -C <marketplace-clone> remote get-url origin`), the marketplace `source.url` in the host's plugin config/registry, and the install URLs documented in the consuming project's README each help identify the canonical repo. Cross-check them rather than trusting one — a README can be stale or absent, and an install/marketplace URL can differ from the canonical source repo (e.g. http vs ssh, or a mirror).
60
60
  - **Author in one standing checkout — clone once, reuse.** Clone that URL a single time (or reuse an existing checkout) and reuse it for every later extraction. Do not clone per change.
61
61
  - **Never edit the consumption copies.** The plugin cache and the marketplace-managed clone are overwritten on update; edits there are lost and never reach the repo. Author only in your own checkout, then land through the normal review/MR gates.
62
- - **Isolate each change with a worktree, not a clone.** Make every extraction in a dedicated worktree off the standing checkout — worktrees share the object store and are not full repo copies — and remove it once the change lands (`git worktree remove`). This meets the concurrent-session isolation rule without proliferating repo copies.
62
+ - **Isolate each change with a worktree, not a clone.** Make every extraction in a dedicated worktree off the standing checkout — worktrees share the object store and are not full repo copies — and remove it once the change lands, following the closeout section (`## 收尾`) of `worktree-isolation/references/merge-and-teardown.md`: you must scan its gitignored outputs before removal (`git -C <worktree> status --ignored -s`, which must exit 0) and copy any output that is costly to recreate back to the standing checkout, because `git worktree remove` deletes gitignored files without asking. This meets the concurrent-session isolation rule without proliferating repo copies.
63
63
  - **Accumulation is consumption-side, not authoring.** Plugin caches on some hosts keep one directory per installed version, so old version dirs can linger after updates — unrelated to authoring. Prune them with the host's plugin prune command if disk matters.
64
64
 
65
65
  After landing, the change reaches every install through the normal update path (marketplace refresh + plugin update); the exact update command lives in the consuming project's README, not in the shared skill tree.
@@ -129,6 +129,8 @@ Result inflation 没有 MAST 对应——它是 context / 成本问题,不是
129
129
 
130
130
  **用**:本 skill `skill-extraction-workflow` 自身的 R0 / drafting 类大改动;产品 skill 的 routing 调整。
131
131
 
132
+ **落地形态**:`scripts/skill-paired-eval.py` + `eval/paired-tasks/`(`make eval-paired`)。同一合成 git 世界里交错跑 off / base / candidate 三臂,按世界状态判分,实现上面的冻结任务库、同轮起跑、成本列与逐断言读数;每个检查都带能把它判红的坏轨迹(`--check-oracles`)。它的数不是什么、隔离怎么核验,以脚本头为准。
133
+
132
134
  ### 3.2 Golden trace(中量)
133
135
 
134
136
  为每个 stable skill 沉淀 1-2 个 **golden agent trace**:
@@ -768,3 +768,12 @@ The pending classification above is superseded by the executed source comparison
768
768
  | The receipt lock check signals from the writer's own lock call, so no wait decides its result | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | updated | Owner key `code-review/SKILL.md`. Supersedes the previous row's claim of scheduling independence: the next delta review found that a writer paused between signalling its start and reaching the lock lets both mutants pass once the one-second wait expires. The writer now signals from inside its exclusive lock call on the receipt directory, and the holder rewrites the count only after that signal; unlock and cleanup run in `finally`, the writer is a daemon and every wait is bounded. Against copies of the controller, each run with and without a forced two-second pause in the writer: the unchanged controller passed 8 of 8; removing the lock and taking a shared lock each failed 8 of 8 on the missing signal; reading the prior receipt before the lock failed 8 of 8 on the count. |
769
769
  | The receipt lock check asserts the exclusive lock at the receipt read and replace instead of racing writers | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | updated | Owner key `code-review/SKILL.md`. Supersedes the two rows above for the check's design. Three successive delta reviews each found a controller variant that a racing-writer check let through: writers that did not overlap, a writer paused past a wait, a lock released before the read. The class was closed by changing the method instead of patching a fourth time. When the controller opens the prior receipt and when it replaces it, the check requires that a second open of the receipt directory cannot take even a shared lock, and then that the count increments the value read. Against copies of the controller, three runs each, the unchanged controller passed; removing the lock, a shared lock, a lock on another descriptor, reading before the lock, unlocking before the read and unlocking before the replace each failed on the recorded lock states. |
770
770
  | The receipt lock check also requires the lock to be held without interruption between the read and the replace, and reads the stored receipt back | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_review_gate.sh | updated | Owner key `code-review/SKILL.md`. The delta review of the previous row's check, the sixth conclusive run in the worktree and the first to return the continuation checkpoint, found that probes at the read and the replace pass a controller that unlocks and relocks in between, and that the check trusted the returned receipt over the stored one. The check now records every lock call during the controller's call and requires none between the read and the replace, and it compares the stored receipt with the returned one. Against copies of the controller, three runs each: the unchanged controller passed; unlocking and relocking between the read and the replace failed on the recorded lock calls, and a stored count or start time that differs from the returned one, applied to the last call only, failed on the stored receipt; the six variants from the previous row still failed. |
771
+ | The always-on pointers to the worktree merge protocol and teardown section resolve to the package reference that now carries that text, the entrypoint sentence that forwards to it is pinned, and the gate still blocks when either side loses its anchor | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_sync_pointers.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. Constructed scenarios, runs recorded in `specs/161-firing-point-loading/evidence/validation.md`: with the merge and teardown sections relocated and the previous registry, `check-sync-pointers.sh` exits 1 on exactly the three worktree pins; with the registry pointed at `worktree-isolation/references/merge-and-teardown.md` it passes; removing the protocol anchor or renaming the teardown heading in that reference blocks the matching pins only. Removing the forwarding sentence from `worktree-isolation/SKILL.md` left the gate green until the new forwarding pair, which then blocks it alone. The suite's fixtures mutate the reference and the forwarding sentence. |
772
+ | Trimming or moving text out of a file first finds its pins by path, every non-Markdown file outside round evidence plus the ledger's file anchors, and runs the fast validator after the edit | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#must search every non-Markdown file in the repository for the path | updated | Owner key `skill-extraction-workflow/SKILL.md`. Recorded incident in `specs/161-firing-point-loading/evidence/validation.md`: a relocation that read only contract anchors and suite assertions moved the shared-branch update bullet out of `worktree-isolation/SKILL.md`, where six register rows anchor their firing paths; the full validator stopped on `register_firing_path_unresolved` and the base tree did not. On the base tree the path search lists the sync registry, the `assert_contains` pin and the sync suite fixtures, and the ledger search lists the six anchors, so the move would have been seen before it was made; the bullet stays in the entrypoint. |
773
+ | A section used only at a later point of the task loads at that point, and text injected there names its canonical section without presenting a partial list as complete | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#must name the canonical section it summarizes | updated | Owner key `skill-extraction-workflow/SKILL.md`. Applied to `worktree-isolation`: entrypoint 38,112 to 13,964 bytes with 81 moved lines verbatim in two references. The post-merge reminder called the integration-branch case the only exception and skipped the ignored-artifact scan; against the previous hook the nine new text assertions in `hooks/test_remind_post_merge_cleanup.sh` fail and the 57 existing probes pass, against the new hook all 67 pass, and ten single mutations each fail the assertions they target. In paired runs of a cleanup task on a synthetic repository, neither the previous nor the relocated skill lost the expensive ignored file or the release branch in five runs each, while runs without the plugin lost them in four and three of five; both recorded in `specs/161-firing-point-loading/evidence/validation.md`. |
774
+ | A worktree closeout that removes the worktree first scans its gitignored outputs from the primary checkout, treats a failed scan as no scan, copies costly outputs out, keeps forced removal and protected branches off the path, and names the canonical teardown section | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/worktree-mechanics.md#Before removing the worktree directory, you must run | updated | Owner key `product-rd-workflow/SKILL.md`. The Closeout Cleanup listed only `git worktree remove`, `git worktree prune` and `git branch -d`, and `git worktree remove` without `--force` deletes gitignored files and exits 0. Each of the nine product delivery rows of the teardown pin family in the shared implementation-gates fixture fails alone against the text before this round and passes on the new one, and the teardown pin walk reds each under an applied mutation. In paired cleanup runs without the plugin, with the recipe committed into a synthetic repository, agents following the previous text lost the expensive ignored file in 2 of 5 runs and stopped to ask in the other 3, while agents following the new text scanned first, kept the file and finished in 5 of 5; with the plugin loaded, both texts kept the file in 5 of 5. Recorded in `specs/162-teardown-recipe-guard/evidence/validation.md`. |
775
+ | Every worker worktree removal in a delegated handoff runs the ignored-output scan first, requires it to succeed, copies costly outputs out, and records the scan in the cleanup proof | `multi-agent-delegation` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/multi-agent-delegation/references/multi-agent-delegation-playbook.md#Every worktree removal below must run the pre-removal scan | updated | Owner key `multi-agent-delegation/SKILL.md`. Handoff step 5 removed the worker worktree on three paths, one of them for a clean worktree, with no scan, so gitignored results a worker produced were deleted with it. Each of the five delegation rows of the teardown pin family fails alone against the text before this round and passes on the new one, and the teardown pin walk reds each under an applied mutation; recorded in `specs/162-teardown-recipe-guard/evidence/validation.md`. |
776
+ | A Markdown surface that restates worktree removal carries the ignored-output scan with its exit-0 requirement on the same line and names the canonical teardown reference, and one added later without them fails the shared fixture | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. Against the tree before this round the sweep names exactly the extraction lifecycle note and the product delivery closeout (no scan) and the worktree handbook (no pointer); on the new tree it passes. `test_teardown_guard_pins.sh` plants decoy surfaces and failure probes and each reds the sweep for the stated reason (no scan, no exit-0 requirement on the scan line, no pointer or a file-name-only pointer, a new top-level directory, a file name starting with a newline, an unreadable file, directory or index, a canonical file hidden from git, a missing repository, a failed classification), while a compliant surface, a package-relative pointer inside the canonical package, a prune-only mention, ignored paths, round records, evaluation inputs and a register row stay green; nineteen sabotaged copies of the fixture each make the walk fail. The sweep keys on the command, so a removal described only in prose is pinned per surface; recorded in `specs/162-teardown-recipe-guard/evidence/validation.md`. |
777
+ | Moving a section retargets every Markdown pointer that names it by its old location, found by searching Markdown for the moved section's name, because no gate resolves those pointers | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/attention-budget-ratchet.md#you must search the Markdown for the name of each moved section | updated | Owner key `skill-extraction-workflow/SKILL.md`. Recorded miss in `specs/162-teardown-recipe-guard/evidence/validation.md`: after round 161 (`specs/161-firing-point-loading/`) moved the worktree merge and teardown sections into references, the worktree handbook still named the entrypoint for three of them. On that tree the path search the recipe asked for, over non-Markdown files, returns registries and suites only, while searching Markdown outside `specs/` for each moved section's name returns the three handbook lines. |
778
+ | A behavior-shaping skill change is measured before it lands with paired runs on synthetic tasks (no plugin, a base export, a candidate export) graded by the world state the agent leaves; each run's isolation is read from its own events and an instruction-file canary, and every task check is proven able to fail by a bad trajectory | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/skill-paired-eval.py | updated | Owner key `skill-extraction-workflow/SKILL.md`. §3.1 of `harness-patterns-and-eval.md` had no tool; rounds 161 and 162 measured with private scripts that passed their whole environment to the tested agent, and in round 162 the no-plugin arm read the installed plugin cache and followed it (`specs/162-teardown-recipe-guard/evidence/validation.md`). `skill-paired-eval.py` strips inherited `CLAUDE*` and `GIT_*` variables, gives each run a private copy of the frozen export and checks its init event for that path, and calibrates a per-sample instruction-file canary that appears with project instructions allowed and stayed absent under the run flags. Each of 108 applied mutations to its graders, isolation checks, guards and batch lifecycle reds the named test in `test_skill_paired_eval.py` on a disposable copy while an unrelated test stays green, and every check of the four tasks in `eval/paired-tasks/` fails on a bad trajectory. First batch, 54 samples on Opus 5.5 comparing the published 0.18.11, main and no plugin: with either plugin the release branch was kept in 5 of 5 cleanup runs against 0 of 5 without it, and the edit on the default branch went into a worktree in 5 of 5 against 0 of 5 (p = 0.008 each); no base-candidate comparison separated, and every sample passed the structural isolation checks. That batch disabled session persistence, so the plugin's hooks that read the session transcript did not run in it; the tool now keeps persistence on and invalidates a run without a transcript. Recorded in `specs/163-paired-behavior-eval/evidence/validation.md`. |
779
+ | A session nothing re-invokes ends at the stop and the host stops its background tasks seconds later, so the rule that awaiting your own work is not a stop condition needs a firing point that reads the running tasks from the Stop hook's input instead of the transcript: one block per still-running task, for SDK entrypoints only | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#blocks such a stop once per background task still running | updated | Owner key `product-rd-workflow/SKILL.md`; changed reference `product-rd-workflow/references/pre-final-continuation-gate.md`. In a paired batch of an earlier round, six of seven background reviews started by plugin-arm headless runs were killed after the turn ended, leaving three changes uncommitted and three reviews unread, while the plugin's transcript-reading reminder could not run. In a pre-registered headless probe with session persistence on, the plugin from `main` lost the work in 3 of 3 runs, its background task killed, and the plugin with the guard in none, waiting and then finishing. The guard's suite has 30 cases, and each of 22 applied mutations turns its named case red. Recorded in `specs/164-headless-background-stop/evidence/validation.md`. |