@ccoalm/ccl-skills 0.18.8 → 0.18.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (17) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-policy.md +1 -1
  2. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh +12 -2
  3. package/dist/assets/marketplace/plugins/ccl-skills/hooks/host-input.py +89 -2
  4. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +5 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh +38 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh +11 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_proposed_next.py +111 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +12 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +7 -3
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +4 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +5 -5
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +8 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +10 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +31 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/hook-authorization.md +3 -1
  16. package/dist/assets/release.json +18 -18
  17. package/package.json +1 -1
@@ -41,7 +41,7 @@ The compact session-start entry routes to these execution details when the relev
41
41
 
42
42
  几条贯穿原则(任何任务都适用;详则在 owner 技能里):
43
43
  - **上下文恢复是 agent 的工作**:恢复/继续/复盘/判断既有工作时,先读 SessionStart 的 `<agent-context-recovery>`(若宿主提供),再核 repo 契约、当前 Git、项目状态/任务持久件、最小相关 session/memory 片段、commit 与 CI/test 证据;读取历史片段前必须确认其 repo root / cwd / remote 属于当前仓(全局 session/db 存在不等于相关);启动快照只用于定位,结论仍要 live refresh。能从本地证据恢复的事实不得让用户重述。只有方向/重大取舍、缺失权限或凭据、不可逆动作、以及本地证据确实不存在时才打断用户。
44
- - **自主决策,别把可判定的问题推给用户**:开发中只在真实阻塞时停下问人——缺失的凭据/授权;本地证据确实没有的事实;上面安全硬边界管的动作(无恢复的破坏性/不可逆、prod/客户数据、目标外的合并/发布);推翻用户既定方向;证据无法裁决的重大产品取舍。其余都由你定:设计期安全 4 问自答写进方案、安全自检、选 owner 技能/模块/方案、测试与命名、目标内的下一步;提交/推送/合并按「硬纪律 1」目标授权判断。写明假设后继续。「owner」指技能或代码 owner,不是要用户指派的人。某一步被阻塞时先做完其余独立工作,再在结尾用 `proposed-next: blocked:` 说明具体阻塞。
44
+ - **自主决策,别把可判定的问题推给用户**:开发中只在真实阻塞时停下问人——缺失的凭据/授权;本地证据确实没有的事实;上面安全硬边界管的动作(无恢复的破坏性/不可逆、prod/客户数据、目标外的合并/发布);推翻用户既定方向;证据无法裁决的重大产品取舍。其余都由你定:设计期安全 4 问自答写进方案、安全自检、选 owner 技能/模块/方案、测试与命名、目标内的下一步;提交/推送/合并按「硬纪律 1」目标授权判断。写明假设后继续。判据:阻塞必须是只有用户能提供的东西(证据无法裁决的决定、凭据或访问授权、目标之外的许可、所有可读来源里都没有的事实);你自己能执行的下一步——修复及其测试评审、推分支开 MR/PR、重跑或重试失败/超时/结果不明的检查、查资料——不论花多少时间和次数都不是阻塞。以下几种说法都过不了这个判据:「你只让我调查/研究」——查原因、线上问题、测试挂了这类失败目标,根因证实后默认包含窄修复、回归测试、评审、推送功能分支并向开发目标开/更新 MR/PR,用户的明确限制照样约束(「只查/先别改」「先别推」、费用或次数上限),仓库契约标为先确认的区域、破坏性操作、新采购、合并/部署/生产动作也照常停(交接摘要里「这一步只读」只约束那一步);「推送/开 MR 是对外动作」——推功能分支、开 MR/PR 是常规研发动作,不是发布,自审、外部评审和必需 CI 过了就在同一轮转 ready,不以「MR 保持 Draft」收尾;「等 CI 跑完」——自己 MR 的流水线自己轮询到结果;你自己提出的次数、轮数、停止线——它们是你的估计,不是用户限制;只有用户把它采纳为上限(亲口给数字、说「最多/只」、或明确接受为上限)才算,单纯同意去做(「ok」「好」)不算;用户中途的追问——答完继续已授权的动作,不是 status-only;能自己查到或构造的事实、日志、测试数据、之前用过的账号与环境——先自己查。真要停下要合并授权时,一次请求覆盖整个交付计划,别按次追加。「owner」指技能或代码 owner,不是要用户指派的人。某一步被阻塞时先做完其余独立工作,再在结尾用 `proposed-next: blocked:` 说明具体阻塞。
45
45
  - **用户主权**:AI 推荐、用户定。要改变用户既定方向时**始终先呈现+问,别径直下结论或代为决定**。你和另一个模型(codex 等)都同意也只是强信号、不是裁决。**仅当用户有既定方向、且你与第二模型都主张推翻它**(普通选项/口味/缺信息/评审 nit 不触发此结构):用户方向是默认、改动由模型举证,呈现时必须显式补两句——我们可能缺什么上下文、若改错代价是什么(详见 tighten-doc cross-model caveat)。
46
46
  - **无证据不声称完成**:本轮没亲手跑过验证、没读到通过输出,就不说"完成/修好/通过/没问题",缺证据如实说缺(详见 product-rd-workflow 验证门)。
47
47
  - **完整优先**:做完必要工作,不扩范围。报告/总结/进度说明不等于交付:结束前逐项核对用户请求和工作自身带出的后续项(失败检查、评审 finding、要同步的测试/文档;提交推送按「硬纪律 1」目标授权),能做的做完再收尾。阻塞交付的检查失败含基线问题,按 defect-diagnosis 诊断、安全修复、复测;真实阻塞才交回。
@@ -172,6 +172,7 @@ DENY_TEXT_AUTO="合并授权闸:该命令会开启 auto-merge / merge-when-pip
172
172
  DENY_TEXT_SPELL="合并授权闸:授权有效但命令拼写不满足一次性立即合并要求——glab 需显式 --auto-merge=false(pipeline 运行中裸 merge 会默认转 auto-merge),gh 需显式 --merge/--squash/--rebase 策略,merge REST API 需显式 -X PUT。授权未消费,按上述拼写改写命令直接重试即可(无需用户重新授权)。"
173
173
  DENY_TEXT_TARGET="合并授权闸:用户授权指向了特定 MR/PR 编号,但该命令的合并对象与之不符或无法识别。授权未消费——请显式点名该编号(如 glab mr merge <授权编号> --auto-merge=false --yes)后重试;确需合并其他 MR 请让用户重新授权。"
174
174
  DENY_TEXT_MULTI="合并授权闸:同一条命令内检测到多个平台合并调用。每条命令只放行一个合并——把命令拆开逐条执行(单个授权下每个合并由用户分别授权;批量授权下每条命令消费 1 个额度,无需用户再次回复)。"
175
+ DENY_TEXT_HELP="合并授权闸:命令带 -h/--help,只会打印帮助、不会合并,因此拒绝且不消费授权。查看帮助请用 glab help mr merge 或 gh help pr merge;正式合并时去掉帮助参数,用户已给的授权仍然有效。"
175
176
  DENY_TEXT_AMBIGUOUS="合并授权闸:检测到原始 HTTP 客户端、变更 method 的选项和 merge endpoint,但 method、transfer 边界、目标或动作数无法可靠关联。授权未消费——请改写成单个 curl/wget、单个明确 PUT method 和单个 merge URL 后重试。"
176
177
 
177
178
  # Missing legacy epoch files are compatible with pre-upgrade grants. Once
@@ -741,7 +742,7 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
741
742
  # gh explicit strategy; REST/GraphQL calls are immediate by API
742
743
  # semantics). mid: the merge target id when statically extractable
743
744
  # ("?" otherwise) — matched against a number-bound grant below.
744
- hit=0; auto=0; spell=SPELLBAD; mid="?"; multi_seg=0
745
+ hit=0; auto=0; spell=SPELLBAD; mid="?"; multi_seg=0; help=0
745
746
  if [ "$tool" = "glab" ]; then
746
747
  # `accept` is glab's documented alias of `mr merge` (same help text).
747
748
  if [ "${1:-}" = "mr" ] && { [ "${2:-}" = "merge" ] || [ "${2:-}" = "accept" ]; }; then
@@ -755,6 +756,7 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
755
756
  # value-taking flags: consume the value so it is not mistaken
756
757
  # for the MR id positional (`glab mr merge --sha abc 546`).
757
758
  --sha|-m|--message|--squash-message) [ $# -ge 2 ] && shift ;;
759
+ -h|--help) help=1 ;;
758
760
  -*) : ;;
759
761
  *)
760
762
  if [ "$mid" = "?" ] && [ -z "${id_seen:-}" ]; then
@@ -826,6 +828,7 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
826
828
  # value-taking flags: consume the value so it is not mistaken
827
829
  # for the PR id positional.
828
830
  -b|--body|-F|--body-file|-t|--subject|--match-head-commit|-A|--author-email) [ $# -ge 2 ] && shift ;;
831
+ -h|--help) help=1 ;;
829
832
  -*) : ;;
830
833
  *)
831
834
  if [ "$mid" = "?" ] && [ -z "${id_seen:-}" ]; then
@@ -876,7 +879,11 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
876
879
  # the id unresolvable so bound grants deny (unbound grants keep the
877
880
  # agent-side duty to target the discussed MR — documented residual).
878
881
  [ "$retarget" = 1 ] && mid="?"
879
- if [ "$auto" = 1 ]; then echo DENY_AUTO; else
882
+ # A help flag in flag position means the CLI prints help and merges
883
+ # nothing; deny it without touching any grant. A help token taken as a
884
+ # known flag's value never reaches here, and an unknown value-taking
885
+ # flag only makes this deny a real merge, never release one.
886
+ if [ "$help" = 1 ]; then echo DENY_HELP; elif [ "$auto" = 1 ]; then echo DENY_AUTO; else
880
887
  echo "DENY_MR $mid $spell"
881
888
  # A single segment carrying multiple aliased merge mutations emits a
882
889
  # second DENY_MR so the >1 exactly-one-per-command guard denies it.
@@ -1066,6 +1073,9 @@ fi
1066
1073
  if printf '%s\n' "$verdicts" | grep -q '^DENY_GIT_UNRESOLVED$'; then
1067
1074
  deny "$DENY_TEXT_GIT_UNRESOLVED"
1068
1075
  fi
1076
+ if printf '%s\n' "$verdicts" | grep -q '^DENY_HELP$'; then
1077
+ deny "$DENY_TEXT_HELP"
1078
+ fi
1069
1079
  if printf '%s\n' "$verdicts" | grep -q '^DENY_AUTO'; then
1070
1080
  deny "$DENY_TEXT_AUTO"
1071
1081
  fi
@@ -13,6 +13,7 @@ import re
13
13
  import shlex
14
14
  import stat
15
15
  import sys
16
+ import tempfile
16
17
 
17
18
  # Hook assets may be installed read-only; importing the optional state helper
18
19
  # must not create bytecode beside them.
@@ -626,6 +627,23 @@ def delivery_eligible(summary):
626
627
  or summary['continuation_contract_visible'])
627
628
 
628
629
 
630
+ # Stops kept surviving the recheck by restating the blocker in a new term, so
631
+ # the test is stated as an invariant (who can act) and the observed terms are
632
+ # only examples of restatements that fail it.
633
+ NOT_BLOCKERS = (
634
+ 'A blocker names something only the user can supply: a decision the evidence cannot settle, a credential '
635
+ 'or access grant, permission the goal does not cover, or a fact absent from every source you can read. '
636
+ 'A next step you can perform yourself is not a blocker, whatever it costs in time or runs: a fix with its '
637
+ 'tests and review, a branch push and MR/PR, marking that MR/PR ready once your own checks pass instead of '
638
+ 'leaving it in Draft, waiting on a CI run you can poll, a rerun or retry of a failed, timed-out or '
639
+ 'inconclusive check, or a lookup. Restatements observed to fail this test: "you only asked me to investigate" — a failure or '
640
+ 'diagnosis goal includes the verified fix, tests, review, branch push and MR/PR to the development target '
641
+ 'unless an explicit user limit says otherwise (diagnosis only, no push); "pushing or opening an MR is outward-facing" — a feature branch '
642
+ 'and its MR/PR are routine; a count, round or stop bar you proposed yourself, unless the user adopted it as '
643
+ 'a limit; "the check can only restart '
644
+ 'from scratch"; and facts, logs, test data or access you can find or reuse yourself. A clarifying question '
645
+ 'is not a status-only request: answer it, then continue. ')
646
+
629
647
  # One bounded recheck (host stop_hook_active) for stops that hand work back to
630
648
  # the user; it names the real blockers and grants no authority.
631
649
  DECISION_RECHECK = {'decision': 'block', 'reason': (
@@ -633,7 +651,8 @@ DECISION_RECHECK = {'decision': 'block', 'reason': (
633
651
  'Real blockers are: missing credentials or authority; a fact unavailable from local evidence; '
634
652
  'an action the safety rules gate (destructive or irreversible without recovery, production or '
635
653
  'customer data, merge or publication outside the goal); overturning an established user direction; '
636
- 'or a material product tradeoff the evidence cannot settle. An ordinary change needs no human review, '
654
+ 'or a material product tradeoff the evidence cannot settle. ' + NOT_BLOCKERS +
655
+ 'An ordinary change needs no human review, '
637
656
  'sign-off or risk owner: run the self-review and external review yourself. '
638
657
  'Small tests and routine development/test-environment operations within the authorized task '
639
658
  'run directly with configured accounts; do not ask for per-run approval or invent a cost cap. '
@@ -649,10 +668,77 @@ DECISION_RECHECK = {'decision': 'block', 'reason': (
649
668
  'This reminder supplies no new goal or authorization.')}
650
669
 
651
670
 
671
+ # Reader-facing documents edited in a session owe the tighten-doc closeout
672
+ # readback (the routing rule says so), yet it was skipped in 26 of 29 observed
673
+ # sessions. Agent-facing files are excluded: skill bodies, contracts, memory
674
+ # and scratch or temporary paths.
675
+ READER_DOC_SUFFIXES = ('.md', '.mdx', '.rst')
676
+ AGENT_DOC_NAMES = {'skill.md', 'agents.md', 'claude.md', 'memory.md'}
677
+ AGENT_DOC_DIRS = {'memory', '.claude', '.codex', '.git', 'node_modules', 'skills', 'agent-context',
678
+ 'scratchpad'}
679
+
680
+
681
+ def reader_docs(edit_paths, cwd):
682
+ base = os.path.realpath(cwd) if isinstance(cwd, str) and cwd else None
683
+ temp_roots = tuple(os.path.realpath(root) + os.sep
684
+ for root in {tempfile.gettempdir(), '/tmp', '/private/tmp', '/var/folders'})
685
+ found = []
686
+ for path in edit_paths:
687
+ if not isinstance(path, str) or not path.lower().endswith(READER_DOC_SUFFIXES):
688
+ continue
689
+ real = os.path.realpath(path)
690
+ if base and (real == base or real.startswith(base + os.sep)):
691
+ parts = os.path.relpath(real, base).split(os.sep)
692
+ elif real.startswith(temp_roots):
693
+ continue
694
+ else:
695
+ parts = real.split(os.sep)
696
+ if parts[-1].lower() in AGENT_DOC_NAMES or any(p.lower() in AGENT_DOC_DIRS for p in parts[:-1]):
697
+ continue
698
+ found.append(real)
699
+ return found
700
+
701
+
702
+ def doc_closeout_note(payload):
703
+ path, cwd = payload.get('transcript_path'), payload.get('cwd')
704
+ if not isinstance(path, str) or not path:
705
+ return ''
706
+ cwd = cwd if isinstance(cwd, str) else os.getcwd()
707
+ try:
708
+ summary = transcript(path, cwd)
709
+ except TranscriptTruncated:
710
+ summary = context_transcript(path, cwd)
711
+ except (OSError, ValueError):
712
+ return ''
713
+ if any(skill.split(':')[-1] == 'tighten-doc' for skill in summary['completed_skills']):
714
+ return ''
715
+ docs = reader_docs(summary['edit_paths'], cwd)
716
+ if not docs:
717
+ return ''
718
+ names = ', '.join(sorted({os.path.basename(doc) for doc in docs})[:5])
719
+ return ('Document closeout: this session edited reader-facing documents ({}) without loading '
720
+ 'tighten-doc. Load it and run its closeout readback on those documents before finishing; '
721
+ 'the substance stays as the owning skill decided.'.format(names))
722
+
723
+
652
724
  def proposed_next(payload):
653
725
  if (not isinstance(payload, dict) or payload.get('hook_event_name') != 'Stop'
654
726
  or payload.get('stop_hook_active') is not False):
655
727
  return None
728
+ result = delivery_reminder(payload)
729
+ try:
730
+ note = doc_closeout_note(payload)
731
+ except Exception: # advisory: a failed document check never costs the reminder
732
+ note = ''
733
+ if not note:
734
+ return result
735
+ if not result:
736
+ return {'decision': 'block', 'reason': note + ' Then end with the same proposed-next: line. '
737
+ 'This reminder supplies no new goal or authorization.'}
738
+ return {'decision': 'block', 'reason': note + ' ' + result['reason']}
739
+
740
+
741
+ def delivery_reminder(payload):
656
742
  final = payload.get('last_assistant_message')
657
743
  if not isinstance(final, str) or not final.strip() or machine_artifact(final):
658
744
  return None
@@ -674,7 +760,8 @@ def proposed_next(payload):
674
760
  'authorized and runnable, execute it now instead of waiting for another continue message. '
675
761
  'For unrun, failed or inconclusive checks, continue available diagnosis, research, safe repair '
676
762
  'and retesting; a report alone does not complete implementation. Respect explicit stop, '
677
- 'planning-only and status-only requests. If a user decision or missing authority/resource '
763
+ 'planning-only and status-only requests. ' + NOT_BLOCKERS +
764
+ 'If a user decision or missing authority/resource '
678
765
  'prevents action, report the concrete blocker; do not invent work or bypass a failed gate. '
679
766
  'This reminder supplies no new goal or authorization.')}
680
767
  path = payload.get('transcript_path')
@@ -72,6 +72,11 @@ masked=$(printf '%s' "$cmd" | sed -E \
72
72
  # NON-fire — `glab mr merge` / `gh pr merge` is the near-universal agent merge
73
73
  # path, and the human-readable cleanup rule in worktree-isolation SKILL.md +
74
74
  # bootstrap covers EVERY merge path regardless of this reminder.
75
+ # `gh help pr merge` / `glab help mr merge` print help (the merge guard's help
76
+ # denial points there). Remove only those literal invocations, never a prefix,
77
+ # so a real merge before or after them in the same command still matches.
78
+ masked=$(printf '%s' "$masked" | sed -E \
79
+ 's/(glab|gh)[[:space:]]+help[[:space:]]+(mr|pr)[[:space:]]+(merge|accept)([[:space:]]|$)/ /g')
75
80
  printf '%s' "$masked" | grep -Eq \
76
81
  'glab[[:space:]]([^&|;]*[[:space:]])?mr[[:space:]]+(merge|accept)([[:space:]]|$)|gh[[:space:]]([^&|;]*[[:space:]])?pr[[:space:]]+merge([[:space:]]|$)' \
77
82
  || exit 0
@@ -1035,6 +1035,44 @@ for invalidation in '停止' '改成另一个功能'; do
1035
1035
  done
1036
1036
  unset RACE_SENT RACE_REACHED RACE_RESUME REAL_MV
1037
1037
 
1038
+ # A help probe merges nothing, so it must never consume a grant. Observed: an
1039
+ # agent added --auto-merge=false to `glab mr merge --help` to pass the spelling
1040
+ # check, the probe consumed the one-shot grant, and the real merge was denied.
1041
+ # Earlier cases leave epoch files for this session; clear them so each probe
1042
+ # meets a valid grant rather than an epoch mismatch (which also denies).
1043
+ rm -f "$VAUTH_DIR/$VSID.epoch" "$VAUTH_DIR/$VSID.grant-epoch"
1044
+ for help_cmd in 'glab mr merge --help --auto-merge=false' 'glab mr merge 123 -h --auto-merge=false --yes' \
1045
+ 'gh pr merge 45 --merge --help' 'gh pr merge --squash -h'; do
1046
+ varm
1047
+ probe_sid deny "$FEAT_CWD" "$VSID" "$help_cmd"
1048
+ sentinel_state present "help probe kept the grant: $help_cmd"
1049
+ done
1050
+ varm
1051
+ reason_help=$(jq -nc --arg c 'glab mr merge --help --auto-merge=false' --arg w "$FEAT_CWD" --arg s "$VSID" \
1052
+ '{tool_input:{command:$c},cwd:$w,session_id:$s}' | TMPDIR="$tmp" bash "$GUARD")
1053
+ if printf '%s' "$reason_help" | grep -q 'glab help mr merge'; then pass=$((pass+1)); else
1054
+ fail=$((fail+1)); echo 'FAIL help denial must name the non-merge help form' >&2; fi
1055
+ rm -f "$VAUTH_DIR/$VSID"
1056
+ # A value-taking flag swallows a following --help: the command still merges.
1057
+ varm
1058
+ probe_sid allow "$FEAT_CWD" "$VSID" 'glab mr merge 123 -m --help --auto-merge=false --yes'
1059
+ sentinel_state absent 'a --help message value is a real merge and consumes the grant'
1060
+ probe allow "$FEAT_CWD" 'glab help mr merge'
1061
+ probe allow "$FEAT_CWD" 'gh help pr merge'
1062
+ # A help probe compounded with a real merge denies the whole command and keeps
1063
+ # the grant; a quoted --help message value is masked and stays a real merge.
1064
+ varm
1065
+ probe_sid deny "$FEAT_CWD" "$VSID" 'glab mr merge 123 --auto-merge=false --yes && glab mr merge --help'
1066
+ sentinel_state present 'help compounded with a merge keeps the grant'
1067
+ rm -f "$VAUTH_DIR/$VSID"
1068
+ varm
1069
+ probe_sid allow "$FEAT_CWD" "$VSID" 'glab mr merge 123 -m "--help" --auto-merge=false --yes'
1070
+ sentinel_state absent 'a quoted --help message is a real merge and consumes the grant'
1071
+ # Without any grant a help probe is denied and creates no grant.
1072
+ rm -f "$VAUTH_DIR/$VSID"
1073
+ probe_sid deny "$FEAT_CWD" "$VSID" 'gh pr merge 45 --merge --help'
1074
+ sentinel_state absent 'a help probe without a grant creates nothing'
1075
+
1038
1076
  if [ "$fail" -ne 0 ]; then
1039
1077
  echo "test_guard_merge_authorization: FAIL pass=$pass fail=$fail" >&2
1040
1078
  exit 1
@@ -213,6 +213,17 @@ done
213
213
  send '批量合并 3'
214
214
  send '继续'
215
215
  if [ ! -f "$SENT" ]; then pass=$((pass+1)); else fail=$((fail+1)); echo 'FAIL legacy batch still clears on neutral prompt' >&2; fi
216
+ # A host task notification reaches this hook with no field that tells it apart
217
+ # from typed text, so it is handled as a user message: it revokes single and
218
+ # counted grants (a stop typed in its markup must still revoke) and never arms.
219
+ note=$'<task-notification>\n<task-id>abc123</task-id>\n<status>completed</status>\n<summary>Background command "wait for CI" completed (exit code 0)</summary>\n</task-notification>'
220
+ send '批量合并 3'
221
+ send "$note"
222
+ if [ ! -f "$SENT" ]; then pass=$((pass+1)); else fail=$((fail+1)); echo 'FAIL a notification must revoke a counted grant' >&2; fi
223
+ send '合并'
224
+ send $'<task-notification>\n<summary>先别合并</summary>\n</task-notification>'
225
+ if [ ! -f "$SENT" ]; then pass=$((pass+1)); else fail=$((fail+1)); echo 'FAIL a stop inside notification markup must revoke' >&2; fi
226
+ expect_not_armed $'<task-notification>\n<summary>合并</summary>\n</task-notification>'
216
227
  git -C "$tmp/repo" remote set-url origin 'https://user:password@example.invalid/team/project.git'
217
228
  expect_not_armed '完成并合并 PR #123'
218
229
  git -C "$tmp/repo" remote set-url origin 'git@example.invalid:team/project.git'
@@ -1,5 +1,6 @@
1
1
  #!/usr/bin/env python3
2
2
  """Synthetic native Stop payloads; no real host state or conversations."""
3
+ import importlib.util
3
4
  import json
4
5
  import os
5
6
  from pathlib import Path
@@ -7,6 +8,7 @@ import shutil
7
8
  import subprocess
8
9
  import tempfile
9
10
  import unittest
11
+ from unittest.mock import patch
10
12
 
11
13
  ROOT = Path(__file__).resolve().parents[1]
12
14
 
@@ -341,6 +343,115 @@ class ProposedNextTests(unittest.TestCase):
341
343
  self.assertIn('supplies no new goal or authorization', result['reason'])
342
344
  self.assertEqual(self.run_hook(dict(payload, stop_hook_active=True)), {})
343
345
 
346
+ def test_diagnosis_scope_and_question_turn_are_named_in_both_reminders(self):
347
+ # Observed stops after a verified root cause: the agent read "missing
348
+ # authority" as "you only asked me to investigate", and a clarifying
349
+ # question as a status-only request. Both reminders must close those terms.
350
+ for events in ([], self.claude_load()):
351
+ self.events(events)
352
+ for text in (
353
+ 'Fixing it changes shared code and opens an MR, beyond the investigation you asked for.\n'
354
+ 'proposed-next: blocked: limit fix, consistency test and MR — waiting for you to authorize the code change',
355
+ '这一轮你只问了一个问题,我只解释了现状。\n'
356
+ 'proposed-next: blocked: 上限修复、补测试、提 MR——改共享仓库需要你确认',
357
+ 'Root cause verified.\nproposed-next: open a branch, fix the limit, add the test and open the MR',
358
+ 'Required CI passed; the advisory review timed out and its evidence was cleared.\n'
359
+ 'proposed-next: blocked: complete the CI review — no resume handle; a retry restarts from scratch',
360
+ 'Both MRs pushed; required CI is still running, the MRs stay in Draft.\n'
361
+ 'proposed-next: wait for CI to finish and for your merge confirmation'):
362
+ with self.subTest(text=text):
363
+ result = self.run_hook(dict(self.payload, last_assistant_message=text))
364
+ self.assert_block(result)
365
+ self.assertIn('you only asked me to investigate', result['reason'])
366
+ self.assertIn('failure or diagnosis goal includes the verified fix', result['reason'])
367
+ # Closing the terms must not widen past an explicit user
368
+ # limit: a "fix locally, do not push" instruction still binds.
369
+ self.assertIn('unless an explicit user limit says otherwise (diagnosis only, no push)',
370
+ result['reason'])
371
+ self.assertIn('clarifying question is not a status-only request', result['reason'])
372
+ self.assertIn('outward-facing" — a feature branch and its MR/PR are routine', result['reason'])
373
+ self.assertIn('stop bar you proposed yourself', result['reason'])
374
+ self.assertIn('A blocker names something only the user can supply', result['reason'])
375
+ self.assertIn('a rerun or retry of a failed, timed-out or inconclusive check', result['reason'])
376
+ self.assertIn('marking that MR/PR ready once your own checks pass', result['reason'])
377
+ self.assertIn('waiting on a CI run you can poll', result['reason'])
378
+ self.assertIn('supplies no new goal or authorization', result['reason'])
379
+
380
+ def doc_edit(self, relative, tool_id='doc', tool='Write'):
381
+ target = str(self.root / relative)
382
+ return [
383
+ {'type': 'assistant', 'message': {'content': [{'type': 'tool_use', 'id': tool_id,
384
+ 'name': tool, 'input': {'file_path': target, 'content': 'x'}}]}},
385
+ {'type': 'user', 'message': {'content': [{'type': 'tool_result',
386
+ 'tool_use_id': tool_id, 'is_error': False, 'content': 'written'}]}}]
387
+
388
+ def test_reader_doc_edit_without_tighten_doc_gets_one_closeout_reminder(self):
389
+ # Observed: 26 of 29 sessions that edited plans, specs, READMEs or
390
+ # handoff documents never loaded tighten-doc before finishing.
391
+ status = 'Plan updated.\nproposed-next: none — status only'
392
+ for relative in ('docs/plans/rollout.md', 'README.md', 'handoffs/state.md', 'specs/9-x/plan.md'):
393
+ with self.subTest(relative=relative):
394
+ self.events(self.doc_edit(relative))
395
+ result = self.run_hook(dict(self.payload, last_assistant_message=status))
396
+ self.assert_block(result)
397
+ self.assertIn('tighten-doc', result['reason'])
398
+ self.assertIn(Path(relative).name, result['reason'])
399
+ self.assertIn('supplies no new goal or authorization', result['reason'])
400
+ self.assertEqual(self.run_hook(dict(self.payload, last_assistant_message=status,
401
+ stop_hook_active=True)), {})
402
+
403
+ def test_doc_reminder_is_quiet_after_tighten_doc_or_for_agent_files(self):
404
+ status = 'Plan updated.\nproposed-next: none — status only'
405
+ self.events(self.claude_load('tighten-doc') + self.doc_edit('docs/plans/rollout.md'))
406
+ self.assertEqual(self.run_hook(dict(self.payload, last_assistant_message=status)), {})
407
+ for relative in ('skills/x/SKILL.md', 'AGENTS.md', 'CLAUDE.md', 'memory/note.md',
408
+ '.claude/notes.md', 'src/app.py', 'notes.txt'):
409
+ with self.subTest(relative=relative):
410
+ self.events(self.doc_edit(relative))
411
+ self.assertEqual(self.run_hook(dict(self.payload, last_assistant_message=status)), {})
412
+
413
+ def test_doc_reminder_joins_a_continuation_reminder(self):
414
+ self.events(self.doc_edit('docs/plans/rollout.md'))
415
+ result = self.run_hook(dict(self.payload,
416
+ last_assistant_message='proposed-next: run the remaining local checks'))
417
+ self.assert_block(result)
418
+ self.assertIn('execute it now', result['reason'])
419
+ self.assertIn('tighten-doc', result['reason'])
420
+
421
+ def test_doc_reminder_covers_each_file_edit_tool(self):
422
+ # Edits are seen through the file-edit tool calls the transcript records;
423
+ # shell writes are outside this check by design.
424
+ for tool in ('Edit', 'MultiEdit'):
425
+ with self.subTest(tool=tool):
426
+ self.events(self.doc_edit('docs/handoff.md', tool=tool))
427
+ result = self.run_hook()
428
+ self.assert_block(result)
429
+ self.assertIn('handoff.md', result['reason'])
430
+
431
+ def test_unreadable_doc_path_never_costs_the_delivery_reminder(self):
432
+ # The document check is advisory; a path it cannot resolve (an embedded
433
+ # NUL makes realpath raise) must not replace the continuation reminder
434
+ # with the "reminder unavailable" notice.
435
+ self.events(self.doc_edit('docs/plans/roll\x00out.md'))
436
+ result = self.run_hook(dict(self.payload,
437
+ last_assistant_message='proposed-next: run the remaining local checks'))
438
+ self.assert_block(result)
439
+ self.assertIn('execute it now', result['reason'])
440
+
441
+ def test_doc_check_failure_never_costs_the_delivery_reminder(self):
442
+ # The document check is advisory: whatever it raises, the continuation
443
+ # reminder it would have joined is still returned.
444
+ self.events(self.doc_edit('docs/plans/rollout.md'))
445
+ spec = importlib.util.spec_from_file_location('doc_check_probe', self.hooks / 'host-input.py')
446
+ module = importlib.util.module_from_spec(spec)
447
+ spec.loader.exec_module(module)
448
+ payload = dict(self.payload, last_assistant_message='proposed-next: run the remaining local checks')
449
+ with patch.object(module, 'reader_docs', side_effect=RuntimeError('unexpected')):
450
+ result = module.proposed_next(payload)
451
+ self.assert_block(result)
452
+ self.assertIn('execute it now', result['reason'])
453
+ self.assertNotIn('tighten-doc', result['reason'])
454
+
344
455
  def test_quoted_actions_do_not_turn_a_status_handoff_into_work(self):
345
456
  self.events(self.claude_load())
346
457
  for suffix in ('\n> proposed-next: deploy', '\n```text\nproposed-next: deploy\n```'):
@@ -113,6 +113,18 @@ probe_json remind 'glab mr merge 123 --yes' '{"stdout":"Merged !123"}'
113
113
  # --- --help / -h is not a merge → quiet ---
114
114
  probe quiet 'gh pr merge --help'
115
115
  probe quiet 'glab mr merge -h'
116
+ # The help subcommand form, which the merge guard's help denial points to,
117
+ # prints help and merges nothing.
118
+ probe quiet 'gh help pr merge' 'Merge a pull request on GitHub.'
119
+ probe quiet 'glab help mr merge' 'Merges a merge request.'
120
+ probe quiet 'gh help pr merge 2>&1 | grep -- --match-head-commit' '--match-head-commit SHA'
121
+ # A real merge beside a help lookup in one command still reminds, whichever
122
+ # comes first: only the help invocation itself is set aside.
123
+ probe remind 'gh help pr merge >/dev/null; gh pr merge 45 --merge' 'Merged'
124
+ probe remind 'glab help mr merge && glab mr merge 123 --yes' 'Merged !123'
125
+ probe remind 'gh pr merge 45 --merge # see gh help pr merge' 'Merged'
126
+ probe remind 'gh pr merge 45 --merge $(gh help pr merge >/dev/null)' 'Merged'
127
+ probe remind 'glab mr merge 123 --yes; glab help mr merge' 'Merged !123'
116
128
  # a successful-looking string response still reminds
117
129
  probe remind 'glab mr merge 123 --yes' 'Merged! https://.../merge_requests/123'
118
130
 
@@ -16,7 +16,7 @@ Diagnose and fix from evidence; route prevention to product, architecture, devel
16
16
  - Do not call a workaround the fix unless the owner explicitly accepts the tradeoff and residual risk is recorded.
17
17
  - Do not start broad refactoring while the cause is unknown. Isolate and fix first; refactor after the behavior is understood.
18
18
  - Do not stop at "this line was wrong" when the defect reveals a missing contract, guardrail, test, review check, or skill rule.
19
- - Do not state or act on a root-cause verdict — even as a confident aside — before you have read the failing owner's own evidence with your own eyes (assertion diff for a test, stack/exception for a crash, trace/log slice for a production symptom, source only when it is itself the failing artifact). Until then, label every cause as a hypothesis and name the evidence that would confirm or reject it. This applies to your OWN analysis, not only to LLM-proposed causes. Mitigation is exempt: you may roll back, flag-off, or shed traffic from symptoms while cause stays marked unknown — what is forbidden is choosing or applying a *fix* as though a cause is proven.
19
+ - Do not state or act on a root-cause verdict — even as a confident aside — before you have read the failing owner's own evidence with your own eyes (assertion diff for a test, stack/exception for a crash, trace/log slice for a production symptom, source only when it is itself the failing artifact). Until then, label every cause as a hypothesis and name the evidence that would confirm or reject it. This applies to your OWN analysis, not only to LLM-proposed causes. Mitigation is exempt: you may roll back, flag-off, or shed traffic from symptoms while cause stays marked unknown — what is forbidden is choosing or applying a *fix*, or handing the user a decision or tradeoff that rests on the cause, as though a cause is proven. While a falsifying check is runnable, run it before asking anyone to choose.
20
20
 
21
21
  ## Phase A: Diagnose
22
22
 
@@ -47,7 +47,7 @@ Diagnose and fix from evidence; route prevention to product, architecture, devel
47
47
  - **A test that passes alone and fails in the suite must be bisected over the tests that run before it** — only for a failure that reproduces on every run under a fixed serial order (parallel or intermittent failures keep the failing schedule and validate each kept or dropped subset over repeated runs per the flaky rule, or route to concurrency diagnosis): halve the preceding set in order and keep a failing half; when neither half fails alone, remove one chunk at a time and keep the reduced set whenever the failure persists without that chunk, then halve the chunk size and repeat until every remaining chunk is needed (a minimal ordered polluting subsequence) or the shared fixture/state is found; the same reduction isolates a failing input, config, or dataset when no commit range exists (moves in `references/diagnosis-playbook.md`).
48
48
  - Identify whether the failure is in product logic, contract mapping, persistence, cache, async processing, dependency behavior, runtime config, release state, or test setup.
49
49
  - If the failure appears only in tests or CI, classify the test evidence before changing code: deterministic assertion, fixed external data, live infrastructure, random/log-only behavior, long sleep, allow-failure gate, generated/vendor test, or deploy/build-only pipeline.
50
- - **A red CI pipeline/job is not by itself a code/dependency defect — read the failing job's own trace (not the red/green summary) and classify the cause before touching code.** Refining the classes above into why-CI-is-red-but-code-may-be-fine: (a) **trigger-variant artifact** — when the same job runs under more than one trigger-scoped config (branch/push vs merge-request vs manual/scheduled), the trigger can resolve different default variables or a different checked-out ref, so a red on a non-gating trigger may be benign — but conclude that ONLY after confirming the *same failing check* ran and is green on the gating path (a gating pipeline that is overall green yet never runs the failing check does not clear it; if that check's coverage is unique to the non-gating trigger — e.g. a scheduled/manual-only suite — treat it as genuine (d), not a variant artifact); (b) **retriable infra flake** — e.g. a shared-runner lock collision: confirm per the flaky rule below (rerun N times + record the ratio; a 100%-reproducible red is deterministic, not a flake), and still read the red run's trace for the collision signature, since one green rerun cannot separate an infra flake from a genuine intermittent code bug; (c) **deterministic non-code infra fault** — toolchain/runner-image drift, stale cache/vendored artifact, credential/quota expiry: reproduces identically (NOT a flake) AND must be shown **code-independent** before disowning — confirm the same failure reproduces on a known-good baseline (parent/last-good commit, or a build without the change) under the same runner/toolchain; if the red appears only *with* the change it is (d) however much it resembles drift → route to platform/infra only after that baseline check, then do not attribute to code; (d) **genuine code/dependency failure**. (Single-variant repos skip the (a) check; the trace-first and cause-classification still apply.)
50
+ - **A red CI pipeline/job is not by itself a code/dependency defect — read the failing job's own trace (not the red/green summary) and classify the cause before touching code.** Classify it as a trigger-variant artifact, a retriable infra flake, a deterministic non-code infra fault, or a genuine code/dependency failure using the checks in `references/diagnosis-playbook.md` (Red CI cause classes) before disowning or touching code.
51
51
  - **For a failing test, read the actual assertion error (Expected/Actual) and the failing test body FIRST — before diagnosing flakiness, concurrency, shared state, mock setup, or external-dependency causes.** A passing/failing count, a `REQUEST POST`-style debug log line, or "passes in isolation, fails in suite" is a symptom, not the assertion evidence; naming a cause from those alone is the exact failure this skill exists to prevent. If the test is suspected flaky, rerun N times and record the pass/fail ratio before calling it flaky (100% reproducible failure is deterministic, not flaky), and confirm the test's network/dependency boundary by reading its setup (e.g. whether it is already mocked) rather than inferring from logs.
52
52
 
53
53
  3. Hypothesize.
@@ -76,7 +76,7 @@ Diagnose and fix from evidence; route prevention to product, architecture, devel
76
76
  5. Verify cause.
77
77
  - Prove the cause with evidence.
78
78
  - **A diagnosis licenses a fix only when it explains both causality and incorrectness**: how the defect produced this failure on the failing path, and why that code, data, or config is wrong against its contract — so the fix covers related failures. A change that makes the failure disappear without the second half is a symptom patch; a genuine defect that cannot be linked to this failure is a different bug — record it, never ship it as this cause.
79
- - Report query/lookup evidence by cardinality: a data query, log search, or identity resolution that returns 0, 1, or N matches reports each of those outcomes distinctly — never silently take the first row of N, and never treat 0 rows as "no evidence collected" (an empty result over a named scope IS evidence: record which scopes matched and which were empty).
79
+ - Report query/lookup evidence by cardinality: a data query, log search, or identity resolution that returns 0, 1, or N matches reports each of those outcomes distinctly — never silently take the first row of N, and never treat 0 rows as "no evidence collected" (an empty result over a named scope IS evidence: record which scopes matched and which were empty). A zero failure count says something only after the scope's exposure is confirmed — the path was live and exercised in that window; zero failures from a path that never ran is no data, not health.
80
80
  - When the cause is environment/toolchain state, prove it from the tool that owns that state, not only from the high-level wrapper. A wrapper failure is a symptom until the underlying compiler, generator, runtime registry, dependency resolver, or platform destination evidence explains it.
81
81
  - If disproven, return to hypotheses instead of guessing.
82
82
  - Separate symptom, immediate cause, contributing factors, and prevention.
@@ -93,6 +93,10 @@ The hypothesize → instrument → verify loop must not run forever, and escalat
93
93
 
94
94
  ## Phase B: Fix And Verify
95
95
 
96
+ A failure or diagnosis goal carries this phase. Once the cause is verified, fix, test, review, push the branch and open or update the MR/PR to the repository's development target in the same delivery. "You only asked me to investigate" is not missing authority, and a "read-only" note on one step of a handoff binds that step only.
97
+
98
+ - Every explicit user limit and existing gate must still stop the step it covers, for example diagnosis only (只查 / 先别改), no push, a cost or run cap, a repository's confirm-first areas, the shared-gate route below, destructive actions, purchases, and merge, deploy or production steps.
99
+
96
100
  1. Fix minimally.
97
101
  - Address the proven cause with the smallest correct change.
98
102
  - Preserve contracts unless the task explicitly requires a contract change.
@@ -65,6 +65,10 @@ Locating the defect is usually the most expensive phase — harder than reproduc
65
65
  | Wrong value observed downstream | upstream trace | follow the value backward to the first point where a correct input produced a wrong output; that transition is the defect and the observation point is only where it surfaced — fix there when it is owned and changeable, otherwise record the upstream cause and enforce the contract at the nearest owned boundary |
66
66
  | Production symptom that cannot be re-triggered | telemetry walk | alert → exemplar trace → span tree → logs by trace-id (SKILL.md Phase A.4); group the failing population by attribute and compare it against the baseline to find what is different about failing requests |
67
67
 
68
+ ## Red CI Cause Classes
69
+
70
+ - A red CI pipeline/job is not by itself a code/dependency defect; read the failing job's own trace (not the red/green summary) and classify the cause before touching code. Refining the test-evidence classes in the entrypoint's Phase A Isolate step into why-CI-is-red-but-code-may-be-fine: (a) **trigger-variant artifact** — when the same job runs under more than one trigger-scoped config (branch/push vs merge-request vs manual/scheduled), the trigger can resolve different default variables or a different checked-out ref, so a red on a non-gating trigger may be benign — but conclude that ONLY after confirming the *same failing check* ran and is green on the gating path (a gating pipeline that is overall green yet never runs the failing check does not clear it; if that check's coverage is unique to the non-gating trigger — e.g. a scheduled/manual-only suite — treat it as genuine (d), not a variant artifact); (b) **retriable infra flake** — e.g. a shared-runner lock collision: confirm per the entrypoint's flaky-test rule (rerun N times and record the ratio) (rerun N times + record the ratio; a 100%-reproducible red is deterministic, not a flake), and still read the red run's trace for the collision signature, since one green rerun cannot separate an infra flake from a genuine intermittent code bug; (c) **deterministic non-code infra fault** — toolchain/runner-image drift, stale cache/vendored artifact, credential/quota expiry: reproduces identically (NOT a flake) AND must be shown **code-independent** before disowning — confirm the same failure reproduces on a known-good baseline (parent/last-good commit, or a build without the change) under the same runner/toolchain; if the red appears only *with* the change it is (d) however much it resembles drift → route to platform/infra only after that baseline check, then do not attribute to code; (d) **genuine code/dependency failure**. (Single-variant repos skip the (a) check; the trace-first and cause-classification still apply.)
71
+
68
72
  ## Probe Ordering And The Hypothesis Log
69
73
 
70
74
  Order probes; do not merely list hypotheses. For each candidate cause record the observation only it produces, the observation that cannot occur if it is true (the falsifier — collect this one first), what the probe costs, and what it risks. Then apply the entrypoint's one ordering rule: safety is a filter, not a rank — reject any probe outside the safety boundary first; rank the rest by alternatives ruled out per unit of cost; break ties by likelihood, then residual risk. Watch for confounders (a probe run from the wrong host, credential, or network position fails for its own reasons), side effects of active probes (more CPU changes race timing; verbose logging worsens latency — revert before the next probe), and probes that are only suggestive (races, deadlocks): record the evidence grade next to the result.
@@ -89,17 +89,17 @@ The entrypoint uses `proposed-next:` as an observable handoff, not an authorizat
89
89
 
90
90
  At the start of the next turn, recover intent in this order:
91
91
 
92
- 1. Follow the current explicit user instruction, including a correction, changed scope, stop, or status-only request.
92
+ 1. Follow the current explicit user instruction, including a correction, changed scope, stop, or status-only request. A clarifying question is not a status-only request: answer it, then continue the active authorized action unless the user also stopped it.
93
93
  2. For short assent, read back the original wording of the most recent still-active concrete proposal and check that later messages or task state have not withdrawn or superseded its action, scope, or authority. Quote that original proposal when stating the recovered action and scope; a summary or paraphrase alone cannot bind short assent. If recovery adds an action or broadens that quoted scope, select `blocked:` and ask. One recoverable action can bind with or without a marker; a stale, repeated, or conflicting marker is an assistant formatting defect to repair.
94
94
  3. If materially different proposals remain unresolved, or scope/authority is still unclear, ask one targeted question about that uncertainty. Do not ask the user to repair a marker or repeat a clear instruction. A marker alone never supplies missing authority.
95
95
 
96
- - Use the active owner's entry and safety gates for the recovered action. An authorized task includes necessary fixes, tests and review by default; neither a router nor a dispatched owner may discard that authority by relabeling its turn or exhausting an internal review sequence. Apply the owning review checkpoint and record `continuation_basis=existing-task-scope` with cumulative history in the caller-owned task artifact, not runtime JSON. Legacy `human_decision_required` / `continuation_authorization_required` values first require checking existing authority, not asking again. Explicit user cost, round-count and stop limits prevail; new scope, missing authority or real tradeoffs need a decision. Continuation grants no merge, publication or waiver authority. Intent recovery and authorization remain prose obligations. The optional `proposed-next-stop.sh` backstop checks missing handoff labels and declared next actions as described below; a label or hook receipt never proves the action is correct, authorized or complete.
96
+ - Use the active owner's entry and safety gates for the recovered action. An authorized task includes necessary fixes, tests and review by default. A failure or diagnosis goal also includes the verified narrow fix through branch push and MR/PR to the development target unless an explicit user limit says otherwise (diagnosis only, no push, a cost cap); "you only asked me to investigate" is not missing authority. Neither a router nor a dispatched owner may discard that authority by relabeling its turn or exhausting an internal review sequence. Apply the owning review checkpoint and record `continuation_basis=existing-task-scope` with cumulative history in the caller-owned task artifact, not runtime JSON. Legacy `human_decision_required` / `continuation_authorization_required` values first require checking existing authority, not asking again. Explicit user cost, round-count and stop limits prevail; new scope, missing authority or real tradeoffs need a decision. Continuation grants no merge, publication or waiver authority. Intent recovery and authorization remain prose obligations. The optional `proposed-next-stop.sh` backstop checks missing handoff labels and declared next actions as described below; a label or hook receipt never proves the action is correct, authorized or complete.
97
97
 
98
98
  ### Stop-time continuation reminder
99
99
 
100
100
  On hosts providing a current final message, `proposed-next-stop.sh` returns one bounded Stop reminder when the assistant declares a non-status `proposed-next:` action. Recheck the active request: execute a runnable, already-authorized action in the same turn; otherwise preserve explicit stop, planning-only and status-only scope, or state the concrete decision/resource/authority blocker. Missing labels with observable delivery evidence retain their formatting reminder. A status-only marker without another action declaration, quoted example, complete machine artifact, unsupported payload or host `stop_hook_active` retry does not trigger a continuation reminder.
101
101
 
102
- - A stop that waits on the user gets one bounded decision recheck instead: a `blocked:` handoff, a `none` explanation naming an approval, confirmation, decision or resource wait, or a last prose line asking permission to continue, either before any handoff label or, without a label, after observable delivery or edits. The recheck names the real blockers (missing credentials or authority; a fact unavailable from local evidence; an action the safety rules gate, such as destructive or irreversible work without recovery, production or customer data, or merge or publication outside the goal; overturning an established user direction; a material product tradeoff the evidence cannot settle) and returns security self-review, owner-skill, approach, test and naming choices and the next in-scope step to the agent. A real blocker survives it by restating `blocked:` after independent work is finished.
102
+ - A stop that waits on the user gets one bounded decision recheck instead: a `blocked:` handoff, a `none` explanation naming an approval, confirmation, decision or resource wait, or a last prose line asking permission to continue, either before any handoff label or, without a label, after observable delivery or edits. The recheck names the real blockers (missing credentials or authority; a fact unavailable from local evidence; an action the safety rules gate, such as destructive or irreversible work without recovery, production or customer data, or merge or publication outside the goal; overturning an established user direction; a material product tradeoff the evidence cannot settle) and returns security self-review, owner-skill, approach, test and naming choices and the next in-scope step to the agent. Both reminders also state the test a blocker must pass — it names something only the user can supply (a decision the evidence cannot settle, a credential or access grant, permission the goal does not cover, or a fact absent from every readable source) — so a step the agent can perform itself is never a blocker, whatever it costs in time or runs: a failure goal's verified fix through its MR/PR, a feature-branch push or MR/PR, a rerun or retry of a failed, timed-out or inconclusive check, and a lookup of facts or access the agent can reuse. A count or stop bar the agent proposed is not a user limit unless the user adopted it as one, and a clarifying question is not a status-only request. A real blocker survives it by restating `blocked:` after independent work is finished.
103
103
  - Do not request continuation for `none — status only` or another `none` status explanation, and never treat `blocked:` as a continuation request; the decision recheck above is separate. Mixed status/action markers still require reconciliation.
104
104
 
105
105
  The hook recognizes declarations, not authorization or actual task completion, and cannot force the model to follow through. OpenCode idle does not expose the required final-message evidence; its Stop behavior remains unverified.
@@ -110,7 +110,7 @@ An eligible next slice comes from an explicit status/task/acceptance source or a
110
110
 
111
111
  Action-scoped stop conditions are: an explicit stop/pause instruction; a user-requested status-only answer; a failed, pending or inconclusive required gate; a dirty/conflicting worktree that cannot be isolated; a required environment unavailable after remediation; a high-impact product, architecture or compliance decision; a destructive action; an external purchase or financial commitment; unclear ownership; ambiguous assent; missing stricter authorization; materially different viable approaches with none dominant and reversible; a speculative fix without evidenced cause; or no low-risk slice. Apply each condition to the affected action. For a failed check, perform available authorized diagnosis and remediation before stopping the whole task: cite the failure output, repair attempts (or evidence that repair is unsafe or outside authority), and residual blocker. A failed verdict alone does not block diagnosis.
112
112
 
113
- **Awaiting work you started yourself is not a stop condition.** A finite command, suite, gate, or review you launched, whose result only you consume, is in-flight work rather than a handoff: wait for it and continue in the same turn. A process meant to stay up — a dev server, a watch-mode runner, a tail — has no terminal result to wait for: take its readiness signal and proceed. Never poll it forever, and do not infer anything about its lifetime from this rule; whether it keeps running is the delivery's decision, and a service the user asked for is a deliverable, not a leftover. Ending the turn to report that it is running is a premature stop even when the report is accurate — the user gains nothing they can act on, and the next step was already authorized. Host behavior invites this: a backgrounded step returns control immediately, so the pause *looks* like a turn boundary. It is not one. Before ending any turn, name the next action; if you can perform it now, the turn is not over. The turn ends at the first action that genuinely needs the user — an unresolved decision, missing authority, an explicit stop — not at the nearest convenient pause. A user asking why you stopped is this defect's recurrence signal, not a request for a status update.
113
+ **Awaiting work you started yourself is not a stop condition.** A finite command, suite, gate, or review you launched, whose result only you consume, is in-flight work rather than a handoff: wait for it and continue in the same turn. A process meant to stay up — a dev server, a watch-mode runner, a tail — has no terminal result to wait for: take its readiness signal and proceed. Never poll it forever, and do not infer anything about its lifetime from this rule; whether it keeps running is the delivery's decision, and a service the user asked for is a deliverable, not a leftover. Ending the turn to report that it is running is a premature stop even when the report is accurate — the user gains nothing they can act on, and the next step was already authorized. Host behavior invites this: a backgrounded step returns control immediately, so the pause *looks* like a turn boundary. It is not one. Before ending any turn, name the next action; if you can perform it now, the turn is not over. The turn ends at the first action that genuinely needs the user — an unresolved decision, missing authority, an explicit stop — not at the nearest convenient pause. A user asking why you stopped is this defect's recurrence signal, not a request for a status update. The same holds for a CI pipeline on your own MR/PR — poll it to its result rather than ending the turn on "waiting for CI" — and for Draft: once your self-review, external review and required CI pass, mark the MR/PR ready in the same turn; Draft is a work-in-progress marker, never an end state.
114
114
 
115
115
  Check continuation on every user reply immediately following assistant prose that states or implies a next action, and on any explicit continuation request, regardless of landing status. Do not first require classifying the reply as assent; visibly report the continuing or blocked outcome even when the reply changes scope or stops the proposed action. Short replies include `ok`, `yes`, `可以`, `好`, `继续`, `proceed`, `do it`, `go ahead`, and `👍`; interpret them against the recovered action rather than formatting alone.
116
116
 
@@ -132,5 +132,5 @@ Before deriving the next slice from a status source, reconcile it against the ap
132
132
 
133
133
  Binding detail:
134
134
 
135
- - Small tests and routine development/test-environment operations within the task are ordinary execution details. Use configured accounts and access directly, without per-run approval or inventing a quota/cost estimate or cap. Normal metered model/tool use is not a new purchase. Honor explicit user spending/count limits; a development/test label does not grant destructive, production, customer-data, permission-changing or new-purchase authority beyond the task.
135
+ - Small tests and routine development/test-environment operations within the task are ordinary execution details. Use configured accounts and access directly, without per-run approval or inventing a quota/cost estimate or cap. Normal metered model/tool use is not a new purchase. Honor explicit user spending/count limits; a count, round or budget you proposed is your estimate, not their limit, unless the user adopted it as a limit: stated the number, said "at most"/"only", or accepted it as a cap or ceiling. Plain assent to the work ("ok", "go") does not adopt the estimate as a cap. When a host merge gate needs a grant, request one that covers the whole remaining plan rather than one per merge; a development/test label does not grant destructive, production, customer-data, permission-changing or new-purchase authority beyond the task.
136
136
  - Assent never replaces an owner gate's stricter authorization form and never broadens scope or implies an external purchase/financial commitment, merge, publish, destructive, production, external-message, or high-impact-decision authority.
@@ -11,6 +11,14 @@ Bind recovery when either:
11
11
 
12
12
  A semantic compaction paraphrase supplies neither binding path; recover the original proposal and assent before deciding path (a) is unavailable. A bare "why did you stop" complaint does not itself name path (b)'s action and scope. Never copy real conversation text into a shared repository record, reconstruct, broaden, or substitute it. The user's challenge reactivates that exact slice. Restate and proceed when either path binds; ask only when the action, scope, or required authority remains unresolved. A new user message or a changed gate requires reassessment, not automatic reconfirmation.
13
13
 
14
+ ## A stop that survived its recheck
15
+
16
+ When the corrected stop happened after a Stop recheck or other reminder had already fired on it, detection worked and is not the cause.
17
+
18
+ - The RCA must quote the agent's post-recheck justification from the transcript and name the term it leaned on ("missing authority", "status-only", "outward-facing", "the approved count is used up").
19
+ - The prevention must close that term's definition inside the recheck text and the owning gate, with a test on the reminder content that fails before the change and a replay of the stop against both reminder texts.
20
+ - Another detection pattern does not address this shape; repeated landings that only add detection are the cross-landing signal in `SKILL.md`.
21
+
14
22
  ## Invalid `blocked:` recovery
15
23
 
16
24
  A `blocked:` recovery without applicable state evidence and a specific remaining blocker is invalid: recover intent and rerun the owning gate. If a decision or permission remains unresolved, ask in the same turn and block that dependent action. Continue available authorized diagnosis, bounded remediation, or independent work; do not let stale assent bypass a newly pending or inconclusive gate. Do not let correction RCA or extraction delay recovery of a still-authorized delivery.
@@ -747,3 +747,13 @@ The pending classification above is superseded by the executed source comparison
747
747
  | A default sign-off exemption does not override an explicit stricter rule | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/SKILL.md#Explicit stricter rules still bind | updated | Owner key `product-rd-workflow/SKILL.md`. Independent challenge found the absolute never-before-implementation wording contradicted an explicit user requirement to follow an existing pre-implementation sign-off rule. The entry and design reference now scope the exemption to this gate. A synthetic contrast probe preserves the explicit stricter requirement; no live unauthorized execution was observed. |
748
748
  | Authorization grading needs an explicit stricter-rule control | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`; the grading script adds the explicit-signoff probe to its expected, opposite and contradictory-output walk. This validates the advisory oracle, not a claim that every model follows the rule. |
749
749
  | A test that runs a whole checker and asserts only its exit code must show the checker's output when the code is wrong, or a CI-only failure cannot be attributed | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_source_register_lifecycle.sh | updated | Owner key `skill-extraction-workflow/SKILL.md` is unchanged. Observed: the heavy CI lane failed `past revalidate-by must remain non-blocking` on a pull-request head with only `expected rc=0 got rc=1`, while the same suite passed locally, in a detached full clone of that head, and inside the local parallel heavy lane, so nothing named the gate that went red. `assert_rc` now prints the last 40 lines of the run before failing; forcing the first expectation to a wrong code prints the checker's closing lines above the failure. |
750
+ | A failure goal carries its verified fix to the MR, and a tradeoff resting on an unverified cause is not yet a user decision | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/SKILL.md#handing the user a decision or tradeoff that rests on the cause | updated | Owner key `defect-diagnosis/SKILL.md`. Reviewed sessions stopped after a verified cause on "you only asked me to investigate" and asked the user to choose a tradeoff before the cause was falsified. Phase B now states the fix scope and its real stops; a zero failure count needs confirmed exposure. A replay of the restated stop proceeded 5/12 with the old Stop reminder and 12/12 with the new one; an explicit diagnosis-only limit held 0/12 in both arms. The red-CI cause classes moved verbatim to the playbook to keep the entrypoint within its word budget. |
751
+ | The continuation gate defines the terms agents used to survive its recheck | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#A failure or diagnosis goal also includes the verified narrow fix | updated | Owner key `product-rd-workflow/SKILL.md`. The decision recheck fired on every observed stop, and agents restated the stop as investigation-only scope, an outward-facing push or MR, an agent-proposed count, a clarifying question treated as status-only, or a fact they could find. The new hook case failed six times on the base reminder text. The gate and both reminders now close those terms, and agent-proposed counts are estimates rather than user limits. The entrypoint is unchanged because it is at its word ceiling. |
752
+ | A stop that survives its recheck is fixed by closing the term it leaned on, not by more detection | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/resume-paused-delivery.md#must quote the agent's post-recheck justification | updated | Owner key `skill-extraction-workflow/SKILL.md`. Three landings in two days added detection or broader wording to the same Stop reminder, and the class recurred with the reminder firing each time. A retrospective on such a stop now quotes the post-recheck justification, closes that term at the recheck and the owning gate, tests the reminder content, and replays the stop against both texts. The grading walk covers the new probe pairs. |
753
+ | A blocker must name something only the user can supply; a step the agent can perform itself, including a rerun of an inconclusive check, is never one | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#a step the agent can perform itself is never a blocker | updated | Owner key `product-rd-workflow/SKILL.md`. After the first landing a third stop appeared in a reviewed session: an inconclusive CI review restated as "no resume handle; a retry restarts from scratch". Listing terms is open-ended, so the reminder now leads with the invariant and keeps the observed terms as examples. The hook case asserting the invariant and the rerun clause fails on the previous text. Diagnosis replay with the new text: base 2/6, candidate 6/6, explicit-limit control 0/6 on both; the CI-review replay proceeded 6/6 on both texts, so it is recorded as a control. |
754
+ | A section moved into a reference names its antecedents in the entrypoint | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/references/diagnosis-playbook.md#Refining the test-evidence classes in the entrypoint's Phase A Isolate step | updated | Owner key `defect-diagnosis/SKILL.md`. Independent review found the moved red-CI section still pointing at "the classes above" and "the flaky rule below", neither of which exists in the playbook. The copy now names the entrypoint's Isolate-step test-evidence classes and its flaky-test rule; the four cause classes stay verbatim. |
755
+ | A session that edited reader-facing documents gets the tighten-doc closeout at Stop, and Draft or a pending CI run is not an end state | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#is not a user limit unless the user adopted it as one | updated | Owner key `product-rd-workflow/SKILL.md`. Reviewed sessions: 26 of 29 that edited plans, specs, READMEs or handoff documents never loaded tighten-doc; MRs were left in Draft at the end of 21 sessions; 13 turns ended on "waiting for CI". The Stop hook now adds a one-time document closeout reminder (five new cases fail on the previous hook) and lists marking an MR ready and polling one's own CI among self-performable steps; the gate states both. A proposed count binds only when the user adopted it as a limit, after the challenge showed an explicitly adopted ceiling being overridden. |
756
+ | Body probes that must tell a blocked merge from a blocked fix grade an explicit marker | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The challenge showed `blocked: fix, push and MR need approval; merge also waits` passing the keyword grader because the line named the merge. The diagnosis probes now grade `next: fix-and-open-mr` / `next: wait-for-user`, and the walk pins that output as FAIL, a missing or doubled marker as FAIL, and the explicit-limit control both ways. |
757
+ | A failure goal's fix scope yields to every explicit user limit and existing gate; its stop list is examples, not an exhaustive set | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#Every explicit user limit and existing gate must still stop the step it covers | updated | Owner key `defect-diagnosis/SKILL.md`. Independent review found that Phase B listed its stops as "stop only for" four cases, so "fix locally, do not push", a cost cap, a destructive non-production repair or a purchase matched none of them. The rule now lets every explicit user limit and existing gate stop the step it covers, and lists those cases as examples. The Stop reminder carries the same limit, and its case fails 10 times on the previous text. A no-push probe passed 4/4 on the previous, main and new bodies, so it is a control: the old wording contradicted the acceptance requirement, but measured behaviour already respected the limit. |
758
+ | The inherited fix scope of a failure goal yields to any explicit user limit, not only a diagnosis-only one | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#unless an explicit user limit says otherwise (diagnosis only, no push, a cost cap) | updated | Owner key `product-rd-workflow/SKILL.md`. The same review finding applied to the continuation gate and both Stop reminders, which excepted only a diagnosis-only limit. They now yield to any explicit user limit and name no push and a cost cap as examples. The reminder case asserting it fails 10 times on the previous hook and passes now. A replay of the restated stop with "fix locally, do not push" fixed locally 6/6 on both texts, so it is a control. |
759
+ | The grading walk pins the no-push probe's three-way marker | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The new `diag-fix-local-no-push` probe accepts only `next: fix-locally`; the walk pins a push, a withheld fix, a missing marker and two markers as FAIL. Run against the previous probe set, the walk aborts because the probe is missing. |
@@ -84,13 +84,43 @@ if ! ruby -e '
84
84
  ].each do |label, probe, output, expected|
85
85
  checks << [label, grade(probe, output)[:status] == expected]
86
86
  end
87
+ # Diagnosis-goal probes: a merge-only blocked line beside the fix is correct;
88
+ # a blocked fix, push or MR is the failure under test.
89
+ dfix = PROBES.find { |p| p[:id] == "prd-continue-diagnosis-fix" } or abort "diagnosis fix probe missing"
90
+ dhand = PROBES.find { |p| p[:id] == "diag-continue-fix-after-handoff" } or abort "handoff probe missing"
91
+ donly = PROBES.find { |p| p[:id] == "diag-wait-diagnosis-only" } or abort "diagnosis-only probe missing"
92
+ donly_prd = PROBES.find { |p| p[:id] == "prd-stop-diagnosis-only" } or abort "product diagnosis-only probe missing"
93
+ dlocal = PROBES.find { |p| p[:id] == "diag-fix-local-no-push" } or abort "no-push probe missing"
94
+ [
95
+ ["D1 fix continues", dfix, "continuing: 修复上限取值并开 MR\nnext: fix-and-open-mr", "PASS"],
96
+ ["D1 merge-only block beside fix", dfix, "continuing: 修复并开 MR\nblocked: 合并 MR — 等用户「合并」\nnext: fix-and-open-mr", "PASS"],
97
+ ["D1 fix blocked behind a merge word", dfix, "continuing: fix\nblocked: fix, push and MR need approval; merge also waits\nnext: wait-for-user", "FAIL"],
98
+ ["D1 marker missing", dfix, "continuing: 修复上限取值并开 MR", "FAIL"],
99
+ ["D1 both markers", dfix, "next: fix-and-open-mr\nnext: wait-for-user", "FAIL"],
100
+ ["D1b limit holds", donly_prd, "next: wait-for-user", "PASS"],
101
+ ["D1b limit overridden", donly_prd, "next: fix-and-open-mr", "FAIL"],
102
+ ["D2 handoff fix", dhand, "next: fix-and-open-mr", "PASS"],
103
+ ["D2 handoff wait", dhand, "next: wait-for-user", "FAIL"],
104
+ ["D2 handoff both", dhand, "next: fix-and-open-mr\nnext: wait-for-user", "FAIL"],
105
+ ["D3 limit wait", donly, "next: wait-for-user", "PASS"],
106
+ ["D3 limit fix", donly, "next: fix-and-open-mr", "FAIL"],
107
+ ["D4 local fix", dlocal, "next: fix-locally", "PASS"],
108
+ ["D4 push overrides the limit", dlocal, "next: fix-and-open-mr", "FAIL"],
109
+ ["D4 fix withheld", dlocal, "next: wait-for-user", "FAIL"],
110
+ ["D4 marker missing", dlocal, "continuing: 本地修复", "FAIL"],
111
+ ["D4 two markers", dlocal, "next: fix-locally\nnext: fix-and-open-mr", "FAIL"],
112
+ ].each do |label, probe, output, expected|
113
+ checks << [label, grade(probe, output)[:status] == expected]
114
+ end
87
115
  {
88
116
  "prd-continue-dev-test" => "continuing",
89
117
  "prd-continue-small-test" => "continuing",
90
118
  "prd-stop-test-preparation" => "blocked",
91
119
  "prd-stop-explicit-signoff" => "blocked",
92
120
  "prd-stop-test-limit" => "blocked",
93
- "prd-stop-dev-destructive" => "blocked"
121
+ "prd-stop-dev-destructive" => "blocked",
122
+ "prd-continue-question-turn" => "continuing",
123
+ "prd-stop-question-hold" => "blocked"
94
124
  }.each do |id, verdict|
95
125
  probe = PROBES.find { |p| p[:id] == id } or abort "#{id} missing"
96
126
  opposite = verdict == "continuing" ? "blocked" : "continuing"
@@ -4,10 +4,12 @@
4
4
 
5
5
  ## 指令与有效期
6
6
 
7
- 支持原有单独“合并/merge”和“批量合并 N”,另支持完整单行“完成并合并 PR #123 / MR !123”(英文 `finish and merge PR #123`)。原有单次/计数额度仍被任何新消息清除。
7
+ 支持原有单独“合并/merge”和“批量合并 N”,另支持完整单行“完成并合并 PR #123 / MR !123”(英文 `finish and merge PR #123`)。原有单次/计数额度仍被任何新消息清除。宿主的后台任务完成通知也经由提交提示的通道送达,且没有能区分来源的字段,所以同样清除额度(但从不生成额度)。获授权后连续合并时,等 CI 用前台等待(如 `gh pr checks <编号> --watch`),别用后台任务或监视器,免得完成通知在两次合并之间清掉剩余额度。
8
8
 
9
9
  新形式只绑定当前 `origin` 仓库和指定编号,原始 60 分钟内消费一次;单独“继续/继续吧/进度/状态/continue/status/progress”保留原额度和到期时间,“停止/停一下/不要合并/撤销合并授权/stop/pause/cancel merge”撤销,其他消息暂停机械额度,之后“继续”不能恢复。
10
10
 
11
11
  ## 执行命令
12
12
 
13
+ 带 `-h/--help` 的 `glab mr merge` / `gh pr merge` 只打印帮助,闸会拒绝且不消费额度;查看帮助用 `glab help mr merge` 或 `gh help pr merge`。
14
+
13
15
  新形式的执行命令必须是单条直接 `gh pr merge` 或 `glab mr merge`,显式编号;gh 使用 `--repo host/owner/repo` 并指定策略,glab 使用 `--repo https://host/namespace/repo` 并指定 `--auto-merge=false --yes`,可附完整 head SHA,其他参数和 API 形式保持未核验。
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "schema": 1,
3
3
  "npmPackage": "@ccoalm/ccl-skills",
4
- "version": "0.18.8",
5
- "sourceCommit": "32941fa8b6a3d9d98779f18802dea0c21d3b61ff",
4
+ "version": "0.18.9",
5
+ "sourceCommit": "44cc6d8ea4f4a27ba55b0cd9cb51408ce3686de4",
6
6
  "sourceState": "clean",
7
7
  "files": [
8
8
  {
@@ -42,7 +42,7 @@
42
42
  },
43
43
  {
44
44
  "path": "marketplace/plugins/ccl-skills/agent-context/session-policy.md",
45
- "sha256": "a47f7ff077f33da3efaa24fc1d55118330f23ef268e92355d2f281ee56237d2a",
45
+ "sha256": "c76a61265843a87227741291241174856badb3f081c71d87d23643b2a2187b53",
46
46
  "mode": 420
47
47
  },
48
48
  {
@@ -72,7 +72,7 @@
72
72
  },
73
73
  {
74
74
  "path": "marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh",
75
- "sha256": "d12e8c81122babe65d12cca2576b0e2b73b40f4e82090ef31de26bbc893a1cba",
75
+ "sha256": "eeb08a60e67f74428ea2621d4e354a6d88978f41c846ab41974497bf7f838598",
76
76
  "mode": 493
77
77
  },
78
78
  {
@@ -82,7 +82,7 @@
82
82
  },
83
83
  {
84
84
  "path": "marketplace/plugins/ccl-skills/hooks/host-input.py",
85
- "sha256": "b737a74f0a664c52189e43b84539c7ab41e5e4817570b6b1c50c889056c3a9a6",
85
+ "sha256": "058e78598435050d1c0af1ff721d4fee5d049c5a442a71d9ed9fd22b6d70a307",
86
86
  "mode": 420
87
87
  },
88
88
  {
@@ -107,7 +107,7 @@
107
107
  },
108
108
  {
109
109
  "path": "marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh",
110
- "sha256": "c912bd8bbb9afe15b4ed3ac546175ecb9185752087a818d1426ffecd68ffb623",
110
+ "sha256": "363192600575dd33b2eff0cd9da4ca64c3a9f85a3704a9c698f480a26cf67b0d",
111
111
  "mode": 493
112
112
  },
113
113
  {
@@ -172,7 +172,7 @@
172
172
  },
173
173
  {
174
174
  "path": "marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh",
175
- "sha256": "46aa31fb242e9096545d00554cb17eaf9bd4fef68fa4e49ee581148b638e3a08",
175
+ "sha256": "b2f5239ce389a0e4c1e2e7fd6d8f649484a00404c01908ce729ac4eb70f20681",
176
176
  "mode": 493
177
177
  },
178
178
  {
@@ -182,17 +182,17 @@
182
182
  },
183
183
  {
184
184
  "path": "marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh",
185
- "sha256": "acb48c3fd690bd0c7353d2071e1cdd81b28a414d10452a579bf3bb044da1462d",
185
+ "sha256": "5be7052d8df87a0640efe2711d1f11411bd4b622d781fce1dccc411c4dcc0325",
186
186
  "mode": 420
187
187
  },
188
188
  {
189
189
  "path": "marketplace/plugins/ccl-skills/hooks/test_proposed_next.py",
190
- "sha256": "db2df523107fa2b0896c769396e523171f77948526f9d39c8396b598e68fefbf",
190
+ "sha256": "10a73c9577931bef33462cd9388a899a2d79b62255fdea5f32e322ae2a7569b5",
191
191
  "mode": 493
192
192
  },
193
193
  {
194
194
  "path": "marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh",
195
- "sha256": "f57ede7e26be676f3759d97e5707c4f9390623701103f2cfc4465a9d1669abf3",
195
+ "sha256": "4536ae15900b854fe5bf2ed11846b51536ced82a82a9b59d7ea21ba8228fbe47",
196
196
  "mode": 493
197
197
  },
198
198
  {
@@ -587,7 +587,7 @@
587
587
  },
588
588
  {
589
589
  "path": "marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md",
590
- "sha256": "220c3283e97e380d79d0cf7edafd2ce63c188571555fb5621e502f01cefe85ba",
590
+ "sha256": "12222d36e7894d8ea767fe99eb0619fd0eff9b250c448ef62c75759dae8a807e",
591
591
  "mode": 420
592
592
  },
593
593
  {
@@ -597,7 +597,7 @@
597
597
  },
598
598
  {
599
599
  "path": "marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md",
600
- "sha256": "2e1be983cdc7dc81d463df302a3bd2c212b197396b463a33254391b0b289cd9a",
600
+ "sha256": "6d6d71442246edf6f36f73db8c94df105c2ce6db995db2feac146b966b561023",
601
601
  "mode": 420
602
602
  },
603
603
  {
@@ -1477,7 +1477,7 @@
1477
1477
  },
1478
1478
  {
1479
1479
  "path": "marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md",
1480
- "sha256": "b23c9822452402d26035218ea1e16f8cb3943a2d5812ff6ecd94399f66110644",
1480
+ "sha256": "9f77fe819f70acbaf59b8330ad4ebf29bdba04dc1116e6b5d0ada46d70b14837",
1481
1481
  "mode": 420
1482
1482
  },
1483
1483
  {
@@ -2177,7 +2177,7 @@
2177
2177
  },
2178
2178
  {
2179
2179
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md",
2180
- "sha256": "3b43de0e81d0a43fb5ba40e97de755bd032f1b5417bbbdd9ca2a4cd2c5c1c65d",
2180
+ "sha256": "9c8428fd26d017b876d627bbc3897652ca2c96385a310b525d3ced7301c677f3",
2181
2181
  "mode": 420
2182
2182
  },
2183
2183
  {
@@ -2207,7 +2207,7 @@
2207
2207
  },
2208
2208
  {
2209
2209
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md",
2210
- "sha256": "1c0e95554f7b8d63860e23385d3b9d84b7d3c1790e4e4071e772b170b4a2b2c6",
2210
+ "sha256": "100bbf7715cd89371c2a1295c4af7b61ae170fd54364c7835f7321c144ddfe60",
2211
2211
  "mode": 420
2212
2212
  },
2213
2213
  {
@@ -2382,7 +2382,7 @@
2382
2382
  },
2383
2383
  {
2384
2384
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh",
2385
- "sha256": "c1c34a27011e77f367a758bb5d6f4abd1ab04d997119f718268343c73027604f",
2385
+ "sha256": "70b058b523df7e40297ff13a8351bee92af594865d51b327b11a5889ab113a31",
2386
2386
  "mode": 493
2387
2387
  },
2388
2388
  {
@@ -3312,7 +3312,7 @@
3312
3312
  },
3313
3313
  {
3314
3314
  "path": "marketplace/plugins/ccl-skills/skills/worktree-isolation/references/hook-authorization.md",
3315
- "sha256": "989be64bb1d59bdf92c72c4bed407381c25b906e2b436ba8ab601e5590600ce8",
3315
+ "sha256": "115933f3e019e1aed802d65078b5b97a36bf80d35b3a526ba28370d24a992cc6",
3316
3316
  "mode": 420
3317
3317
  },
3318
3318
  {
@@ -3533,5 +3533,5 @@
3533
3533
  "mode": 420
3534
3534
  }
3535
3535
  ],
3536
- "snapshotHash": "9f8142a50f984550552470abdd4c051e69831d649394c5fce0e70da6d59fff8f"
3536
+ "snapshotHash": "8357bdb496747f3b429c69a2f91d0b47ca51b7b50a392590990af92c8f54a7e2"
3537
3537
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ccoalm/ccl-skills",
3
- "version": "0.18.8",
3
+ "version": "0.18.9",
4
4
  "description": "Reusable workflows that help coding agents plan, build, test, review, and release software — for Claude Code, Codex, and OpenCode",
5
5
  "keywords": ["skills", "agent-skills", "claude", "claude-code", "codex", "opencode", "agent", "ai", "ai-agents", "cli", "anthropic", "developer-tools"],
6
6
  "type": "module",