@ccoalm/ccl-skills 0.18.7 → 0.18.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +69 -2
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-policy.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh +12 -2
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/host-input.py +89 -2
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_proposed_next.py +111 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +7 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +5 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +31 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/hook-authorization.md +3 -1
- package/dist/assets/release.json +31 -21
- package/dist/auto-update.d.ts +26 -0
- package/dist/auto-update.js +551 -0
- package/dist/cli.d.ts +2 -0
- package/dist/cli.js +29 -3
- package/dist/opencode-adapter.js +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -60,7 +60,7 @@ ccl-skills uninstall # preview
|
|
|
60
60
|
ccl-skills uninstall --yes # remove host assets
|
|
61
61
|
```
|
|
62
62
|
|
|
63
|
-
Limit
|
|
63
|
+
Limit these operations to one host with `--host claude`, `--host codex`, or `--host opencode`. Add `--json` for machine-readable output.
|
|
64
64
|
|
|
65
65
|
For `--host codex`, unreadable public plugin state returns exit `3` with `host-state-unknown`; a missing CLI or failed capability probe returns `4`. If that host failure occurs with a pending journal, recovery is deferred: exit `5` with `partial-journal` retains the journal and records `details.hostFailure`. Restore the CLI or readable public plugin state, then rerun the command. Other outcomes can share these exit codes, so inspect the JSON status as well.
|
|
66
66
|
|
|
@@ -70,9 +70,76 @@ Codex `doctor` also reads the native hook inventory. When every expected package
|
|
|
70
70
|
|
|
71
71
|
After `ccl-skills uninstall --yes`, remove the CLI package itself with `npm uninstall --global @ccoalm/ccl-skills` if it is no longer needed.
|
|
72
72
|
|
|
73
|
+
## Automatic CCL updates
|
|
74
|
+
|
|
75
|
+
On macOS, explicitly enable daily CCL updates for one host:
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
# Existing Codex Git plugin; Codex is the default host
|
|
79
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update enable
|
|
80
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update status --json
|
|
81
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update disable
|
|
82
|
+
|
|
83
|
+
# Existing npm-managed OpenCode skills, plugin and hooks
|
|
84
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update enable --host opencode
|
|
85
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update status --host opencode --json
|
|
86
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update disable --host opencode
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
With the CLI installed globally, replace `npx --yes @ccoalm/ccl-skills@latest` with `ccl-skills`. Each command acts on one host and executes directly; `--yes` and `--host all` are unsupported. Installation never enables scheduling. Separate launchd jobs run every 24 hours while logged in and when loaded at login. A non-root macOS user is required; other platforms return exit `4` with `unsupported-platform`.
|
|
90
|
+
|
|
91
|
+
### Codex
|
|
92
|
+
|
|
93
|
+
The target is `ccl-skills@ccl-skills` in the current `CODEX_HOME`, or `~/.codex` when unset. Its marketplace and plugin must use the canonical `ccoalm/ccl-skills` Git repository over HTTPS or SSH, and the plugin must be enabled. Codex must support `--no-daemon` and plugin JSON output.
|
|
94
|
+
|
|
95
|
+
Each run checks public plugin state before and between these steps:
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
codex --no-daemon plugin marketplace upgrade ccl-skills
|
|
99
|
+
codex --no-daemon plugin add ccl-skills@ccl-skills
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
An unsuccessful upgrade prevents the add step. Each update command has a two-minute limit; timeout terminates its child process group. Source changes are checked at those boundaries; Codex does not provide an atomic compare-and-update operation. Restart Codex to load changes. Hook trust follows Codex's normal policy.
|
|
103
|
+
|
|
104
|
+
Codex npm installations use a local snapshot under `ccl-skills-npm`, which this scheduler refuses. Keep those current with `ccl-skills update --yes`. To migrate, preview removal with `ccl-skills uninstall --host codex`, then repeat with `--yes`. Register the Git installation explicitly:
|
|
105
|
+
|
|
106
|
+
```bash
|
|
107
|
+
codex --no-daemon plugin marketplace add https://github.com/ccoalm/ccl-skills.git
|
|
108
|
+
codex --no-daemon plugin add ccl-skills@ccl-skills
|
|
109
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update enable
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
### OpenCode
|
|
113
|
+
|
|
114
|
+
OpenCode requires an existing healthy npm-managed CCL installation with bundled assets. The scheduler updates only its CCL skills, native plugin and hook runtime under `~/.config/opencode`. It preserves OpenCode itself, application configuration, authentication, unrelated files and `~/.agents`. Restart OpenCode to load changes.
|
|
115
|
+
|
|
116
|
+
Each run downloads `@ccoalm/ccl-skills@latest` from the public npm registry into a fresh private directory. It disables lifecycle scripts and package self-update, validates package identity, then invokes the fetched CLI's OpenCode doctor and update commands. It refuses missing installations, source overrides and locally changed managed files. A newer installed version is retained. Success requires a healthy receipt matching the fetched package version and bundled source.
|
|
117
|
+
|
|
118
|
+
Source-checkout installs are unsupported. To migrate, first back up their CCL files and receipt, move only those verified CCL-owned files out of OpenCode's shared directories, then explicitly run the npm installer:
|
|
119
|
+
|
|
120
|
+
```bash
|
|
121
|
+
npx --yes @ccoalm/ccl-skills@latest install --host opencode
|
|
122
|
+
npx --yes @ccoalm/ccl-skills@latest doctor --host opencode
|
|
123
|
+
npx --yes @ccoalm/ccl-skills@latest auto-update enable --host opencode
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
The scheduler never adopts or deletes a source-copy installation. Resolve any collision or drift reported by the installer before enabling it.
|
|
127
|
+
|
|
128
|
+
Package download has a three-minute limit; asset update has a two-minute limit. Cancellation sends SIGINT and allows ten seconds for rollback before forcing termination. A forced termination reports unknown finality; inspect `doctor --host opencode` before retrying. Fetch and preflight failures leave installed assets intact.
|
|
129
|
+
|
|
130
|
+
### Status and recovery
|
|
131
|
+
|
|
132
|
+
State lives in `$CODEX_HOME/ccl-skills-auto-update` for Codex and `~/.config/opencode/ccl-skills-auto-update` for OpenCode. Each profile has its own file in `~/Library/LaunchAgents`. The saved runner survives removal of an npx download and keeps the executable paths and profile selected at enable time. Those executables must remain installed. Re-enabling retains the saved runner; upgrading the npm CLI does not replace it.
|
|
133
|
+
|
|
134
|
+
Commands receive fixed HOME/PATH plus CODEX_HOME and `GIT_TERMINAL_PROMPT=0` for Codex, or the self-update/notifier opt-outs for OpenCode. npm uses private empty config files and cache. Credentials and source/registry overrides from the enabling shell are not persisted; macOS may add runtime variables.
|
|
135
|
+
|
|
136
|
+
`status` checks live registration and the latest bounded `last-run.json` record, which excludes raw command output. Failed or interrupted runs, registration errors and ownership errors return exit `5`. Restore the reported dependency and retry `enable` or `disable`. Overlapping runs are skipped. `disable` removes only that host's owned schedule and retains its runner, state and log. Disable first if removing the CLI should also stop updates.
|
|
137
|
+
|
|
138
|
+
SIGINT and SIGTERM release owned locks, including between commands. A crash or forced kill can leave a `run.lock` or `manage.lock`. A lock at least 15 minutes old, or dated at least 15 minutes into the future, reports failure even if its PID is alive. The scheduler never removes these locks automatically. Inspect the recorded PID and launchd job; remove only that lock after confirming no updater owns it. Never remove a running updater's lock.
|
|
139
|
+
|
|
73
140
|
## Update notice
|
|
74
141
|
|
|
75
|
-
|
|
142
|
+
The npm package does not update itself automatically. An interactive run prints a one-line notice on **stderr** when a newer version exists, at most once a day per version:
|
|
76
143
|
|
|
77
144
|
```
|
|
78
145
|
Update available: 0.2.0 -> 0.3.0
|
|
@@ -41,7 +41,7 @@ The compact session-start entry routes to these execution details when the relev
|
|
|
41
41
|
|
|
42
42
|
几条贯穿原则(任何任务都适用;详则在 owner 技能里):
|
|
43
43
|
- **上下文恢复是 agent 的工作**:恢复/继续/复盘/判断既有工作时,先读 SessionStart 的 `<agent-context-recovery>`(若宿主提供),再核 repo 契约、当前 Git、项目状态/任务持久件、最小相关 session/memory 片段、commit 与 CI/test 证据;读取历史片段前必须确认其 repo root / cwd / remote 属于当前仓(全局 session/db 存在不等于相关);启动快照只用于定位,结论仍要 live refresh。能从本地证据恢复的事实不得让用户重述。只有方向/重大取舍、缺失权限或凭据、不可逆动作、以及本地证据确实不存在时才打断用户。
|
|
44
|
-
- **自主决策,别把可判定的问题推给用户**:开发中只在真实阻塞时停下问人——缺失的凭据/授权;本地证据确实没有的事实;上面安全硬边界管的动作(无恢复的破坏性/不可逆、prod/客户数据、目标外的合并/发布);推翻用户既定方向;证据无法裁决的重大产品取舍。其余都由你定:设计期安全 4 问自答写进方案、安全自检、选 owner 技能/模块/方案、测试与命名、目标内的下一步;提交/推送/合并按「硬纪律 1
|
|
44
|
+
- **自主决策,别把可判定的问题推给用户**:开发中只在真实阻塞时停下问人——缺失的凭据/授权;本地证据确实没有的事实;上面安全硬边界管的动作(无恢复的破坏性/不可逆、prod/客户数据、目标外的合并/发布);推翻用户既定方向;证据无法裁决的重大产品取舍。其余都由你定:设计期安全 4 问自答写进方案、安全自检、选 owner 技能/模块/方案、测试与命名、目标内的下一步;提交/推送/合并按「硬纪律 1」目标授权判断。写明假设后继续。判据:阻塞必须是只有用户能提供的东西(证据无法裁决的决定、凭据或访问授权、目标之外的许可、所有可读来源里都没有的事实);你自己能执行的下一步——修复及其测试评审、推分支开 MR/PR、重跑或重试失败/超时/结果不明的检查、查资料——不论花多少时间和次数都不是阻塞。以下几种说法都过不了这个判据:「你只让我调查/研究」——查原因、线上问题、测试挂了这类失败目标,根因证实后默认包含窄修复、回归测试、评审、推送功能分支并向开发目标开/更新 MR/PR,用户的明确限制照样约束(「只查/先别改」「先别推」、费用或次数上限),仓库契约标为先确认的区域、破坏性操作、新采购、合并/部署/生产动作也照常停(交接摘要里「这一步只读」只约束那一步);「推送/开 MR 是对外动作」——推功能分支、开 MR/PR 是常规研发动作,不是发布,自审、外部评审和必需 CI 过了就在同一轮转 ready,不以「MR 保持 Draft」收尾;「等 CI 跑完」——自己 MR 的流水线自己轮询到结果;你自己提出的次数、轮数、停止线——它们是你的估计,不是用户限制;只有用户把它采纳为上限(亲口给数字、说「最多/只」、或明确接受为上限)才算,单纯同意去做(「ok」「好」)不算;用户中途的追问——答完继续已授权的动作,不是 status-only;能自己查到或构造的事实、日志、测试数据、之前用过的账号与环境——先自己查。真要停下要合并授权时,一次请求覆盖整个交付计划,别按次追加。「owner」指技能或代码 owner,不是要用户指派的人。某一步被阻塞时先做完其余独立工作,再在结尾用 `proposed-next: blocked:` 说明具体阻塞。
|
|
45
45
|
- **用户主权**:AI 推荐、用户定。要改变用户既定方向时**始终先呈现+问,别径直下结论或代为决定**。你和另一个模型(codex 等)都同意也只是强信号、不是裁决。**仅当用户有既定方向、且你与第二模型都主张推翻它**(普通选项/口味/缺信息/评审 nit 不触发此结构):用户方向是默认、改动由模型举证,呈现时必须显式补两句——我们可能缺什么上下文、若改错代价是什么(详见 tighten-doc cross-model caveat)。
|
|
46
46
|
- **无证据不声称完成**:本轮没亲手跑过验证、没读到通过输出,就不说"完成/修好/通过/没问题",缺证据如实说缺(详见 product-rd-workflow 验证门)。
|
|
47
47
|
- **完整优先**:做完必要工作,不扩范围。报告/总结/进度说明不等于交付:结束前逐项核对用户请求和工作自身带出的后续项(失败检查、评审 finding、要同步的测试/文档;提交推送按「硬纪律 1」目标授权),能做的做完再收尾。阻塞交付的检查失败含基线问题,按 defect-diagnosis 诊断、安全修复、复测;真实阻塞才交回。
|
|
@@ -172,6 +172,7 @@ DENY_TEXT_AUTO="合并授权闸:该命令会开启 auto-merge / merge-when-pip
|
|
|
172
172
|
DENY_TEXT_SPELL="合并授权闸:授权有效但命令拼写不满足一次性立即合并要求——glab 需显式 --auto-merge=false(pipeline 运行中裸 merge 会默认转 auto-merge),gh 需显式 --merge/--squash/--rebase 策略,merge REST API 需显式 -X PUT。授权未消费,按上述拼写改写命令直接重试即可(无需用户重新授权)。"
|
|
173
173
|
DENY_TEXT_TARGET="合并授权闸:用户授权指向了特定 MR/PR 编号,但该命令的合并对象与之不符或无法识别。授权未消费——请显式点名该编号(如 glab mr merge <授权编号> --auto-merge=false --yes)后重试;确需合并其他 MR 请让用户重新授权。"
|
|
174
174
|
DENY_TEXT_MULTI="合并授权闸:同一条命令内检测到多个平台合并调用。每条命令只放行一个合并——把命令拆开逐条执行(单个授权下每个合并由用户分别授权;批量授权下每条命令消费 1 个额度,无需用户再次回复)。"
|
|
175
|
+
DENY_TEXT_HELP="合并授权闸:命令带 -h/--help,只会打印帮助、不会合并,因此拒绝且不消费授权。查看帮助请用 glab help mr merge 或 gh help pr merge;正式合并时去掉帮助参数,用户已给的授权仍然有效。"
|
|
175
176
|
DENY_TEXT_AMBIGUOUS="合并授权闸:检测到原始 HTTP 客户端、变更 method 的选项和 merge endpoint,但 method、transfer 边界、目标或动作数无法可靠关联。授权未消费——请改写成单个 curl/wget、单个明确 PUT method 和单个 merge URL 后重试。"
|
|
176
177
|
|
|
177
178
|
# Missing legacy epoch files are compatible with pre-upgrade grants. Once
|
|
@@ -741,7 +742,7 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
|
|
|
741
742
|
# gh explicit strategy; REST/GraphQL calls are immediate by API
|
|
742
743
|
# semantics). mid: the merge target id when statically extractable
|
|
743
744
|
# ("?" otherwise) — matched against a number-bound grant below.
|
|
744
|
-
hit=0; auto=0; spell=SPELLBAD; mid="?"; multi_seg=0
|
|
745
|
+
hit=0; auto=0; spell=SPELLBAD; mid="?"; multi_seg=0; help=0
|
|
745
746
|
if [ "$tool" = "glab" ]; then
|
|
746
747
|
# `accept` is glab's documented alias of `mr merge` (same help text).
|
|
747
748
|
if [ "${1:-}" = "mr" ] && { [ "${2:-}" = "merge" ] || [ "${2:-}" = "accept" ]; }; then
|
|
@@ -755,6 +756,7 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
|
|
|
755
756
|
# value-taking flags: consume the value so it is not mistaken
|
|
756
757
|
# for the MR id positional (`glab mr merge --sha abc 546`).
|
|
757
758
|
--sha|-m|--message|--squash-message) [ $# -ge 2 ] && shift ;;
|
|
759
|
+
-h|--help) help=1 ;;
|
|
758
760
|
-*) : ;;
|
|
759
761
|
*)
|
|
760
762
|
if [ "$mid" = "?" ] && [ -z "${id_seen:-}" ]; then
|
|
@@ -826,6 +828,7 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
|
|
|
826
828
|
# value-taking flags: consume the value so it is not mistaken
|
|
827
829
|
# for the PR id positional.
|
|
828
830
|
-b|--body|-F|--body-file|-t|--subject|--match-head-commit|-A|--author-email) [ $# -ge 2 ] && shift ;;
|
|
831
|
+
-h|--help) help=1 ;;
|
|
829
832
|
-*) : ;;
|
|
830
833
|
*)
|
|
831
834
|
if [ "$mid" = "?" ] && [ -z "${id_seen:-}" ]; then
|
|
@@ -876,7 +879,11 @@ printf '%s\n' "$masked" | tr ';|&(){}' '\n' | while IFS= read -r seg; do
|
|
|
876
879
|
# the id unresolvable so bound grants deny (unbound grants keep the
|
|
877
880
|
# agent-side duty to target the discussed MR — documented residual).
|
|
878
881
|
[ "$retarget" = 1 ] && mid="?"
|
|
879
|
-
|
|
882
|
+
# A help flag in flag position means the CLI prints help and merges
|
|
883
|
+
# nothing; deny it without touching any grant. A help token taken as a
|
|
884
|
+
# known flag's value never reaches here, and an unknown value-taking
|
|
885
|
+
# flag only makes this deny a real merge, never release one.
|
|
886
|
+
if [ "$help" = 1 ]; then echo DENY_HELP; elif [ "$auto" = 1 ]; then echo DENY_AUTO; else
|
|
880
887
|
echo "DENY_MR $mid $spell"
|
|
881
888
|
# A single segment carrying multiple aliased merge mutations emits a
|
|
882
889
|
# second DENY_MR so the >1 exactly-one-per-command guard denies it.
|
|
@@ -1066,6 +1073,9 @@ fi
|
|
|
1066
1073
|
if printf '%s\n' "$verdicts" | grep -q '^DENY_GIT_UNRESOLVED$'; then
|
|
1067
1074
|
deny "$DENY_TEXT_GIT_UNRESOLVED"
|
|
1068
1075
|
fi
|
|
1076
|
+
if printf '%s\n' "$verdicts" | grep -q '^DENY_HELP$'; then
|
|
1077
|
+
deny "$DENY_TEXT_HELP"
|
|
1078
|
+
fi
|
|
1069
1079
|
if printf '%s\n' "$verdicts" | grep -q '^DENY_AUTO'; then
|
|
1070
1080
|
deny "$DENY_TEXT_AUTO"
|
|
1071
1081
|
fi
|
|
@@ -13,6 +13,7 @@ import re
|
|
|
13
13
|
import shlex
|
|
14
14
|
import stat
|
|
15
15
|
import sys
|
|
16
|
+
import tempfile
|
|
16
17
|
|
|
17
18
|
# Hook assets may be installed read-only; importing the optional state helper
|
|
18
19
|
# must not create bytecode beside them.
|
|
@@ -626,6 +627,23 @@ def delivery_eligible(summary):
|
|
|
626
627
|
or summary['continuation_contract_visible'])
|
|
627
628
|
|
|
628
629
|
|
|
630
|
+
# Stops kept surviving the recheck by restating the blocker in a new term, so
|
|
631
|
+
# the test is stated as an invariant (who can act) and the observed terms are
|
|
632
|
+
# only examples of restatements that fail it.
|
|
633
|
+
NOT_BLOCKERS = (
|
|
634
|
+
'A blocker names something only the user can supply: a decision the evidence cannot settle, a credential '
|
|
635
|
+
'or access grant, permission the goal does not cover, or a fact absent from every source you can read. '
|
|
636
|
+
'A next step you can perform yourself is not a blocker, whatever it costs in time or runs: a fix with its '
|
|
637
|
+
'tests and review, a branch push and MR/PR, marking that MR/PR ready once your own checks pass instead of '
|
|
638
|
+
'leaving it in Draft, waiting on a CI run you can poll, a rerun or retry of a failed, timed-out or '
|
|
639
|
+
'inconclusive check, or a lookup. Restatements observed to fail this test: "you only asked me to investigate" — a failure or '
|
|
640
|
+
'diagnosis goal includes the verified fix, tests, review, branch push and MR/PR to the development target '
|
|
641
|
+
'unless an explicit user limit says otherwise (diagnosis only, no push); "pushing or opening an MR is outward-facing" — a feature branch '
|
|
642
|
+
'and its MR/PR are routine; a count, round or stop bar you proposed yourself, unless the user adopted it as '
|
|
643
|
+
'a limit; "the check can only restart '
|
|
644
|
+
'from scratch"; and facts, logs, test data or access you can find or reuse yourself. A clarifying question '
|
|
645
|
+
'is not a status-only request: answer it, then continue. ')
|
|
646
|
+
|
|
629
647
|
# One bounded recheck (host stop_hook_active) for stops that hand work back to
|
|
630
648
|
# the user; it names the real blockers and grants no authority.
|
|
631
649
|
DECISION_RECHECK = {'decision': 'block', 'reason': (
|
|
@@ -633,7 +651,8 @@ DECISION_RECHECK = {'decision': 'block', 'reason': (
|
|
|
633
651
|
'Real blockers are: missing credentials or authority; a fact unavailable from local evidence; '
|
|
634
652
|
'an action the safety rules gate (destructive or irreversible without recovery, production or '
|
|
635
653
|
'customer data, merge or publication outside the goal); overturning an established user direction; '
|
|
636
|
-
'or a material product tradeoff the evidence cannot settle.
|
|
654
|
+
'or a material product tradeoff the evidence cannot settle. ' + NOT_BLOCKERS +
|
|
655
|
+
'An ordinary change needs no human review, '
|
|
637
656
|
'sign-off or risk owner: run the self-review and external review yourself. '
|
|
638
657
|
'Small tests and routine development/test-environment operations within the authorized task '
|
|
639
658
|
'run directly with configured accounts; do not ask for per-run approval or invent a cost cap. '
|
|
@@ -649,10 +668,77 @@ DECISION_RECHECK = {'decision': 'block', 'reason': (
|
|
|
649
668
|
'This reminder supplies no new goal or authorization.')}
|
|
650
669
|
|
|
651
670
|
|
|
671
|
+
# Reader-facing documents edited in a session owe the tighten-doc closeout
|
|
672
|
+
# readback (the routing rule says so), yet it was skipped in 26 of 29 observed
|
|
673
|
+
# sessions. Agent-facing files are excluded: skill bodies, contracts, memory
|
|
674
|
+
# and scratch or temporary paths.
|
|
675
|
+
READER_DOC_SUFFIXES = ('.md', '.mdx', '.rst')
|
|
676
|
+
AGENT_DOC_NAMES = {'skill.md', 'agents.md', 'claude.md', 'memory.md'}
|
|
677
|
+
AGENT_DOC_DIRS = {'memory', '.claude', '.codex', '.git', 'node_modules', 'skills', 'agent-context',
|
|
678
|
+
'scratchpad'}
|
|
679
|
+
|
|
680
|
+
|
|
681
|
+
def reader_docs(edit_paths, cwd):
|
|
682
|
+
base = os.path.realpath(cwd) if isinstance(cwd, str) and cwd else None
|
|
683
|
+
temp_roots = tuple(os.path.realpath(root) + os.sep
|
|
684
|
+
for root in {tempfile.gettempdir(), '/tmp', '/private/tmp', '/var/folders'})
|
|
685
|
+
found = []
|
|
686
|
+
for path in edit_paths:
|
|
687
|
+
if not isinstance(path, str) or not path.lower().endswith(READER_DOC_SUFFIXES):
|
|
688
|
+
continue
|
|
689
|
+
real = os.path.realpath(path)
|
|
690
|
+
if base and (real == base or real.startswith(base + os.sep)):
|
|
691
|
+
parts = os.path.relpath(real, base).split(os.sep)
|
|
692
|
+
elif real.startswith(temp_roots):
|
|
693
|
+
continue
|
|
694
|
+
else:
|
|
695
|
+
parts = real.split(os.sep)
|
|
696
|
+
if parts[-1].lower() in AGENT_DOC_NAMES or any(p.lower() in AGENT_DOC_DIRS for p in parts[:-1]):
|
|
697
|
+
continue
|
|
698
|
+
found.append(real)
|
|
699
|
+
return found
|
|
700
|
+
|
|
701
|
+
|
|
702
|
+
def doc_closeout_note(payload):
|
|
703
|
+
path, cwd = payload.get('transcript_path'), payload.get('cwd')
|
|
704
|
+
if not isinstance(path, str) or not path:
|
|
705
|
+
return ''
|
|
706
|
+
cwd = cwd if isinstance(cwd, str) else os.getcwd()
|
|
707
|
+
try:
|
|
708
|
+
summary = transcript(path, cwd)
|
|
709
|
+
except TranscriptTruncated:
|
|
710
|
+
summary = context_transcript(path, cwd)
|
|
711
|
+
except (OSError, ValueError):
|
|
712
|
+
return ''
|
|
713
|
+
if any(skill.split(':')[-1] == 'tighten-doc' for skill in summary['completed_skills']):
|
|
714
|
+
return ''
|
|
715
|
+
docs = reader_docs(summary['edit_paths'], cwd)
|
|
716
|
+
if not docs:
|
|
717
|
+
return ''
|
|
718
|
+
names = ', '.join(sorted({os.path.basename(doc) for doc in docs})[:5])
|
|
719
|
+
return ('Document closeout: this session edited reader-facing documents ({}) without loading '
|
|
720
|
+
'tighten-doc. Load it and run its closeout readback on those documents before finishing; '
|
|
721
|
+
'the substance stays as the owning skill decided.'.format(names))
|
|
722
|
+
|
|
723
|
+
|
|
652
724
|
def proposed_next(payload):
|
|
653
725
|
if (not isinstance(payload, dict) or payload.get('hook_event_name') != 'Stop'
|
|
654
726
|
or payload.get('stop_hook_active') is not False):
|
|
655
727
|
return None
|
|
728
|
+
result = delivery_reminder(payload)
|
|
729
|
+
try:
|
|
730
|
+
note = doc_closeout_note(payload)
|
|
731
|
+
except Exception: # advisory: a failed document check never costs the reminder
|
|
732
|
+
note = ''
|
|
733
|
+
if not note:
|
|
734
|
+
return result
|
|
735
|
+
if not result:
|
|
736
|
+
return {'decision': 'block', 'reason': note + ' Then end with the same proposed-next: line. '
|
|
737
|
+
'This reminder supplies no new goal or authorization.'}
|
|
738
|
+
return {'decision': 'block', 'reason': note + ' ' + result['reason']}
|
|
739
|
+
|
|
740
|
+
|
|
741
|
+
def delivery_reminder(payload):
|
|
656
742
|
final = payload.get('last_assistant_message')
|
|
657
743
|
if not isinstance(final, str) or not final.strip() or machine_artifact(final):
|
|
658
744
|
return None
|
|
@@ -674,7 +760,8 @@ def proposed_next(payload):
|
|
|
674
760
|
'authorized and runnable, execute it now instead of waiting for another continue message. '
|
|
675
761
|
'For unrun, failed or inconclusive checks, continue available diagnosis, research, safe repair '
|
|
676
762
|
'and retesting; a report alone does not complete implementation. Respect explicit stop, '
|
|
677
|
-
'planning-only and status-only requests.
|
|
763
|
+
'planning-only and status-only requests. ' + NOT_BLOCKERS +
|
|
764
|
+
'If a user decision or missing authority/resource '
|
|
678
765
|
'prevents action, report the concrete blocker; do not invent work or bypass a failed gate. '
|
|
679
766
|
'This reminder supplies no new goal or authorization.')}
|
|
680
767
|
path = payload.get('transcript_path')
|
|
@@ -72,6 +72,11 @@ masked=$(printf '%s' "$cmd" | sed -E \
|
|
|
72
72
|
# NON-fire — `glab mr merge` / `gh pr merge` is the near-universal agent merge
|
|
73
73
|
# path, and the human-readable cleanup rule in worktree-isolation SKILL.md +
|
|
74
74
|
# bootstrap covers EVERY merge path regardless of this reminder.
|
|
75
|
+
# `gh help pr merge` / `glab help mr merge` print help (the merge guard's help
|
|
76
|
+
# denial points there). Remove only those literal invocations, never a prefix,
|
|
77
|
+
# so a real merge before or after them in the same command still matches.
|
|
78
|
+
masked=$(printf '%s' "$masked" | sed -E \
|
|
79
|
+
's/(glab|gh)[[:space:]]+help[[:space:]]+(mr|pr)[[:space:]]+(merge|accept)([[:space:]]|$)/ /g')
|
|
75
80
|
printf '%s' "$masked" | grep -Eq \
|
|
76
81
|
'glab[[:space:]]([^&|;]*[[:space:]])?mr[[:space:]]+(merge|accept)([[:space:]]|$)|gh[[:space:]]([^&|;]*[[:space:]])?pr[[:space:]]+merge([[:space:]]|$)' \
|
|
77
82
|
|| exit 0
|
|
@@ -1035,6 +1035,44 @@ for invalidation in '停止' '改成另一个功能'; do
|
|
|
1035
1035
|
done
|
|
1036
1036
|
unset RACE_SENT RACE_REACHED RACE_RESUME REAL_MV
|
|
1037
1037
|
|
|
1038
|
+
# A help probe merges nothing, so it must never consume a grant. Observed: an
|
|
1039
|
+
# agent added --auto-merge=false to `glab mr merge --help` to pass the spelling
|
|
1040
|
+
# check, the probe consumed the one-shot grant, and the real merge was denied.
|
|
1041
|
+
# Earlier cases leave epoch files for this session; clear them so each probe
|
|
1042
|
+
# meets a valid grant rather than an epoch mismatch (which also denies).
|
|
1043
|
+
rm -f "$VAUTH_DIR/$VSID.epoch" "$VAUTH_DIR/$VSID.grant-epoch"
|
|
1044
|
+
for help_cmd in 'glab mr merge --help --auto-merge=false' 'glab mr merge 123 -h --auto-merge=false --yes' \
|
|
1045
|
+
'gh pr merge 45 --merge --help' 'gh pr merge --squash -h'; do
|
|
1046
|
+
varm
|
|
1047
|
+
probe_sid deny "$FEAT_CWD" "$VSID" "$help_cmd"
|
|
1048
|
+
sentinel_state present "help probe kept the grant: $help_cmd"
|
|
1049
|
+
done
|
|
1050
|
+
varm
|
|
1051
|
+
reason_help=$(jq -nc --arg c 'glab mr merge --help --auto-merge=false' --arg w "$FEAT_CWD" --arg s "$VSID" \
|
|
1052
|
+
'{tool_input:{command:$c},cwd:$w,session_id:$s}' | TMPDIR="$tmp" bash "$GUARD")
|
|
1053
|
+
if printf '%s' "$reason_help" | grep -q 'glab help mr merge'; then pass=$((pass+1)); else
|
|
1054
|
+
fail=$((fail+1)); echo 'FAIL help denial must name the non-merge help form' >&2; fi
|
|
1055
|
+
rm -f "$VAUTH_DIR/$VSID"
|
|
1056
|
+
# A value-taking flag swallows a following --help: the command still merges.
|
|
1057
|
+
varm
|
|
1058
|
+
probe_sid allow "$FEAT_CWD" "$VSID" 'glab mr merge 123 -m --help --auto-merge=false --yes'
|
|
1059
|
+
sentinel_state absent 'a --help message value is a real merge and consumes the grant'
|
|
1060
|
+
probe allow "$FEAT_CWD" 'glab help mr merge'
|
|
1061
|
+
probe allow "$FEAT_CWD" 'gh help pr merge'
|
|
1062
|
+
# A help probe compounded with a real merge denies the whole command and keeps
|
|
1063
|
+
# the grant; a quoted --help message value is masked and stays a real merge.
|
|
1064
|
+
varm
|
|
1065
|
+
probe_sid deny "$FEAT_CWD" "$VSID" 'glab mr merge 123 --auto-merge=false --yes && glab mr merge --help'
|
|
1066
|
+
sentinel_state present 'help compounded with a merge keeps the grant'
|
|
1067
|
+
rm -f "$VAUTH_DIR/$VSID"
|
|
1068
|
+
varm
|
|
1069
|
+
probe_sid allow "$FEAT_CWD" "$VSID" 'glab mr merge 123 -m "--help" --auto-merge=false --yes'
|
|
1070
|
+
sentinel_state absent 'a quoted --help message is a real merge and consumes the grant'
|
|
1071
|
+
# Without any grant a help probe is denied and creates no grant.
|
|
1072
|
+
rm -f "$VAUTH_DIR/$VSID"
|
|
1073
|
+
probe_sid deny "$FEAT_CWD" "$VSID" 'gh pr merge 45 --merge --help'
|
|
1074
|
+
sentinel_state absent 'a help probe without a grant creates nothing'
|
|
1075
|
+
|
|
1038
1076
|
if [ "$fail" -ne 0 ]; then
|
|
1039
1077
|
echo "test_guard_merge_authorization: FAIL pass=$pass fail=$fail" >&2
|
|
1040
1078
|
exit 1
|
|
@@ -213,6 +213,17 @@ done
|
|
|
213
213
|
send '批量合并 3'
|
|
214
214
|
send '继续'
|
|
215
215
|
if [ ! -f "$SENT" ]; then pass=$((pass+1)); else fail=$((fail+1)); echo 'FAIL legacy batch still clears on neutral prompt' >&2; fi
|
|
216
|
+
# A host task notification reaches this hook with no field that tells it apart
|
|
217
|
+
# from typed text, so it is handled as a user message: it revokes single and
|
|
218
|
+
# counted grants (a stop typed in its markup must still revoke) and never arms.
|
|
219
|
+
note=$'<task-notification>\n<task-id>abc123</task-id>\n<status>completed</status>\n<summary>Background command "wait for CI" completed (exit code 0)</summary>\n</task-notification>'
|
|
220
|
+
send '批量合并 3'
|
|
221
|
+
send "$note"
|
|
222
|
+
if [ ! -f "$SENT" ]; then pass=$((pass+1)); else fail=$((fail+1)); echo 'FAIL a notification must revoke a counted grant' >&2; fi
|
|
223
|
+
send '合并'
|
|
224
|
+
send $'<task-notification>\n<summary>先别合并</summary>\n</task-notification>'
|
|
225
|
+
if [ ! -f "$SENT" ]; then pass=$((pass+1)); else fail=$((fail+1)); echo 'FAIL a stop inside notification markup must revoke' >&2; fi
|
|
226
|
+
expect_not_armed $'<task-notification>\n<summary>合并</summary>\n</task-notification>'
|
|
216
227
|
git -C "$tmp/repo" remote set-url origin 'https://user:password@example.invalid/team/project.git'
|
|
217
228
|
expect_not_armed '完成并合并 PR #123'
|
|
218
229
|
git -C "$tmp/repo" remote set-url origin 'git@example.invalid:team/project.git'
|
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
#!/usr/bin/env python3
|
|
2
2
|
"""Synthetic native Stop payloads; no real host state or conversations."""
|
|
3
|
+
import importlib.util
|
|
3
4
|
import json
|
|
4
5
|
import os
|
|
5
6
|
from pathlib import Path
|
|
@@ -7,6 +8,7 @@ import shutil
|
|
|
7
8
|
import subprocess
|
|
8
9
|
import tempfile
|
|
9
10
|
import unittest
|
|
11
|
+
from unittest.mock import patch
|
|
10
12
|
|
|
11
13
|
ROOT = Path(__file__).resolve().parents[1]
|
|
12
14
|
|
|
@@ -341,6 +343,115 @@ class ProposedNextTests(unittest.TestCase):
|
|
|
341
343
|
self.assertIn('supplies no new goal or authorization', result['reason'])
|
|
342
344
|
self.assertEqual(self.run_hook(dict(payload, stop_hook_active=True)), {})
|
|
343
345
|
|
|
346
|
+
def test_diagnosis_scope_and_question_turn_are_named_in_both_reminders(self):
|
|
347
|
+
# Observed stops after a verified root cause: the agent read "missing
|
|
348
|
+
# authority" as "you only asked me to investigate", and a clarifying
|
|
349
|
+
# question as a status-only request. Both reminders must close those terms.
|
|
350
|
+
for events in ([], self.claude_load()):
|
|
351
|
+
self.events(events)
|
|
352
|
+
for text in (
|
|
353
|
+
'Fixing it changes shared code and opens an MR, beyond the investigation you asked for.\n'
|
|
354
|
+
'proposed-next: blocked: limit fix, consistency test and MR — waiting for you to authorize the code change',
|
|
355
|
+
'这一轮你只问了一个问题,我只解释了现状。\n'
|
|
356
|
+
'proposed-next: blocked: 上限修复、补测试、提 MR——改共享仓库需要你确认',
|
|
357
|
+
'Root cause verified.\nproposed-next: open a branch, fix the limit, add the test and open the MR',
|
|
358
|
+
'Required CI passed; the advisory review timed out and its evidence was cleared.\n'
|
|
359
|
+
'proposed-next: blocked: complete the CI review — no resume handle; a retry restarts from scratch',
|
|
360
|
+
'Both MRs pushed; required CI is still running, the MRs stay in Draft.\n'
|
|
361
|
+
'proposed-next: wait for CI to finish and for your merge confirmation'):
|
|
362
|
+
with self.subTest(text=text):
|
|
363
|
+
result = self.run_hook(dict(self.payload, last_assistant_message=text))
|
|
364
|
+
self.assert_block(result)
|
|
365
|
+
self.assertIn('you only asked me to investigate', result['reason'])
|
|
366
|
+
self.assertIn('failure or diagnosis goal includes the verified fix', result['reason'])
|
|
367
|
+
# Closing the terms must not widen past an explicit user
|
|
368
|
+
# limit: a "fix locally, do not push" instruction still binds.
|
|
369
|
+
self.assertIn('unless an explicit user limit says otherwise (diagnosis only, no push)',
|
|
370
|
+
result['reason'])
|
|
371
|
+
self.assertIn('clarifying question is not a status-only request', result['reason'])
|
|
372
|
+
self.assertIn('outward-facing" — a feature branch and its MR/PR are routine', result['reason'])
|
|
373
|
+
self.assertIn('stop bar you proposed yourself', result['reason'])
|
|
374
|
+
self.assertIn('A blocker names something only the user can supply', result['reason'])
|
|
375
|
+
self.assertIn('a rerun or retry of a failed, timed-out or inconclusive check', result['reason'])
|
|
376
|
+
self.assertIn('marking that MR/PR ready once your own checks pass', result['reason'])
|
|
377
|
+
self.assertIn('waiting on a CI run you can poll', result['reason'])
|
|
378
|
+
self.assertIn('supplies no new goal or authorization', result['reason'])
|
|
379
|
+
|
|
380
|
+
def doc_edit(self, relative, tool_id='doc', tool='Write'):
|
|
381
|
+
target = str(self.root / relative)
|
|
382
|
+
return [
|
|
383
|
+
{'type': 'assistant', 'message': {'content': [{'type': 'tool_use', 'id': tool_id,
|
|
384
|
+
'name': tool, 'input': {'file_path': target, 'content': 'x'}}]}},
|
|
385
|
+
{'type': 'user', 'message': {'content': [{'type': 'tool_result',
|
|
386
|
+
'tool_use_id': tool_id, 'is_error': False, 'content': 'written'}]}}]
|
|
387
|
+
|
|
388
|
+
def test_reader_doc_edit_without_tighten_doc_gets_one_closeout_reminder(self):
|
|
389
|
+
# Observed: 26 of 29 sessions that edited plans, specs, READMEs or
|
|
390
|
+
# handoff documents never loaded tighten-doc before finishing.
|
|
391
|
+
status = 'Plan updated.\nproposed-next: none — status only'
|
|
392
|
+
for relative in ('docs/plans/rollout.md', 'README.md', 'handoffs/state.md', 'specs/9-x/plan.md'):
|
|
393
|
+
with self.subTest(relative=relative):
|
|
394
|
+
self.events(self.doc_edit(relative))
|
|
395
|
+
result = self.run_hook(dict(self.payload, last_assistant_message=status))
|
|
396
|
+
self.assert_block(result)
|
|
397
|
+
self.assertIn('tighten-doc', result['reason'])
|
|
398
|
+
self.assertIn(Path(relative).name, result['reason'])
|
|
399
|
+
self.assertIn('supplies no new goal or authorization', result['reason'])
|
|
400
|
+
self.assertEqual(self.run_hook(dict(self.payload, last_assistant_message=status,
|
|
401
|
+
stop_hook_active=True)), {})
|
|
402
|
+
|
|
403
|
+
def test_doc_reminder_is_quiet_after_tighten_doc_or_for_agent_files(self):
|
|
404
|
+
status = 'Plan updated.\nproposed-next: none — status only'
|
|
405
|
+
self.events(self.claude_load('tighten-doc') + self.doc_edit('docs/plans/rollout.md'))
|
|
406
|
+
self.assertEqual(self.run_hook(dict(self.payload, last_assistant_message=status)), {})
|
|
407
|
+
for relative in ('skills/x/SKILL.md', 'AGENTS.md', 'CLAUDE.md', 'memory/note.md',
|
|
408
|
+
'.claude/notes.md', 'src/app.py', 'notes.txt'):
|
|
409
|
+
with self.subTest(relative=relative):
|
|
410
|
+
self.events(self.doc_edit(relative))
|
|
411
|
+
self.assertEqual(self.run_hook(dict(self.payload, last_assistant_message=status)), {})
|
|
412
|
+
|
|
413
|
+
def test_doc_reminder_joins_a_continuation_reminder(self):
|
|
414
|
+
self.events(self.doc_edit('docs/plans/rollout.md'))
|
|
415
|
+
result = self.run_hook(dict(self.payload,
|
|
416
|
+
last_assistant_message='proposed-next: run the remaining local checks'))
|
|
417
|
+
self.assert_block(result)
|
|
418
|
+
self.assertIn('execute it now', result['reason'])
|
|
419
|
+
self.assertIn('tighten-doc', result['reason'])
|
|
420
|
+
|
|
421
|
+
def test_doc_reminder_covers_each_file_edit_tool(self):
|
|
422
|
+
# Edits are seen through the file-edit tool calls the transcript records;
|
|
423
|
+
# shell writes are outside this check by design.
|
|
424
|
+
for tool in ('Edit', 'MultiEdit'):
|
|
425
|
+
with self.subTest(tool=tool):
|
|
426
|
+
self.events(self.doc_edit('docs/handoff.md', tool=tool))
|
|
427
|
+
result = self.run_hook()
|
|
428
|
+
self.assert_block(result)
|
|
429
|
+
self.assertIn('handoff.md', result['reason'])
|
|
430
|
+
|
|
431
|
+
def test_unreadable_doc_path_never_costs_the_delivery_reminder(self):
|
|
432
|
+
# The document check is advisory; a path it cannot resolve (an embedded
|
|
433
|
+
# NUL makes realpath raise) must not replace the continuation reminder
|
|
434
|
+
# with the "reminder unavailable" notice.
|
|
435
|
+
self.events(self.doc_edit('docs/plans/roll\x00out.md'))
|
|
436
|
+
result = self.run_hook(dict(self.payload,
|
|
437
|
+
last_assistant_message='proposed-next: run the remaining local checks'))
|
|
438
|
+
self.assert_block(result)
|
|
439
|
+
self.assertIn('execute it now', result['reason'])
|
|
440
|
+
|
|
441
|
+
def test_doc_check_failure_never_costs_the_delivery_reminder(self):
|
|
442
|
+
# The document check is advisory: whatever it raises, the continuation
|
|
443
|
+
# reminder it would have joined is still returned.
|
|
444
|
+
self.events(self.doc_edit('docs/plans/rollout.md'))
|
|
445
|
+
spec = importlib.util.spec_from_file_location('doc_check_probe', self.hooks / 'host-input.py')
|
|
446
|
+
module = importlib.util.module_from_spec(spec)
|
|
447
|
+
spec.loader.exec_module(module)
|
|
448
|
+
payload = dict(self.payload, last_assistant_message='proposed-next: run the remaining local checks')
|
|
449
|
+
with patch.object(module, 'reader_docs', side_effect=RuntimeError('unexpected')):
|
|
450
|
+
result = module.proposed_next(payload)
|
|
451
|
+
self.assert_block(result)
|
|
452
|
+
self.assertIn('execute it now', result['reason'])
|
|
453
|
+
self.assertNotIn('tighten-doc', result['reason'])
|
|
454
|
+
|
|
344
455
|
def test_quoted_actions_do_not_turn_a_status_handoff_into_work(self):
|
|
345
456
|
self.events(self.claude_load())
|
|
346
457
|
for suffix in ('\n> proposed-next: deploy', '\n```text\nproposed-next: deploy\n```'):
|
|
@@ -113,6 +113,18 @@ probe_json remind 'glab mr merge 123 --yes' '{"stdout":"Merged !123"}'
|
|
|
113
113
|
# --- --help / -h is not a merge → quiet ---
|
|
114
114
|
probe quiet 'gh pr merge --help'
|
|
115
115
|
probe quiet 'glab mr merge -h'
|
|
116
|
+
# The help subcommand form, which the merge guard's help denial points to,
|
|
117
|
+
# prints help and merges nothing.
|
|
118
|
+
probe quiet 'gh help pr merge' 'Merge a pull request on GitHub.'
|
|
119
|
+
probe quiet 'glab help mr merge' 'Merges a merge request.'
|
|
120
|
+
probe quiet 'gh help pr merge 2>&1 | grep -- --match-head-commit' '--match-head-commit SHA'
|
|
121
|
+
# A real merge beside a help lookup in one command still reminds, whichever
|
|
122
|
+
# comes first: only the help invocation itself is set aside.
|
|
123
|
+
probe remind 'gh help pr merge >/dev/null; gh pr merge 45 --merge' 'Merged'
|
|
124
|
+
probe remind 'glab help mr merge && glab mr merge 123 --yes' 'Merged !123'
|
|
125
|
+
probe remind 'gh pr merge 45 --merge # see gh help pr merge' 'Merged'
|
|
126
|
+
probe remind 'gh pr merge 45 --merge $(gh help pr merge >/dev/null)' 'Merged'
|
|
127
|
+
probe remind 'glab mr merge 123 --yes; glab help mr merge' 'Merged !123'
|
|
116
128
|
# a successful-looking string response still reminds
|
|
117
129
|
probe remind 'glab mr merge 123 --yes' 'Merged! https://.../merge_requests/123'
|
|
118
130
|
|
|
@@ -16,7 +16,7 @@ Diagnose and fix from evidence; route prevention to product, architecture, devel
|
|
|
16
16
|
- Do not call a workaround the fix unless the owner explicitly accepts the tradeoff and residual risk is recorded.
|
|
17
17
|
- Do not start broad refactoring while the cause is unknown. Isolate and fix first; refactor after the behavior is understood.
|
|
18
18
|
- Do not stop at "this line was wrong" when the defect reveals a missing contract, guardrail, test, review check, or skill rule.
|
|
19
|
-
- Do not state or act on a root-cause verdict — even as a confident aside — before you have read the failing owner's own evidence with your own eyes (assertion diff for a test, stack/exception for a crash, trace/log slice for a production symptom, source only when it is itself the failing artifact). Until then, label every cause as a hypothesis and name the evidence that would confirm or reject it. This applies to your OWN analysis, not only to LLM-proposed causes. Mitigation is exempt: you may roll back, flag-off, or shed traffic from symptoms while cause stays marked unknown — what is forbidden is choosing or applying a *fix
|
|
19
|
+
- Do not state or act on a root-cause verdict — even as a confident aside — before you have read the failing owner's own evidence with your own eyes (assertion diff for a test, stack/exception for a crash, trace/log slice for a production symptom, source only when it is itself the failing artifact). Until then, label every cause as a hypothesis and name the evidence that would confirm or reject it. This applies to your OWN analysis, not only to LLM-proposed causes. Mitigation is exempt: you may roll back, flag-off, or shed traffic from symptoms while cause stays marked unknown — what is forbidden is choosing or applying a *fix*, or handing the user a decision or tradeoff that rests on the cause, as though a cause is proven. While a falsifying check is runnable, run it before asking anyone to choose.
|
|
20
20
|
|
|
21
21
|
## Phase A: Diagnose
|
|
22
22
|
|
|
@@ -47,7 +47,7 @@ Diagnose and fix from evidence; route prevention to product, architecture, devel
|
|
|
47
47
|
- **A test that passes alone and fails in the suite must be bisected over the tests that run before it** — only for a failure that reproduces on every run under a fixed serial order (parallel or intermittent failures keep the failing schedule and validate each kept or dropped subset over repeated runs per the flaky rule, or route to concurrency diagnosis): halve the preceding set in order and keep a failing half; when neither half fails alone, remove one chunk at a time and keep the reduced set whenever the failure persists without that chunk, then halve the chunk size and repeat until every remaining chunk is needed (a minimal ordered polluting subsequence) or the shared fixture/state is found; the same reduction isolates a failing input, config, or dataset when no commit range exists (moves in `references/diagnosis-playbook.md`).
|
|
48
48
|
- Identify whether the failure is in product logic, contract mapping, persistence, cache, async processing, dependency behavior, runtime config, release state, or test setup.
|
|
49
49
|
- If the failure appears only in tests or CI, classify the test evidence before changing code: deterministic assertion, fixed external data, live infrastructure, random/log-only behavior, long sleep, allow-failure gate, generated/vendor test, or deploy/build-only pipeline.
|
|
50
|
-
- **A red CI pipeline/job is not by itself a code/dependency defect — read the failing job's own trace (not the red/green summary) and classify the cause before touching code.**
|
|
50
|
+
- **A red CI pipeline/job is not by itself a code/dependency defect — read the failing job's own trace (not the red/green summary) and classify the cause before touching code.** Classify it as a trigger-variant artifact, a retriable infra flake, a deterministic non-code infra fault, or a genuine code/dependency failure using the checks in `references/diagnosis-playbook.md` (Red CI cause classes) before disowning or touching code.
|
|
51
51
|
- **For a failing test, read the actual assertion error (Expected/Actual) and the failing test body FIRST — before diagnosing flakiness, concurrency, shared state, mock setup, or external-dependency causes.** A passing/failing count, a `REQUEST POST`-style debug log line, or "passes in isolation, fails in suite" is a symptom, not the assertion evidence; naming a cause from those alone is the exact failure this skill exists to prevent. If the test is suspected flaky, rerun N times and record the pass/fail ratio before calling it flaky (100% reproducible failure is deterministic, not flaky), and confirm the test's network/dependency boundary by reading its setup (e.g. whether it is already mocked) rather than inferring from logs.
|
|
52
52
|
|
|
53
53
|
3. Hypothesize.
|
|
@@ -76,7 +76,7 @@ Diagnose and fix from evidence; route prevention to product, architecture, devel
|
|
|
76
76
|
5. Verify cause.
|
|
77
77
|
- Prove the cause with evidence.
|
|
78
78
|
- **A diagnosis licenses a fix only when it explains both causality and incorrectness**: how the defect produced this failure on the failing path, and why that code, data, or config is wrong against its contract — so the fix covers related failures. A change that makes the failure disappear without the second half is a symptom patch; a genuine defect that cannot be linked to this failure is a different bug — record it, never ship it as this cause.
|
|
79
|
-
- Report query/lookup evidence by cardinality: a data query, log search, or identity resolution that returns 0, 1, or N matches reports each of those outcomes distinctly — never silently take the first row of N, and never treat 0 rows as "no evidence collected" (an empty result over a named scope IS evidence: record which scopes matched and which were empty).
|
|
79
|
+
- Report query/lookup evidence by cardinality: a data query, log search, or identity resolution that returns 0, 1, or N matches reports each of those outcomes distinctly — never silently take the first row of N, and never treat 0 rows as "no evidence collected" (an empty result over a named scope IS evidence: record which scopes matched and which were empty). A zero failure count says something only after the scope's exposure is confirmed — the path was live and exercised in that window; zero failures from a path that never ran is no data, not health.
|
|
80
80
|
- When the cause is environment/toolchain state, prove it from the tool that owns that state, not only from the high-level wrapper. A wrapper failure is a symptom until the underlying compiler, generator, runtime registry, dependency resolver, or platform destination evidence explains it.
|
|
81
81
|
- If disproven, return to hypotheses instead of guessing.
|
|
82
82
|
- Separate symptom, immediate cause, contributing factors, and prevention.
|
|
@@ -93,6 +93,10 @@ The hypothesize → instrument → verify loop must not run forever, and escalat
|
|
|
93
93
|
|
|
94
94
|
## Phase B: Fix And Verify
|
|
95
95
|
|
|
96
|
+
A failure or diagnosis goal carries this phase. Once the cause is verified, fix, test, review, push the branch and open or update the MR/PR to the repository's development target in the same delivery. "You only asked me to investigate" is not missing authority, and a "read-only" note on one step of a handoff binds that step only.
|
|
97
|
+
|
|
98
|
+
- Every explicit user limit and existing gate must still stop the step it covers, for example diagnosis only (只查 / 先别改), no push, a cost or run cap, a repository's confirm-first areas, the shared-gate route below, destructive actions, purchases, and merge, deploy or production steps.
|
|
99
|
+
|
|
96
100
|
1. Fix minimally.
|
|
97
101
|
- Address the proven cause with the smallest correct change.
|
|
98
102
|
- Preserve contracts unless the task explicitly requires a contract change.
|
|
@@ -65,6 +65,10 @@ Locating the defect is usually the most expensive phase — harder than reproduc
|
|
|
65
65
|
| Wrong value observed downstream | upstream trace | follow the value backward to the first point where a correct input produced a wrong output; that transition is the defect and the observation point is only where it surfaced — fix there when it is owned and changeable, otherwise record the upstream cause and enforce the contract at the nearest owned boundary |
|
|
66
66
|
| Production symptom that cannot be re-triggered | telemetry walk | alert → exemplar trace → span tree → logs by trace-id (SKILL.md Phase A.4); group the failing population by attribute and compare it against the baseline to find what is different about failing requests |
|
|
67
67
|
|
|
68
|
+
## Red CI Cause Classes
|
|
69
|
+
|
|
70
|
+
- A red CI pipeline/job is not by itself a code/dependency defect; read the failing job's own trace (not the red/green summary) and classify the cause before touching code. Refining the test-evidence classes in the entrypoint's Phase A Isolate step into why-CI-is-red-but-code-may-be-fine: (a) **trigger-variant artifact** — when the same job runs under more than one trigger-scoped config (branch/push vs merge-request vs manual/scheduled), the trigger can resolve different default variables or a different checked-out ref, so a red on a non-gating trigger may be benign — but conclude that ONLY after confirming the *same failing check* ran and is green on the gating path (a gating pipeline that is overall green yet never runs the failing check does not clear it; if that check's coverage is unique to the non-gating trigger — e.g. a scheduled/manual-only suite — treat it as genuine (d), not a variant artifact); (b) **retriable infra flake** — e.g. a shared-runner lock collision: confirm per the entrypoint's flaky-test rule (rerun N times and record the ratio) (rerun N times + record the ratio; a 100%-reproducible red is deterministic, not a flake), and still read the red run's trace for the collision signature, since one green rerun cannot separate an infra flake from a genuine intermittent code bug; (c) **deterministic non-code infra fault** — toolchain/runner-image drift, stale cache/vendored artifact, credential/quota expiry: reproduces identically (NOT a flake) AND must be shown **code-independent** before disowning — confirm the same failure reproduces on a known-good baseline (parent/last-good commit, or a build without the change) under the same runner/toolchain; if the red appears only *with* the change it is (d) however much it resembles drift → route to platform/infra only after that baseline check, then do not attribute to code; (d) **genuine code/dependency failure**. (Single-variant repos skip the (a) check; the trace-first and cause-classification still apply.)
|
|
71
|
+
|
|
68
72
|
## Probe Ordering And The Hypothesis Log
|
|
69
73
|
|
|
70
74
|
Order probes; do not merely list hypotheses. For each candidate cause record the observation only it produces, the observation that cannot occur if it is true (the falsifier — collect this one first), what the probe costs, and what it risks. Then apply the entrypoint's one ordering rule: safety is a filter, not a rank — reject any probe outside the safety boundary first; rank the rest by alternatives ruled out per unit of cost; break ties by likelihood, then residual risk. Watch for confounders (a probe run from the wrong host, credential, or network position fails for its own reasons), side effects of active probes (more CPU changes race timing; verbose logging worsens latency — revert before the next probe), and probes that are only suggestive (races, deadlocks): record the evidence grade next to the result.
|