@ccoalm/ccl-skills 0.18.5 → 0.18.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-policy.md +5 -4
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +11 -9
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/host-input.py +135 -11
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-review-covers-head.sh +34 -9
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/skill-extraction-gate-stop.sh +14 -3
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/skill-loading.py +39 -2
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_host_input.py +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_proposed_next.py +180 -3
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_review_covers_head.sh +19 -3
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_skill_loading.py +77 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +4 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +11 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +65 -11
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +158 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +39 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +48 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +5 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +4 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/review-reception.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-sync-pointers.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +11 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +4 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_register_pending_exclusion.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_route_drift.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +4 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_source_register_lifecycle.sh +9 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_sync_pointers.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_resolution.sh +32 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_surface_binding.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_prose_target.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_locale_independent_gates.sh +264 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_loading_budget.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +2 -0
- package/dist/assets/release.json +61 -56
- package/package.json +1 -1
|
@@ -64,6 +64,9 @@ base_url = "https://example.invalid"
|
|
|
64
64
|
[future_tools]
|
|
65
65
|
enabled = true
|
|
66
66
|
|
|
67
|
+
[watch]
|
|
68
|
+
enabled = true
|
|
69
|
+
|
|
67
70
|
[permission]
|
|
68
71
|
deny = ["Read"]
|
|
69
72
|
|
|
@@ -164,6 +167,7 @@ cat >"$WORK/bin/kimi" <<'KIMI_STUB'
|
|
|
164
167
|
#!/usr/bin/env bash
|
|
165
168
|
set -u
|
|
166
169
|
state="$REVIEW_WRAPPER_TEST_STATE"
|
|
170
|
+
printf '%s\t%s\n' invocation "${KIMI_CODE_WATCH-}" >>"$state/kimi_watch_all"
|
|
167
171
|
if [ "${1:-}" = --version ]; then
|
|
168
172
|
if [ "${STUB_BEHAVIOR:-}" = version_hang ]; then
|
|
169
173
|
trap '' TERM
|
|
@@ -176,6 +180,23 @@ fi
|
|
|
176
180
|
if [ "${1:-}" = doctor ]; then
|
|
177
181
|
touch "$state/kimi_doctor_checked"
|
|
178
182
|
printf '%s\n' "$*" >"$state/kimi_doctor_args"
|
|
183
|
+
printf '%s\t%s\n' doctor "${KIMI_CODE_WATCH-}" >>"$state/kimi_watch"
|
|
184
|
+
# Keep every validated config so a retry can be compared against the attempt
|
|
185
|
+
# it replaced, rather than only asserted about one property at a time.
|
|
186
|
+
doctor_n=$(( $(cat "$state/kimi_doctor_count" 2>/dev/null || echo 0) + 1 ))
|
|
187
|
+
printf '%s' "$doctor_n" >"$state/kimi_doctor_count"
|
|
188
|
+
if [ -n "${3:-}" ] && [ -f "$3" ]; then
|
|
189
|
+
cp "$3" "$state/kimi_doctor_config.$doctor_n"
|
|
190
|
+
fi
|
|
191
|
+
if [ "${STUB_BEHAVIOR:-pass}" = doctor_reject_all ]; then
|
|
192
|
+
printf '%s\n' 'config rejected' >&2
|
|
193
|
+
exit 1
|
|
194
|
+
fi
|
|
195
|
+
if [ "${STUB_BEHAVIOR:-pass}" = doctor_reject_watch ] \
|
|
196
|
+
&& [ -n "${3:-}" ] && [ -f "$3" ] && grep -q '^\[watch\]' "$3"; then
|
|
197
|
+
printf '%s\n' 'unknown config table: watch' >&2
|
|
198
|
+
exit 1
|
|
199
|
+
fi
|
|
179
200
|
exit 0
|
|
180
201
|
fi
|
|
181
202
|
touch "$state/kimi_invoked"
|
|
@@ -199,6 +220,12 @@ while [ "$#" -gt 0 ]; do
|
|
|
199
220
|
*) shift ;;
|
|
200
221
|
esac
|
|
201
222
|
done
|
|
223
|
+
if [[ "$prompt" = "No-tools capability probe."* ]]; then
|
|
224
|
+
watch_phase=capability
|
|
225
|
+
else
|
|
226
|
+
watch_phase=formal
|
|
227
|
+
fi
|
|
228
|
+
printf '%s\t%s\n' "$watch_phase" "${KIMI_CODE_WATCH-}" >>"$state/kimi_watch"
|
|
202
229
|
printf '%s' "$has_stream" >"$state/kimi_stream"
|
|
203
230
|
printf '%s' "$has_model" >"$state/kimi_model_override"
|
|
204
231
|
printf '%s' "$has_skills_dir" >"$state/kimi_skills_dir"
|
|
@@ -213,7 +240,9 @@ if [ -f "$KIMI_CODE_HOME/config.toml" ]; then
|
|
|
213
240
|
grep -q '/trusted/local/hook-marker' "$KIMI_CODE_HOME/config.toml" && touch "$state/kimi_hooks_preserved"
|
|
214
241
|
grep -q 'enabled = \["\*"\]' "$KIMI_CODE_HOME/config.toml" \
|
|
215
242
|
&& touch "$state/kimi_no_tools_configured"
|
|
216
|
-
|
|
243
|
+
grep -c '^# Generated by code-review' "$KIMI_CODE_HOME/config.toml" \
|
|
244
|
+
>"$state/kimi_generated_marker_count"
|
|
245
|
+
python3 - "$KIMI_CODE_HOME/config.toml" "$state/kimi_runtime_config_mode" "$state/kimi_dotted_config_preserved" "$state/kimi_watch_disabled" "$state/kimi_watch_absent" <<'PY'
|
|
217
246
|
import os, sys, tomllib
|
|
218
247
|
from pathlib import Path
|
|
219
248
|
|
|
@@ -222,6 +251,10 @@ with open(sys.argv[2], "w", encoding="utf-8") as stream:
|
|
|
222
251
|
config = tomllib.loads(Path(sys.argv[1]).read_text(encoding="utf-8"))
|
|
223
252
|
if config.get("providers") and config.get("models"):
|
|
224
253
|
Path(sys.argv[3]).touch()
|
|
254
|
+
if config.get("watch") == {"enabled": False}:
|
|
255
|
+
Path(sys.argv[4]).touch()
|
|
256
|
+
if "watch" not in config:
|
|
257
|
+
Path(sys.argv[5]).touch()
|
|
225
258
|
PY
|
|
226
259
|
fi
|
|
227
260
|
if [ -f "$KIMI_CODE_HOME/mcp.json" ]; then
|
|
@@ -258,6 +291,58 @@ if [[ "$prompt" = "No-tools capability probe."* ]]; then
|
|
|
258
291
|
printf '%s\n' 'EMFILE: too many open files, watch' >&2
|
|
259
292
|
exit 1
|
|
260
293
|
fi
|
|
294
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_quota ]; then
|
|
295
|
+
printf '%s\n' "provider.auth_error: 403 You've reached your weekly (7-day) usage limit" >&2
|
|
296
|
+
exit 1
|
|
297
|
+
fi
|
|
298
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_auth ]; then
|
|
299
|
+
printf '%s\n' 'provider.auth_error: 403 authentication required' >&2
|
|
300
|
+
exit 1
|
|
301
|
+
fi
|
|
302
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_auth_quota_mention ]; then
|
|
303
|
+
printf '%s\n' 'provider.auth_error: 403 authentication required; quota metadata unavailable' >&2
|
|
304
|
+
exit 1
|
|
305
|
+
fi
|
|
306
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_auth_limit ]; then
|
|
307
|
+
printf '%s\n' 'provider.auth_error: 403 authentication required; usage limit metadata unavailable' >&2
|
|
308
|
+
exit 1
|
|
309
|
+
fi
|
|
310
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_auth_word ]; then
|
|
311
|
+
printf '%s\n' 'tool surface rejected; authentication metadata unavailable' >&2
|
|
312
|
+
exit 1
|
|
313
|
+
fi
|
|
314
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_offset ]; then
|
|
315
|
+
printf '%s\n' 'provider failed at byte offset 4290 and line 14030' >&2
|
|
316
|
+
exit 1
|
|
317
|
+
fi
|
|
318
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_exact_offset ]; then
|
|
319
|
+
printf '%s\n' 'provider failed at byte offset 429 and line 403' >&2
|
|
320
|
+
exit 1
|
|
321
|
+
fi
|
|
322
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_rate_limit_403 ]; then
|
|
323
|
+
printf '%s\n' 'provider.auth_error: 403 rate limit exceeded' >&2
|
|
324
|
+
exit 1
|
|
325
|
+
fi
|
|
326
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_http_quota ]; then
|
|
327
|
+
printf '%s\n' 'provider error: HTTP 429 Too Many Requests' >&2
|
|
328
|
+
exit 1
|
|
329
|
+
fi
|
|
330
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_decode_offset ]; then
|
|
331
|
+
printf '%s\n' 'provider failed to decode offset 429' >&2
|
|
332
|
+
exit 1
|
|
333
|
+
fi
|
|
334
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_auth_weekly_metadata ]; then
|
|
335
|
+
printf '%s\n' 'provider.auth_error: 403 authentication required; weekly usage limit metadata unavailable' >&2
|
|
336
|
+
exit 1
|
|
337
|
+
fi
|
|
338
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_auth_too_many ]; then
|
|
339
|
+
printf '%s\n' 'provider.auth_error: HTTP 429 Too Many Requests' >&2
|
|
340
|
+
exit 1
|
|
341
|
+
fi
|
|
342
|
+
if [ "${STUB_BEHAVIOR:-pass}" = capability_forbidden ]; then
|
|
343
|
+
printf '%s\n' 'provider error: 403 Forbidden' >&2
|
|
344
|
+
exit 1
|
|
345
|
+
fi
|
|
261
346
|
if [ "${STUB_BEHAVIOR:-pass}" = capability_missing ]; then
|
|
262
347
|
printf '%s\n' '{"role":"meta","type":"system.version","version":"future"}'
|
|
263
348
|
exit 0
|
|
@@ -546,6 +631,7 @@ content = {
|
|
|
546
631
|
"foreign_tool": clean,
|
|
547
632
|
"nested_tool": clean,
|
|
548
633
|
"unknown_event": clean,
|
|
634
|
+
"doctor_reject_watch": clean,
|
|
549
635
|
"multi_message": clean,
|
|
550
636
|
"mcp_bad_chunk": clean,
|
|
551
637
|
"mcp_eof_confirmation": clean,
|
|
@@ -1321,15 +1407,37 @@ for client in kimi codex; do
|
|
|
1321
1407
|
done
|
|
1322
1408
|
done
|
|
1323
1409
|
|
|
1324
|
-
|
|
1410
|
+
rm -f "$WORK/state/kimi_watch" "$WORK/state/kimi_watch_all" "$WORK/state/kimi_watch_disabled"
|
|
1411
|
+
out="$(KIMI_CODE_WATCH=1 run_kimi pass)"; rc=$?
|
|
1325
1412
|
check "Kimi clean result passes with fixed Moonshot attribution" \
|
|
1326
|
-
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ "$(field concern_results.0.concern "$out")" = correctness ] && [ "$(field reviewer_family "$out")" = moonshot ] && [ "$(field provider "$out")" = kimi-cli ] && [ "$(field model "$out")" = None ] && [ "$(dir_mode "$WORK/kimi-source/config.toml")" = 400 ] && [ "$(cat "$WORK/state/kimi_runtime_config_mode")" = 600 ]'
|
|
1413
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ "$(field concern_results.0.concern "$out")" = correctness ] && [ "$(field reviewer_family "$out")" = moonshot ] && [ "$(field provider "$out")" = kimi-cli ] && [ "$(field model "$out")" = None ] && [ "$(dir_mode "$WORK/kimi-source/config.toml")" = 400 ] && [ "$(cat "$WORK/state/kimi_runtime_config_mode")" = 600 ] && [ "$(wc -l < "$WORK/state/kimi_watch" | tr -d " ")" = 3 ] && [ "$(grep -c "^doctor[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && [ "$(grep -c "^capability[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && [ "$(grep -c "^formal[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && [ "$(wc -l < "$WORK/state/kimi_watch_all" | tr -d " ")" = 3 ] && ! grep -qEv "^invocation[[:space:]]0$" "$WORK/state/kimi_watch_all" && [ -e "$WORK/state/kimi_watch_disabled" ]'
|
|
1414
|
+
# The retry's config must differ from the attempt it replaced by exactly the
|
|
1415
|
+
# watcher table: asserting only that watch is absent would also pass a retry
|
|
1416
|
+
# that dropped or widened something else.
|
|
1417
|
+
retry_config_is_watch_only_delta() {
|
|
1418
|
+
python3 -c '
|
|
1419
|
+
import sys
|
|
1420
|
+
first = open(sys.argv[1], encoding="utf-8").read()
|
|
1421
|
+
retry = open(sys.argv[2], encoding="utf-8").read()
|
|
1422
|
+
sys.exit(0 if first.replace("[watch]\nenabled = false\n", "", 1) == retry else 1)
|
|
1423
|
+
' "$1" "$2"
|
|
1424
|
+
}
|
|
1425
|
+
rm -f "$WORK/state/kimi_watch" "$WORK/state/kimi_watch_all" "$WORK/state/kimi_watch_disabled" "$WORK/state/kimi_watch_absent" "$WORK/state/kimi_doctor_count" "$WORK/state/kimi_doctor_config."*
|
|
1426
|
+
out="$(KIMI_CODE_WATCH=1 run_kimi doctor_reject_watch)"; rc=$?
|
|
1427
|
+
# The watcher guard is belt on top of the per-invocation override, which the
|
|
1428
|
+
# runtime reads ahead of the config. A runtime that does not know the table may
|
|
1429
|
+
# not cost the whole lane: drop the table, keep the override, stay admitted.
|
|
1430
|
+
check "Kimi keeps the lane when the runtime rejects only the watcher table" \
|
|
1431
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ -e "$WORK/state/kimi_watch_absent" ] && [ ! -e "$WORK/state/kimi_watch_disabled" ] && [ "$(tr -d " " < "$WORK/state/kimi_generated_marker_count")" = 1 ] && retry_config_is_watch_only_delta "$WORK/state/kimi_doctor_config.1" "$WORK/state/kimi_doctor_config.2" && [ "$(grep -c "^doctor[[:space:]]0$" "$WORK/state/kimi_watch")" = 2 ] && [ "$(grep -c "^capability[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && [ "$(grep -c "^formal[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && ! grep -qEv "^invocation[[:space:]]0$" "$WORK/state/kimi_watch_all"'
|
|
1432
|
+
out="$(run_kimi doctor_reject_all)"; rc=$?
|
|
1433
|
+
check "Kimi still fails closed when no generated config is accepted" \
|
|
1434
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_packet_only_config_unrecognized ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1327
1435
|
# Admission is version-neutral: deliberately unparseable version output must not
|
|
1328
1436
|
# trigger --version or block a capable runtime.
|
|
1329
|
-
rm -f "$WORK/state/kimi_invoked" "$WORK/state/kimi_version_checked" "$WORK/state/kimi_no_tools_configured" "$WORK/state/kimi_no_tools_policy_verified"
|
|
1437
|
+
rm -f "$WORK/state/kimi_invoked" "$WORK/state/kimi_version_checked" "$WORK/state/kimi_no_tools_configured" "$WORK/state/kimi_no_tools_policy_verified" "$WORK/state/kimi_watch_all"
|
|
1330
1438
|
out="$(KIMI_STUB_VERSION='not-a-version' run_kimi pass)"; rc=$?
|
|
1331
1439
|
check "Kimi admission depends on runtime capability, not a version string" \
|
|
1332
|
-
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ ! -e "$WORK/state/kimi_version_checked" ] && [ -e "$WORK/state/kimi_doctor_checked" ] && grep -q "^doctor config /.*config.toml$" "$WORK/state/kimi_doctor_args" && [ -e "$WORK/state/kimi_invoked" ] && [ -e "$WORK/state/kimi_no_tools_configured" ] && [ -e "$WORK/state/kimi_no_tools_policy_verified" ]'
|
|
1440
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ ! -e "$WORK/state/kimi_version_checked" ] && [ -e "$WORK/state/kimi_doctor_checked" ] && grep -q "^doctor config /.*config.toml$" "$WORK/state/kimi_doctor_args" && [ -e "$WORK/state/kimi_invoked" ] && [ -e "$WORK/state/kimi_no_tools_configured" ] && [ -e "$WORK/state/kimi_no_tools_policy_verified" ] && [ "$(wc -l < "$WORK/state/kimi_watch_all" | tr -d " ")" = 3 ] && ! grep -qEv "^invocation[[:space:]]0$" "$WORK/state/kimi_watch_all"'
|
|
1333
1441
|
rm -f "$WORK/state/kimi_invoked"
|
|
1334
1442
|
out="$(run_kimi pass claude "$WORK/kimi-source" "$WORK/nul.patch")"; rc=$?
|
|
1335
1443
|
check "Kimi rejects a NUL-bearing diff before inference" \
|
|
@@ -1346,6 +1454,47 @@ check "Kimi rejects any tool exposed during the no-tools probe" \
|
|
|
1346
1454
|
out="$(run_kimi capability_emfile)"; rc=$?
|
|
1347
1455
|
check "Kimi classifies probe-time EMFILE as a local client failure" \
|
|
1348
1456
|
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_host_resource_exhausted ] && [ "$(field reason_code "$out")" = client_unavailable ] && [ "$(field cascade_eligible "$out")" = True ]'
|
|
1457
|
+
# Provider prose is not sub-classified: these shapes each broke a predecessor
|
|
1458
|
+
# predicate, and all of them now land in the one capability class.
|
|
1459
|
+
out="$(run_kimi capability_quota)"; rc=$?
|
|
1460
|
+
check "Kimi does not sub-classify provider prose: quota" \
|
|
1461
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1462
|
+
out="$(run_kimi capability_auth)"; rc=$?
|
|
1463
|
+
check "Kimi does not sub-classify provider prose: auth" \
|
|
1464
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1465
|
+
out="$(run_kimi capability_auth_quota_mention)"; rc=$?
|
|
1466
|
+
check "Kimi does not sub-classify provider prose: auth quota mention" \
|
|
1467
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1468
|
+
out="$(run_kimi capability_auth_limit)"; rc=$?
|
|
1469
|
+
check "Kimi does not sub-classify provider prose: auth limit" \
|
|
1470
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1471
|
+
out="$(run_kimi capability_auth_word)"; rc=$?
|
|
1472
|
+
check "Kimi does not sub-classify provider prose: auth word" \
|
|
1473
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1474
|
+
out="$(run_kimi capability_offset)"; rc=$?
|
|
1475
|
+
check "Kimi does not sub-classify provider prose: offset" \
|
|
1476
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1477
|
+
out="$(run_kimi capability_exact_offset)"; rc=$?
|
|
1478
|
+
check "Kimi does not sub-classify provider prose: exact offset" \
|
|
1479
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1480
|
+
out="$(run_kimi capability_rate_limit_403)"; rc=$?
|
|
1481
|
+
check "Kimi does not sub-classify provider prose: rate limit 403" \
|
|
1482
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1483
|
+
out="$(run_kimi capability_http_quota)"; rc=$?
|
|
1484
|
+
check "Kimi does not sub-classify provider prose: http quota" \
|
|
1485
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1486
|
+
out="$(run_kimi capability_forbidden)"; rc=$?
|
|
1487
|
+
check "Kimi does not sub-classify provider prose: forbidden" \
|
|
1488
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1489
|
+
out="$(run_kimi capability_decode_offset)"; rc=$?
|
|
1490
|
+
check "Kimi does not sub-classify provider prose: decode offset" \
|
|
1491
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1492
|
+
out="$(run_kimi capability_auth_weekly_metadata)"; rc=$?
|
|
1493
|
+
check "Kimi does not sub-classify provider prose: auth weekly metadata" \
|
|
1494
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1495
|
+
out="$(run_kimi capability_auth_too_many)"; rc=$?
|
|
1496
|
+
check "Kimi does not sub-classify provider prose: auth too many" \
|
|
1497
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = kimi_tool_capability_unverified ] && [ "$(field reason_code "$out")" = capability_missing ] && [ "$(field cascade_eligible "$out")" = True ] && [ "$(field transport_exit_code "$out")" = 1 ]'
|
|
1349
1498
|
probe_started=$SECONDS
|
|
1350
1499
|
out="$(REVIEW_TEST_TIMEOUT=5 run_kimi capability_hang)"; rc=$?
|
|
1351
1500
|
probe_elapsed=$((SECONDS - probe_started))
|
|
@@ -1406,10 +1555,13 @@ out="$(REVIEW_TEST_TIMEOUT=180 REVIEW_TEST_PATH_PREFIX="$WORK/timeout-probe-bin"
|
|
|
1406
1555
|
check "Kimi caps argv-exposed inline review at 120 seconds" \
|
|
1407
1556
|
'[ "$rc" = 0 ] && grep -q -- "--kill-after=1s 120s " "$WORK/state/timeout_args"'
|
|
1408
1557
|
rm -f "$WORK/state/timeout_args"
|
|
1409
|
-
|
|
1558
|
+
rm -f "$WORK/state/kimi_watch" "$WORK/state/kimi_watch_all" "$WORK/state/kimi_watch_disabled"
|
|
1559
|
+
out="$(KIMI_CODE_WATCH=1 REVIEW_TEST_TIMEOUT=180 REVIEW_TEST_PATH_PREFIX="$WORK/timeout-probe-bin" run_kimi pass claude "$WORK/kimi-source" "$WORK/agent-template.patch")"; rc=$?
|
|
1410
1560
|
formal_timeout="$(tail -n 1 "$WORK/state/timeout_args" | awk '{ value=$2; sub(/s$/, "", value); print value }')"
|
|
1411
1561
|
check "Kimi gives private MCP review only the remaining controller budget" \
|
|
1412
1562
|
'[ "$rc" = 0 ] && [ "$formal_timeout" -ge 1 ] && [ "$formal_timeout" -le 180 ]'
|
|
1563
|
+
check "Kimi sets the watcher override on private MCP formal delivery" \
|
|
1564
|
+
'[ "$(wc -l < "$WORK/state/kimi_watch" | tr -d " ")" = 3 ] && [ "$(grep -c "^doctor[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && [ "$(grep -c "^capability[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && [ "$(grep -c "^formal[[:space:]]0$" "$WORK/state/kimi_watch")" = 1 ] && [ "$(wc -l < "$WORK/state/kimi_watch_all" | tr -d " ")" = 3 ] && ! grep -qEv "^invocation[[:space:]]0$" "$WORK/state/kimi_watch_all" && [ -e "$WORK/state/kimi_watch_disabled" ]'
|
|
1413
1565
|
rm -f "$WORK/state/timeout_args"
|
|
1414
1566
|
out="$(REVIEW_TEST_TIMEOUT=10 REVIEW_TEST_PATH_PREFIX="$WORK/timeout-probe-bin" run_kimi capability_delay claude "$WORK/kimi-source" "$WORK/agent-template.patch")"; rc=$?
|
|
1415
1567
|
formal_timeout="$(tail -n 1 "$WORK/state/timeout_args" | awk '{ value=$2; sub(/s$/, "", value); print value }')"
|
|
@@ -832,11 +832,49 @@ class CompletionFindingDispositionTest(unittest.TestCase):
|
|
|
832
832
|
"failure_path": "A distinct synthetic failure on the same line."})
|
|
833
833
|
return subprocess.CompletedProcess(command, 0, json.dumps(payload).encode("utf-8"), b"")
|
|
834
834
|
|
|
835
|
+
base_repo: Path | None = None
|
|
836
|
+
|
|
835
837
|
def arguments(self, mode: str) -> list[str]:
|
|
836
|
-
|
|
838
|
+
if self.base_repo is not None:
|
|
839
|
+
candidate = ["--cwd", str(self.base_repo), "--base", self.base_commit]
|
|
840
|
+
else:
|
|
841
|
+
candidate = ["--cwd", str(self.root), "--diff-file", str(self.packet)]
|
|
842
|
+
return ["--mode", mode, *candidate,
|
|
837
843
|
"--implementer-family", "openai", "--review-plan-file", str(self.plan),
|
|
838
844
|
"--challenge-budget", "1"]
|
|
839
845
|
|
|
846
|
+
def test_source_refuted_completion_records_a_passed_receipt(self) -> None:
|
|
847
|
+
# The chain the pull-request reminder reads: findings recorded on a
|
|
848
|
+
# committed whole-worktree candidate, then disposed through completion.
|
|
849
|
+
repo = self.root / "repo"
|
|
850
|
+
repo.mkdir()
|
|
851
|
+
def git(*command: str) -> str:
|
|
852
|
+
return subprocess.run(["git", "-C", str(repo), "-c", "user.email=t@example.invalid",
|
|
853
|
+
"-c", "user.name=t", *command],
|
|
854
|
+
capture_output=True, check=True, text=True).stdout.strip()
|
|
855
|
+
git("init", "-q")
|
|
856
|
+
(repo / "x").write_text("a\n", encoding="utf-8")
|
|
857
|
+
git("add", "x")
|
|
858
|
+
git("commit", "-q", "-m", "base")
|
|
859
|
+
self.base_commit = git("rev-parse", "HEAD")
|
|
860
|
+
(repo / "x").write_text("b\n", encoding="utf-8")
|
|
861
|
+
git("commit", "-q", "-am", "candidate")
|
|
862
|
+
head = git("rev-parse", "HEAD")
|
|
863
|
+
self.base_repo = repo
|
|
864
|
+
self.record_receipts()
|
|
865
|
+
receipt = repo / ".git" / "ccl-code-review" / "last-review.json"
|
|
866
|
+
before = json.loads(receipt.read_text(encoding="utf-8"))
|
|
867
|
+
self.assertEqual((before["mode"], before["status"], before["head"], before["worktree_clean"]),
|
|
868
|
+
("challenge", "findings", head, True))
|
|
869
|
+
|
|
870
|
+
code, result = self.complete()
|
|
871
|
+
|
|
872
|
+
self.assertEqual(code, 0, result)
|
|
873
|
+
self.assertEqual(result["status"], "passed", result)
|
|
874
|
+
after = json.loads(receipt.read_text(encoding="utf-8"))
|
|
875
|
+
self.assertEqual((after["mode"], after["status"], after["head"], after["worktree_clean"]),
|
|
876
|
+
("complete", "passed", head, True))
|
|
877
|
+
|
|
840
878
|
def invoke(self, mode: str, extra: list[str]) -> tuple[int, dict]:
|
|
841
879
|
output = io.StringIO()
|
|
842
880
|
with (mock.patch.object(review_gate, "run", side_effect=self.wrapper_result),
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh
CHANGED
|
@@ -4182,10 +4182,38 @@ module.signal_reviewer_process_group(
|
|
|
4182
4182
|
)
|
|
4183
4183
|
tree_process.communicate(timeout=2)
|
|
4184
4184
|
tree_child_gone = False
|
|
4185
|
-
|
|
4185
|
+
# The killed grandchild is reparented to PID 1. Where PID 1 never reaps
|
|
4186
|
+
# orphans (some containers), it stays a zombie that still answers signal 0,
|
|
4187
|
+
# so a zombie counts as gone: it has exited and holds no pipe or lock.
|
|
4188
|
+
def exited(pid):
|
|
4186
4189
|
try:
|
|
4187
|
-
module.os.kill(
|
|
4190
|
+
module.os.kill(pid, 0)
|
|
4188
4191
|
except ProcessLookupError:
|
|
4192
|
+
return True
|
|
4193
|
+
try:
|
|
4194
|
+
with open(f"/proc/{pid}/stat", encoding="utf-8") as stat:
|
|
4195
|
+
return stat.read().rsplit(")", 1)[1].split()[0] == "Z"
|
|
4196
|
+
except FileNotFoundError:
|
|
4197
|
+
# procfs may be absent (macOS), or the PID may have exited between
|
|
4198
|
+
# probes. Only a second failed liveness check proves the latter.
|
|
4199
|
+
try:
|
|
4200
|
+
module.os.kill(pid, 0)
|
|
4201
|
+
except ProcessLookupError:
|
|
4202
|
+
return True
|
|
4203
|
+
return False
|
|
4204
|
+
except OSError:
|
|
4205
|
+
return False
|
|
4206
|
+
from unittest.mock import patch, mock_open
|
|
4207
|
+
# Negative control: a live PID must remain live even on hosts without procfs.
|
|
4208
|
+
assert not exited(module.os.getpid())
|
|
4209
|
+
with patch("builtins.open", side_effect=FileNotFoundError):
|
|
4210
|
+
assert not exited(module.os.getpid()), "missing procfs is not process-exit evidence"
|
|
4211
|
+
with patch.object(module.os, "kill", side_effect=ProcessLookupError):
|
|
4212
|
+
assert exited(42), "a missing PID must count as exited without relying on PID reuse timing"
|
|
4213
|
+
with patch.object(module.os, "kill"), patch("builtins.open", mock_open(read_data="42 (worker) Z 1")):
|
|
4214
|
+
assert exited(42), "a Linux zombie has exited even when signal 0 succeeds"
|
|
4215
|
+
for _ in range(100):
|
|
4216
|
+
if exited(tree_child_pid):
|
|
4189
4217
|
tree_child_gone = True
|
|
4190
4218
|
break
|
|
4191
4219
|
time.sleep(0.01)
|
|
@@ -4863,6 +4891,24 @@ out="$(REVIEW_GATE_TEST_STATE="$WORK/state" "$WORK/harness/scripts/review_gate.s
|
|
|
4863
4891
|
check "a bare --diff-file review records no receipt" \
|
|
4864
4892
|
'[ "$rc" = 0 ] && json_fields "$out" status=passed && [ ! -e "$contract_repo/.git/ccl-code-review" ]'
|
|
4865
4893
|
|
|
4894
|
+
# The completion checkpoint is what the pull-request reminder reads as disposed:
|
|
4895
|
+
# on the same whole-worktree candidate it replaces the review's receipt with a
|
|
4896
|
+
# passed one; a checkpoint that fails leaves the receipt as it was.
|
|
4897
|
+
reset_case passed unavailable unavailable
|
|
4898
|
+
out="$(run_contract_gate --mode review)"; rc=$?
|
|
4899
|
+
printf '%s\n' "$out" >"$WORK/contract-completion-review.json"
|
|
4900
|
+
check "the review before a completion checkpoint records its own receipt" \
|
|
4901
|
+
'[ "$rc" = 0 ] && [ "$(jq -r .mode "$receipt_file")" = review ]'
|
|
4902
|
+
reset_case passed unavailable unavailable
|
|
4903
|
+
out="$(run_contract_gate --mode complete --completion-review-result-file "$WORK/contract-completion-review.json")"; rc=$?
|
|
4904
|
+
check "a successful completion checkpoint records a passed receipt for the HEAD it bound" \
|
|
4905
|
+
'[ "$rc" = 0 ] && json_fields "$out" mode=complete status=passed && [ "$(jq -r .mode "$receipt_file")" = complete ] && [ "$(jq -r .status "$receipt_file")" = passed ] && [ "$(jq -r .head "$receipt_file")" = "$(git -C "$contract_repo" rev-parse HEAD)" ]'
|
|
4906
|
+
cp "$receipt_file" "$WORK/receipt-before-failed-completion"
|
|
4907
|
+
reset_case passed unavailable unavailable
|
|
4908
|
+
out="$(run_contract_gate --mode complete --review-plan-file "$WORK/changed-review-plan.json" --completion-review-result-file "$WORK/contract-completion-review.json")"; rc=$?
|
|
4909
|
+
check "a failed completion checkpoint leaves the receipt untouched" \
|
|
4910
|
+
'[ "$rc" = 2 ] && cmp -s "$receipt_file" "$WORK/receipt-before-failed-completion"'
|
|
4911
|
+
|
|
4866
4912
|
swap_out="$(PYTHONPATH="$WORK/harness/scripts" python3 - "$WORK" <<'PY' 2>&1
|
|
4867
4913
|
import os, sys, review_gate
|
|
4868
4914
|
root = os.path.realpath(os.path.join(sys.argv[1], "gitdir-swap"))
|
|
@@ -125,7 +125,7 @@ At each stage boundary, walk the per-stage entry-state enumeration in [Stage-Ent
|
|
|
125
125
|
- **Changed candidate = refreshed row + fresh full-scope rerun** — any post-review change, tests/docs included, mechanically triggers a full-scope rerun; the author cannot narrow its scope or call the change immaterial.
|
|
126
126
|
- **Independent gate surfacing basics = process defect** — repair the self-review/deterministic-gate loop before rerunning, and findings still require disposition, not waiver.
|
|
127
127
|
- **Green-tests-alone merge is the same defect** as reaching implementation with only spec plus plan.
|
|
128
|
-
- **Human/team sign-off** —
|
|
128
|
+
- **Human/team sign-off** — this gate alone: high-risk money/permission/data or breaking external APIs, before merge/launch, never before implementation. Explicit stricter rules still bind.
|
|
129
129
|
- **Floor** — reaching implementation with only spec plus plan on triggered work is a process defect, not a shortcut.
|
|
130
130
|
- For any multi-step request that combines assessment, fixes, and verification, produce a task plan before editing code (required fields in `references/delivery-lifecycle.md` §Plan Authoring).
|
|
131
131
|
- **Concurrent-session isolation**: when more than one session/agent/work-line may edit the same repository, give each line its own `git worktree` (or separate clone) on a unique per-line branch before editing — never stash another line's uncommitted changes, host-install symlinks into shared repos count as shared-tree edits, if isolation was skipped do not commit unreviewed shared changes to dodge clobber, and run the pre-merge freshness gate before merging back (recipe: `worktree-isolation`; mechanics: `references/worktree-mechanics.md`).
|
|
@@ -164,8 +164,8 @@ At each stage boundary, walk the per-stage entry-state enumeration in [Stage-Ent
|
|
|
164
164
|
- High-risk workflows cannot be accepted by happy-path tests alone. Require a risk scenario matrix and replayable incident drills for the relevant classes: duplicate money/quota/write side effects, permission uncertainty, tenant/user data isolation, AI provider/model failure, user repeated submission or unclear final state, and traceable incident explanation.
|
|
165
165
|
- For UI backed by APIs or generated content, require `testing-strategy` to produce evidence that covers rendered states, contract/error handling, and one real visible flow where feasible. Do not let ideal mocked data stand in for runtime integration evidence.
|
|
166
166
|
- When live infrastructure is required, keep it explicit and separate from default fast tests.
|
|
167
|
-
- After code/test edits, self-
|
|
168
|
-
-
|
|
167
|
+
- After code/test edits, deep self-review in the main agent; then, for external independent review, never a human one, invoke `code-review` automatically before completion. Persist review status; timeout, empty output or inconclusive results remain pending, never passed.
|
|
168
|
+
- Review runs take bounded diff input plus the repository's `AGENTS.md` or equivalent rules, skip generated/docs noise unless targeted, and keep timeout/inconclusive results in a durable pending record.
|
|
169
169
|
- For product/spec normalization, standards-to-health-gate work, or any change that edits a shared deterministic gate/verifier (workspace verifier, conformance script, contract-coverage gate, status-source validator, CI harness, continuation-state checker, or a cross-repo contract/status/version/release/compatibility coordination surface), this workflow classifies the artifact before implementation — `spec/plan`, `gate design`, `gate implementation`, `status sync`, or `runtime/code` — without delegating that decision (do not delegate the spec-vs-plan-vs-code decision to `feature-risk-router`), then routes every shared-gate change through `feature-risk-router` and applies its `shared-gate` decision before shared branch push or MR merge; a recorded independent adversarial review names concrete objections, their disposition, and the reviewer/tool identity — prefer the session's review/challenge skill, otherwise a ccl-owned independent review (external tools supplement only; same-agent inline prose review only for explicitly low-risk, non-cross-boundary work). Rule/scope/failure/completion semantics changes require a concrete repo-local persistent artifact before editing; a `gate implementation` runs the plan/status verifier(s) before implementation and before claiming the plan active — an explicit status-source validator takes precedence, otherwise run every authoritative non-alias verifier or record why each is not applicable; a verifier gap or unavailable required review/challenge stays `interim` / pending-review. Details live in `references/shared-gate-artifact-classification.md`.
|
|
170
170
|
- Do not use landing labels without matching evidence. `landed` requires the relevant local commit or persisted artifact; `MR-ready` requires branch, push, review artifact, known CI/pipeline status when applicable, known mergeability when applicable, and review status that matches reality; `release-ready` requires the relevant release checks, rollback/mitigation, and runtime verification evidence; `shared-status-ready` requires the owning status or product document to match the real branch/MR/review/verification state. For local-only or exploratory slices, report the actual uncommitted/unpushed state and use a local status label instead of treating MR evidence as mandatory.
|
|
171
171
|
- When the delivery changes shared product status, roadmap, verification state, or cross-repository readiness, update the owning product/status document in the same delivery batch after the code MR lands. The **agent-consumed status-doc rule** is stricter: such a tracker is part of the work itself and must match the final handed-off/green/shared state — before merge and after reviewer edits/squash/rebase/platform merge, re-validate it against the final diff/ref/CI and point it at the final landed ref (`references/status-tracker-sync.md`).
|
|
@@ -187,13 +187,13 @@ At each stage boundary, walk the per-stage entry-state enumeration in [Stage-Ent
|
|
|
187
187
|
|
|
188
188
|
### Pre-Final Continuation Gate
|
|
189
189
|
|
|
190
|
-
Run this gate before finalizing a product R&D turn after any delivery slice lands, on every user reply immediately following an assistant message that states or implies a next action, and on any explicit continuation request.
|
|
190
|
+
Run this gate before finalizing a product R&D turn after any delivery slice lands, on every user reply immediately following an assistant message that states or implies a next action, and on any explicit continuation request. Recover the intended action from the current request and relevant conversation; a malformed or absent marker alone never selects `blocked:`. Report, on its own line, `continuing: <action and scope>` for work you will execute, or `blocked: <action and scope> — <specific blocker>` when no authorized work can proceed; `proposed-next:` never replaces it. A blocked dependent action may remain pending while a different authorized action continues. Apply landing checks only to actual landing claims; an already-authorized local investigation does not require inventing a prior landing. Details: `references/pre-final-continuation-gate.md` (Gate triggers and outcome contract).
|
|
191
191
|
|
|
192
192
|
1. Confirm the landing state from real evidence (local branch, remote sync, MR/review artifact, CI/pipeline when applicable, review/challenge status when required, status-doc sync, dirty worktree), proving the landing before reading any document (for the remote-backed default, fetch/update the target ref from its remote immediately before classifying the slice landed) per `references/pre-final-continuation-gate.md` (Landing-state proof); content/tree/patch equivalence never by itself proves a slice landed.
|
|
193
193
|
2. Inspect the current product/status source of truth, issue list, repo-local next-step artifact, unresolved acceptance item, or direct user continuation instruction for the next implied slice. **Reconcile it against the current branch/MR/merge/CI/tag state from step 1 before deriving: contradicting reality means stale — stop, repair the status source first, and do NOT derive from the stale source or a git-log/grep scan** (`references/pre-final-continuation-gate.md` §Status-source reconciliation).
|
|
194
194
|
- **Deferred-evidence continuation check (`DFE-CONT`).** When real/runtime evidence is due (named by an acceptance item, status source, landing-evidence row, required gate, user correction, or because it is the behavior's only meaningful proof) yet deferred, blocked after remediation, skipped at finalization, or replaced by local/mock verification. Report deferred real evidence as `interim`/outstanding; do NOT report the turn complete while it is outstanding. A local/mock substitution is terminal only when a cited **non-agent** anchor — **agent-authored or agent-co-edited status/router/gate/handoff text never satisfies this** — names the same evidence, declares the deferral terminal, and carries the outstanding command/source forward for the active slice/ref. Never add verifier/config/test hardening motivated only by missing deferred evidence; never auto-continue past the pending gate. **Load `references/pre-final-continuation-gate.md` before treating any deferral as terminal** — it owns the valid/invalid-anchor list and hardening boundary.
|
|
195
195
|
- **Affirmative-assent binding rule** lives in `references/pre-final-continuation-gate.md` §Assent binding — load it when recovering a short reply. Bind to the current explicit request or one recoverable concrete proposal, including an unmarked proposal; preserve its scope and existing authority. Ask only if action, scope, or required authority remains unresolved after recovery. A status remark or output marker cannot substitute for a proposal or permission; self-classifying the reply or marker away is never an exit from carrying out an already-clear request.
|
|
196
|
-
3. Continue
|
|
196
|
+
3. Continue an owned, verifiable, low-risk slice from the active task/status/acceptance source or user continuation, within scope and authority; apply `references/pre-final-continuation-gate.md`. Necessary fixes, tests and review inherit task authorization. Reviewer-budget flags require a method checkpoint and cumulative history, not renewed permission. Small tests and routine development/test-environment operations use configured access directly: no per-run approval or invented cost cap. Developer-self-use metered accounts are not new purchases; explicit user limits and high-impact boundaries still govern.
|
|
197
197
|
4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, a concrete blocker for the affected action, or no safe authorized work remains. Block materially differing viable approaches (none dominant-and-reversible) and a fix lacking evidenced cause; load `references/pre-final-continuation-gate.md` for the full stop conditions. **Scope each blocker to its dependent action or claim.** An unproven cause blocks the speculative patch, not available diagnosis; a pending gate blocks dependent landing/completion, not authorized remediation or independent work. Before ending, perform in-scope diagnosis, owner discovery, remediation or independent work, and poll any finite step you started to its result, never reporting it as running. Quality-gate failures require diagnosis and available related behavior-preserving cleanup before escalation; preserve readability and compatibility, never game counters (`references/refactoring-discipline.md`). Never bypass the blocked gate, invent a pass, widen scope, or substitute unrelated hardening. With one dominant reversible approach and no applicable stop condition, do not stop at a recommendation: deliver a tested reviewable draft.
|
|
198
198
|
5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form and classifies the turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
|
|
199
199
|
6. **Assent-outcome closeout check.** Every user reply immediately following an assistant message that states or implies a next action requires a visible `continuing:` or `blocked:` outcome before finalizing, even if the reply is not classified as assent; every explicit continuation request does too. Missing markers never waive it. Reconcile the current request, original proposal, scope/authority changes, tool/output evidence, and remaining blockers. Respect a current explicit stop or status-only request; name that reason in the blocked outcome without executing the prior proposal. Otherwise `continuing:` must be followed by execution in the same turn; a promised next step is not execution. If part remains blocked, report its pending state and independent work performed. A status-only handoff cannot discharge an unexecuted accepted action. Repair marker formatting; for short assent, if the original proposal cannot be recovered verbatim, select `blocked:` and ask. Formatting never requires clarification. Do not silently drop an accepted action or claim a pending gate passed.
|
|
@@ -55,7 +55,7 @@ Merging triggered-work implementation on green-tests-alone, with the adversarial
|
|
|
55
55
|
|
|
56
56
|
## Human/team sign-off
|
|
57
57
|
|
|
58
|
-
(b) human/team sign-off before
|
|
58
|
+
(b) This gate requires human/team sign-off before merge or launch of high-risk money/permission/data paths and breaking contract/API changes with an external consumer (reviewed interface diff plus compatibility decision). It adds no sign-off prerequisite to local, reversible implementation; explicit stricter user, repository or team rules still bind. A backward-compatible addition outside those high-risk paths needs no human sign-off under this gate; recorded deep self-review and external independent review cover it. Read-only investigation or an isolated disposable prototype that will not be merged, reused, or launched is exempt from this gate; promoting it to a merge or launch reruns the gate.
|
|
59
59
|
|
|
60
60
|
## Floor
|
|
61
61
|
|
|
@@ -99,7 +99,8 @@ At the start of the next turn, recover intent in this order:
|
|
|
99
99
|
|
|
100
100
|
On hosts providing a current final message, `proposed-next-stop.sh` returns one bounded Stop reminder when the assistant declares a non-status `proposed-next:` action. Recheck the active request: execute a runnable, already-authorized action in the same turn; otherwise preserve explicit stop, planning-only and status-only scope, or state the concrete decision/resource/authority blocker. Missing labels with observable delivery evidence retain their formatting reminder. A status-only marker without another action declaration, quoted example, complete machine artifact, unsupported payload or host `stop_hook_active` retry does not trigger a continuation reminder.
|
|
101
101
|
|
|
102
|
-
-
|
|
102
|
+
- A stop that waits on the user gets one bounded decision recheck instead: a `blocked:` handoff, a `none` explanation naming an approval, confirmation, decision or resource wait, or a last prose line asking permission to continue, either before any handoff label or, without a label, after observable delivery or edits. The recheck names the real blockers (missing credentials or authority; a fact unavailable from local evidence; an action the safety rules gate, such as destructive or irreversible work without recovery, production or customer data, or merge or publication outside the goal; overturning an established user direction; a material product tradeoff the evidence cannot settle) and returns security self-review, owner-skill, approach, test and naming choices and the next in-scope step to the agent. A real blocker survives it by restating `blocked:` after independent work is finished.
|
|
103
|
+
- Do not request continuation for `none — status only` or another `none` status explanation, and never treat `blocked:` as a continuation request; the decision recheck above is separate. Mixed status/action markers still require reconciliation.
|
|
103
104
|
|
|
104
105
|
The hook recognizes declarations, not authorization or actual task completion, and cannot force the model to follow through. OpenCode idle does not expose the required final-message evidence; its Stop behavior remains unverified.
|
|
105
106
|
|
|
@@ -113,7 +114,7 @@ Action-scoped stop conditions are: an explicit stop/pause instruction; a user-re
|
|
|
113
114
|
|
|
114
115
|
Check continuation on every user reply immediately following assistant prose that states or implies a next action, and on any explicit continuation request, regardless of landing status. Do not first require classifying the reply as assent; visibly report the continuing or blocked outcome even when the reply changes scope or stops the proposed action. Short replies include `ok`, `yes`, `可以`, `好`, `继续`, `proceed`, `do it`, `go ahead`, and `👍`; interpret them against the recovered action rather than formatting alone.
|
|
115
116
|
|
|
116
|
-
- Select `continuing: <action and scope>` when that action is clear and authorized, then execute it in the same turn. A tool call and its result or a produced artifact establish execution; the label alone does not.
|
|
117
|
+
- Select `continuing: <action and scope>` when that action is clear and authorized, then execute it in the same turn. Write either outcome on its own line; a `proposed-next:` label never replaces the outcome line. A tool call and its result or a produced artifact establish execution; the label alone does not.
|
|
117
118
|
- A blocked patch, review, or landing does not block every action. Keep that dependent action/claim pending while continuing available diagnosis, bounded remediation, monitoring of the existing live handle, or independent accepted work. These paths retain their own scope and permission checks; they cannot bypass the blocked gate or substitute unrelated hardening for missing evidence.
|
|
118
119
|
- A failed quality gate calls for a repair that preserves its purpose. Before asking the user to choose a workaround, inspect and perform a safe structural cleanup necessary for the authorized delivery when available, including baseline failures that block it, then rerun the gate and affected tests. Follow [refactoring discipline](refactoring-discipline.md#responding-to-quality-gates): preserve behavior, compatibility and readability; do not shrink identifiers or necessary comments, weaken a baseline or rewrite history solely to make the counter pass. If no safe in-scope repair remains, report the evidence and the actual decision needed.
|
|
119
120
|
- Independent work must neither depend on the pending verdict nor modify the candidate being evaluated. Name the pending gate and the independence basis when continuing. A candidate-changing fix is remediation, not independent work: let the existing run reach a terminal state, then refresh affected evidence and re-enter the owning gate. The deferred-evidence hardening prohibition still applies.
|
|
@@ -131,5 +132,5 @@ Before deriving the next slice from a status source, reconcile it against the ap
|
|
|
131
132
|
|
|
132
133
|
Binding detail:
|
|
133
134
|
|
|
134
|
-
-
|
|
135
|
+
- Small tests and routine development/test-environment operations within the task are ordinary execution details. Use configured accounts and access directly, without per-run approval or inventing a quota/cost estimate or cap. Normal metered model/tool use is not a new purchase. Honor explicit user spending/count limits; a development/test label does not grant destructive, production, customer-data, permission-changing or new-purchase authority beyond the task.
|
|
135
136
|
- Assent never replaces an owner gate's stricter authorization form and never broadens scope or implies an external purchase/financial commitment, merge, publish, destructive, production, external-message, or high-impact-decision authority.
|
|
@@ -11,7 +11,7 @@ Use this when receiving review feedback from a human reviewer, AI reviewer, CI a
|
|
|
11
11
|
- Push back with technical evidence when a suggestion is incorrect or harmful.
|
|
12
12
|
- If your pushback was wrong, correct it factually and move on: retract only after re-verifying against the codebase or the authority that settles the point (requirement owner, product decision, observed runtime behavior) — state the concrete counter-evidence; reviewer insistence alone is never grounds, and an inconclusive re-check goes to the risk owner rather than holding the pushback — except on a scope-adding finding, which follows the mirror-case procedure below instead. Once re-verification or that risk owner establishes the pushback was wrong, retract and implement — no long apology, no defending the earlier pushback.
|
|
13
13
|
- Implement one coherent review item or group at a time, then verify.
|
|
14
|
-
- Partition findings before fixing: a mechanical finding (bug, missing check, wrong value) goes on the fix list; a design-level finding — one questioning a mechanism's cost, operability, trust-model fit, or existence — is a risk-owner decision item (keep / delete / narrow / replace) to surface BEFORE investing hardening rounds in the questioned mechanism. Hardening first and deciding later pays the cost twice — once to build, once to tear down. The mirror case — a finding that adds scope or abstraction beyond the agreed deliverable (a new capability or abstraction layer — not a guard, error path, or rollback that a current acceptance point or observed hard constraint already requires for behavior already in scope; a remedy that itself introduces a new capability or abstraction stays scope-adding however it is labelled; an arguable classification is settled by running the test named next, never by defaulting either way) — is never an automatic build: answer it against the structural-minimality test in `implementation-completeness-and-minimality.md` — name the current acceptance point or observed hard constraint that needs it, or state that there is none — and reply with that answer before deciding. No locator means the answer is decline, not build: building it anyway first requires the risk owner to change the governing acceptance requirement. that evidence-backed answer is the receiver's to give and the reviewer's to confirm, and a rejected answer returns the item to the receiver to rebuild the evidence rather than to a risk owner — on this question the risk owner's role is to change the governing acceptance requirement, not to override the reviewer's rejection. When one finding is both design-level and scope-adding — it questions an existing mechanism AND proposes a new capability or abstraction, whether that wraps, supplements, or replaces the mechanism — the two paths govern different objects and both run: the minimality test decides the proposed addition (receiver answers, reviewer confirms), while the questioned mechanism's keep / delete / narrow / replace stays the risk owner's. Neither substitutes for the other, and neither decision alone authorizes the addition.
|
|
14
|
+
- Partition findings before fixing: a mechanical finding (bug, missing check, wrong value) goes on the fix list; a design-level finding — one questioning a mechanism's cost, operability, trust-model fit, or existence — is a risk-owner decision item (keep / delete / narrow / replace) to surface BEFORE investing hardening rounds in the questioned mechanism. Hardening first and deciding later pays the cost twice — once to build, once to tear down. In this file the risk owner is you, deciding from evidence and recording the reason; it is the user only when the decision would change accepted scope or an established user direction, or touches a money, permission, data-loss, breaking-contract or release path — never a human review or sign-off for an ordinary change. The mirror case — a finding that adds scope or abstraction beyond the agreed deliverable (a new capability or abstraction layer — not a guard, error path, or rollback that a current acceptance point or observed hard constraint already requires for behavior already in scope; a remedy that itself introduces a new capability or abstraction stays scope-adding however it is labelled; an arguable classification is settled by running the test named next, never by defaulting either way) — is never an automatic build: answer it against the structural-minimality test in `implementation-completeness-and-minimality.md` — name the current acceptance point or observed hard constraint that needs it, or state that there is none — and reply with that answer before deciding. No locator means the answer is decline, not build: building it anyway first requires the risk owner to change the governing acceptance requirement. that evidence-backed answer is the receiver's to give and the reviewer's to confirm, and a rejected answer returns the item to the receiver to rebuild the evidence rather than to a risk owner — on this question the risk owner's role is to change the governing acceptance requirement, not to override the reviewer's rejection. When one finding is both design-level and scope-adding — it questions an existing mechanism AND proposes a new capability or abstraction, whether that wraps, supplements, or replaces the mechanism — the two paths govern different objects and both run: the minimality test decides the proposed addition (receiver answers, reviewer confirms), while the questioned mechanism's keep / delete / narrow / replace stays the risk owner's. Neither substitutes for the other, and neither decision alone authorizes the addition.
|
|
15
15
|
- Avoid performative agreement. Technical correctness matters more than sounding agreeable: agree only after verifying — the reply shows what you checked and what you found, and praise or thanks offered in place of that evidence is the failure ("you're absolutely right" / "great point" are the common forms).
|
|
16
16
|
|
|
17
17
|
## Response Pattern
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md
CHANGED
|
@@ -133,7 +133,7 @@ Before changing architecture guidance, contracts, service boundaries, diagrams,
|
|
|
133
133
|
- High-risk operations require a resilience contract: fail-closed policy, idempotency strategy, durable status, reconciliation or repair path, trace/request id propagation, user/support explanation surface, and proof that fallback/degradation cannot bypass authorization, tenant/user isolation, quota, audit, or data-retention controls.
|
|
134
134
|
- High-risk context resolution must reject missing tenant, actor, subject, or resource scope instead of falling back to default identities. Durable side effects need atomic audit/outbox evidence or an explicit reconciliation/repair workflow.
|
|
135
135
|
- Python AI/RAG service hosts must separate service wiring from inference design. Model routing, prompt policy, retrieval design, evaluation, and replay belong to `llm-inference-integration`.
|
|
136
|
-
- Generated API clients and generated protobuf code are output surfaces; do not hand-edit them. Generated migrations are drafts
|
|
136
|
+
- Generated API clients and generated protobuf code are output surfaces; do not hand-edit them. Generated migrations are drafts: review them before landing, and get human sign-off only before a destructive one runs on shared data.
|
|
137
137
|
|
|
138
138
|
## Reference Loading
|
|
139
139
|
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md
CHANGED
|
@@ -50,7 +50,7 @@ Turn observed experience into reusable skills without business-specific details.
|
|
|
50
50
|
- **A retrospective over a LARGE multi-batch / multi-phase session has a second axis beyond the per-delivery chain: distinct lesson-TYPE axes that must each be covered or explicitly marked `no-new-lesson` — (a) per-artifact CONTENT lessons (the specific bug / contract value / domain rule; for research/writing/design programs this axis is the METHOD/CRAFT — how the work was done well), (b) PROGRAM/PROCESS lessons (how the multi-batch effort was structured and driven), (c) WORKFLOW/META lessons (did the retro or extraction itself recur shallow, under-trigger, or stop at the most salient content lesson), and (d) SUSTAIN lessons (what went RIGHT and how the next run reuses it — counts only with mechanism + non-luck evidence + owner routing; **axes (a)-craft and (d) read from the produced-artifact class — an enumeration driven by correction turns cannot reach them and will come back falsely empty**). Landing only the loudest content lesson and declaring the session "fully summarized / 复盘完成" is incomplete.** "LARGE" is not a vibe — it fires when the session already carries a coverage/program structure: a source register or named batch-progress standard was applied, OR the work spanned multiple explicit phases/batches/verticals. A user re-ask after a "done" claim = same-scope correction signal — classify first; never manufacture a lesson. Per-axis detail, re-ask classification, DO-CONFIRM card, `covered-through` watermark: `references/source-to-skill-extraction.md` (Task Retrospective Extraction).
|
|
51
51
|
- **A long operational delivery session also needs a separate non-lesson delivery-state axis (in addition to the content/program/meta lesson axes above).** When the session changed operational delivery state across multiple repositories, branches, MRs, pipelines, releases, or deployable artifacts, the source register must carry that axis — changed artifact set, branch/worktree state, remote/MR state, CI or local verification state, cancelled/retried pipeline state, unresolved risks, and the next concrete action — before "whole-session retro complete" is claimed; closeout records either the axis rows' locator (sanitized labels in the shared landing, real per-repo evidence in scratch/private archive) or `artifact/status axis: not-applicable` with a reason. If required rows are absent, the retro can be reported only as `interim`, even when the extracted lesson text is correct. The row-family fields and closeout-evidence forms: `references/source-to-skill-extraction.md` (Task Retrospective Extraction).
|
|
52
52
|
- A blocked verification item is not closed by naming the blockage. Before marking a test, device, browser, service, credential, or environment layer unavailable, attempt the normal remediation path for that layer — restart the client daemon, run the documented fallback (full rung list and sandbox-denial triage: the next ladder). Only record `unavailable` after remediation fails, with command evidence, residual risk, and the next concrete unblock action.
|
|
53
|
-
- **A landed CONCLUSION is a hypothesis until an operation that could have falsified it has been run** —
|
|
53
|
+
- **A landed CONCLUSION is a hypothesis until an operation that could have falsified it has been run** — a claim that a tool / capability / lane is unavailable, impossible, or must permanently fail-closed ("fail-closed is the safe default" does not waive the in-env attempt), a DIAGNOSIS of why an observed failure happened, and a **MEASUREMENT reported as a finding about the subject** — run the falsifying operation first — exercise the suspected mechanism on the failing path for an observation only IT predicts, or build a paired control differing in exactly ONE variable — **only within existing sandbox/permission, non-destructive, synthetic-target, and credential-safety boundaries** (never unsafe mutation, prod/live credentials, secret-bearing state, or a permission-boundary bypass — trading this rule for the security/authority/data-loss axis). Where no safe attempt is available after remediation the record is `pending` with remediation and residual risk — never `unavailable`, `fail-closed`, or a stated cause — and an unfalsified cause is `hypothesis`, kept off shared surfaces, because withdrawing a landed cause costs more than testing it. **When REVIEWING a change that asserts impossibility/unavailability, rests on a diagnosis, or reports a measurement, independently run the same falsification first** — an inherited "it can't be done", "this is why it broke" or "the tool measured N" is hypothesis-grade (see the named-convention primary-source re-verify rule). Forms and failure shapes: `references/validation-and-landing.md` (Behavioral Validation).
|
|
54
54
|
- A blocked source read is not closed by naming the blockage. If Figma, code, document, API, or repository reads time out, return partial output, or fail transport, switch to a smaller or different read strategy before extracting rules — and when the source is **missing rather than unreadable**, change WHERE you enumerate instead. Both ladders: `references/source-to-skill-extraction.md#blocked-verification-and-source-read-remediation`. Failed or timed-out reads do not count as coverage.
|
|
55
55
|
- **Large reads can lose the middle with no reliable signal — the trigger is read-OUTPUT size, so chunk proactively.** A read whose OUTPUT exceeds ~256 lines / ~10 KiB can be silently head+tail truncated (no marker guaranteed), so a single `cat`/whole-file read does not count as coverage even when it returns no error. Whenever you need a **complete** view — whole-file coverage, a no-findings/absence claim, or a load-bearing section read — chunk it under **both ~200 lines AND ~8 KiB** and confirm a mid-file section was ingested. Detail: `references/source-to-skill-extraction.md#read-in-chunks-large-reads-lose-the-middle`.
|
|
56
56
|
- Think across the full delivery lifecycle before editing: product intent, design/UX, implementation, debugging, test strategy, launch acceptance, iteration feedback, team onboarding, and normal users without source access. A rule that improves only one slice while leaving another slice ambiguous is incomplete or belongs in a narrower skill.
|
|
@@ -61,6 +61,10 @@
|
|
|
61
61
|
|
|
62
62
|
**地板管「能不能动手」,不等于「点估计已经准到能和阈值比」**:10 次有效观测在 70–85% 区间的抽样误差约 ±15–20 个点。实测形态(115 轮,同一个候选、同一条用例 `ab-b5`):一次 10 副本得 5/10(50%),紧接着 20 副本得 17/20(85%),合并 22/30(73%)——单看前者会判成「稳定失败」并据以改描述,单看后者会判成「健康」。所以**任何按阈值分档的判断(稳定失败 / 边缘 / 抖动)必须读合并观测,落在约 45–75% 之间的读数在 10 副本下不构成判定**,要么补到 30 次以上,要么如实记成「区间未定」。同理,改前/改后的差值也按合并观测比:本轮那条 8/10→5/10 的「回归」在 24/30 vs 22/30 下相差两次命中,不可分。
|
|
63
63
|
|
|
64
|
+
**被测模型也是测量对象的一部分,不只是评分工具**:路由由读取技能清单的那个模型决定,规则执行由运行该技能的那个模型决定;两个 runner 默认的廉价档(`claude-haiku-4-5`)只适合**筛查**候选。凡据评测结果做的技能决定——改 description、改正文规则、撤回改动、判「测不出改进」——都必须在**实际运行该技能的模型档**上测得(以 `--model` 显式指定,JSON `model_source: explicit`),合并观测地板同上;默认档的读数只能写成「候选」。runner 在落到默认档时打印 `router_model_default` / `subject_model_default` 提示。实测形态:同一组继续闸探针在默认廉价档上 10 副本读数互相矛盾,并据此撤回了一处改动;被纠正后改在部署档上重测,才得到可用于决定的基线。
|
|
65
|
+
|
|
66
|
+
**被测进程必须与本机插件隔离**:`claude --print` 会加载用户级配置,包括已安装插件的 hooks——SessionStart 注入路由块、UserPromptSubmit 注入任务入口、Stop hook 可以把被评分的最终回答替换成对它自身提醒的回复。两个 runner 因此以 `--settings '{"disableAllHooks":true}'` 调用被测模型(`--bare` 同样跳过 hooks,但只接受 API key 认证);bootstrap 只经 `--with-bootstrap` 这一受控臂进入测量。实测:未隔离时一次调用触发 12 个 hook 事件,同一档模型三次全量读数在 13/29–22/29 之间摆动,失败回答里出现了只属于 Stop 提醒的内容。隔离之前的历史读数混入了这部分上下文,不能与隔离后的读数直接比较。
|
|
67
|
+
|
|
64
68
|
1. 动任何 description 之前必须先跑 **≥10 轮有效观测**的稳定性基线,把稳定失败与抖动分开;抖动不得作为修改依据(grader 超时/不可解析轮不算有效观测,须补跑)。
|
|
65
69
|
2. 改后通过数必须在**最终措辞**上重测:中间稿的通过数在措辞再变的那一刻作废,不得挪用到最终候选的证据里。
|
|
66
70
|
**`newly_failed` 是候选,不是回归判定。** runner 的 `--baseline` diff 在三副本下按保守共识判 status,于是一次孤立偏离就把一条用例记进 `newly_failed`;而这个集合**每跑一次就换一批**。实测(115 轮,同一条分支上四次全量三副本运行):`{route-opencode-project-config, route-nodejs-arch}`、`{skip-pytest-cmd, ctrl-ai-risk, miss-refactor-python-unqualified, route-nodejs-arch}`、`{mem-api-log-redact, route-nodejs-arch}`——除 `route-nodejs-arch` 外每一条只出现过一次、再未复现,逐条做成对 20 副本探针后**无一可归因于该轮改动**(两例两臂分布完全相同,一例两臂都红,一例合并后相差两次命中)。所以:`newly_failed` 的每一条都要按「同一用例、改前/改后两棵树、合并 ≥20 次观测」复测才能称为回归,不得直接写进轮记录当回归清单;同样地,不得因为它每轮都有内容就把整轮判红。
|
|
@@ -20,6 +20,7 @@ For the specific owner-dispatch case, the firing point is now **mechanically enf
|
|
|
20
20
|
|
|
21
21
|
- `owner-dispatch` PreToolUse/Stop hooks (`hooks/owner-dispatch-guard.sh`, `hooks/owner-dispatch-stop.sh`) gate the first product-code edit and session close.
|
|
22
22
|
- **Subagent extension** (delegated workers are a separate firing surface — SessionStart routing is NOT inherited by subagents, so a cold worker never sees the gate): `SubagentStart` (`hooks/subagent-start.sh`) injects a slim self-gating routing pointer, and `SubagentStop` reuses `owner-dispatch-stop.sh` (PreToolUse already fires inside subagents) with `agent_id`-scoped, actor-precise markers/cap so the invoke-owner backstop applies one level down for a worker's **path-attributable** gated Edit/Write (Bash-only and missing-baseline cases stay advisory under the session Stop / CI backstop, not SubagentStop). This closes the recurrence where `multi-agent-delegation` had to *manually* inject owners into every worker prompt (`skills/multi-agent-delegation/SKILL.md` Core Rules) and the user reminded when it was forgotten.
|
|
23
|
+
- **Extraction owner at the first shared-skill edit:** the default source-edit checkpoint (`hooks/skill-loading.py`) recognises a target on a ccl-skills checkout's `skills/`, `hooks/` or root `scripts/` surface (markdown included there only; the marker is probed at every such component, so a checkout under an ancestor carrying one of those names still counts; plugin-cache and `.codex` install copies excluded — the stop backstop's scope, which this round aligned) and, while `skill-extraction-workflow` is not loaded in the current context, names it and the charter-before-editing step in its one replan. Observed miss it answers: a shared-gate bug routed to its debugging owner satisfied the generic checkpoint, and the extraction backstop in `hooks/skill-extraction-gate-stop.sh` fired only at Stop, after commit, push and pull request. Limits: one replan per actor/context, Edit/Write/patch tools only (Bash-written edits still reach only the Stop backstop), and it cannot prove the charter was written.
|
|
23
24
|
- `scripts/owner-dispatch/owner-dispatch.sh ci` is the host-agnostic merge backstop (engine: `scripts/owner-dispatch/owner-dispatch.sh`; rationale + safety posture: `scripts/owner-dispatch/README.md`).
|
|
24
25
|
|
|
25
26
|
This is the worked example of moving an already-precise recognition-dependent gate ONTO the transition. **Honesty preserved — enforcement is partial, not total:** opt-in per product repo (`.owner-dispatch.json`), default `ask`, fail-open, Claude-hook-hard / Codex-advisory, in-session Bash-write detection is best-effort (the un-bypassable layer is CI, and only when the CI job itself is enforced), and the whole thing is a no-op on the generic Agent-Skills channel. So it raises enforcement where the plugin/hook + CI layers load, but the closeout gate + user-signal escalation remain the backstop everywhere they do not. The opt-in gap itself is now narrowed by a **closeout-acquire check** in `product-rd-workflow` (a gated multi-owner delivery whose repo lacks `.owner-dispatch.json` installs the backstop or records exempt at closeout) — so "silently absent" is caught, though still recognition-dependent at closeout, not a hard gate.
|