oh-my-customcode 1.1.62 → 1.1.63
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli/index.js +1 -1
- package/dist/index.js +1 -1
- package/package.json +1 -1
- package/templates/.claude/hooks/scripts/audit-log.sh +24 -4
- package/templates/.claude/hooks/scripts/failure-ledger.sh +6 -1
- package/templates/.claude/hooks/scripts/git-delegation-guard.sh +12 -2
- package/templates/.claude/hooks/scripts/model-escalation-advisor.sh +14 -3
- package/templates/.claude/hooks/scripts/playwright-compress.sh +13 -5
- package/templates/.claude/hooks/scripts/r007-r008-drift-advisor.sh +3 -0
- package/templates/.claude/hooks/scripts/secret-filter.sh +14 -10
- package/templates/.claude/hooks/scripts/stuck-detector.sh +37 -3
- package/templates/.claude/hooks/scripts/task-outcome-recorder.sh +99 -19
- package/templates/.claude/rules/MUST-completion-verification.md +6 -0
- package/templates/.claude/rules/MUST-orchestrator-coordination.md +2 -0
- package/templates/.claude/rules/MUST-tool-identification.md +11 -0
- package/templates/.claude/rules/SHOULD-verification-ladder.md +2 -0
- package/templates/.claude/skills/fsd/SKILL.md +1 -0
- package/templates/.claude/skills/pipeline/workflows/auto-dev.yaml +5 -1
- package/templates/manifest.json +1 -1
- package/templates/workflows/auto-dev.yaml +5 -1
package/dist/cli/index.js
CHANGED
package/dist/index.js
CHANGED
package/package.json
CHANGED
|
@@ -20,14 +20,34 @@ printf '%s' "$input" | jq -e 'type=="object"' >/dev/null 2>&1 || exit 0
|
|
|
20
20
|
# Extract fields from hook input
|
|
21
21
|
tool_name=$(printf '%s\n' "$input" | jq -r '.tool_name // "unknown"')
|
|
22
22
|
file_path=$(printf '%s\n' "$input" | jq -r '.tool_input.file_path? // .tool_input.command? // ""' | head -c 200)
|
|
23
|
-
|
|
24
|
-
|
|
23
|
+
# Agent identity and model. PostToolUse has NO top-level `model` at all, and its
|
|
24
|
+
# top-level `agent_type` is present only when the hook fires from inside a subagent
|
|
25
|
+
# (CC 2.1.259 schema: "Present when the hook fires from within a subagent") — so on
|
|
26
|
+
# the main thread both reads resolved to "unknown" on every single entry.
|
|
27
|
+
# Measured replacements (#1656 C): the Agent tool's result object carries
|
|
28
|
+
# `resolvedModel` (892 occurrences in this project's transcript corpus, e.g.
|
|
29
|
+
# "claude-sonnet-5" / "claude-opus-5[1m]") and, on the synchronous shape, `agentType`;
|
|
30
|
+
# the spawn arguments carry `subagent_type`.
|
|
31
|
+
# `?` keeps a scalar-shaped tool_input/tool_response from aborting under `set -euo pipefail`.
|
|
32
|
+
agent_type=$(printf '%s\n' "$input" | jq -r '
|
|
33
|
+
(.agent_type?
|
|
34
|
+
// .tool_input?.subagent_type?
|
|
35
|
+
// .tool_response?.agentType?
|
|
36
|
+
// "unknown") | tostring
|
|
37
|
+
' 2>/dev/null) || agent_type="unknown"
|
|
38
|
+
model=$(printf '%s\n' "$input" | jq -r '
|
|
39
|
+
(.model?
|
|
40
|
+
// .tool_response?.resolvedModel?
|
|
41
|
+
// "unknown") | tostring
|
|
42
|
+
' 2>/dev/null) || model="unknown"
|
|
25
43
|
# Outcome signal. PostToolUse carries the result under `tool_response`, not
|
|
26
44
|
# `tool_output` (measured 2026-09-03: 1764/1764 PostToolUse payloads had
|
|
27
45
|
# `tool_response`, 0/1764 had `tool_output`), so the old `.tool_output.is_error`
|
|
28
46
|
# read was always absent and every entry was logged as outcome=success.
|
|
29
|
-
#
|
|
30
|
-
# `.tool_response.interrupted
|
|
47
|
+
# Failure signals read below: `.tool_response.is_error` (generic) and
|
|
48
|
+
# `.tool_response.interrupted`. The latter is present on every Bash response but
|
|
49
|
+
# was never observed true in this corpus (0/1555), so it is defensive coverage —
|
|
50
|
+
# not a signal measured to fire.
|
|
31
51
|
# jq's `//` treats both null and false as absent, so this chain is an OR:
|
|
32
52
|
# the result is true iff at least one signal is true.
|
|
33
53
|
is_error=$(printf '%s\n' "$input" | jq -r '
|
|
@@ -52,6 +52,9 @@ ts=$(date -u +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || echo "")
|
|
|
52
52
|
# (r007-r008-drift-advisor.sh의 `.role` vs `.message.role`과 동일 계열).
|
|
53
53
|
# .tool_error / .tool_response.* fallback은 스키마 변화에 대한 방어로만 남긴다 —
|
|
54
54
|
# tool_response가 문자열인 경우 인덱싱 에러로 레코드가 통째로 유실되므로 type 검사로 감싼다.
|
|
55
|
+
# .tool_input도 같은 이유로 type 검사를 씌운다 (#1656 A 실측): 스칼라/배열 tool_input이
|
|
56
|
+
# 오면 jq가 인덱싱 에러로 죽고 `2>/dev/null || true` 때문에 **조용히 레코드 전체가
|
|
57
|
+
# 유실**됐다 — rc는 0이고 stderr도 비어 있어 크래시가 아니라 데이터 손실로 나타난다.
|
|
55
58
|
#
|
|
56
59
|
# 단일 라인(<1KB) append 이므로 O_APPEND 원자성에 기대어 병렬 에이전트 환경에서도 안전.
|
|
57
60
|
printf '%s' "$input" \
|
|
@@ -61,7 +64,9 @@ printf '%s' "$input" \
|
|
|
61
64
|
session: (.session_id // ""),
|
|
62
65
|
cwd: $cwd,
|
|
63
66
|
tool: (.tool_name // "unknown"),
|
|
64
|
-
target: ((.tool_input
|
|
67
|
+
target: ((if (.tool_input | type) == "object"
|
|
68
|
+
then (.tool_input.command // .tool_input.file_path // "")
|
|
69
|
+
else "" end) | tostring | .[0:160]),
|
|
65
70
|
interrupt: (.is_interrupt == true),
|
|
66
71
|
err: ((.error
|
|
67
72
|
// .tool_error
|
|
@@ -14,8 +14,18 @@ command -v jq >/dev/null 2>&1 || exit 0
|
|
|
14
14
|
|
|
15
15
|
input=$(cat)
|
|
16
16
|
|
|
17
|
-
|
|
18
|
-
|
|
17
|
+
# Guard: non-object stdin (non-JSON / JSON array / empty) — swallow and exit 0
|
|
18
|
+
# instead of reflecting garbage back on stdout. The sibling hooks (audit-log.sh,
|
|
19
|
+
# model-escalation-advisor.sh, task-outcome-recorder.sh) all carry this guard;
|
|
20
|
+
# this one did not, so a non-object payload was echoed through verbatim. (#1656 F)
|
|
21
|
+
printf '%s' "$input" | jq -e 'type=="object"' >/dev/null 2>&1 || exit 0
|
|
22
|
+
|
|
23
|
+
# `?` suppresses jq's "Cannot index <type> with string" when tool_input arrives as a
|
|
24
|
+
# scalar/array instead of an object (measured #1656 A: the bare form leaked a jq error to
|
|
25
|
+
# stderr on every non-object tool_input). 2>/dev/null + `|| echo ""` keep the hook silent
|
|
26
|
+
# and advisory-only even if jq itself fails for an unrelated reason.
|
|
27
|
+
agent_type=$(printf '%s\n' "$input" | jq -r '.tool_input.subagent_type? // ""' 2>/dev/null || echo "")
|
|
28
|
+
prompt=$(printf '%s\n' "$input" | jq -r '.tool_input.prompt? // ""' 2>/dev/null || echo "")
|
|
19
29
|
|
|
20
30
|
# R010 violation tracking file (PPID-scoped for session persistence)
|
|
21
31
|
VIOLATION_FILE="/tmp/.claude-r010-violations-${PPID}"
|
|
@@ -34,13 +34,22 @@ CONSECUTIVE_THRESHOLD=3
|
|
|
34
34
|
COOLDOWN=5
|
|
35
35
|
|
|
36
36
|
# Count failures for this agent type
|
|
37
|
+
#
|
|
38
|
+
# `grep -c` COUNT HYGIENE (#1656 E): a no-match `grep -c` prints `0` on stdout AND
|
|
39
|
+
# exits 1, so the old `$(grep -c ... || echo "0")` appended a SECOND zero and the
|
|
40
|
+
# variable became the two-line string "0\n0". Every later `[ "$count" -ge N ]` then
|
|
41
|
+
# failed with "integer expression expected" on stderr — noise indistinguishable
|
|
42
|
+
# from a real hook fault, on the exact path (zero failures) that is the common case.
|
|
43
|
+
# `|| true` keeps grep's own `0` and `${x:-0}` covers an unreadable file.
|
|
37
44
|
agent_failures=0
|
|
38
45
|
if [ -n "$agent_type" ] && [ "$agent_type" != "unknown" ]; then
|
|
39
|
-
agent_failures=$(grep -c "\"agent_type\":\"${agent_type}\".*\"outcome\":\"failure\"" "$OUTCOME_FILE" 2>/dev/null ||
|
|
46
|
+
agent_failures=$(grep -c "\"agent_type\":\"${agent_type}\".*\"outcome\":\"failure\"" "$OUTCOME_FILE" 2>/dev/null || true)
|
|
47
|
+
agent_failures=${agent_failures:-0}
|
|
40
48
|
fi
|
|
41
49
|
|
|
42
50
|
# Count consecutive failures (tail entries)
|
|
43
|
-
consecutive_failures=$(tail -${CONSECUTIVE_THRESHOLD} "$OUTCOME_FILE" 2>/dev/null | grep -c '"outcome":"failure"' 2>/dev/null ||
|
|
51
|
+
consecutive_failures=$(tail -${CONSECUTIVE_THRESHOLD} "$OUTCOME_FILE" 2>/dev/null | grep -c '"outcome":"failure"' 2>/dev/null || true)
|
|
52
|
+
consecutive_failures=${consecutive_failures:-0}
|
|
44
53
|
|
|
45
54
|
# Escalation path
|
|
46
55
|
# NOTE: Agent tool `model` param is an enum of exactly 4 values: sonnet | opus | haiku | fable.
|
|
@@ -90,7 +99,9 @@ fi
|
|
|
90
99
|
|
|
91
100
|
# De-escalation check
|
|
92
101
|
if [ "$current_model" != "haiku" ] && [ "$current_model" != "inherit" ] && [ "$current_model" != "" ]; then
|
|
93
|
-
|
|
102
|
+
# Same `grep -c` hygiene as above (#1656 E).
|
|
103
|
+
recent_successes=$(tail -${COOLDOWN} "$OUTCOME_FILE" 2>/dev/null | grep -c '"outcome":"success"' 2>/dev/null || true)
|
|
104
|
+
recent_successes=${recent_successes:-0}
|
|
94
105
|
|
|
95
106
|
if [ "$recent_successes" -ge "$COOLDOWN" ]; then
|
|
96
107
|
lower_model=""
|
|
@@ -12,11 +12,19 @@ input=$(cat)
|
|
|
12
12
|
# Hooks must never crash (R021); jq parse errors would otherwise abort under `set -e`. (#1650)
|
|
13
13
|
printf '%s' "$input" | jq -e 'type=="object"' >/dev/null 2>&1 || exit 0
|
|
14
14
|
# PostToolUse carries the tool result under `tool_response`, not `tool_output`
|
|
15
|
-
# (measured
|
|
16
|
-
#
|
|
17
|
-
# "" and this hook never compressed anything.
|
|
18
|
-
#
|
|
19
|
-
#
|
|
15
|
+
# (measured at v1.1.62: every PostToolUse record in this project's transcripts
|
|
16
|
+
# carried `tool_response` and none carried `tool_output`), so the previous
|
|
17
|
+
# `.tool_output` read always yielded "" and this hook never compressed anything.
|
|
18
|
+
# MCP responses arrive either as a bare string or as
|
|
19
|
+
# `{content: [{type:"text", text:...}]}`; `.tool_output` is kept as a fallback
|
|
20
|
+
# for events still using the older shape.
|
|
21
|
+
#
|
|
22
|
+
# MATCHER-SCOPED: `.stdout` and `.file.content` are read here only because this
|
|
23
|
+
# hook's matcher is `mcp__playwright__.*|mcp__claude-in-chrome__.*`, so those
|
|
24
|
+
# branches can never fire on a real Bash or Read result. This is a lossy hook —
|
|
25
|
+
# it REPLACES `.tool_response` with a Haiku summary — so widening the matcher to
|
|
26
|
+
# cover Bash/Read without first dropping those two branches would silently
|
|
27
|
+
# destroy genuine command output and file contents. (#1656 F)
|
|
20
28
|
tool_output=$(printf '%s\n' "$input" | jq -r '
|
|
21
29
|
[
|
|
22
30
|
(.tool_response? | if type == "string" then . else empty end),
|
|
@@ -52,6 +52,9 @@
|
|
|
52
52
|
# with no tool_result block) ends a turn.
|
|
53
53
|
# * `thinking` blocks are interleaved with text/tool_use and never carry an R008 prefix;
|
|
54
54
|
# they are filtered out before analysis.
|
|
55
|
+
# * NOTE (#1654, v1.1.63): narration 블록은 트랜스크립트에 type:"thinking"(signature 라벨
|
|
56
|
+
# narration)으로 직렬화되므로 아래 select(.type? != "thinking")이 함께 배제한다 — 의도된
|
|
57
|
+
# 동작: narration에는 R007/R008 마커가 실리지 않는다(실측 47블록 0건).
|
|
55
58
|
#
|
|
56
59
|
# ── R008 verdict: TURN-LEVEL COUNTING, not block adjacency (#1563 찐빠 #1) ─────────────
|
|
57
60
|
# R008 (`.claude/rules/MUST-tool-identification.md`) says, verbatim:
|
|
@@ -12,8 +12,8 @@ command -v jq >/dev/null 2>&1 || exit 0
|
|
|
12
12
|
|
|
13
13
|
input=$(cat)
|
|
14
14
|
|
|
15
|
-
# Non-object stdin guard (#1650 B)
|
|
16
|
-
#
|
|
15
|
+
# Non-object stdin guard (#1650 B). PostToolUse stdin is always an object; a
|
|
16
|
+
# bare-string stdin is rejected by design, not because it is known to be empty.
|
|
17
17
|
printf '%s' "$input" | jq -e 'type=="object"' >/dev/null 2>&1 || exit 0
|
|
18
18
|
|
|
19
19
|
tool_name=$(printf '%s\n' "$input" | jq -r '.tool_name? // "unknown"')
|
|
@@ -21,19 +21,23 @@ tool_name=$(printf '%s\n' "$input" | jq -r '.tool_name? // "unknown"')
|
|
|
21
21
|
# Collect every text-bearing field of the payload into one scan buffer.
|
|
22
22
|
#
|
|
23
23
|
# PostToolUse carries the result under `tool_response`, NOT `tool_output`
|
|
24
|
-
# (measured
|
|
25
|
-
#
|
|
24
|
+
# (measured at v1.1.62 against this project's session transcripts: every
|
|
25
|
+
# PostToolUse record carried `tool_response` and none carried `tool_output`;
|
|
26
26
|
# the CC 2.1.259 embedded hook reference documents `"tool_response": {...}
|
|
27
27
|
# // PostToolUse only`). Reading `.tool_output.output` therefore always yielded
|
|
28
28
|
# "" and the scan below never ran — every AWS key, private key and PAT passed
|
|
29
29
|
# through unflagged.
|
|
30
30
|
#
|
|
31
|
-
#
|
|
32
|
-
#
|
|
33
|
-
#
|
|
34
|
-
#
|
|
35
|
-
#
|
|
36
|
-
#
|
|
31
|
+
# Text-bearing fields, by tool. This hook's matcher is `Bash|Read|Grep`, so only
|
|
32
|
+
# the first two rows fire in production; the rest are defensive so a matcher
|
|
33
|
+
# widening does not silently reintroduce the blind spot above.
|
|
34
|
+
# Bash .tool_response.stdout / .stderr (measured, in matcher)
|
|
35
|
+
# Read .tool_response.file.content (measured, in matcher)
|
|
36
|
+
# Write .tool_response.content (measured, out of matcher)
|
|
37
|
+
# Agent .tool_response.prompt / .output (measured, out of matcher)
|
|
38
|
+
# MCP .tool_response (string)
|
|
39
|
+
# or .tool_response.content[].text (defensive — unmeasured;
|
|
40
|
+
# MCP corpus 0 records)
|
|
37
41
|
# `.tool_output.output` and a string-shaped `.tool_output` are kept as
|
|
38
42
|
# fallbacks for events that still use the older shape (e.g. SubagentStop).
|
|
39
43
|
#
|
|
@@ -207,6 +207,9 @@ _strip_heredoc_bodies() {
|
|
|
207
207
|
# A literal single quote cannot be written inside the bracket expression of
|
|
208
208
|
# the quote-counting expansion below, so hold one in a variable.
|
|
209
209
|
local sq_char="'"
|
|
210
|
+
# Same expression the extraction below used to hand to "sed -n -E", now fed
|
|
211
|
+
# to bash's own "=~" (see the delimiter scan for why it moved).
|
|
212
|
+
local delim_re="^.*[^<]<<-?[[:space:]]*[\"']?([A-Za-z_][A-Za-z0-9_]*)[\"']?.*\$"
|
|
210
213
|
while IFS= read -r line; do
|
|
211
214
|
if [ "$in_body" -eq 1 ]; then
|
|
212
215
|
# "<<-" allows leading tabs before the terminator; accept leading
|
|
@@ -253,8 +256,18 @@ _strip_heredoc_bodies() {
|
|
|
253
256
|
printf '%s' "${out}__QUOTED_HEREDOC_OPENER__"$'\n'
|
|
254
257
|
return 0
|
|
255
258
|
fi
|
|
256
|
-
|
|
257
|
-
|
|
259
|
+
# bash's "=~" runs the SAME POSIX ERE the "sed -n -E" here used to run —
|
|
260
|
+
# both engines are leftmost-longest, so the greedy "^.*" still settles on
|
|
261
|
+
# the LAST "<<" (verified against sed over 251 cases, #1656 E). Two forks
|
|
262
|
+
# per opener line are saved, but the reason it moved is correctness: sed
|
|
263
|
+
# aborts with "RE error: illegal byte sequence" on a line carrying invalid
|
|
264
|
+
# UTF-8, and under "set -o pipefail" that non-zero pipeline killed the
|
|
265
|
+
# whole hook. "=~" simply reports no match there.
|
|
266
|
+
if [[ "$line" =~ $delim_re ]]; then
|
|
267
|
+
delim="${BASH_REMATCH[1]}"
|
|
268
|
+
else
|
|
269
|
+
delim=""
|
|
270
|
+
fi
|
|
258
271
|
if [ -n "$delim" ]; then
|
|
259
272
|
in_body=1
|
|
260
273
|
fi
|
|
@@ -292,7 +305,28 @@ is_readonly_bash_command() {
|
|
|
292
305
|
return 0
|
|
293
306
|
fi
|
|
294
307
|
|
|
295
|
-
|
|
308
|
+
# sed applies "^" and "$" once PER LINE, so this trims every line of a
|
|
309
|
+
# multi-line command — it is NOT the whole-string parameter-expansion trim
|
|
310
|
+
# used in Step 6 and in _strip_heredoc_bodies, which is why it survived
|
|
311
|
+
# #1650 A. The loop below reproduces the per-line form exactly: each line is
|
|
312
|
+
# trimmed and re-terminated, then the trailing newlines are dropped the same
|
|
313
|
+
# way the "$( )" around the old pipeline dropped them. Two forks are saved,
|
|
314
|
+
# but the reason it changed (#1656 E) is correctness: sed aborts a multi-line
|
|
315
|
+
# command with "RE error: illegal byte sequence" at the first invalid UTF-8
|
|
316
|
+
# byte, and under "set -o pipefail" that non-zero pipeline killed the hook
|
|
317
|
+
# outright — so a heredoc body carrying raw bytes could stop the classifier
|
|
318
|
+
# before it ever reached a trailing "rm -rf". Parameter expansion has no
|
|
319
|
+
# locale dependency.
|
|
320
|
+
local trimmed="" tline
|
|
321
|
+
while IFS= read -r tline; do
|
|
322
|
+
tline="${tline#"${tline%%[![:space:]]*}"}"
|
|
323
|
+
tline="${tline%"${tline##*[![:space:]]}"}"
|
|
324
|
+
trimmed="${trimmed}${tline}"$'\n'
|
|
325
|
+
done <<< "$cmd"
|
|
326
|
+
while [ "${trimmed%$'\n'}" != "$trimmed" ]; do
|
|
327
|
+
trimmed="${trimmed%$'\n'}"
|
|
328
|
+
done
|
|
329
|
+
cmd="$trimmed"
|
|
296
330
|
if [ -z "$cmd" ]; then
|
|
297
331
|
echo "false"
|
|
298
332
|
return 0
|
|
@@ -5,9 +5,33 @@ set -euo pipefail
|
|
|
5
5
|
command -v jq >/dev/null 2>&1 || exit 0
|
|
6
6
|
|
|
7
7
|
# Task/Agent Outcome Recorder Hook
|
|
8
|
-
# Trigger:
|
|
9
|
-
# Purpose: Record
|
|
8
|
+
# Trigger: SubagentStop (settings.json wires this script to that event and no other)
|
|
9
|
+
# Purpose: Record agent outcomes for model escalation decisions
|
|
10
10
|
# Protocol: stdin JSON -> process -> stdout pass-through, exit 0 always
|
|
11
|
+
#
|
|
12
|
+
# MEASURED SubagentStop payload (CC 2.1.259 embedded hook schema; #1656 B):
|
|
13
|
+
# session_id, transcript_path, cwd, prompt_id?, permission_mode?, effort?,
|
|
14
|
+
# hook_event_name:"SubagentStop", stop_hook_active, agent_id,
|
|
15
|
+
# agent_transcript_path, agent_type,
|
|
16
|
+
# last_assistant_message?, background_tasks?, session_crons?
|
|
17
|
+
#
|
|
18
|
+
# Note what the event does NOT carry: no `tool_input`, no `tool_output`, no
|
|
19
|
+
# `model`, no `description`, no `prompt`, and no error signal of any kind. The
|
|
20
|
+
# selectors below therefore read the top-level fields first and fall back to
|
|
21
|
+
# `last_assistant_message` — the one text-bearing field SubagentStop provides —
|
|
22
|
+
# for the description and skill-name extraction. The `tool_input`/`tool_output`
|
|
23
|
+
# branches are retained as fallbacks so the script stays correct if it is ever
|
|
24
|
+
# re-wired to PostToolUse.
|
|
25
|
+
#
|
|
26
|
+
# This project's transcript corpus (889 files) holds ZERO SubagentStop hook
|
|
27
|
+
# records — SubagentStop output is not attached to the parent transcript — so
|
|
28
|
+
# the shape above is schema-derived, not corpus-derived.
|
|
29
|
+
#
|
|
30
|
+
# Outcome caveat: SubagentStop exposes no failure signal, so entries recorded
|
|
31
|
+
# from that event are always outcome=success. There is no working hand-off for
|
|
32
|
+
# that gap today: SubagentStop payload carries no failure signal;
|
|
33
|
+
# subagent-failure-advisor.sh currently reads `.tool_output.is_error`, which is
|
|
34
|
+
# absent here, so failure detection on this path is 0 until #1656 G is resolved.
|
|
11
35
|
|
|
12
36
|
input=$(cat)
|
|
13
37
|
|
|
@@ -15,11 +39,24 @@ input=$(cat)
|
|
|
15
39
|
# jq's parse error propagates through `set -euo pipefail` as rc=5 without this.
|
|
16
40
|
printf '%s' "$input" | jq -e 'type=="object"' >/dev/null 2>&1 || exit 0
|
|
17
41
|
|
|
18
|
-
# Extract
|
|
19
|
-
#
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
42
|
+
# Extract agent info. Measured SubagentStop fields come first; the PostToolUse
|
|
43
|
+
# `tool_input.*` shape is kept as a fallback.
|
|
44
|
+
# `?` suppresses "Cannot index string with ..." when tool_input arrives as a scalar;
|
|
45
|
+
# `tostring` normalises a scalar (e.g. a numeric agent_type) instead of aborting.
|
|
46
|
+
agent_type=$(printf '%s\n' "$input" | jq -r \
|
|
47
|
+
'(.agent_type? // .tool_input?.subagent_type? // "unknown") | tostring' 2>/dev/null) \
|
|
48
|
+
|| agent_type="unknown"
|
|
49
|
+
# SubagentStop carries no `model`; this resolves to "inherit" there by design.
|
|
50
|
+
model=$(printf '%s\n' "$input" | jq -r \
|
|
51
|
+
'(.model? // .tool_input?.model? // "inherit") | tostring' 2>/dev/null) || model="inherit"
|
|
52
|
+
# `strings` drops non-string shapes so an object-valued field degrades to "" rather
|
|
53
|
+
# than dumping raw JSON into the description.
|
|
54
|
+
description=$(printf '%s\n' "$input" | jq -r '
|
|
55
|
+
[ (.tool_input?.description? | strings),
|
|
56
|
+
(.description? | strings),
|
|
57
|
+
(.last_assistant_message? | strings) ]
|
|
58
|
+
| map(select(. != "")) | first // ""
|
|
59
|
+
' 2>/dev/null | head -c 200) || description=""
|
|
23
60
|
|
|
24
61
|
# Extract skill name from description or prompt
|
|
25
62
|
skill_name=""
|
|
@@ -28,18 +65,35 @@ skill_name=""
|
|
|
28
65
|
if echo "$description" | grep -qiE '(skill:|routing|→.*skill)'; then
|
|
29
66
|
skill_name=$(echo "$description" | grep -oiE '[a-z]+-[a-z]+(-[a-z]+)*-?(routing|skill|practices|detection|decomposition|orchestration|pipeline|guards|cycle|plan|review|refactor|publish|version|audit|exec|analyze|bundle|report|setup|watch|lists|status|help|save|recall)' | head -1 || true)
|
|
30
67
|
fi
|
|
31
|
-
# Fallback: check prompt
|
|
68
|
+
# Fallback: check the prompt / last assistant message for a "Skill: {name}" pattern.
|
|
69
|
+
# SubagentStop has no `prompt`, so `last_assistant_message` is the live source here.
|
|
32
70
|
if [ -z "$skill_name" ]; then
|
|
33
|
-
prompt=$(printf '%s\n' "$input" | jq -r '
|
|
34
|
-
|
|
71
|
+
prompt=$(printf '%s\n' "$input" | jq -r '
|
|
72
|
+
[ (.tool_input?.prompt? | strings),
|
|
73
|
+
(.last_assistant_message? | strings) ]
|
|
74
|
+
| map(select(. != "")) | first // ""
|
|
75
|
+
' 2>/dev/null | head -c 500) || prompt=""
|
|
76
|
+
# POSIX character classes, not `\s`: BSD sed (macOS, this repo's runtime) does not
|
|
77
|
+
# implement the GNU `\s` escape, so `s/[Ss]kill:\s*//` stripped only the colon and
|
|
78
|
+
# left a leading space on every extracted skill name. Same family as the BSD `\?`
|
|
79
|
+
# gap recorded in R005. (#1656 B)
|
|
80
|
+
skill_name=$(echo "$prompt" | grep -oiE 'Skill:[[:space:]]*[a-z]+-[a-z]+(-[a-z]+)*' | sed 's/[Ss]kill:[[:space:]]*//' | head -1 || true)
|
|
35
81
|
fi
|
|
36
82
|
|
|
37
|
-
# Determine outcome
|
|
38
|
-
is_error
|
|
83
|
+
# Determine outcome. SubagentStop carries no error signal, so this is always false
|
|
84
|
+
# there; `.tool_response.is_error` is the measured PostToolUse field (`.tool_output`
|
|
85
|
+
# is the legacy shape) and both are kept for re-wiring safety.
|
|
86
|
+
is_error=$(printf '%s\n' "$input" | jq -r '
|
|
87
|
+
(.tool_response?.is_error? // .tool_output?.is_error? // false) | tostring
|
|
88
|
+
' 2>/dev/null) || is_error="false"
|
|
39
89
|
|
|
40
90
|
if [ "$is_error" = "true" ]; then
|
|
41
91
|
outcome="failure"
|
|
42
|
-
error_summary=$(printf '%s\n' "$input" | jq -r '
|
|
92
|
+
error_summary=$(printf '%s\n' "$input" | jq -r '
|
|
93
|
+
[ (.tool_response?.output? | strings),
|
|
94
|
+
(.tool_output?.output? | strings) ]
|
|
95
|
+
| map(select(. != "")) | first // ""
|
|
96
|
+
' 2>/dev/null | head -c 200) || error_summary=""
|
|
43
97
|
else
|
|
44
98
|
outcome="success"
|
|
45
99
|
error_summary=""
|
|
@@ -50,11 +104,29 @@ OUTCOME_FILE="/tmp/.claude-task-outcomes-${PPID}"
|
|
|
50
104
|
TASK_COUNT_FILE="/tmp/.claude-task-count-${PPID}"
|
|
51
105
|
|
|
52
106
|
# --- Pattern Detection ---
|
|
53
|
-
# Priority: skill-specific patterns > parallel >
|
|
54
|
-
|
|
107
|
+
# Priority: skill-specific patterns > parallel > agent-count inference > default.
|
|
108
|
+
#
|
|
109
|
+
# INPUT SCOPE (#1656 D): the cascade below reads ONLY `tool_input.description` —
|
|
110
|
+
# the spawn argument the orchestrator wrote — never `$description`, which now
|
|
111
|
+
# falls back to `last_assistant_message`, i.e. free-form prose the subagent wrote
|
|
112
|
+
# about itself. A closing summary that merely says "ran the parallel review" or
|
|
113
|
+
# "orchestrator finished" would otherwise be recorded as
|
|
114
|
+
# pattern_used=parallel/orchestrator, fabricating a workflow shape out of English.
|
|
115
|
+
# SubagentStop carries no `tool_input`, so on that event this source is empty and
|
|
116
|
+
# the pattern degrades to the session-level agent-count signal (a real, non-prose
|
|
117
|
+
# signal) and then to "unknown" — an honest absence rather than a default
|
|
118
|
+
# "sequential" that reads as a measurement.
|
|
119
|
+
pattern_source=$(printf '%s\n' "$input" | jq -r '
|
|
120
|
+
(.tool_input?.description? | strings) // ""
|
|
121
|
+
' 2>/dev/null | head -c 200) || pattern_source=""
|
|
122
|
+
|
|
123
|
+
if [ -n "$pattern_source" ]; then
|
|
124
|
+
pattern="sequential"
|
|
125
|
+
else
|
|
126
|
+
pattern="unknown"
|
|
127
|
+
fi
|
|
55
128
|
|
|
56
|
-
|
|
57
|
-
desc_lower=$(echo "$description" | tr '[:upper:]' '[:lower:]')
|
|
129
|
+
desc_lower=$(printf '%s' "$pattern_source" | tr '[:upper:]' '[:lower:]')
|
|
58
130
|
|
|
59
131
|
if echo "$desc_lower" | grep -qE '(evaluator.optimizer|evaluator_optimizer)'; then
|
|
60
132
|
pattern="evaluator-optimizer"
|
|
@@ -87,9 +159,17 @@ if [ -f "$AGENT_START_FILE" ]; then
|
|
|
87
159
|
fi
|
|
88
160
|
fi
|
|
89
161
|
|
|
90
|
-
# Append JSON line entry
|
|
162
|
+
# Append JSON line entry.
|
|
163
|
+
# `-c` (compact) is REQUIRED, not cosmetic: this file is consumed as JSONL by
|
|
164
|
+
# eval-core's outcome-parser (`content.split('\n')` + `JSON.parse(line)`) and by
|
|
165
|
+
# model-escalation-advisor.sh (`grep -c '"agent_type":"X".*"outcome":"failure"'`,
|
|
166
|
+
# `tail -N`), and the ring buffer below trims by `wc -l`. A pretty-printed entry
|
|
167
|
+
# spans 11 lines (measured: 9 fields plus the braces), so it broke every one of
|
|
168
|
+
# those readers and let `tail -50` slice
|
|
169
|
+
# an object in half. Sibling recorders (agent-start-recorder.sh, stuck-detector.sh)
|
|
170
|
+
# already use `jq -cn`. (#1656 B)
|
|
91
171
|
timestamp=$(date -u +%Y-%m-%dT%H:%M:%SZ)
|
|
92
|
-
entry=$(jq -
|
|
172
|
+
entry=$(jq -cn \
|
|
93
173
|
--arg ts "$timestamp" \
|
|
94
174
|
--arg agent "$agent_type" \
|
|
95
175
|
--arg model "$model" \
|
|
@@ -388,6 +388,12 @@ This applies when a change touches a field that participates in an override/prec
|
|
|
388
388
|
|--------------|----------|
|
|
389
389
|
| Plan a provider/endpoint switch as N commands without reading the config's override chain | Read the full config schema (which field wins, defaults, inheritance) → enumerate EVERY field the switch touches (incl. base_url) → then plan |
|
|
390
390
|
|
|
391
|
+
**훅 스크립트 각도 — stdin 필드 형상은 실측 후 편집 (Origin: #1658 #1, v1.1.62)**: 훅 스크립트가 읽는 stdin 필드(`tool_input`/`tool_response`/`agent_id` 등)를 편집·가드·억제하기 전에 **실제 페이로드 형상을 실측**한다 — 트랜스크립트의 `attachment.type=="hook_success"` 레코드에서 `attachment.stdout`이 pass-through 훅이 되돌린 stdin 원문이며, CC 바이너리 내장 훅 문서로 교차검증한다. 처방("가드 추가·`?` 억제")만 위임하면 선택자 결함 위에 가드를 얹어 마지막 실패 신호까지 지운다. 실증: v1.1.62에서 `secret-filter.sh`가 PostToolUse에 존재하지 않는 `tool_output`(0/1764)을 읽어 실제 페이로드를 한 번도 스캔하지 않던 선재 결함 위에 `?` 억제가 추가됐고(rc=5 신호 소멸), 적대적 리뷰가 실측 형상 재현으로 FAIL 판정해 `tool_response`(1764/1764)로 교체했다. 훅 편집 위임서 표준 문안: "스크립트가 읽는 stdin 필드는 `hook_success` 레코드로 형상 실측 후 편집".
|
|
392
|
+
|
|
393
|
+
| Anti-pattern | Required |
|
|
394
|
+
|--------------|----------|
|
|
395
|
+
| 훅 위임서에 "가드 추가·`?` 억제" 처방만 전달 | 읽는 필드의 실제 형상(`hook_success` stdin 원문 + 바이너리 훅 문서) 실측을 위임서 완료 조건에 포함 |
|
|
396
|
+
|
|
391
397
|
Sibling discipline to Read-Before-Characterize (that rule governs diagnosis — don't label before reading; this one governs edit-planning completeness — enumerate every interdependent field before editing). Cross-ref: R023 (verification ladder — config completeness is a Tier-1 deterministic pre-check).
|
|
392
398
|
|
|
393
399
|
### Degraded-Output Re-Verification Gate (529 / buffering)
|
|
@@ -390,6 +390,8 @@ Origin: #1595 #2 (v1.1.48 세션 — `git checkout -b release/v1.1.48 develop`
|
|
|
390
390
|
| 파일 소유권만 고지하고 동일 검증 명령을 각 에이전트 완료 조건에 넣어 병렬 발주 | 검증 명령 공유를 고지하거나 검증을 오케스트레이터가 직렬 1회로 회수 |
|
|
391
391
|
| 공유 `$TMPDIR`에 고정 경로로 임시 파일을 쓰고 그 디렉토리를 전수 계수 | 에이전트별 고유 경로 사용 + 그 경로만 계수 |
|
|
392
392
|
|
|
393
|
+
**순차 위임도 고지 대상 — 직전 완료 변경분의 출처 (Origin: #1658 #4)**: 병렬 형제뿐 아니라 **이미 워킹트리에 있는 미커밋 변경분의 출처**(직전에 완료한 에이전트와 그 담당 파일)를 위임서에 한 줄 고지한다. 고지가 없으면 후속 에이전트가 `git diff`에 보이는 타 변경분을 "형제 담당분"으로 오귀속해 서술한다(v1.1.62 세션 3건 — 행동에는 영향 없었으나 보고가 오염). 문안: "워킹트리의 미커밋 변경 중 X·Y는 직전 에이전트 [N]의 완료분이다 — 건드리지 말고 보고에서도 네 변경분과 구분하라."
|
|
394
|
+
|
|
393
395
|
> Origin: #1518 (찐빠 #3 — 미고지 git 에이전트가 형제를 "외부 프로세스"로 오귀속; 같은 세션에서 고지한 4개 구현 에이전트는 전원 정확히 구분 보고 — 대조 실증). Cross-ref: R009 (병렬 실행 조건).
|
|
394
396
|
|
|
395
397
|
> Origin 보강: #1598 — 파일이 완전 disjoint한 병렬 배치에서 위양성 4종 발생(judge.sh 테스트 7건 ENOENT: 두 테스트가 tracked `verdict-schema.json`을 cp→rm→복구 / reviewers.sh 타임아웃 테스트 간헐 실패: CPU 포화 / "임시 파일 누수 1건" 오측정: 형제 잔여물, 격리 셔임 재측정 시 0). **3종의 원인은 오케스트레이터가 위임서에 넣은 완료 조건 자체였다** — 형제 고지의 결함이 아니라 고지 항목의 누락이다.
|
|
@@ -125,6 +125,17 @@ Origin: #1595 #5 (v1.1.48 세션 — R008 위반 3건이 단일 턴에 집중. t
|
|
|
125
125
|
> **v2.1.174+**: Fixed the Workflow tool's `agent()` subagents missing per-agent attribution headers. Workflow-spawned subagents now carry attribution consistent with R008 — when authoring Workflow scripts, each `agent()` call is attributed like a direct Agent tool spawn. Align Workflow orchestration with the R008 `[agent][model] → Tool:` identification discipline: a Workflow `agent()` fan-out should still be reasoned about with the same per-agent identification model as parallel Agent tool spawns.
|
|
126
126
|
-->
|
|
127
127
|
|
|
128
|
+
## announce와 헤더는 narration이 아니라 visible text 블록으로 (Origin: #1654)
|
|
129
|
+
|
|
130
|
+
모델 출력에는 `text` 블록과 **narration 블록**(트랜스크립트에 `type:"thinking"` + signature 라벨 `narration`으로 직렬화되는 사용자향 짧은 산문)이 있고, 한 API 메시지에는 **둘 중 하나만** 실린다(v1.1.61~62 세션 실측: 115메시지 중 공존 0). 도구 호출 턴을 narration 요약 한 문장("…했습니다. 이제 …하겠습니다")으로 시작하면 R007 헤더와 R008 접두사는 **어디에도 남지 않는다** — 실측: narration 47블록에 R007 헤더 0건, 대괄호 번호 항목 0건, Tool 표기 0건(조사 문장 인용 제외). advisor는 `type != "thinking"` 필터로 narration을 배제하므로 이 턴들은 전부 누락으로 계상되며, 실제로 v1.1.61 세션 advisory 18건은 **전부 진양성**이었다(직렬화 유실 가설은 advisor가 메시지 직후에 판정했다는 사실로 배제됨).
|
|
131
|
+
|
|
132
|
+
| Anti-pattern | Required |
|
|
133
|
+
|--------------|----------|
|
|
134
|
+
| 도구 호출 턴을 짧은 요약 산문만으로 시작(narration 채널로 흐름) | 헤더(`┌─ Agent:` 또는 단축 헤더)와 Core Rule 접두사를 **text 블록**으로 명시 — 산문 요약은 그 뒤에 |
|
|
135
|
+
| "announce를 썼다"는 기억으로 advisory를 오탐으로 가정 | 트랜스크립트의 `text` 블록에서 마커를 실측(R020 Self-Violation Counting) |
|
|
136
|
+
|
|
137
|
+
Iteration 1(Agent 스폰 15메시지 전부 narration)과 Iteration 2(7메시지 text)의 대비는 계수 도구 결함이 아니라 출력 채널 선택의 차이였다. 채널 선택 요인은 미귀속이다.
|
|
138
|
+
|
|
128
139
|
## Tier-3 Interaction Tool Prefix (MANDATORY)
|
|
129
140
|
|
|
130
141
|
R008 "every tool call" applies to Tier-3 interaction tools too — NOT only file/exec tools. Applying the Core Rule prefix form (에이전트·모델 대괄호 다음 화살표와 Tool 표기) to Agent/Bash/Read while omitting it on `AskUserQuestion`, `TodoWrite`, `EnterPlanMode`, etc. is a violation.
|
|
@@ -113,6 +113,8 @@ Origin: #1455 #1 (Session 127 회고 찐빠 #1) — cc-release-monitor PR #1449
|
|
|
113
113
|
|--------------|----------|
|
|
114
114
|
| "변경분 영향 범위"만 보고 검증 항목을 정해 위임 → CI 전용 잡(lint 등) 누락 | 워크플로 잡 목록을 하한선으로 삼아 로컬 대응 명령을 완료 조건에 열거 |
|
|
115
115
|
|
|
116
|
+
**상한선 — 오케스트레이터 사전 실측 항목은 재실행 금지 (Origin: #1655 제안 2)**: 하한선이 CI 잡 목록이라면 상한선은 "오케스트레이터가 이미 실측한 항목"이다. 검증 위임서에는 사전 실측 결과(테스트 pass/fail, lint·typecheck exit, 스크립트 exit, 미러 md5)를 **재실행 금지 목록**으로 열거하고 재실측 대상만 지정한다 — 서브에이전트가 전체 테스트를 재실행하면 턴 예산이 소진돼 판정 없이 절단된다. 실증: v1.1.61 세션 mgr-sauron 1차 위임(금지 목록 없음)은 `bun test` 재실행으로 25턴 절단, 2차 위임(금지 목록 + 15턴 내 판정 명시)은 16 tool_uses로 PASS 완주(R020 maxTurns 절단 누적 7건째). auto-dev.yaml deep-verify 스텝 description에 같은 문안이 배선돼 있다.
|
|
117
|
+
|
|
116
118
|
Origin: #1574 (v1.1.44 세션 — 병렬 위임 3건 모두 `bun run lint`를 누락해 verify-build halt, 수정 에이전트 1회 추가 발주). 기존 `feedback_delegation_verify_scope_by_impact`("영향 범위 기준")의 하한선을 명문화한 것이다. Cross-reference: R020(완료 검증 — 선언 전 실제 게이트 통과 확인), R017(커밋 전 검증 게이트).
|
|
117
119
|
|
|
118
120
|
## Conditional-Output Verification — Positive/Negative Pair Mandate (Origin: #1563 #2)
|
|
@@ -131,6 +131,7 @@ Each iteration operates under full project rules — no relaxation because FSD i
|
|
|
131
131
|
| R015 (intent persistence) | If user explicitly defers a specific PR this session, honor that deferral — do not retry it. |
|
|
132
132
|
| R017 (sync verification) | mgr-sauron passes required before any commit. |
|
|
133
133
|
| R020 (completion verification) | Each release verified via `npm view`, `gh release view`, closed issues before `[Done]`. PR merges verified via `gh pr view` ground-truth before declaring iteration complete. |
|
|
134
|
+
| cost-cap advisory (`cost-cap-advisor.sh`, Conversation Block) | FSD는 사용자가 명시 호출한 다중 릴리즈 루프라 세션 비용이 상한(기본 $5)을 필연적으로 넘는다. advisory 발화는 루프를 멈추는 신호가 아니며, 오케스트레이터는 발화 사실과 누적 비용을 반복 경계(homework 게이트)에서 사용자에게 고지한다. 상한을 올리려면 `CLAUDE_COST_CAP`(v1.1.61 세션 실측: 착수 직후 $5.38에서 발화). |
|
|
134
135
|
|
|
135
136
|
`/homework` runs as a **retrospective gate** between iterations — findings go through `omcustom-feedback`'s Phase 4A confirmation gate. The loop does NOT skip homework on the grounds that it is "automated". If homework requires user confirmation (e.g., to file a feedback issue), the loop pauses and waits.
|
|
136
137
|
|
|
@@ -441,7 +441,7 @@ steps:
|
|
|
441
441
|
|
|
442
442
|
- name: deep-verify
|
|
443
443
|
skill: deep-verify
|
|
444
|
-
description: "Multi-angle release quality verification — self-review checklist if docs-only; mgr-sauron R017 + core self-check if lite. MULTI-PHASE: 스킬을 그대로 spawn하지 말고 R020 「위임 경계를 Phase 개수로 설계」에 따라 단일 목표 위임으로 분할해 순차 발주한다 (#1595 #4). lite 분할 표준(#1652 #3-4): (1) mgr-sauron R017 구조 검증 단일 목표 위임 1건 + (2) 변경 성격별 적대적 리뷰 단일 목표 위임 1건 — 스크립트 변경이면 실행 재현 기반 adversarial-review, 룰/스킬/yaml 텍스트 변경이면 문구 정합·배선 리뷰. 근거: v1.1.59/60 두 반복 연속 적대적 리뷰가 신규 회귀(M-3/M-4, heredoc 위조)를 실행 재현으로 포착."
|
|
444
|
+
description: "Multi-angle release quality verification — self-review checklist if docs-only; mgr-sauron R017 + core self-check if lite. MULTI-PHASE: 스킬을 그대로 spawn하지 말고 R020 「위임 경계를 Phase 개수로 설계」에 따라 단일 목표 위임으로 분할해 순차 발주한다 (#1595 #4). lite 분할 표준(#1652 #3-4): (1) mgr-sauron R017 구조 검증 단일 목표 위임 1건 + (2) 변경 성격별 적대적 리뷰 단일 목표 위임 1건 — 스크립트 변경이면 실행 재현 기반 adversarial-review, 룰/스킬/yaml 텍스트 변경이면 문구 정합·배선 리뷰. 근거: v1.1.59/60 두 반복 연속 적대적 리뷰가 신규 회귀(M-3/M-4, heredoc 위조)를 실행 재현으로 포착. 검증 위임 표준 문안: 오케스트레이터가 이미 실측한 항목(bun test·lint·typecheck·template-sync·wiki-sync·validate-docs·미러 md5)은 위임서에 재실행 금지 목록으로 열거하고 재실측 대상만 지정한다 — v1.1.61 세션에서 금지 목록 없는 sauron 위임이 bun test를 재실행하다 25턴 절단됐고, 금지 목록을 명시한 재위임은 16 tool_uses로 완주했다(R020 maxTurns 절단 누적 7건째)."
|
|
445
445
|
depends_on: verify-build
|
|
446
446
|
|
|
447
447
|
- name: release
|
|
@@ -585,6 +585,10 @@ steps:
|
|
|
585
585
|
- If failures: diagnose, fix, re-verify
|
|
586
586
|
3. For npm projects with auto-tag.yml: MANDATORY additional check:
|
|
587
587
|
gh run list --workflow auto-tag.yml --limit 1 --json conclusion,displayTitle
|
|
588
|
+
⚠ auto-tag.yml run의 `headSha`는 **release 브랜치 head(PR head SHA)**이지 머지 커밋이 아니다 — 폴링 종료
|
|
589
|
+
조건을 머지 커밋 SHA로 걸면 영원히 미일치해 타임아웃까지 전량 소진한다(v1.1.61 세션 실측: 40회/10분 전량
|
|
590
|
+
소진, v1.1.62에서 PR head로 정정 → 22회 정상 종료). release.yml run은 `displayTitle`에 PR 번호/브랜치명이
|
|
591
|
+
들어가므로 그것으로 식별한다.
|
|
588
592
|
→ conclusion MUST be "success" before declaring release complete.
|
|
589
593
|
If auto-tag.yml concluded "failure": diagnose root cause. Do NOT retry with manual git tag.
|
|
590
594
|
Common causes:
|
package/templates/manifest.json
CHANGED
|
@@ -441,7 +441,7 @@ steps:
|
|
|
441
441
|
|
|
442
442
|
- name: deep-verify
|
|
443
443
|
skill: deep-verify
|
|
444
|
-
description: "Multi-angle release quality verification — self-review checklist if docs-only; mgr-sauron R017 + core self-check if lite. MULTI-PHASE: 스킬을 그대로 spawn하지 말고 R020 「위임 경계를 Phase 개수로 설계」에 따라 단일 목표 위임으로 분할해 순차 발주한다 (#1595 #4). lite 분할 표준(#1652 #3-4): (1) mgr-sauron R017 구조 검증 단일 목표 위임 1건 + (2) 변경 성격별 적대적 리뷰 단일 목표 위임 1건 — 스크립트 변경이면 실행 재현 기반 adversarial-review, 룰/스킬/yaml 텍스트 변경이면 문구 정합·배선 리뷰. 근거: v1.1.59/60 두 반복 연속 적대적 리뷰가 신규 회귀(M-3/M-4, heredoc 위조)를 실행 재현으로 포착."
|
|
444
|
+
description: "Multi-angle release quality verification — self-review checklist if docs-only; mgr-sauron R017 + core self-check if lite. MULTI-PHASE: 스킬을 그대로 spawn하지 말고 R020 「위임 경계를 Phase 개수로 설계」에 따라 단일 목표 위임으로 분할해 순차 발주한다 (#1595 #4). lite 분할 표준(#1652 #3-4): (1) mgr-sauron R017 구조 검증 단일 목표 위임 1건 + (2) 변경 성격별 적대적 리뷰 단일 목표 위임 1건 — 스크립트 변경이면 실행 재현 기반 adversarial-review, 룰/스킬/yaml 텍스트 변경이면 문구 정합·배선 리뷰. 근거: v1.1.59/60 두 반복 연속 적대적 리뷰가 신규 회귀(M-3/M-4, heredoc 위조)를 실행 재현으로 포착. 검증 위임 표준 문안: 오케스트레이터가 이미 실측한 항목(bun test·lint·typecheck·template-sync·wiki-sync·validate-docs·미러 md5)은 위임서에 재실행 금지 목록으로 열거하고 재실측 대상만 지정한다 — v1.1.61 세션에서 금지 목록 없는 sauron 위임이 bun test를 재실행하다 25턴 절단됐고, 금지 목록을 명시한 재위임은 16 tool_uses로 완주했다(R020 maxTurns 절단 누적 7건째)."
|
|
445
445
|
depends_on: verify-build
|
|
446
446
|
|
|
447
447
|
- name: release
|
|
@@ -585,6 +585,10 @@ steps:
|
|
|
585
585
|
- If failures: diagnose, fix, re-verify
|
|
586
586
|
3. For npm projects with auto-tag.yml: MANDATORY additional check:
|
|
587
587
|
gh run list --workflow auto-tag.yml --limit 1 --json conclusion,displayTitle
|
|
588
|
+
⚠ auto-tag.yml run의 `headSha`는 **release 브랜치 head(PR head SHA)**이지 머지 커밋이 아니다 — 폴링 종료
|
|
589
|
+
조건을 머지 커밋 SHA로 걸면 영원히 미일치해 타임아웃까지 전량 소진한다(v1.1.61 세션 실측: 40회/10분 전량
|
|
590
|
+
소진, v1.1.62에서 PR head로 정정 → 22회 정상 종료). release.yml run은 `displayTitle`에 PR 번호/브랜치명이
|
|
591
|
+
들어가므로 그것으로 식별한다.
|
|
588
592
|
→ conclusion MUST be "success" before declaring release complete.
|
|
589
593
|
If auto-tag.yml concluded "failure": diagnose root cause. Do NOT retry with manual git tag.
|
|
590
594
|
Common causes:
|