@ccoalm/ccl-skills 0.1.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +93 -16
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +249 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +394 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +50 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +59 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +16 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +391 -14
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +74 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +74 -33
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +11 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +421 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +576 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +176 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_lanes.sh +101 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +6 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +1 -1
- package/dist/assets/release.json +45 -20
- package/package.json +1 -1
|
@@ -0,0 +1,394 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
# Proves an aborted test_review_gate.sh leaves no reviewer wrapper behind, by both of
|
|
3
|
+
# the mechanisms that keep that true.
|
|
4
|
+
#
|
|
5
|
+
# The suite's reviewer wrappers are deliberately TERM-immune and the controller starts
|
|
6
|
+
# them with start_new_session=True, so they sit in their own session where no signal
|
|
7
|
+
# aimed at the suite's process group can reach them. That makes the controller the only
|
|
8
|
+
# party that reaps them on the happy path. Both legs below remove the controller FIRST —
|
|
9
|
+
# the way a tree-kill, a CI cancellation, or a host suspend does — and then abort the
|
|
10
|
+
# suite two different ways:
|
|
11
|
+
#
|
|
12
|
+
# leg 1 SIGTERM the suite: its cleanup trap runs and must reap the abandoned wrapper.
|
|
13
|
+
# leg 2 SIGKILL the suite: nothing can trap that, so no reaper runs at all and the
|
|
14
|
+
# wrapper must die on the fixture's own lifetime bound instead.
|
|
15
|
+
#
|
|
16
|
+
# Leg 2 exists because leg 1 alone stays green if the fixture's bound were restored to
|
|
17
|
+
# an unbounded loop: the trap would still reap it. It was an unbounded loop that turned
|
|
18
|
+
# this defect into a wrapper observed alive for 39 hours with ppid=1, ignoring SIGTERM.
|
|
19
|
+
#
|
|
20
|
+
# Each leg runs the suite under a private TMPDIR so every process this probe may signal
|
|
21
|
+
# is identified by a path no other run can share. Nothing here matches on the shared
|
|
22
|
+
# `review-gate-test.` prefix: a concurrent lane running the same suite would be killed
|
|
23
|
+
# by that, which is precisely the class scripts/test_lane_isolation.py keeps out of the
|
|
24
|
+
# parallel lanes.
|
|
25
|
+
set -uo pipefail
|
|
26
|
+
|
|
27
|
+
DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)"
|
|
28
|
+
SUITE="$DIR/test_review_gate.sh"
|
|
29
|
+
# How long an abandoned wrapper may still be alive after the suite is gone.
|
|
30
|
+
GRACE="${ABORT_LEAK_PROBE_GRACE:-30}"
|
|
31
|
+
# Bound on reaching the target case. The suite reaches it in ~2min on an idle host; the
|
|
32
|
+
# headroom is for a slow runner, NOT for sharing one — this probe runs in its own job
|
|
33
|
+
# precisely because a shared runner stretches that round past any budget.
|
|
34
|
+
REACH="${ABORT_LEAK_PROBE_REACH:-420}"
|
|
35
|
+
# Leg 2 shortens the fixture's own bound so the leg does not have to wait out the 300s
|
|
36
|
+
# default; leg 2's runtime is essentially this number, because the leg ends by watching
|
|
37
|
+
# the abandoned wrapper reach it. The constraint is the budget each case gives ITS
|
|
38
|
+
# WRAPPER — the `--timeout` the controller passes down, which is 5s for the claude hang
|
|
39
|
+
# case this probe aborts in and under 12s for the fallback hang cases that run before it
|
|
40
|
+
# (per those cases' own assertions) — NOT the case's `--total-timeout`, which bounds the
|
|
41
|
+
# whole gate rather than the wrapper. 25 keeps ~2x margin over the largest of those while
|
|
42
|
+
# taking 20s off the job: measured on CI, this job was the longest branch at 213s against
|
|
43
|
+
# 188s for the previous longest, so that 20s is the difference between adding a critical
|
|
44
|
+
# path and landing inside the existing one.
|
|
45
|
+
LEG2_HANG_BOUND="${ABORT_LEAK_PROBE_HANG_BOUND:-25}"
|
|
46
|
+
# Which leg(s) to run: `1`, `2`, or `all`. Each leg needs its OWN full run of the suite —
|
|
47
|
+
# one abort must be trappable and the other must not — and that pair, not the leg-2 wait,
|
|
48
|
+
# is what makes this the longest CI branch. Measured: cutting the leg-2 bound 45s -> 25s
|
|
49
|
+
# moved the job 213s -> 214s, i.e. not at all. So CI runs the legs as two parallel jobs
|
|
50
|
+
# and each stays well under the next-longest branch; `all` remains the local default.
|
|
51
|
+
ABORT_LEAK_PROBE_LEG="${ABORT_LEAK_PROBE_LEG:-all}"
|
|
52
|
+
case "$ABORT_LEAK_PROBE_LEG" in
|
|
53
|
+
1|2|all) : ;;
|
|
54
|
+
*) echo "ABORT_LEAK_PROBE_LEG must be 1, 2, or all" >&2; exit 2 ;;
|
|
55
|
+
esac
|
|
56
|
+
# Which fixture stub to orphan. The suite generates TWO of them — the claude stub and the
|
|
57
|
+
# shared candidate stub the fallback clients run — and they carry the same hang behavior
|
|
58
|
+
# and the same lifetime bound. A probe pinned to one leaves the other's bound revertible
|
|
59
|
+
# with the probe still green, so CI points its two jobs at different stubs.
|
|
60
|
+
ABORT_LEAK_PROBE_CLIENT="${ABORT_LEAK_PROBE_CLIENT:-claude}"
|
|
61
|
+
case "$ABORT_LEAK_PROBE_CLIENT" in
|
|
62
|
+
claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh ;;
|
|
63
|
+
fallback) PROBE_BEHAVIOR_FILE=kimi_behavior; PROBE_WRAPPER=kimi_review.sh ;;
|
|
64
|
+
*) echo "ABORT_LEAK_PROBE_CLIENT must be claude or fallback" >&2; exit 2 ;;
|
|
65
|
+
esac
|
|
66
|
+
|
|
67
|
+
fails=0
|
|
68
|
+
check() {
|
|
69
|
+
if eval "$2"; then echo "ok - $1"; else echo "FAIL - $1"; fails=$((fails+1)); fi
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
PROBE_TMP=""
|
|
73
|
+
PROBE_TMP_REAL=""
|
|
74
|
+
suite_pid=""
|
|
75
|
+
suite_pgid=""
|
|
76
|
+
suite_start=""
|
|
77
|
+
wrapper_pid=""
|
|
78
|
+
wrapper_pgid=""
|
|
79
|
+
|
|
80
|
+
# Every process a leg starts carries that leg's private tmp path in argv. The needles go
|
|
81
|
+
# through the environment, not `awk -v`: `ps -e` lists this very awk, and an argv-passed
|
|
82
|
+
# needle makes the scanner match itself.
|
|
83
|
+
probe_processes() {
|
|
84
|
+
[ -n "$PROBE_TMP_REAL" ] || return 0
|
|
85
|
+
ps -eo pid=,stat=,command= 2>/dev/null |
|
|
86
|
+
PROBE_A="$PROBE_TMP/" PROBE_B="$PROBE_TMP_REAL/" \
|
|
87
|
+
awk 'BEGIN { a = ENVIRON["PROBE_A"]; b = ENVIRON["PROBE_B"] }
|
|
88
|
+
(index($0, a) || index($0, b)) && $2 !~ /Z/ { print $1 }'
|
|
89
|
+
}
|
|
90
|
+
|
|
91
|
+
# Never let the probe become the leak it tests for — and never on someone else's
|
|
92
|
+
# processes: `suite_pgid` and any pid read earlier are cached numbers the OS recycles,
|
|
93
|
+
# so ownership is re-proved against the live process table at the moment of signalling.
|
|
94
|
+
# A group is signalled only while it still contains a process carrying this probe's
|
|
95
|
+
# private tmp path; everything else is signalled by pid, which `probe_processes` has
|
|
96
|
+
# just re-matched by that same path.
|
|
97
|
+
probe_group_is_ours() {
|
|
98
|
+
ps -eo pgid=,stat=,command= 2>/dev/null |
|
|
99
|
+
PROBE_A="$PROBE_TMP/" PROBE_B="$PROBE_TMP_REAL/" PROBE_PGID="$1" \
|
|
100
|
+
awk 'BEGIN { a = ENVIRON["PROBE_A"]; b = ENVIRON["PROBE_B"]; g = ENVIRON["PROBE_PGID"]; found = 1 }
|
|
101
|
+
$1 == g && $2 !~ /Z/ && (index($0, a) || index($0, b)) { found = 0; exit }
|
|
102
|
+
END { exit found }'
|
|
103
|
+
}
|
|
104
|
+
# The suite's pid and group id were captured minutes earlier and the OS recycles both,
|
|
105
|
+
# so every signal aimed at them re-proves identity first: the pid must still be running
|
|
106
|
+
# THIS suite script and must still lead the group we recorded. Validating the group
|
|
107
|
+
# through a live member we can name is what makes the group signal safe — a
|
|
108
|
+
# tmp-path-membership test would not do here, because by design the controller (the one
|
|
109
|
+
# group member carrying that path) has just been killed.
|
|
110
|
+
suite_is_ours() {
|
|
111
|
+
case "${suite_pid:-}" in ''|*[!0-9]*) return 1 ;; esac
|
|
112
|
+
case "${suite_pgid:-}" in ''|*[!0-9]*) return 1 ;; esac
|
|
113
|
+
kill -0 "$suite_pid" 2>/dev/null || return 1
|
|
114
|
+
case "$(ps -o command= -p "$suite_pid" 2>/dev/null || true)" in
|
|
115
|
+
*"$SUITE"*) : ;;
|
|
116
|
+
*) return 1 ;;
|
|
117
|
+
esac
|
|
118
|
+
[ "$(ps -o pgid= -p "$suite_pid" 2>/dev/null | tr -d ' ')" = "$suite_pgid" ] || return 1
|
|
119
|
+
# Start time, not just argv and group: with job control the suite's pgid equals its
|
|
120
|
+
# own pid, so those two conditions collapse into "this pid is running this script" —
|
|
121
|
+
# which a recycled pid on another run of the SAME suite would also satisfy. The launch
|
|
122
|
+
# instant is what separates our process from its namesake.
|
|
123
|
+
[ -n "$suite_start" ] || return 1
|
|
124
|
+
[ "$(ps -o lstart= -p "$suite_pid" 2>/dev/null || true)" = "$suite_start" ] || return 1
|
|
125
|
+
return 0
|
|
126
|
+
}
|
|
127
|
+
signal_suite() {
|
|
128
|
+
suite_is_ours || return 0
|
|
129
|
+
kill -"$1" -"$suite_pgid" 2>/dev/null || true
|
|
130
|
+
kill -"$1" "$suite_pid" 2>/dev/null || true
|
|
131
|
+
}
|
|
132
|
+
cleanup() {
|
|
133
|
+
local pid probe_cmd
|
|
134
|
+
signal_suite KILL
|
|
135
|
+
for pid in $(probe_processes); do
|
|
136
|
+
# Re-check immediately before signalling: a pid read even a moment ago can already be
|
|
137
|
+
# a stranger's. This narrows the window to the width of one `ps`; a shell cannot close
|
|
138
|
+
# it entirely, because there is no portable kill-by-handle, and that residual TOCTOU
|
|
139
|
+
# is a documented boundary rather than an oversight.
|
|
140
|
+
# Match BOTH spellings, like probe_processes does: on a host where the temp dir is a
|
|
141
|
+
# symlink (macOS /var -> /private/var) a process can carry either form, and a check
|
|
142
|
+
# that knows only one silently skips a process this probe started.
|
|
143
|
+
probe_cmd="$(ps -o command= -p "$pid" 2>/dev/null || true)"
|
|
144
|
+
case "$probe_cmd" in
|
|
145
|
+
*"$PROBE_TMP/"*|*"$PROBE_TMP_REAL/"*) : ;;
|
|
146
|
+
*) continue ;;
|
|
147
|
+
esac
|
|
148
|
+
probe_group_is_ours "$pid" && kill -KILL -"$pid" 2>/dev/null || true
|
|
149
|
+
kill -KILL "$pid" 2>/dev/null || true
|
|
150
|
+
done
|
|
151
|
+
[ -z "$PROBE_TMP" ] || rm -rf "$PROBE_TMP"
|
|
152
|
+
}
|
|
153
|
+
trap cleanup EXIT
|
|
154
|
+
trap 'exit 130' INT
|
|
155
|
+
trap 'exit 143' TERM HUP
|
|
156
|
+
|
|
157
|
+
# Selecting the target reads the suite's OWN state rather than pattern-matching an argv
|
|
158
|
+
# tail. Two earlier selectors failed here and each failure is the reason for this shape:
|
|
159
|
+
# picking by lifetime chose one of the behaviors that merely sleeps and then exits on its
|
|
160
|
+
# own (`quota_slow` 2s, `passed_slow` 5s), so the probe was green against a broken suite;
|
|
161
|
+
# and matching `--focus process-group-timeout` in the command line depends on how ps
|
|
162
|
+
# renders a long argv — it found the case when the probe ran alone and found nothing at
|
|
163
|
+
# all under `make`, burning the whole reach budget twice. The behavior file is written by
|
|
164
|
+
# the suite before it starts the case, so "a hang case is running now" is a fact to read,
|
|
165
|
+
# not a shape to guess. It also matches the FIRST hang case rather than the last, so the
|
|
166
|
+
# probe reaches its target sooner.
|
|
167
|
+
#
|
|
168
|
+
# The wrapper and the `( ... ) &` child it backgrounds share a command line, so "the
|
|
169
|
+
# first matching ps row" is whichever the kernel lists first. Picking the child would
|
|
170
|
+
# make the next step read the WRAPPER as the controller — and the wrapper's own command
|
|
171
|
+
# line satisfies the tmp-path validation, so that mistake would sail through and the leg
|
|
172
|
+
# would kill the wrong process. Select by ancestry: the wrapper is the match whose parent
|
|
173
|
+
# is the controller.
|
|
174
|
+
suite_work_dir() {
|
|
175
|
+
local dir
|
|
176
|
+
for dir in "$PROBE_TMP_REAL"/review-gate-test.*; do
|
|
177
|
+
[ -d "$dir" ] && printf '%s\n' "$dir" && return 0
|
|
178
|
+
done
|
|
179
|
+
return 1
|
|
180
|
+
}
|
|
181
|
+
live_wrapper() {
|
|
182
|
+
local work behavior pid parent_cmd
|
|
183
|
+
work="$(suite_work_dir)" || return 0
|
|
184
|
+
behavior="$(cat "$work/state/$PROBE_BEHAVIOR_FILE" 2>/dev/null || true)"
|
|
185
|
+
[ "$behavior" = "hang" ] || return 0
|
|
186
|
+
for pid in $(ps -eo pid=,command= 2>/dev/null |
|
|
187
|
+
grep -e "$PROBE_TMP/" -e "$PROBE_TMP_REAL/" -F |
|
|
188
|
+
grep -F "$PROBE_WRAPPER" |
|
|
189
|
+
grep -v '[g]rep' |
|
|
190
|
+
awk '{print $1}'); do
|
|
191
|
+
parent_cmd="$(ps -o command= -p "$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ')" 2>/dev/null || true)"
|
|
192
|
+
case "$parent_cmd" in
|
|
193
|
+
*review_gate.py*) printf '%s\n' "$pid"; return 0 ;;
|
|
194
|
+
esac
|
|
195
|
+
done
|
|
196
|
+
return 0
|
|
197
|
+
}
|
|
198
|
+
|
|
199
|
+
# Runs the suite until the target wrapper is in flight, then removes its controller so
|
|
200
|
+
# the wrapper is orphaned with no reaper but the suite itself. Sets wrapper_pid.
|
|
201
|
+
# $1 is the value for the fixture's lifetime bound, or empty for its default.
|
|
202
|
+
orphan_target_wrapper() {
|
|
203
|
+
local hang_bound="$1" candidate controller_pid controller_ok deadline suite_died
|
|
204
|
+
PROBE_TMP="$(mktemp -d "${TMPDIR:-/tmp}/review-gate-abort-probe.XXXXXX")" || return 1
|
|
205
|
+
PROBE_TMP_REAL="$(cd "$PROBE_TMP" && pwd -P)" || return 1
|
|
206
|
+
wrapper_pid=""
|
|
207
|
+
|
|
208
|
+
# Job control so the suite gets its own process group and the probe can signal that
|
|
209
|
+
# group without signalling itself.
|
|
210
|
+
set -m
|
|
211
|
+
TMPDIR="$PROBE_TMP" REVIEW_GATE_TEST_HANG_SECONDS="$hang_bound" \
|
|
212
|
+
bash "$SUITE" >/dev/null 2>&1 &
|
|
213
|
+
suite_pid=$!
|
|
214
|
+
suite_pgid="$(ps -o pgid= -p "$suite_pid" 2>/dev/null | tr -d ' ')"
|
|
215
|
+
suite_start="$(ps -o lstart= -p "$suite_pid" 2>/dev/null || true)"
|
|
216
|
+
|
|
217
|
+
deadline=$(( $(date +%s) + REACH ))
|
|
218
|
+
suite_died=0
|
|
219
|
+
while [ "$(date +%s)" -lt "$deadline" ]; do
|
|
220
|
+
candidate="$(live_wrapper)"
|
|
221
|
+
if [ -n "$candidate" ]; then wrapper_pid="$candidate"; break; fi
|
|
222
|
+
kill -0 "$suite_pid" 2>/dev/null || { suite_died=1; break; }
|
|
223
|
+
sleep 0.1
|
|
224
|
+
done
|
|
225
|
+
if [ -z "$wrapper_pid" ]; then
|
|
226
|
+
# Say WHICH failure this is. "unreached" alone cannot distinguish a suite that ran
|
|
227
|
+
# fine but never got to a hang case inside the budget from a suite that died early —
|
|
228
|
+
# and a run of this probe under `make` failed exactly here with no way to tell them
|
|
229
|
+
# apart, which is why the distinction is printed rather than inferred.
|
|
230
|
+
if [ "$suite_died" = 1 ]; then
|
|
231
|
+
wait "$suite_pid" 2>/dev/null
|
|
232
|
+
printf 'abort_leak_probe_unreached: the suite exited (status %s) before any hang case was observed\n' "$?" >&2
|
|
233
|
+
else
|
|
234
|
+
printf 'abort_leak_probe_unreached: %ss budget elapsed with the suite still running; last behavior=%s\n' \
|
|
235
|
+
"$REACH" "$(cat "$(suite_work_dir 2>/dev/null)/state/claude_behavior" 2>/dev/null || printf unknown)" >&2
|
|
236
|
+
fi
|
|
237
|
+
return 1
|
|
238
|
+
fi
|
|
239
|
+
|
|
240
|
+
# Validate the target before signalling it, never after. Between selecting the wrapper
|
|
241
|
+
# and reading its parent the wrapper can be reaped and its pid reused, so this value is
|
|
242
|
+
# untrustworthy by construction: it can be 1 (already orphaned) or an unrelated
|
|
243
|
+
# process. `kill -KILL 1` is harmless for an unprivileged user but not for a root
|
|
244
|
+
# container, and no probe should be the thing that finds that out.
|
|
245
|
+
controller_pid="$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')"
|
|
246
|
+
controller_ok=0
|
|
247
|
+
case "$controller_pid" in
|
|
248
|
+
''|*[!0-9]*) : ;;
|
|
249
|
+
*)
|
|
250
|
+
if [ "$controller_pid" -gt 1 ] \
|
|
251
|
+
&& [ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "$controller_pid" ] \
|
|
252
|
+
&& { case "$(ps -o command= -p "$controller_pid" 2>/dev/null || true)" in
|
|
253
|
+
*"$PROBE_TMP/"*|*"$PROBE_TMP_REAL/"*) true ;; *) false ;; esac; }; then
|
|
254
|
+
controller_ok=1
|
|
255
|
+
fi
|
|
256
|
+
;;
|
|
257
|
+
esac
|
|
258
|
+
if [ "$controller_ok" != 1 ]; then
|
|
259
|
+
printf 'abort leak setup: refusing to signal pid=%s (not this run'"'"'s controller)\n' \
|
|
260
|
+
"${controller_pid:-none}" >&2
|
|
261
|
+
return 1
|
|
262
|
+
fi
|
|
263
|
+
# Order matters: the controller must die BEFORE the suite, or its own timeout path
|
|
264
|
+
# reaps the wrapper and the leg proves nothing about the suite. `kill` returning is
|
|
265
|
+
# not that proof: confirm here, once, that the controller is really gone and the
|
|
266
|
+
# wrapper really was reparented, so BOTH legs inherit a verified precondition instead
|
|
267
|
+
# of each re-deriving it (leg 2 had no such check, and its later kill of the suite
|
|
268
|
+
# group could have removed the controller itself — making every remaining assertion
|
|
269
|
+
# pass without the required ordering ever holding).
|
|
270
|
+
wrapper_pgid="$(ps -o pgid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')"
|
|
271
|
+
kill -KILL "$controller_pid" 2>/dev/null || true
|
|
272
|
+
deadline=$(( $(date +%s) + 10 ))
|
|
273
|
+
while :; do
|
|
274
|
+
if ! kill -0 "$controller_pid" 2>/dev/null ||
|
|
275
|
+
[ -n "$(ps -o stat= -p "$controller_pid" 2>/dev/null | grep Z || true)" ]; then
|
|
276
|
+
break
|
|
277
|
+
fi
|
|
278
|
+
if [ "$(date +%s)" -ge "$deadline" ]; then
|
|
279
|
+
echo "abort leak setup: controller $controller_pid survived SIGKILL" >&2
|
|
280
|
+
return 1
|
|
281
|
+
fi
|
|
282
|
+
sleep 0.5
|
|
283
|
+
done
|
|
284
|
+
while :; do
|
|
285
|
+
[ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "1" ] && return 0
|
|
286
|
+
kill -0 "$wrapper_pid" 2>/dev/null || {
|
|
287
|
+
echo "abort leak setup: wrapper $wrapper_pid died with its controller" >&2
|
|
288
|
+
return 1
|
|
289
|
+
}
|
|
290
|
+
if [ "$(date +%s)" -ge "$deadline" ]; then
|
|
291
|
+
echo "abort leak setup: wrapper $wrapper_pid was never reparented" >&2
|
|
292
|
+
return 1
|
|
293
|
+
fi
|
|
294
|
+
sleep 0.5
|
|
295
|
+
done
|
|
296
|
+
}
|
|
297
|
+
|
|
298
|
+
await_suite_exit() {
|
|
299
|
+
local deadline
|
|
300
|
+
deadline=$(( $(date +%s) + GRACE ))
|
|
301
|
+
while kill -0 "$suite_pid" 2>/dev/null; do
|
|
302
|
+
[ "$(date +%s)" -ge "$deadline" ] && break
|
|
303
|
+
sleep 0.5
|
|
304
|
+
done
|
|
305
|
+
kill -0 "$suite_pid" 2>/dev/null && return 1
|
|
306
|
+
return 0
|
|
307
|
+
}
|
|
308
|
+
|
|
309
|
+
# The verdict is about the GROUP, not the leader. The wrapper backgrounds a TERM-immune
|
|
310
|
+
# child in that same group; a cleanup that kills the session leader but misses the child
|
|
311
|
+
# leaves the leak in place while "the wrapper is gone" reads true — and the probe's own
|
|
312
|
+
# cleanup would then erase the very evidence the leg exists to surface. Residue under the
|
|
313
|
+
# leg's private tmp path is checked too, so a survivor that left the group still counts.
|
|
314
|
+
wrapper_group_members_alive() {
|
|
315
|
+
[ -n "$wrapper_pgid" ] || return 0
|
|
316
|
+
ps -eo pid=,pgid=,stat= 2>/dev/null |
|
|
317
|
+
awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ { print $1 }'
|
|
318
|
+
}
|
|
319
|
+
wrapper_gone_within() {
|
|
320
|
+
local deadline members residue
|
|
321
|
+
deadline=$(( $(date +%s) + $1 ))
|
|
322
|
+
while :; do
|
|
323
|
+
members="$(wrapper_group_members_alive | tr '\n' ' ')"
|
|
324
|
+
residue="$(probe_processes | tr '\n' ' ')"
|
|
325
|
+
case "$members$residue" in
|
|
326
|
+
*[0-9]*) : ;;
|
|
327
|
+
*) return 0 ;;
|
|
328
|
+
esac
|
|
329
|
+
[ "$(date +%s)" -ge "$deadline" ] && break
|
|
330
|
+
sleep 1
|
|
331
|
+
done
|
|
332
|
+
printf 'abort leak: group members alive: %s | private-path residue: %s\n' \
|
|
333
|
+
"${members:-none}" "${residue:-none}" >&2
|
|
334
|
+
return 1
|
|
335
|
+
}
|
|
336
|
+
|
|
337
|
+
diagnose_wrapper() {
|
|
338
|
+
printf 'abort leak diagnostic (%s): wrapper=%s %s\n' "$1" "$wrapper_pid" \
|
|
339
|
+
"$(ps -o pid=,ppid=,etime=,stat= -p "$wrapper_pid" 2>/dev/null | tr -s ' ')" >&2
|
|
340
|
+
}
|
|
341
|
+
|
|
342
|
+
# ---- leg 1: a trappable abort — the suite's own cleanup must reap the wrapper --------
|
|
343
|
+
if [ "$ABORT_LEAK_PROBE_LEG" = 1 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
|
|
344
|
+
leg1_orphaned=0
|
|
345
|
+
orphan_target_wrapper "" && leg1_orphaned=1
|
|
346
|
+
check "leg1: the controller was removed and the wrapper reparented to init" '[ "$leg1_orphaned" = 1 ]'
|
|
347
|
+
if [ "$leg1_orphaned" = 1 ]; then
|
|
348
|
+
signal_suite TERM
|
|
349
|
+
leg1_suite_gone=0; await_suite_exit && leg1_suite_gone=1
|
|
350
|
+
check "leg1: the aborted suite exits" '[ "$leg1_suite_gone" = 1 ]'
|
|
351
|
+
leg1_reaped=0; wrapper_gone_within "$GRACE" && leg1_reaped=1
|
|
352
|
+
[ "$leg1_reaped" = 1 ] || diagnose_wrapper leg1
|
|
353
|
+
check "leg1: a terminated suite reaps the reviewer wrapper it started" '[ "$leg1_reaped" = 1 ]'
|
|
354
|
+
fi
|
|
355
|
+
cleanup; PROBE_TMP=""; PROBE_TMP_REAL=""; suite_pgid=""
|
|
356
|
+
fi
|
|
357
|
+
|
|
358
|
+
# ---- leg 2: an untrappable abort — only the fixture's own bound can end the wrapper --
|
|
359
|
+
if [ "$ABORT_LEAK_PROBE_LEG" = 2 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
|
|
360
|
+
leg2_orphaned=0
|
|
361
|
+
orphan_target_wrapper "$LEG2_HANG_BOUND" && leg2_orphaned=1
|
|
362
|
+
check "leg2: the controller was removed and the wrapper reparented to init" '[ "$leg2_orphaned" = 1 ]'
|
|
363
|
+
if [ "$leg2_orphaned" = 1 ]; then
|
|
364
|
+
signal_suite KILL
|
|
365
|
+
leg2_suite_gone=0; await_suite_exit && leg2_suite_gone=1
|
|
366
|
+
check "leg2: the killed suite exits without running any cleanup" '[ "$leg2_suite_gone" = 1 ]'
|
|
367
|
+
# Precondition: with no trap able to run, the wrapper must still be here. If it is
|
|
368
|
+
# already gone, something else reaped it and the bound was never exercised.
|
|
369
|
+
# `kill -0` succeeds for a zombie, and the verdict below excludes zombies — so a wrapper
|
|
370
|
+
# that had already died would satisfy "still alive" here and "gone" there, passing the
|
|
371
|
+
# leg without the bound ever being exercised. Require a live, non-zombie process.
|
|
372
|
+
leg2_survived_abort=0
|
|
373
|
+
if kill -0 "$wrapper_pid" 2>/dev/null; then
|
|
374
|
+
case "$(ps -o stat= -p "$wrapper_pid" 2>/dev/null || true)" in
|
|
375
|
+
''|*Z*) : ;;
|
|
376
|
+
*) leg2_survived_abort=1 ;;
|
|
377
|
+
esac
|
|
378
|
+
fi
|
|
379
|
+
check "leg2: no reaper ran, so the wrapper is still alive right after the kill" \
|
|
380
|
+
'[ "$leg2_survived_abort" = 1 ]'
|
|
381
|
+
leg2_self_exit=0; wrapper_gone_within "$(( LEG2_HANG_BOUND + GRACE ))" && leg2_self_exit=1
|
|
382
|
+
[ "$leg2_self_exit" = 1 ] || diagnose_wrapper leg2
|
|
383
|
+
check "leg2: an unreaped wrapper still exits on the fixture's own lifetime bound" \
|
|
384
|
+
'[ "$leg2_self_exit" = 1 ]'
|
|
385
|
+
fi
|
|
386
|
+
|
|
387
|
+
fi
|
|
388
|
+
|
|
389
|
+
if [ "$fails" -eq 0 ]; then
|
|
390
|
+
echo "review_gate_abort_leak_ok leg=$ABORT_LEAK_PROBE_LEG client=$ABORT_LEAK_PROBE_CLIENT"
|
|
391
|
+
exit 0
|
|
392
|
+
fi
|
|
393
|
+
echo "review_gate_abort_leak_failed=$fails" >&2
|
|
394
|
+
exit 1
|
|
@@ -69,8 +69,15 @@ The fitness function may be a scalar metric, unit tests / validators / schema ch
|
|
|
69
69
|
|
|
70
70
|
Every material model or prompt change should define:
|
|
71
71
|
|
|
72
|
+
- the decision the eval supports, named before choosing datasets or scorers — this reference's working classes: version/model selection (benchmarking), release/regression gate, capability/unit check, and production monitoring and failure diagnosis (informed by the public offline-vs-online eval-type split per `docs.langchain.com/langsmith/evaluation-types` and success-criteria-first design per `docs.anthropic.com/en/docs/test-and-evaluate/define-success`). Conclusions do not transfer automatically across decision classes — each intended decision must satisfy its own predeclared criteria: a selection benchmark's average is not by itself a release verdict (a release verdict needs regression comparison against the incumbent baseline plus red-line blockers — see Eval Reliability and Freeze Comparison Criteria below), and a monitoring signal is not a version verdict; one combined, predeclared design may serve several decisions when every decision's criteria are declared up front and each is satisfied on its own terms (per the combined-design rule under Freeze Comparison Criteria);
|
|
72
73
|
- dataset or replay source;
|
|
73
74
|
- expected output, rubric, or comparator;
|
|
75
|
+
- run protocol: a declared repetition and uncertainty protocol matched to the decision and to the observed nondeterminism —
|
|
76
|
+
- one run per task only when the evaluated path is deterministic end to end, generation included (a deterministic code grader over stochastic model output does not qualify), or with recorded justification that output variance cannot move the decision;
|
|
77
|
+
- where variance can move an offline or replay-based comparison, resample the same task several times and use question-level averages for aggregate quality metrics (question-level-averaging technique per `arxiv.org/abs/2411.00640`, stated there for chain-of-thought evals; applying it to other variance-sensitive comparisons is this reference's extension), while red-line/blocker checks aggregate conservatively (any hit fails, never averaged away);
|
|
78
|
+
- online production-monitoring decisions observe live traffic through predeclared observation windows with minimum volume, persistence, and hysteresis (per the operational-gate rule under Freeze Comparison Criteria and `inference-capacity-operations.md`);
|
|
79
|
+
- failure-diagnosis decisions reproduce the failure under control: sanitized and isolated replay of recorded inputs, with side effects suppressed or stubbed wherever mutation is possible (reproduction discipline per `defect-diagnosis`); live side-effecting requests are never re-executed;
|
|
80
|
+
- for compared runs, hold the conditions fixed — everything except the declared variable under test, with both values of that variable recorded (model string, prompt version, retrieval index, tool set, and decoding parameters are the usual fixed set);
|
|
74
81
|
- quality metrics such as accuracy, recall, consistency, parse success, or groundedness;
|
|
75
82
|
- runtime metrics such as latency, token usage, success rate, fallback rate, and cost;
|
|
76
83
|
- regression examples and human review notes when judgment is subjective.
|
|
@@ -10,7 +10,8 @@ These are external quality benchmarks, not visual style sources. Do not copy ano
|
|
|
10
10
|
- W3C WCAG 2.2: accessibility criteria for keyboard, focus, contrast, labels, target size, error identification, and consistent navigation.
|
|
11
11
|
- web.dev Core Web Vitals: Largest Contentful Paint, Cumulative Layout Shift, Interaction to Next Paint, and field/lab measurement.
|
|
12
12
|
- Google HEART framework: product-experience metrics across happiness, engagement, adoption, retention, and task success.
|
|
13
|
-
-
|
|
13
|
+
- Apple Human Interface Guidelines (developer.apple.com/design/human-interface-guidelines) and Material Design 3 (m3.material.io): first-party platform specifications for feedback, loading, interaction states, accessibility, and platform conventions — distilled into the Platform Convention Walkthrough section below. Each criterion there is labeled (HIG) or (Material); a single-source criterion is that platform's convention, not a cross-platform standard.
|
|
14
|
+
- Other design-system guidance such as Atlassian Design System: secondary confirmation for touch targets, empty/error states, progressive disclosure, feedback, and content clarity.
|
|
14
15
|
|
|
15
16
|
## Heuristic Review Layer
|
|
16
17
|
|
|
@@ -38,6 +39,54 @@ Before launch or review, verify:
|
|
|
38
39
|
- **Reduced-motion / color-scheme / contrast / transparency design intent must be declared, not auto-derived from `@media` query alone**. For each non-decorative animation, name whether it is *essential* (progress indicator, drag preview, view transition that conveys a state change) or *decorative* (parallax, autoplay carousel, hover bounce); `prefers-reduced-motion: reduce` should remove or replace decorative motion by default and may shorten essential motion but cannot omit feedback. An explicit in-product opt-in (e.g. user setting "Show celebration animation even when system asks for reduced motion") MAY override the default for brand-splash / completion-celebration moments when policy permits, but the override must be opt-in not opt-out and the default behavior must respect the system preference. `prefers-color-scheme` requires the design system to ship Light + Dark token pairs (no missing pair = no dark-mode claim). `prefers-contrast: more` and `prefers-reduced-transparency` are useful supplements where browser support permits but are NOT Baseline yet — do not gate accessibility compliance on them; route them through a separate "increased contrast" theme variant when product needs require it.
|
|
39
40
|
- **APCA (Accessible Perceptual Contrast Algorithm) is the WCAG 3 / Silver candidate contrast method, not a WCAG 2.2 replacement**. WCAG 2.2 SC 1.4.3 / 1.4.11 (4.5:1 body / 3:1 large text + UI components) remains the legal/audit baseline. APCA can be used as a *supplementary* perceptual check (per APCA Bronze Simple Mode, Lc 75 minimum / Lc 90 preferred for body text) when WCAG 2.x mathematical contrast passes but the result looks washed-out, or when designing dark mode where WCAG 2.x ratios systematically over-permit low-readability combinations. Decision matrix: WCAG 2.x fail blocks accessibility/legal compliance claims regardless of APCA result; WCAG 2.x pass + APCA fail is NOT a WCAG failure but should be treated as a readability / product-quality defect — either adjust the color tokens to also pass APCA Bronze, or document the acceptance with a rationale (brand constraint, dark-mode literal preserved). Do not ship a design that passes only APCA but fails WCAG 2.2.
|
|
40
41
|
|
|
42
|
+
## Platform Convention Walkthrough (HIG / Material)
|
|
43
|
+
|
|
44
|
+
Use this section when reviewing or accepting a mobile-platform surface (iOS/iPadOS or Android/Material-based, including Flutter/React Native apps that adopt a platform design language). Each criterion is a pass/fail walkthrough check distilled from the first-party spec named in its label; verify against the rendered surface, not the design file alone. These complement — never replace — the WCAG 2.2 acceptance line above: where a platform minimum is stricter than WCAG (e.g. target size), the platform minimum is the walkthrough bar for that platform. Criteria mirror each source's own normative strength: where the source states a recommendation ("ideally", "consider", "in general", "in most cases"), a recorded, justified exception passes as an exception — only a silent shortfall fails; where the source states a requirement, the failure is unconditional.
|
|
45
|
+
|
|
46
|
+
### State completeness against the platform spec
|
|
47
|
+
|
|
48
|
+
- **Interaction-state matrix is complete for the component class and input modalities (Material).** Every interactive component accounts for each state its Material component class inherits — enabled, plus disabled/hover/focused/pressed/dragged only where the class takes them (action/selection/input components inherit most; app bars, dialogs, menus, navigation components inherit few) — across the input modalities the surface ships on (hover needs a pointer; focused needs a focus-capable input such as keyboard or voice). States the class does not inherit are marked inapplicable, never styled in. Fail: an action component on a keyboard-capable surface styles only enabled and pressed.
|
|
49
|
+
- **Disabled semantics are real, not painted (Material).** A disabled component cannot be focused, dragged, or pressed, and does not change state when tapped or hovered — unrelated explanatory feedback (a tooltip saying why it is disabled) is not prohibited. Components whose class does not take a disabled state in Material — app bars, badges, dialogs, FABs, menus, navigation bar/drawer/rail, sheets, tabs, tooltips — never render a "disabled" look: when a FAB's action is unavailable, remove the FAB rather than disabling it. Fail: a grayed-out FAB or tab sits on screen, or a "disabled" card still accepts a drag.
|
|
50
|
+
- **Each state change is signaled by more than one visual cue (Material).** Material's baseline is two visual indicators per state so state remains perceivable under color-vision or contrast loss; opacity-only or color-only state styling fails.
|
|
51
|
+
- **Transient input states are singletons (Material).** At most one hover, one focus, one pressed, and one dragged state visible at a time in a layout; persistent states (selected, activated) may combine with them on the same element (e.g. a selected chip showing hover).
|
|
52
|
+
- **Feedback reaches people through more than one channel (HIG).** Significant feedback pairs color with text/icon, and sound with haptic where sound is used, so it survives a silenced device, a glance away, or a screen reader. Fail: success/failure conveyed by hue change alone or by sound alone.
|
|
53
|
+
- **Interruption level matches significance (HIG).** Passive status renders in-context near the item it describes (badge, inline line); modal alerts are reserved for critical, ideally actionable information. Fail: routine status delivered as a modal, or a data-loss warning delivered as a passive toast.
|
|
54
|
+
- **Data-loss warnings fire on the unexpected-and-irreversible boundary, both directions (HIG).** Warn before an action whose data loss is unexpected and irreversible; do NOT interpose confirmation when loss is the expected result of the user's own action (e.g. moving a file to trash). Fail in either direction: silent irreversible loss, or confirmation nagging on expected outcomes.
|
|
55
|
+
- **Completion feedback is reserved for significant outcomes; failure feedback is never omitted (HIG).** People expect success, so confirm only payment-grade/significant completions — but every command that cannot be carried out must say so and say why, with the next step. Fail: a no-op button press with no explanation.
|
|
56
|
+
- **Content loading shows something immediately and frees the user (HIG).** HIG scopes this to content/asset loading: placeholder/skeleton content appears at once instead of a blank wait, loading continues in the background so unrelated safe actions stay available, a determinate indicator is used when duration is known and indeterminate only when it is not, and an unavoidably long load gets meaningful interim content. This does not apply to in-flight mutations (payment, deletion, submission): while one is pending, its duplicate or conflicting mutation controls are blocked per the high-risk resilience states in `SKILL.md`, not left available.
|
|
57
|
+
|
|
58
|
+
### Accessibility against the platform spec
|
|
59
|
+
|
|
60
|
+
- **Text scales to 200% without breaking the layout (HIG, stated as "ideally").** Support the platform text-size setting (Dynamic Type on Apple platforms) up to 200% enlargement — HIG's recommended target, so a smaller ceiling passes only as a recorded exception; the walkthrough re-renders key screens at enlarged sizes and checks truncation, overlap, and control reachability. Adoption mechanics (Dynamic Type APIs, per-platform text-scaling behavior) belong to the stack implementation owner, not this walkthrough.
|
|
61
|
+
- **iOS hit regions measure at least 44×44pt (HIG, stated as "a general rule").** Every control's hit region is at least 44×44pt (60×60pt on visionOS), and the padded hit area is what must measure up, not the visual glyph; a smaller region passes only as a recorded exception. The ~12pt padding around bezeled elements and ~24pt around bezel-less ones is HIG's "generally works well" guidance, checked the same way. Stricter than WCAG 2.2's 24px floor — the platform benchmark is the walkthrough bar.
|
|
62
|
+
- **Android touch targets measure at least 48×48dp with 8dp spacing (Material, stated as "consider" / "in most cases").** Touch targets at least 48×48dp with at least 8dp between targets, pointer targets at least 44×44dp, the padded target extending beyond the visual bounds; a shortfall passes only as a recorded exception. Stricter than WCAG 2.2's 24px floor — the platform benchmark is the walkthrough bar.
|
|
63
|
+
- **Contrast meets the W3C-derived platform bar (Material, citing W3C).** Small text at least 4.5:1 against background; large text (14pt bold / 18pt regular and up) and meaningful graphics at least 3:1. Clustered non-text containers (e.g. a button group) need 3:1 container-vs-background. The standalone-prominence exemption (a FAB) applies only to that container-vs-background ratio — the element's own text, icons, focus indicators, and meaningful graphics still need their 3:1/4.5:1 bars; disabled states are the only class exempt from contrast requirements.
|
|
64
|
+
- **Reading and focus order follows content hierarchy (Material).** Screen-reader order follows the top-down source/DOM structure, headings do not skip levels, one H1 per web page, repeated landmarks get unique labels. Fail: visual order diverges from traversal order with no remediation.
|
|
65
|
+
- **Focus is managed across context changes (Material).** Initial focus is defined per screen; opening a dialog moves focus into it; closing returns focus to the element that opened it; a visible focus ring appears on keyboard traversal. Fail: focus lost to page top after a dialog closes.
|
|
66
|
+
- **Labels describe purpose, not appearance, and omit the role (Material).** Icon-only controls, meaningful images, and progress/error cues carry labels naming the action or meaning ("Voice search", not "Microphone"); decorative images are hidden from assistive tech; the role word ("button") never appears inside the label.
|
|
67
|
+
- **Core functionality is never gesture-only (HIG).** Any action in the UI's core functionality or supported task flows that a gesture performs (swipe-to-dismiss, swipe-row actions, custom gestures) is also reachable through a visible onscreen control; an optional convenience gesture duplicating an already-visible control needs no second alternative. Frequent actions use the simplest gesture available, no custom multi-finger requirements.
|
|
68
|
+
- **Keyboard access is complete and system shortcuts stay untouched (HIG).** Core flows complete with the keyboard alone (Full Keyboard Access on Apple platforms), and system-defined keyboard shortcuts are not overridden.
|
|
69
|
+
- **Custom shortcuts default to two-key combinations and are discoverable (Material).** Custom keyboard shortcuts use two or more keys by default — a single-key shortcut needs a remap option, component-focus scoping, or an off switch — and a help surface lists them.
|
|
70
|
+
- **Timed UI does not self-dismiss content people must act on (HIG).** Views and controls that auto-dismiss on a timer are minimized; anything carrying a decision or unfinished reading dismisses by explicit action. Fail: an error toast that disappears before its recovery action can be reached.
|
|
71
|
+
- **Reduce Motion is honored with concrete substitutions (HIG).** When the OS reduce-motion setting is on, decorative/repetitive animation stops by default — the Accessibility Baseline's narrowly scoped, explicitly opt-in in-product override for brand-splash/celebration moments remains valid and is the only exception. For animations that use these effects: springs tighten (no bounce), x/y/z transitions become fades, z-depth and blur animations are avoided, and gesture-driven animation tracks the gesture. Essential status motion (progress) remains. Declare essential-vs-decorative intent per the reduced-motion rule in the Accessibility Baseline above; native OS setting detection routes to `platform-mobile-patterns.md` Mobile Motion Discipline.
|
|
72
|
+
|
|
73
|
+
### Platform conventions walkthrough
|
|
74
|
+
|
|
75
|
+
- **One-handed reachability shapes iPhone layouts (HIG).** Primary and frequent controls live in the middle or bottom of the screen; back-swipe from the edge and list-row swipe actions are preserved, not hijacked by custom edge gestures (Android's equivalent predictive-back geometry routes to `platform-mobile-patterns.md`).
|
|
76
|
+
- **The surface adapts to user-chosen appearance settings (HIG).** A screen passes walkthrough only after being checked under orientation change, Dark Mode, and enlarged Dynamic Type — the platform treats these as user choices the app must follow, not edge cases.
|
|
77
|
+
- **Standard platform components are the default for standard tasks (Material).** Standard platform controls and semantic elements inherit assistive-technology support for free; a custom replacement for a standard task (e.g. a non-standard dialog) carries the burden of extra AT verification before it passes walkthrough.
|
|
78
|
+
- **Contrast and appearance adaptation is verified on the rendered surface (HIG).** Beyond the orientation/Dark Mode/Dynamic Type checks above, the walkthrough gate is observable: the surface renders correctly with the Increase Contrast setting on — text, icons, and state indicators keep sufficient contrast and their meaning. Preferring system-defined colors (whose accessible variants adapt automatically) and familiar system behaviors is non-blocking implementation guidance verified in code review, because two token implementations can render identically.
|
|
79
|
+
|
|
80
|
+
### Deliberately not absorbed from the platform specs
|
|
81
|
+
|
|
82
|
+
Recorded so future rounds do not re-import them; each was read and rejected for walkthrough use:
|
|
83
|
+
|
|
84
|
+
- HIG media-accessibility taxonomy (captions vs subtitles vs audio descriptions vs transcripts) — media-content-type guidance, not a screen-walkthrough criterion; consult the HIG Hearing section directly when shipping media surfaces.
|
|
85
|
+
- HIG platform-capability integrations (Siri/Shortcuts, Switch Control, Voice Control setup, Assistive Access optimization) and watchOS/visionOS-specific rules — implementation- or platform-mode-specific; route to the stack implementation skill if those surfaces enter scope. One exception is retained: the visionOS 60×60pt hit-region figure stays inside the iOS hit-region criterion as informational context only — this walkthrough's scope remains iOS/Android mobile surfaces and does not govern visionOS.
|
|
86
|
+
- Material state-layer token mechanics (fixed opacity percentages, on-color derivation) — design-kit implementation detail owned by token/component references, not an acceptance criterion.
|
|
87
|
+
- Material web-landmark role enumeration (the eight ARIA roles) — imported only as the "landmarks get unique labels" criterion; the full role catalog is reference material, not a checklist.
|
|
88
|
+
- Visual-style content from either spec (Liquid Glass materials, M3 Expressive shapes/motion values) — style adoption is a product decision covered by `platform-mobile-patterns.md` Platform OS Updates; copying platform visual language is already ruled out by "What Not To Absorb" below.
|
|
89
|
+
|
|
41
90
|
## Performance And Perceived Speed
|
|
42
91
|
|
|
43
92
|
Use these checks for AI, feed, media, upload, and document-heavy surfaces:
|
|
@@ -45,6 +45,7 @@ Always include concrete file/line references for code reviews and Figma file/pag
|
|
|
45
45
|
7. **Check accessibility basics**: readable contrast, keyboard/focus where relevant, touch target size, visible labels, reduced-motion risk, safe-area/keyboard behavior on mobile.
|
|
46
46
|
8. **Check visual craft**: anti-slop, product-level identity, spacing rhythm, typography scale, consistent iconography, appropriate density.
|
|
47
47
|
9. **Check serious-domain adaptation** when relevant: source, timestamp, partial data, confirmation, audit labels, and no unsafe optimistic UI.
|
|
48
|
+
10. **Check platform-convention conformance** for iOS/Android app surfaces: run the pass/fail criteria in `external-ui-ux-quality-benchmarks.md` Platform Convention Walkthrough (HIG/Material state completeness, platform accessibility minima, and platform conventions) against the rendered surface.
|
|
48
49
|
|
|
49
50
|
## UI Checks
|
|
50
51
|
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md
CHANGED
|
@@ -160,7 +160,7 @@ Use this skill to turn observed experience into durable agent skills without cop
|
|
|
160
160
|
- The trigger is a correction about *reusable skill/process behavior*, NOT every bug/QA/review nit handled inside its own owner skill. Do not wait for the user to say "沉淀": if the user points out a missed source, missed sibling skill, shallow rule, overclaim, domain leakage, missing trigger, missing verification, repeated correction, **or that this workflow should have been invoked at all (an under-trigger / "should you have used 提炼/复盘" correction, including outside an active extraction)**, run correction RCA, update the smallest owning skill/reference/validator, and verify the prevention point before finalizing the turn.
|
|
161
161
|
- When the user asks whether a lesson was durably landed after a failed extraction, verify the actual skill diff or file content first. Do not answer from memory or intent. If the prevention rule is not present in the owning skill, add it or state that it has not been durably landed.
|
|
162
162
|
- **Consolidate and retire rules; a skill's rule set must not grow monotonically.** Every correction adds a guard, but an N-bullet wall on one theme is itself the over-prescription/unreadability failure, and "just append another bullet" is how it regrows.
|
|
163
|
-
- **
|
|
163
|
+
- **Prose rules compete for a finite attention budget, and joint satisfaction degrades with the number of constraints.** (Predecessor claim — only the most-salient rule applies, the rest dormant — **withdrawn**, unsupported.) **Descriptive, not permissive**: attention limits never excuse a violated rule, and are not a reason to refuse a needed one. (a) **Merge-into-canonical beats append**: appending adds contradiction surface and spends budget. (b) Do not rely on co-resident prose for requirements that must hold JOINTLY — structure them as a **walked enumeration at their firing point**: walking a finite list works where holding a conjunction does not. It is why "the rule was loaded" never predicts "the rule was applied". Detail: `references/external-practice-controls.md#instruction-following-mechanisms`.
|
|
164
164
|
- A rule's TEXT in a `SKILL.md`/reference is **living** — merged, tightened, or retired in place — unlike the `source-register.md` ledger (append-only + supersede-by-pointer; never edit/delete a row). "Land a durable prevention point" (per the always-land rule) is satisfied by **a merge into an existing canonical rule, a validator/checklist gate, or a reference pointer — NOT necessarily a new top-level bullet**: before appending a rule, grep the section for an existing owner of the same failure-class and merge instead (this extends the `keep/merge/discard` conflict rule from *incompatible* to *redundant* rules). A consolidation that rewrites or retires rule text is a non-wording shared-skill change whose **behavioral-evidence row (`semantic-control`) MUST carry a zero-loss obligation map** — every before-obligation maps to surviving or genuinely-subsumed text, none silently dropped; the challenge inspects that map. Dropping a real guard under the banner of "consolidation" is a regression.
|
|
165
165
|
- Wording-level dedup mechanics route to `tighten-doc`; the at-add-time consolidation check, the form-by-failure drafting table, the obligation-table format (incl. the verifiable-survivor-pointer rule), the package/support-file integrity axis, strict "subsumed" criteria, register-row boundary, and audit recovery: `references/rule-consolidation.md`.
|
|
166
166
|
|