@ccoalm/ccl-skills 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +93 -16
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +10 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +249 -5
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate_abort_leak.sh +394 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +2 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +6 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +3 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-capability-composition.md +128 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +35 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +54 -2
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +11 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +7 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +6 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +50 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +1 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +1 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +6 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +1 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +59 -1
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +1 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +3 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +43 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +16 -4
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +391 -14
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +74 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +74 -33
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +11 -4
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +421 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_round_attribution.sh +576 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_source_refuted.sh +176 -0
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_lanes.sh +101 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +6 -2
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +2 -2
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +19 -2
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +1 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +3 -3
  37. package/dist/assets/release.json +63 -33
  38. package/package.json +1 -1
@@ -0,0 +1,394 @@
1
+ #!/usr/bin/env bash
2
+ # Proves an aborted test_review_gate.sh leaves no reviewer wrapper behind, by both of
3
+ # the mechanisms that keep that true.
4
+ #
5
+ # The suite's reviewer wrappers are deliberately TERM-immune and the controller starts
6
+ # them with start_new_session=True, so they sit in their own session where no signal
7
+ # aimed at the suite's process group can reach them. That makes the controller the only
8
+ # party that reaps them on the happy path. Both legs below remove the controller FIRST —
9
+ # the way a tree-kill, a CI cancellation, or a host suspend does — and then abort the
10
+ # suite two different ways:
11
+ #
12
+ # leg 1 SIGTERM the suite: its cleanup trap runs and must reap the abandoned wrapper.
13
+ # leg 2 SIGKILL the suite: nothing can trap that, so no reaper runs at all and the
14
+ # wrapper must die on the fixture's own lifetime bound instead.
15
+ #
16
+ # Leg 2 exists because leg 1 alone stays green if the fixture's bound were restored to
17
+ # an unbounded loop: the trap would still reap it. It was an unbounded loop that turned
18
+ # this defect into a wrapper observed alive for 39 hours with ppid=1, ignoring SIGTERM.
19
+ #
20
+ # Each leg runs the suite under a private TMPDIR so every process this probe may signal
21
+ # is identified by a path no other run can share. Nothing here matches on the shared
22
+ # `review-gate-test.` prefix: a concurrent lane running the same suite would be killed
23
+ # by that, which is precisely the class scripts/test_lane_isolation.py keeps out of the
24
+ # parallel lanes.
25
+ set -uo pipefail
26
+
27
+ DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)"
28
+ SUITE="$DIR/test_review_gate.sh"
29
+ # How long an abandoned wrapper may still be alive after the suite is gone.
30
+ GRACE="${ABORT_LEAK_PROBE_GRACE:-30}"
31
+ # Bound on reaching the target case. The suite reaches it in ~2min on an idle host; the
32
+ # headroom is for a slow runner, NOT for sharing one — this probe runs in its own job
33
+ # precisely because a shared runner stretches that round past any budget.
34
+ REACH="${ABORT_LEAK_PROBE_REACH:-420}"
35
+ # Leg 2 shortens the fixture's own bound so the leg does not have to wait out the 300s
36
+ # default; leg 2's runtime is essentially this number, because the leg ends by watching
37
+ # the abandoned wrapper reach it. The constraint is the budget each case gives ITS
38
+ # WRAPPER — the `--timeout` the controller passes down, which is 5s for the claude hang
39
+ # case this probe aborts in and under 12s for the fallback hang cases that run before it
40
+ # (per those cases' own assertions) — NOT the case's `--total-timeout`, which bounds the
41
+ # whole gate rather than the wrapper. 25 keeps ~2x margin over the largest of those while
42
+ # taking 20s off the job: measured on CI, this job was the longest branch at 213s against
43
+ # 188s for the previous longest, so that 20s is the difference between adding a critical
44
+ # path and landing inside the existing one.
45
+ LEG2_HANG_BOUND="${ABORT_LEAK_PROBE_HANG_BOUND:-25}"
46
+ # Which leg(s) to run: `1`, `2`, or `all`. Each leg needs its OWN full run of the suite —
47
+ # one abort must be trappable and the other must not — and that pair, not the leg-2 wait,
48
+ # is what makes this the longest CI branch. Measured: cutting the leg-2 bound 45s -> 25s
49
+ # moved the job 213s -> 214s, i.e. not at all. So CI runs the legs as two parallel jobs
50
+ # and each stays well under the next-longest branch; `all` remains the local default.
51
+ ABORT_LEAK_PROBE_LEG="${ABORT_LEAK_PROBE_LEG:-all}"
52
+ case "$ABORT_LEAK_PROBE_LEG" in
53
+ 1|2|all) : ;;
54
+ *) echo "ABORT_LEAK_PROBE_LEG must be 1, 2, or all" >&2; exit 2 ;;
55
+ esac
56
+ # Which fixture stub to orphan. The suite generates TWO of them — the claude stub and the
57
+ # shared candidate stub the fallback clients run — and they carry the same hang behavior
58
+ # and the same lifetime bound. A probe pinned to one leaves the other's bound revertible
59
+ # with the probe still green, so CI points its two jobs at different stubs.
60
+ ABORT_LEAK_PROBE_CLIENT="${ABORT_LEAK_PROBE_CLIENT:-claude}"
61
+ case "$ABORT_LEAK_PROBE_CLIENT" in
62
+ claude) PROBE_BEHAVIOR_FILE=claude_behavior; PROBE_WRAPPER=claude_review.sh ;;
63
+ fallback) PROBE_BEHAVIOR_FILE=kimi_behavior; PROBE_WRAPPER=kimi_review.sh ;;
64
+ *) echo "ABORT_LEAK_PROBE_CLIENT must be claude or fallback" >&2; exit 2 ;;
65
+ esac
66
+
67
+ fails=0
68
+ check() {
69
+ if eval "$2"; then echo "ok - $1"; else echo "FAIL - $1"; fails=$((fails+1)); fi
70
+ }
71
+
72
+ PROBE_TMP=""
73
+ PROBE_TMP_REAL=""
74
+ suite_pid=""
75
+ suite_pgid=""
76
+ suite_start=""
77
+ wrapper_pid=""
78
+ wrapper_pgid=""
79
+
80
+ # Every process a leg starts carries that leg's private tmp path in argv. The needles go
81
+ # through the environment, not `awk -v`: `ps -e` lists this very awk, and an argv-passed
82
+ # needle makes the scanner match itself.
83
+ probe_processes() {
84
+ [ -n "$PROBE_TMP_REAL" ] || return 0
85
+ ps -eo pid=,stat=,command= 2>/dev/null |
86
+ PROBE_A="$PROBE_TMP/" PROBE_B="$PROBE_TMP_REAL/" \
87
+ awk 'BEGIN { a = ENVIRON["PROBE_A"]; b = ENVIRON["PROBE_B"] }
88
+ (index($0, a) || index($0, b)) && $2 !~ /Z/ { print $1 }'
89
+ }
90
+
91
+ # Never let the probe become the leak it tests for — and never on someone else's
92
+ # processes: `suite_pgid` and any pid read earlier are cached numbers the OS recycles,
93
+ # so ownership is re-proved against the live process table at the moment of signalling.
94
+ # A group is signalled only while it still contains a process carrying this probe's
95
+ # private tmp path; everything else is signalled by pid, which `probe_processes` has
96
+ # just re-matched by that same path.
97
+ probe_group_is_ours() {
98
+ ps -eo pgid=,stat=,command= 2>/dev/null |
99
+ PROBE_A="$PROBE_TMP/" PROBE_B="$PROBE_TMP_REAL/" PROBE_PGID="$1" \
100
+ awk 'BEGIN { a = ENVIRON["PROBE_A"]; b = ENVIRON["PROBE_B"]; g = ENVIRON["PROBE_PGID"]; found = 1 }
101
+ $1 == g && $2 !~ /Z/ && (index($0, a) || index($0, b)) { found = 0; exit }
102
+ END { exit found }'
103
+ }
104
+ # The suite's pid and group id were captured minutes earlier and the OS recycles both,
105
+ # so every signal aimed at them re-proves identity first: the pid must still be running
106
+ # THIS suite script and must still lead the group we recorded. Validating the group
107
+ # through a live member we can name is what makes the group signal safe — a
108
+ # tmp-path-membership test would not do here, because by design the controller (the one
109
+ # group member carrying that path) has just been killed.
110
+ suite_is_ours() {
111
+ case "${suite_pid:-}" in ''|*[!0-9]*) return 1 ;; esac
112
+ case "${suite_pgid:-}" in ''|*[!0-9]*) return 1 ;; esac
113
+ kill -0 "$suite_pid" 2>/dev/null || return 1
114
+ case "$(ps -o command= -p "$suite_pid" 2>/dev/null || true)" in
115
+ *"$SUITE"*) : ;;
116
+ *) return 1 ;;
117
+ esac
118
+ [ "$(ps -o pgid= -p "$suite_pid" 2>/dev/null | tr -d ' ')" = "$suite_pgid" ] || return 1
119
+ # Start time, not just argv and group: with job control the suite's pgid equals its
120
+ # own pid, so those two conditions collapse into "this pid is running this script" —
121
+ # which a recycled pid on another run of the SAME suite would also satisfy. The launch
122
+ # instant is what separates our process from its namesake.
123
+ [ -n "$suite_start" ] || return 1
124
+ [ "$(ps -o lstart= -p "$suite_pid" 2>/dev/null || true)" = "$suite_start" ] || return 1
125
+ return 0
126
+ }
127
+ signal_suite() {
128
+ suite_is_ours || return 0
129
+ kill -"$1" -"$suite_pgid" 2>/dev/null || true
130
+ kill -"$1" "$suite_pid" 2>/dev/null || true
131
+ }
132
+ cleanup() {
133
+ local pid probe_cmd
134
+ signal_suite KILL
135
+ for pid in $(probe_processes); do
136
+ # Re-check immediately before signalling: a pid read even a moment ago can already be
137
+ # a stranger's. This narrows the window to the width of one `ps`; a shell cannot close
138
+ # it entirely, because there is no portable kill-by-handle, and that residual TOCTOU
139
+ # is a documented boundary rather than an oversight.
140
+ # Match BOTH spellings, like probe_processes does: on a host where the temp dir is a
141
+ # symlink (macOS /var -> /private/var) a process can carry either form, and a check
142
+ # that knows only one silently skips a process this probe started.
143
+ probe_cmd="$(ps -o command= -p "$pid" 2>/dev/null || true)"
144
+ case "$probe_cmd" in
145
+ *"$PROBE_TMP/"*|*"$PROBE_TMP_REAL/"*) : ;;
146
+ *) continue ;;
147
+ esac
148
+ probe_group_is_ours "$pid" && kill -KILL -"$pid" 2>/dev/null || true
149
+ kill -KILL "$pid" 2>/dev/null || true
150
+ done
151
+ [ -z "$PROBE_TMP" ] || rm -rf "$PROBE_TMP"
152
+ }
153
+ trap cleanup EXIT
154
+ trap 'exit 130' INT
155
+ trap 'exit 143' TERM HUP
156
+
157
+ # Selecting the target reads the suite's OWN state rather than pattern-matching an argv
158
+ # tail. Two earlier selectors failed here and each failure is the reason for this shape:
159
+ # picking by lifetime chose one of the behaviors that merely sleeps and then exits on its
160
+ # own (`quota_slow` 2s, `passed_slow` 5s), so the probe was green against a broken suite;
161
+ # and matching `--focus process-group-timeout` in the command line depends on how ps
162
+ # renders a long argv — it found the case when the probe ran alone and found nothing at
163
+ # all under `make`, burning the whole reach budget twice. The behavior file is written by
164
+ # the suite before it starts the case, so "a hang case is running now" is a fact to read,
165
+ # not a shape to guess. It also matches the FIRST hang case rather than the last, so the
166
+ # probe reaches its target sooner.
167
+ #
168
+ # The wrapper and the `( ... ) &` child it backgrounds share a command line, so "the
169
+ # first matching ps row" is whichever the kernel lists first. Picking the child would
170
+ # make the next step read the WRAPPER as the controller — and the wrapper's own command
171
+ # line satisfies the tmp-path validation, so that mistake would sail through and the leg
172
+ # would kill the wrong process. Select by ancestry: the wrapper is the match whose parent
173
+ # is the controller.
174
+ suite_work_dir() {
175
+ local dir
176
+ for dir in "$PROBE_TMP_REAL"/review-gate-test.*; do
177
+ [ -d "$dir" ] && printf '%s\n' "$dir" && return 0
178
+ done
179
+ return 1
180
+ }
181
+ live_wrapper() {
182
+ local work behavior pid parent_cmd
183
+ work="$(suite_work_dir)" || return 0
184
+ behavior="$(cat "$work/state/$PROBE_BEHAVIOR_FILE" 2>/dev/null || true)"
185
+ [ "$behavior" = "hang" ] || return 0
186
+ for pid in $(ps -eo pid=,command= 2>/dev/null |
187
+ grep -e "$PROBE_TMP/" -e "$PROBE_TMP_REAL/" -F |
188
+ grep -F "$PROBE_WRAPPER" |
189
+ grep -v '[g]rep' |
190
+ awk '{print $1}'); do
191
+ parent_cmd="$(ps -o command= -p "$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ')" 2>/dev/null || true)"
192
+ case "$parent_cmd" in
193
+ *review_gate.py*) printf '%s\n' "$pid"; return 0 ;;
194
+ esac
195
+ done
196
+ return 0
197
+ }
198
+
199
+ # Runs the suite until the target wrapper is in flight, then removes its controller so
200
+ # the wrapper is orphaned with no reaper but the suite itself. Sets wrapper_pid.
201
+ # $1 is the value for the fixture's lifetime bound, or empty for its default.
202
+ orphan_target_wrapper() {
203
+ local hang_bound="$1" candidate controller_pid controller_ok deadline suite_died
204
+ PROBE_TMP="$(mktemp -d "${TMPDIR:-/tmp}/review-gate-abort-probe.XXXXXX")" || return 1
205
+ PROBE_TMP_REAL="$(cd "$PROBE_TMP" && pwd -P)" || return 1
206
+ wrapper_pid=""
207
+
208
+ # Job control so the suite gets its own process group and the probe can signal that
209
+ # group without signalling itself.
210
+ set -m
211
+ TMPDIR="$PROBE_TMP" REVIEW_GATE_TEST_HANG_SECONDS="$hang_bound" \
212
+ bash "$SUITE" >/dev/null 2>&1 &
213
+ suite_pid=$!
214
+ suite_pgid="$(ps -o pgid= -p "$suite_pid" 2>/dev/null | tr -d ' ')"
215
+ suite_start="$(ps -o lstart= -p "$suite_pid" 2>/dev/null || true)"
216
+
217
+ deadline=$(( $(date +%s) + REACH ))
218
+ suite_died=0
219
+ while [ "$(date +%s)" -lt "$deadline" ]; do
220
+ candidate="$(live_wrapper)"
221
+ if [ -n "$candidate" ]; then wrapper_pid="$candidate"; break; fi
222
+ kill -0 "$suite_pid" 2>/dev/null || { suite_died=1; break; }
223
+ sleep 0.1
224
+ done
225
+ if [ -z "$wrapper_pid" ]; then
226
+ # Say WHICH failure this is. "unreached" alone cannot distinguish a suite that ran
227
+ # fine but never got to a hang case inside the budget from a suite that died early —
228
+ # and a run of this probe under `make` failed exactly here with no way to tell them
229
+ # apart, which is why the distinction is printed rather than inferred.
230
+ if [ "$suite_died" = 1 ]; then
231
+ wait "$suite_pid" 2>/dev/null
232
+ printf 'abort_leak_probe_unreached: the suite exited (status %s) before any hang case was observed\n' "$?" >&2
233
+ else
234
+ printf 'abort_leak_probe_unreached: %ss budget elapsed with the suite still running; last behavior=%s\n' \
235
+ "$REACH" "$(cat "$(suite_work_dir 2>/dev/null)/state/claude_behavior" 2>/dev/null || printf unknown)" >&2
236
+ fi
237
+ return 1
238
+ fi
239
+
240
+ # Validate the target before signalling it, never after. Between selecting the wrapper
241
+ # and reading its parent the wrapper can be reaped and its pid reused, so this value is
242
+ # untrustworthy by construction: it can be 1 (already orphaned) or an unrelated
243
+ # process. `kill -KILL 1` is harmless for an unprivileged user but not for a root
244
+ # container, and no probe should be the thing that finds that out.
245
+ controller_pid="$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')"
246
+ controller_ok=0
247
+ case "$controller_pid" in
248
+ ''|*[!0-9]*) : ;;
249
+ *)
250
+ if [ "$controller_pid" -gt 1 ] \
251
+ && [ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "$controller_pid" ] \
252
+ && { case "$(ps -o command= -p "$controller_pid" 2>/dev/null || true)" in
253
+ *"$PROBE_TMP/"*|*"$PROBE_TMP_REAL/"*) true ;; *) false ;; esac; }; then
254
+ controller_ok=1
255
+ fi
256
+ ;;
257
+ esac
258
+ if [ "$controller_ok" != 1 ]; then
259
+ printf 'abort leak setup: refusing to signal pid=%s (not this run'"'"'s controller)\n' \
260
+ "${controller_pid:-none}" >&2
261
+ return 1
262
+ fi
263
+ # Order matters: the controller must die BEFORE the suite, or its own timeout path
264
+ # reaps the wrapper and the leg proves nothing about the suite. `kill` returning is
265
+ # not that proof: confirm here, once, that the controller is really gone and the
266
+ # wrapper really was reparented, so BOTH legs inherit a verified precondition instead
267
+ # of each re-deriving it (leg 2 had no such check, and its later kill of the suite
268
+ # group could have removed the controller itself — making every remaining assertion
269
+ # pass without the required ordering ever holding).
270
+ wrapper_pgid="$(ps -o pgid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')"
271
+ kill -KILL "$controller_pid" 2>/dev/null || true
272
+ deadline=$(( $(date +%s) + 10 ))
273
+ while :; do
274
+ if ! kill -0 "$controller_pid" 2>/dev/null ||
275
+ [ -n "$(ps -o stat= -p "$controller_pid" 2>/dev/null | grep Z || true)" ]; then
276
+ break
277
+ fi
278
+ if [ "$(date +%s)" -ge "$deadline" ]; then
279
+ echo "abort leak setup: controller $controller_pid survived SIGKILL" >&2
280
+ return 1
281
+ fi
282
+ sleep 0.5
283
+ done
284
+ while :; do
285
+ [ "$(ps -o ppid= -p "$wrapper_pid" 2>/dev/null | tr -d ' ')" = "1" ] && return 0
286
+ kill -0 "$wrapper_pid" 2>/dev/null || {
287
+ echo "abort leak setup: wrapper $wrapper_pid died with its controller" >&2
288
+ return 1
289
+ }
290
+ if [ "$(date +%s)" -ge "$deadline" ]; then
291
+ echo "abort leak setup: wrapper $wrapper_pid was never reparented" >&2
292
+ return 1
293
+ fi
294
+ sleep 0.5
295
+ done
296
+ }
297
+
298
+ await_suite_exit() {
299
+ local deadline
300
+ deadline=$(( $(date +%s) + GRACE ))
301
+ while kill -0 "$suite_pid" 2>/dev/null; do
302
+ [ "$(date +%s)" -ge "$deadline" ] && break
303
+ sleep 0.5
304
+ done
305
+ kill -0 "$suite_pid" 2>/dev/null && return 1
306
+ return 0
307
+ }
308
+
309
+ # The verdict is about the GROUP, not the leader. The wrapper backgrounds a TERM-immune
310
+ # child in that same group; a cleanup that kills the session leader but misses the child
311
+ # leaves the leak in place while "the wrapper is gone" reads true — and the probe's own
312
+ # cleanup would then erase the very evidence the leg exists to surface. Residue under the
313
+ # leg's private tmp path is checked too, so a survivor that left the group still counts.
314
+ wrapper_group_members_alive() {
315
+ [ -n "$wrapper_pgid" ] || return 0
316
+ ps -eo pid=,pgid=,stat= 2>/dev/null |
317
+ awk -v want="$wrapper_pgid" '$2 == want && $3 !~ /Z/ { print $1 }'
318
+ }
319
+ wrapper_gone_within() {
320
+ local deadline members residue
321
+ deadline=$(( $(date +%s) + $1 ))
322
+ while :; do
323
+ members="$(wrapper_group_members_alive | tr '\n' ' ')"
324
+ residue="$(probe_processes | tr '\n' ' ')"
325
+ case "$members$residue" in
326
+ *[0-9]*) : ;;
327
+ *) return 0 ;;
328
+ esac
329
+ [ "$(date +%s)" -ge "$deadline" ] && break
330
+ sleep 1
331
+ done
332
+ printf 'abort leak: group members alive: %s | private-path residue: %s\n' \
333
+ "${members:-none}" "${residue:-none}" >&2
334
+ return 1
335
+ }
336
+
337
+ diagnose_wrapper() {
338
+ printf 'abort leak diagnostic (%s): wrapper=%s %s\n' "$1" "$wrapper_pid" \
339
+ "$(ps -o pid=,ppid=,etime=,stat= -p "$wrapper_pid" 2>/dev/null | tr -s ' ')" >&2
340
+ }
341
+
342
+ # ---- leg 1: a trappable abort — the suite's own cleanup must reap the wrapper --------
343
+ if [ "$ABORT_LEAK_PROBE_LEG" = 1 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
344
+ leg1_orphaned=0
345
+ orphan_target_wrapper "" && leg1_orphaned=1
346
+ check "leg1: the controller was removed and the wrapper reparented to init" '[ "$leg1_orphaned" = 1 ]'
347
+ if [ "$leg1_orphaned" = 1 ]; then
348
+ signal_suite TERM
349
+ leg1_suite_gone=0; await_suite_exit && leg1_suite_gone=1
350
+ check "leg1: the aborted suite exits" '[ "$leg1_suite_gone" = 1 ]'
351
+ leg1_reaped=0; wrapper_gone_within "$GRACE" && leg1_reaped=1
352
+ [ "$leg1_reaped" = 1 ] || diagnose_wrapper leg1
353
+ check "leg1: a terminated suite reaps the reviewer wrapper it started" '[ "$leg1_reaped" = 1 ]'
354
+ fi
355
+ cleanup; PROBE_TMP=""; PROBE_TMP_REAL=""; suite_pgid=""
356
+ fi
357
+
358
+ # ---- leg 2: an untrappable abort — only the fixture's own bound can end the wrapper --
359
+ if [ "$ABORT_LEAK_PROBE_LEG" = 2 ] || [ "$ABORT_LEAK_PROBE_LEG" = all ]; then
360
+ leg2_orphaned=0
361
+ orphan_target_wrapper "$LEG2_HANG_BOUND" && leg2_orphaned=1
362
+ check "leg2: the controller was removed and the wrapper reparented to init" '[ "$leg2_orphaned" = 1 ]'
363
+ if [ "$leg2_orphaned" = 1 ]; then
364
+ signal_suite KILL
365
+ leg2_suite_gone=0; await_suite_exit && leg2_suite_gone=1
366
+ check "leg2: the killed suite exits without running any cleanup" '[ "$leg2_suite_gone" = 1 ]'
367
+ # Precondition: with no trap able to run, the wrapper must still be here. If it is
368
+ # already gone, something else reaped it and the bound was never exercised.
369
+ # `kill -0` succeeds for a zombie, and the verdict below excludes zombies — so a wrapper
370
+ # that had already died would satisfy "still alive" here and "gone" there, passing the
371
+ # leg without the bound ever being exercised. Require a live, non-zombie process.
372
+ leg2_survived_abort=0
373
+ if kill -0 "$wrapper_pid" 2>/dev/null; then
374
+ case "$(ps -o stat= -p "$wrapper_pid" 2>/dev/null || true)" in
375
+ ''|*Z*) : ;;
376
+ *) leg2_survived_abort=1 ;;
377
+ esac
378
+ fi
379
+ check "leg2: no reaper ran, so the wrapper is still alive right after the kill" \
380
+ '[ "$leg2_survived_abort" = 1 ]'
381
+ leg2_self_exit=0; wrapper_gone_within "$(( LEG2_HANG_BOUND + GRACE ))" && leg2_self_exit=1
382
+ [ "$leg2_self_exit" = 1 ] || diagnose_wrapper leg2
383
+ check "leg2: an unreaped wrapper still exits on the fixture's own lifetime bound" \
384
+ '[ "$leg2_self_exit" = 1 ]'
385
+ fi
386
+
387
+ fi
388
+
389
+ if [ "$fails" -eq 0 ]; then
390
+ echo "review_gate_abort_leak_ok leg=$ABORT_LEAK_PROBE_LEG client=$ABORT_LEAK_PROBE_CLIENT"
391
+ exit 0
392
+ fi
393
+ echo "review_gate_abort_leak_failed=$fails" >&2
394
+ exit 1
@@ -17,6 +17,8 @@ Do not create a service only because:
17
17
  - A team wants a cleaner folder.
18
18
  - A future scale concern is speculative.
19
19
 
20
+ - External grounding (adopted in part): the criteria above take [Bounded Context](https://martinfowler.com/bliki/BoundedContext.html) (Fowler's overview of the central DDD strategic-design pattern; origin credited there to Eric Evans, *Domain-Driven Design*) as an input to boundary drawing — a boundary is where one internally consistent model stops being valid, so boundaries follow model and responsibility lines, never noun or table count — and [Conway's Law](https://martinfowler.com/bliki/ConwaysLaw.html) (Melvin Conway, ["How Do Committees Invent?", 1968](https://www.melconway.com/Home/Committees_Paper.html); Fowler's overview) — a system's structure mirrors the communication structure of the organization that builds it, which is why team ownership and communication structure are a legitimate split input. Two limits: the transactional-data-ownership, scaling, deployment-cadence, and rollback tests above are this skill's own operational criteria, attributed to neither source; and a bounded context does not require its own deployable service — several contexts can ship inside one modular monolith (this skill's monolith-first default), so these sources justify model/ownership boundaries, not a service-per-context split. Borrowed scope: boundary-input criteria only; no claim to the full DDD strategic-design method (context maps, ubiquitous language) or an inverse-Conway process. Sibling: `python-service-architecture/references/architecture-playbook.md` ("Architecture Decisions") carries the same grounding; keep the two in sync.
21
+
20
22
  ## Recommended Service Types
21
23
 
22
24
  ### API Gateway / HTTP Service
@@ -70,3 +70,9 @@
70
70
  - Every ctx key has a typed accessor; bare-string `ctx.Value(...)` returning `any` is an architecture finding, not an idiom to spread. Type assertion lives behind helpers, not in domain code.
71
71
  - For frameworks that propagate metadata over the wire, define which keys travel persistently (every downstream hop forwards them) versus transiently (one hop only). Lane, stress tag, and trace identity are typically persistent; one-off control flags should not be promoted to persistent.
72
72
  - Dual-injection compatibility: when one binary serves multiple transports (TTHeader Thrift + HTTP/2 gRPC), the propagation layer writes to both the framework's persistent value (`metainfo.WithPersistentValue`) and the transport's outgoing metadata so the framework's meta handler picks the right wire format at send time. Application code stays transport-agnostic.
73
+
74
+ ## Topic-extension backlog
75
+
76
+ Entries here are registered candidates, not adopted guidance. Each names the candidate, its evidence status, and the condition that unblocks adoption; the round that evaluates one records keep/narrow/discard against its entry.
77
+
78
+ - **Package-owned runtime-invariant registries (weak-keep candidate).** Observed form, from one agent-native product repository (evolving portfolio): each package contributes its own invariant checks from a companion module, normal entrypoints do not depend on the diagnostics layer, and allowlist/blocklist configuration selects which checks run. Evidence status: weak — single source; whether this generalizes as a cross-cutting pattern, and which reference should own it, is undecided. Evaluate the generalization at the next architecture round touching diagnostics, invariant checking, or startup validation, and record the outcome against this entry. The Python-stack sibling registration lives in `../python-service-architecture/references/packaging-runtime-readiness.md`.
@@ -24,7 +24,8 @@ Use this for product backend work that calls, hosts, evaluates, or operates LLM
24
24
  - Core: model/version registry, prompt lifecycle, request schema, safety boundaries, streaming protocol, retry/fallback, eval datasets, replay/shadow rollout, token/cost metrics, batch/concurrency control, audit trails.
25
25
  - agent-skill-system runtime (progressive-disclosure loading, description-driven skill routing, skill trust/sandbox boundary)
26
26
  - MCP integration (server primitives, server-as-untrusted-domain trust boundary, OAuth 2.1 / audience-binding auth)
27
- - agent command-execution sandbox (OS-sandbox composition, command-policy DSL, approval/escalation state machine, loopback-only egress proxy — see `references/agent-command-sandbox.md`)
27
+ - agent capability composition (seam/provider/model-facing-tool decomposition, reversible registration with scope-owned disposal, policy-plugin-vs-enforcement split, sub-agent provider seam with one-shot/continuable separation — see `references/agent-capability-composition.md`)
28
+ - agent command-execution sandbox (OS-sandbox composition, single-owner policy resolution consumed by every enforcement backend, command-policy DSL, approval/escalation state machine, loopback-only egress proxy — see `references/agent-command-sandbox.md`)
28
29
  - agent session persistence (append-only event-log source of truth, background-writer flush-before-finality, resume/fork with restored token accounting, context-window compaction — see `references/agent-session-persistence.md`)
29
30
  - agent ambient context freshness (diffable world-state sections vs one-shot fragments, comparison-snapshot-vs-rendered-text staleness detection, supersede-stale-in-band-not-retract, self-recognizing injections, derivable baseline + merge-patch re-derive on resume, budget-signaling-vs-compaction-enforcement — see `references/agent-context-freshness.md`)
30
31
  - robust model-driven file-edit (context-anchored hunks over line numbers, graduated-strictness matching, resolve-all-before-write, honest partial-failure reporting — see `references/agent-file-edit-protocol.md`)
@@ -72,6 +73,7 @@ Use this for product backend work that calls, hosts, evaluates, or operates LLM
72
73
  - Trace representative turns through every state transformation: at minimum the normal allow path, deny/ask permission path, resume/recovery path, and dynamic-tool-change path. Each trace must cover user input normalization, context construction, budget or compaction projection, model streaming, tool-call validation, permission decision, tool-result insertion, continuation, terminal stop condition, transcript write, and usage/cost accounting.
73
74
  - Each runtime concern below gets layer-separation, and its gate / assertions / routing live in the named reference (read before gating that concern):
74
75
  - Runtime startup & config bootstrap -> `references/agent-runtime-bootstrap.md`
76
+ - Capability composition — module/plugin boundaries, registration lifecycle, sub-agent providers -> `references/agent-capability-composition.md`
75
77
  - Turn lifecycle — per-turn loop, fg/bg handoff, progress, plan-to-execute, recap -> `references/agent-turn-lifecycle.md`
76
78
  - Session & transport — discovery/fork, history sync, workspace scope, protocol/stdout, settings migration -> `references/agent-session-persistence.md`
77
79
  - Hooks — hook/classifier output trust, hook config control plane, post-turn hooks -> `references/agent-lifecycle-hooks.md`
@@ -0,0 +1,128 @@
1
+ # Agent Capability Composition (Seams, Providers, Reversible Registration)
2
+
3
+ How an agent runtime packages its capabilities so they can be swapped, unloaded, and extended
4
+ without a privileged core: the decomposition unit, the registration lifecycle, and the sub-agent
5
+ provider seam. This is the composition layer *around* the frameworks the sibling references own —
6
+ `agent-tool-dispatch.md` owns how a call reaches a handler; `agent-runtime-bootstrap.md` owns
7
+ config/trust at startup; this reference owns how the handler's capability is packaged, installed,
8
+ and torn down.
9
+
10
+ Patterns observed in one agent-native product repository (an evolving portfolio — treat as one
11
+ industry form, not a standard) and consistent with general plugin-architecture practice. Use when
12
+ designing or reviewing an agent runtime's module boundaries, plugin/extension system, or sub-agent
13
+ integration; skip for a single-purpose agent whose capabilities never vary by deployment.
14
+
15
+ ## 1. A swappable capability is three parts — seam, provider, model-facing tool
16
+
17
+ - When a capability must vary by deployment (local vs sandboxed filesystem, real vs replay model
18
+ transport, in-process vs external sub-agent), decompose it into three separately-packaged roles:
19
+ the **seam** (the service contract: abstract interface + vocabulary types), one or more
20
+ **providers** (implementations of the seam), and the **model-facing tool** (a *consumer* of the
21
+ seam that renders it into the model's tool surface). Consumers — tools, policies, other
22
+ capabilities — depend only on the seam; a provider is selected at assembly time. The tool is not
23
+ the implementation and never reaches around the seam into one.
24
+ - Verify the decomposition by the swap test: a new provider (remote store, different vendor,
25
+ fault-injection fake) must drop in without touching the seam's consumers or the tool. If adding a
26
+ provider forces edits in consumers, the contract leaked implementation detail.
27
+ - Design the seam's contract for **all current consumers**, not the loudest one: a method shaped for
28
+ a single consumer's UI/transport/private need does not belong on the shared contract. A public
29
+ seam method with exactly one internal caller is API surface without a second user — prefer a
30
+ capability closure injected at construction into the one consumer that needs it.
31
+ - Keep role vocabulary honest in reviews: "the tool does X" is a smell when X is storage, policy, or
32
+ enforcement — those belong to a provider or a policy plugin, and the tool only surfaces them.
33
+ (Worked instances of the split: the spill store vs spill policy split in `agent-tool-dispatch.md`
34
+ §Result shaping; the sandbox policy-resolution vs enforcement-backend split in
35
+ `agent-command-sandbox.md` §2.)
36
+
37
+ ## 2. Registration is reversible; teardown is scope disposal, not per-plugin cleanup
38
+
39
+ - Every registration-like effect a capability performs at install time (registering a tool, an
40
+ event listener, a provider, a background worker) must carry its inverse, and the runtime — not
41
+ each plugin author — owns applying inverses on unload. Model installation as effects recorded
42
+ into an **ownership scope**; teardown = dispose the scope, which replays the inverses in reverse
43
+ order. Correct unload is then guaranteed by one abstraction instead of N hand-written uninstall
44
+ paths, and a plugin author cannot forget it.
45
+ - Hot-swap/reload is dispose-old-then-install-new under the same rule — never patch-in-place, which
46
+ accumulates the old registration's residue. The failure half needs a contract too: preflight the
47
+ new provider in an isolated scope *before* disposing the old one where the capability tolerates
48
+ brief coexistence. Where it does not, know what rollback can and cannot promise: registration
49
+ inverses reverse *local* effects only — disposal may have released a lease, port, credential, or
50
+ other external resource that reinstallation cannot reacquire — so a non-coexistent cutover either
51
+ uses an atomic/preflightable swap mechanism, or retains the old install descriptor and
52
+ rollback-critical resources until the new install commits, and when the install fails *and*
53
+ rollback also fails, the seam enters a **typed, surfaced unavailable state** (fail closed), never
54
+ a silent half-registered one. Verify with an install → dispose → re-install cycle asserting no
55
+ duplicate handlers, no leaked listeners, and no orphaned background work — plus failure injection
56
+ through BOTH layers: new provider's install throws (capability ends served by old or new where
57
+ the swap mechanism can guarantee it), and restoration fails too (capability ends in the typed
58
+ unavailable state, observably, rather than asserting never-neither for a path that cannot
59
+ guarantee it).
60
+ - Disposal order matters: dispose in reverse registration order so dependents release before their
61
+ dependencies, and make disposal idempotent so a crash-then-cleanup path cannot double-free.
62
+ - This composes with (does not replace) the trust and generation rules in
63
+ `agent-runtime-bootstrap.md`: a reversible registration that loads untrusted code is still gated
64
+ by trust state, and a reload still invalidates the caches keyed to the old generation.
65
+
66
+ ## 3. Policy plugins decide; providers enforce; absence has a declared default
67
+
68
+ - When a rule about *when/whether* to act is separable from the mechanics of acting (when to spill
69
+ an oversized result, whether an edit target is stale, when to compact), package the rule as a
70
+ swappable **policy plugin** over the capability seam rather than hard-coding it into the
71
+ provider. The provider keeps the atomic enforcement check at the moment of action (freshness,
72
+ no-clobber, bounds) because a policy's observation can be stale by the time the action runs.
73
+ - Declare what happens when the policy plugin is absent, and the declared default must fail toward
74
+ the safe direction for that capability's risk class — for a destructive-capable capability
75
+ (file overwrite, deletion, spend), absence means deny or require explicit approval, never a
76
+ silent degrade to the unguarded behavior. The cautionary shape: a runtime whose read-before-edit
77
+ observation policy is merely a removable plugin, with nothing beneath it, silently degrades to
78
+ unconditional overwrite when the plugin is missing — which is why the enforcement-layer
79
+ freshness/no-clobber check (`agent-command-sandbox.md` §Filesystem write authorization) stays
80
+ mandatory regardless of which policy plugins are installed, and a deployment-level opt-out of a
81
+ safety default must be explicit, never implied by absence.
82
+ - Do not let advisory policy observations become authority: a policy that watches tool traffic to
83
+ derive guard state (which files were read, which calls repeated) produces *hints and gates*, and
84
+ the enforcement layer re-validates at execution time.
85
+
86
+ ## 4. Sub-agent integration is a provider seam, not a hardcoded runner
87
+
88
+ - Expose sub-agent execution through the same seam/provider decomposition (§1), with one shape
89
+ difference: sub-agent providers are **named and co-resident** — a registry of concurrently
90
+ available providers selected per spawn (in-process fork, spawned worker, an external vendor
91
+ agent CLI) — unlike a single-selected executor seam. External vendor agent CLIs integrate as
92
+ first-class providers behind the seam, so consumers do not care whether a child is in-process or
93
+ a different vendor's product.
94
+ - Split the contract by lifecycle: a **one-shot** run (spawn → result) and a **continuable**
95
+ session (spawn → handle → further turns → teardown) are different APIs; do not overload one call
96
+ shape with both. For continuable children, the child-side provider emits only a **creation
97
+ spec**; the parent-side runtime owns handle allocation, turn delivery, and teardown — a provider
98
+ that never sees handles cannot leak or forge them.
99
+ - Make **fulfillment the single publish/ownership-transfer boundary**: a spawned child becomes
100
+ visible to consumers only when the provider fulfills the creation, and ownership (who tears it
101
+ down, whose scope disposes it) transfers exactly there — before fulfillment the provider owns
102
+ cleanup on failure; after, the runtime's scope does (§2). Two owners or zero owners at any point
103
+ is a leak or a double-free.
104
+ - This reference owns the seam shape only. Spawn-depth/live-count caps, capability intersection,
105
+ and cascade-cancel live in `retrieval-agent-safety.md` §Agent SDK Building Blocks; task-state
106
+ finality and recovery live in `agent-task-orchestration.md`; both apply in full to every
107
+ provider, including external-CLI children.
108
+
109
+ ## Non-negotiables
110
+
111
+ - A model-facing tool never depends on a concrete provider; consumers import the seam only.
112
+ - Every registration carries its inverse; unload correctness is owned once by scope disposal, and
113
+ hot-swap is dispose-then-reinstall, never patch-in-place.
114
+ - A policy plugin's absence has a declared, deliberate default; policy observations are advisory
115
+ and the provider re-checks atomically at action time.
116
+ - One-shot and continuable sub-agent APIs stay separate; fulfillment is the only ownership-transfer
117
+ point, and continuable child providers emit creation specs, never handles.
118
+ - These are composition patterns from one industry form — apply the swap/dispose/ownership tests
119
+ above as design gates, but do not cargo-cult the package layout onto a runtime whose capabilities
120
+ genuinely never vary.
121
+
122
+ ## Routing
123
+
124
+ - Tool registry/dispatch mechanics, result shaping, and output spill policy → `agent-tool-dispatch.md`.
125
+ - Startup trust, config generations, and cache invalidation on reload → `agent-runtime-bootstrap.md`.
126
+ - Sandbox policy resolution vs enforcement backends → `agent-command-sandbox.md`.
127
+ - Sub-agent spawn bounds and capability scoping → `retrieval-agent-safety.md`; task control plane → `agent-task-orchestration.md`.
128
+ - Session/scope persistence of what was installed when (for replay) → `agent-session-persistence.md`.
@@ -59,6 +59,41 @@ Express the filesystem/network posture as a closed enum of named profiles. A min
59
59
  Default network to **off** in every write-capable profile. Network is a separate grant from
60
60
  filesystem write (§4).
61
61
 
62
+ **One resolution owner; many enforcement backends.** Resolve "which profile + which writable
63
+ root(s) apply to this call" in exactly one place — a single policy owner that combines the
64
+ deployment default, the session's durable mode override, and any explicitly approved per-call mode
65
+ (explicit approval outranks session override outranks default — but only among profiles the
66
+ deployment/managed policy permits, and only after the §5/§6 authorization decision: a managed hard
67
+ deny or a mandatory sandbox ceiling is non-overridable by per-call approval, per §6's layered
68
+ precedence), and canonicalizes the session's
69
+ working root with filesystem semantics before it becomes the writable root. Every enforcing
70
+ backend — file tools, one-shot shell, persistent interactive sessions (§10) — consumes that one
71
+ resolved mode-and-root result per call, keeping only its platform dialect (§4) local. If each
72
+ backend re-resolves its own mode + root, they drift into a split world where the file layer and
73
+ the shell layer disagree about what is confined — a gap a command that touches both walks straight
74
+ through. One caveat is load-bearing: handing a backend the newly-resolved policy cannot *revoke*
75
+ a kernel sandbox an already-running persistent session (§10) was launched under — those sessions
76
+ bind the resolved policy and roots at creation, and whenever the creation-time grant is not a
77
+ subset of the current grant, or the policy/root generation has changed incompatibly (a swapped
78
+ writable root is not "narrower", yet the old root's access is revoked), the §10
79
+ generation-binding rule applies: tear down and recreate under the new sandbox, or reject the
80
+ command — never dispatch into the stale sandbox. (This is the policy/enforcement split of `agent-capability-composition.md` §3 applied to
81
+ the sandbox.)
82
+
83
+ **Present the resolved policy to the model as a replay-safe snapshot, not a capability inventory.**
84
+ Before each request, the model should see the currently-resolved policy as part of a cache-safe
85
+ runtime-context snapshot that is recorded into model-visible history — so replay can reconstruct
86
+ exactly what the model was told without rewriting the stable system prompt (prompt-cache churn;
87
+ see `llm-client-gateway.md`). Every generation persists for replay, but context construction
88
+ exposes only the *current* snapshot to the model and supersedes prior ones for subsequent
89
+ requests — after a broad-to-narrow policy change, stale snapshots left in the live envelope keep
90
+ advertising permissions and roots the enforcement layer no longer grants (denied attempts,
91
+ needless escalation, stale-path disclosure). Write that policy text as *what the policy means for operations*,
92
+ never as an enumerated capability/tool inventory (the maybe-absent-tool steering anti-pattern in
93
+ `agent-tool-dispatch.md`), and include the instruction that the model should not refuse an action
94
+ merely because policy *might* deny it — attempt the tool and follow the denial/escalation guidance
95
+ (§7) — otherwise the model self-censors more broadly than the policy actually restricts.
96
+
62
97
  **Reads need a confidentiality boundary too, not just writes.** A profile that confines *writes* but
63
98
  allows reading the whole host is still an exfil hole: an auto-approved "safe" read (`cat`, `rg`,
64
99
  `sed`, `find`) can slurp `~/.ssh`, `~/.aws`/cloud-credential files, browser profiles, keychains,