@ccoalm/ccl-skills 0.15.4 → 0.15.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +30 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +187 -29
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +234 -25
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +54 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +1 -1
- package/dist/assets/release.json +13 -13
- package/package.json +1 -1
|
@@ -16,6 +16,36 @@ The budget is a ceiling, not a quota: after a clean or fully source-refuted trac
|
|
|
16
16
|
`autonomous_review_allowed=false`; release/high-risk still requires at least one
|
|
17
17
|
challenge before this early close is eligible.
|
|
18
18
|
|
|
19
|
+
## What `complete` closes, and what it does not
|
|
20
|
+
|
|
21
|
+
`--mode complete` is the checkpoint for a chain whose findings were **shown to
|
|
22
|
+
be wrong**. Every original occurrence must carry a `source_refuted`
|
|
23
|
+
disposition; `unresolved`, `accepted_risk`, `accepted_tradeoff`, and
|
|
24
|
+
`needs_human_decision` are refused, and that refusal is deliberate. The gate
|
|
25
|
+
binds structure and provenance, never authority: it cannot tell a human
|
|
26
|
+
acceptance from an agent that labelled its own findings accepted, so it does
|
|
27
|
+
not let an acceptance close a machine checkpoint.
|
|
28
|
+
|
|
29
|
+
**A chain whose findings are accepted, out of scope, or input defects is not
|
|
30
|
+
stalled — it is simply not closed by this mode.** Such a round ends at its
|
|
31
|
+
`findings` result with a recorded disposition per occurrence, and the round's
|
|
32
|
+
own ledger carries the wider vocabulary. Do not read a refused `complete` as an
|
|
33
|
+
unfinished review; read it as "no refutation was claimed". Reporting the round
|
|
34
|
+
requires the dispositions, not a completion receipt.
|
|
35
|
+
|
|
36
|
+
Two mechanics that cost time when they are discovered by experiment:
|
|
37
|
+
|
|
38
|
+
- **`--stage` and `--risk-tag` must be passed to `complete`, not omitted.** The
|
|
39
|
+
binding predicate compares the prior rounds against the profile derived from
|
|
40
|
+
the arguments given here, so a risk-tagged chain checked without its tags
|
|
41
|
+
fails as an unbound candidate rather than as a mismatch.
|
|
42
|
+
- **A round that edits `skills/code-review/scripts/**` cannot bind its own
|
|
43
|
+
earlier rounds.** `review_controller_sha256` covers every `.py` and `.sh`
|
|
44
|
+
there, so any further edit to the harness mid-round changes the controller
|
|
45
|
+
identity and both chain succession and `complete` refuse the earlier
|
|
46
|
+
receipts. Land every harness edit first, then run review and challenge back
|
|
47
|
+
to back with nothing changed in between.
|
|
48
|
+
|
|
19
49
|
## Plan and owner binding
|
|
20
50
|
|
|
21
51
|
The plan is optional for `review` and `challenge` and required for `complete`.
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh
CHANGED
|
@@ -196,6 +196,142 @@ PROMPT_FILE="$RUN_ROOT/prompt.txt"
|
|
|
196
196
|
SCHEMA_FILE="$RUN_ROOT/schema.json"
|
|
197
197
|
RUN_WORKSPACE="$RUN_ROOT/workspace"
|
|
198
198
|
mkdir -p "$RUN_WORKSPACE"
|
|
199
|
+
# The reviewer runs from a private CODEX_HOME, not the user's. The user's home
|
|
200
|
+
# carries MCP servers -- their own, plus any an installed plugin contributes --
|
|
201
|
+
# and those servers run outside the CLI sandbox, so `--sandbox read-only` and
|
|
202
|
+
# `--disable shell_tool` do not reach them. A tool call completes before
|
|
203
|
+
# `audit_codex` can refuse the verdict, and a server auto-approves itself by
|
|
204
|
+
# declaring `readOnlyHint`, which the CLI trusts, so a tool that executes
|
|
205
|
+
# arbitrary code can be auto-approved while claiming to be read-only. Denying
|
|
206
|
+
# them without naming them was measured and does not work
|
|
207
|
+
# (`apps._default.default_tools_approval_mode` does not override the hint), and
|
|
208
|
+
# naming them cannot work either: an override under `mcp_servers` for a
|
|
209
|
+
# plugin-contributed server builds a transportless entry the CLI rejects
|
|
210
|
+
# outright. So this run gets a home that never had them.
|
|
211
|
+
#
|
|
212
|
+
# Model preferences are carried across explicitly, because this lane is
|
|
213
|
+
# contracted to review on the user's own default model and an empty home
|
|
214
|
+
# silently substitutes the CLI default. That carry-over is an allowlist, and
|
|
215
|
+
# deliberately not a denylist: a key this list has not heard of costs a
|
|
216
|
+
# preference, while a key a denylist has not heard of would let an executable
|
|
217
|
+
# server back in.
|
|
218
|
+
RUNTIME_HOME="$RUN_ROOT/codex-home"
|
|
219
|
+
mkdir -m 700 "$RUNTIME_HOME" \
|
|
220
|
+
|| die_inconclusive runtime_home_unavailable local_tool_failure false
|
|
221
|
+
AUTH_LINK_TARGET=""
|
|
222
|
+
if [ -e "$SOURCE_HOME/auth.json" ]; then
|
|
223
|
+
# A link, not a copy: the CLI refreshes the credential in place, and the
|
|
224
|
+
# rotated token has to land in the user's own file. The link is re-checked
|
|
225
|
+
# after the run, because a replaced link means the credential was written
|
|
226
|
+
# into this run directory instead.
|
|
227
|
+
AUTH_LINK_TARGET="$SOURCE_HOME/auth.json"
|
|
228
|
+
ln -s "$AUTH_LINK_TARGET" "$RUNTIME_HOME/auth.json" \
|
|
229
|
+
|| die_inconclusive runtime_home_auth_link_failed local_tool_failure false
|
|
230
|
+
fi
|
|
231
|
+
if [ -f "$SOURCE_HOME/config.toml" ]; then
|
|
232
|
+
python3 - "$SOURCE_HOME/config.toml" "$RUNTIME_HOME/config.toml" <<'PY_HOME_PREFERENCES' \
|
|
233
|
+
|| die_inconclusive codex_home_preferences_unreadable capability_missing true
|
|
234
|
+
import sys, tomllib
|
|
235
|
+
from pathlib import Path
|
|
236
|
+
|
|
237
|
+
# Model identity only, by KEY. `model_providers` is the exception worth naming:
|
|
238
|
+
# its value is a subtree this list does not inspect, so the allowlist bounds
|
|
239
|
+
# which keys travel, not everything that travels inside them. It is copied from
|
|
240
|
+
# the host's own configuration into a run-scoped home, so it grants a provider
|
|
241
|
+
# definition the host already had; narrowing it is a recorded follow-up.
|
|
242
|
+
# Nothing here can introduce a tool, a server, a hook, or a skill.
|
|
243
|
+
PREFERENCE_KEYS = (
|
|
244
|
+
"model",
|
|
245
|
+
"model_provider",
|
|
246
|
+
"model_providers",
|
|
247
|
+
"model_reasoning_effort",
|
|
248
|
+
"model_reasoning_summary",
|
|
249
|
+
"model_verbosity",
|
|
250
|
+
"service_tier",
|
|
251
|
+
)
|
|
252
|
+
try:
|
|
253
|
+
source = tomllib.loads(Path(sys.argv[1]).read_text(encoding="utf-8"))
|
|
254
|
+
except (OSError, UnicodeError, tomllib.TOMLDecodeError):
|
|
255
|
+
sys.exit(1)
|
|
256
|
+
if not isinstance(source, dict):
|
|
257
|
+
sys.exit(1)
|
|
258
|
+
|
|
259
|
+
|
|
260
|
+
ESCAPES = {"\\": "\\\\", '"': '\\"', "\b": "\\b", "\t": "\\t",
|
|
261
|
+
"\n": "\\n", "\f": "\\f", "\r": "\\r"}
|
|
262
|
+
|
|
263
|
+
|
|
264
|
+
def render_string(value):
|
|
265
|
+
# A basic TOML string cannot carry a literal newline or control character,
|
|
266
|
+
# and a key is a string too: an unquoted `proxy.v1` would silently become a
|
|
267
|
+
# dotted path and rewrite the provider map this run is supposed to copy.
|
|
268
|
+
out = []
|
|
269
|
+
for character in value:
|
|
270
|
+
if character in ESCAPES:
|
|
271
|
+
out.append(ESCAPES[character])
|
|
272
|
+
elif ord(character) < 0x20 or ord(character) == 0x7F:
|
|
273
|
+
out.append("\\u%04X" % ord(character))
|
|
274
|
+
else:
|
|
275
|
+
out.append(character)
|
|
276
|
+
return '"' + "".join(out) + '"'
|
|
277
|
+
|
|
278
|
+
|
|
279
|
+
def render(value):
|
|
280
|
+
if isinstance(value, bool):
|
|
281
|
+
return "true" if value else "false"
|
|
282
|
+
if isinstance(value, (int, float)):
|
|
283
|
+
return repr(value)
|
|
284
|
+
if isinstance(value, str):
|
|
285
|
+
return render_string(value)
|
|
286
|
+
if isinstance(value, list):
|
|
287
|
+
return "[" + ", ".join(render(item) for item in value) + "]"
|
|
288
|
+
if isinstance(value, dict):
|
|
289
|
+
return "{" + ", ".join(
|
|
290
|
+
f"{render_string(key)} = {render(item)}" for key, item in value.items()
|
|
291
|
+
) + "}"
|
|
292
|
+
raise TypeError(value)
|
|
293
|
+
|
|
294
|
+
|
|
295
|
+
# A profile selects the model on many hosts, and the profile table itself is
|
|
296
|
+
# not copied: it can carry approval, sandbox, or server settings this run must
|
|
297
|
+
# not inherit. Resolve the selected profile's model identity into top-level
|
|
298
|
+
# keys instead, so a profile-configured host keeps its own model rather than
|
|
299
|
+
# silently falling back to the CLI default.
|
|
300
|
+
resolved = {key: source[key] for key in PREFERENCE_KEYS if key in source}
|
|
301
|
+
selected = source.get("profile")
|
|
302
|
+
if selected is not None:
|
|
303
|
+
# A selected profile that cannot be resolved is refused, not skipped:
|
|
304
|
+
# falling through would run the review on a different model than the host
|
|
305
|
+
# explicitly asked for, which is the substitution this carry-over exists to
|
|
306
|
+
# prevent.
|
|
307
|
+
profiles = source.get("profiles")
|
|
308
|
+
profile = profiles.get(selected) if isinstance(profiles, dict) and isinstance(selected, str) else None
|
|
309
|
+
if not isinstance(selected, str) or not selected or not isinstance(profile, dict):
|
|
310
|
+
sys.exit(1)
|
|
311
|
+
for key in PREFERENCE_KEYS:
|
|
312
|
+
if key in profile:
|
|
313
|
+
resolved[key] = profile[key]
|
|
314
|
+
lines = []
|
|
315
|
+
try:
|
|
316
|
+
for key in PREFERENCE_KEYS:
|
|
317
|
+
if key in resolved:
|
|
318
|
+
lines.append(f"{render_string(key)} = {render(resolved[key])}")
|
|
319
|
+
except TypeError:
|
|
320
|
+
sys.exit(1)
|
|
321
|
+
Path(sys.argv[2]).write_text("".join(line + "\n" for line in lines), encoding="utf-8")
|
|
322
|
+
PY_HOME_PREFERENCES
|
|
323
|
+
chmod 0600 "$RUNTIME_HOME/config.toml" 2>/dev/null || true
|
|
324
|
+
fi
|
|
325
|
+
if [ "$REVIEW_SKILL_COUNT" -gt 0 ]; then
|
|
326
|
+
# Copied, not linked: the CLI does not follow a symlinked skill directory,
|
|
327
|
+
# so a link here would silently cost the owner-skill binding.
|
|
328
|
+
mkdir -m 700 "$RUNTIME_HOME/skills" \
|
|
329
|
+
|| die_inconclusive runtime_home_unavailable local_tool_failure false
|
|
330
|
+
for review_skill in "${REVIEW_SKILLS[@]}"; do
|
|
331
|
+
cp -R "$INSTALLED_SKILL_REGISTRY_ROOT/$review_skill" "$RUNTIME_HOME/skills/$review_skill" \
|
|
332
|
+
|| die_inconclusive codex_installed_skill_unavailable capability_missing true
|
|
333
|
+
done
|
|
334
|
+
fi
|
|
199
335
|
MODEL=""
|
|
200
336
|
PROVIDER="openai"
|
|
201
337
|
FAMILY="openai"
|
|
@@ -261,34 +397,33 @@ import json, sys
|
|
|
261
397
|
values = [sys.argv[2], "--packet", sys.argv[1], "--sha256", sys.argv[3], "--allow-search"]
|
|
262
398
|
print('mcp_servers={code_review_packet={command=' + json.dumps(sys.executable)
|
|
263
399
|
+ ',args=[' + ','.join(json.dumps(value) for value in values)
|
|
264
|
-
+ '],enabled=true,enabled_tools=["read_packet","search_packet"]
|
|
400
|
+
+ '],enabled=true,enabled_tools=["read_packet","search_packet"]'
|
|
401
|
+
+ ',default_tools_approval_mode="approve"}}')
|
|
265
402
|
PY_MCP_CONFIG
|
|
266
403
|
)" || die_inconclusive packet_config_failed local_tool_failure false
|
|
267
|
-
#
|
|
268
|
-
#
|
|
404
|
+
# Inherited MCP servers are data, not a boundary. Disabling them by name was
|
|
405
|
+
# tried and cannot work: a plugin contributes its server outside `mcp_servers`,
|
|
406
|
+
# so `mcp_servers.<name>={enabled=false}` builds a transportless entry and the
|
|
407
|
+
# CLI refuses the whole configuration -- while leaving it enabled failed an
|
|
408
|
+
# exactly-one-server count. Either branch dead-ended the lane before inference.
|
|
409
|
+
# So this preflight verifies only that the frozen packet server is present and
|
|
410
|
+
# bound to the exact interpreter, script, packet and digest this run created.
|
|
411
|
+
#
|
|
412
|
+
# Accepted residual, measured rather than assumed: other servers stay enabled
|
|
413
|
+
# and CAN execute during a review. `audit_codex` refuses a verdict from any
|
|
414
|
+
# stream containing a foreign mcp_tool_call, but it runs afterwards -- the call
|
|
415
|
+
# has already completed, and a remote write or send cannot be undone by
|
|
416
|
+
# rejecting the verdict. Auto-approval is not a defence either: a server opts
|
|
417
|
+
# itself in by declaring `readOnlyHint` on a tool, which the CLI trusts, so a
|
|
418
|
+
# tool that executes arbitrary code can be auto-approved while claiming to be
|
|
419
|
+
# read-only. Two containment routes that name no server were measured and both
|
|
420
|
+
# failed: a global `apps._default.default_tools_approval_mode` did not override
|
|
421
|
+
# the hint, and `--disable plugins` would disable the reviewer's own installed
|
|
422
|
+
# skill registry, which ships as a plugin. The owner accepted this residual for
|
|
423
|
+
# this round; the route that would close it is a private CODEX_HOME seeded with
|
|
424
|
+
# auth and the registry only, as the Kimi lane already does.
|
|
269
425
|
CODEX_PACKET_CONFIG=(-c "$MCP_CONFIG" -c 'web_search="disabled"' -c 'approval_policy="never"')
|
|
270
|
-
CODEX_HOME="$
|
|
271
|
-
|| die_inconclusive codex_packet_tools_unavailable capability_missing true
|
|
272
|
-
MCP_CONFIG="$(python3 - "$RUN_ROOT/mcp.json" "$MCP_CONFIG" <<'PY_MCP_OVERRIDES'
|
|
273
|
-
import json, sys
|
|
274
|
-
from pathlib import Path
|
|
275
|
-
try:
|
|
276
|
-
rows = json.loads(Path(sys.argv[1]).read_text())
|
|
277
|
-
if not isinstance(rows, list):
|
|
278
|
-
raise ValueError()
|
|
279
|
-
disabled = []
|
|
280
|
-
for row in rows:
|
|
281
|
-
if not isinstance(row, dict) or not isinstance(row.get("name"), str) or not row["name"]:
|
|
282
|
-
raise ValueError()
|
|
283
|
-
if row["name"] != "code_review_packet":
|
|
284
|
-
disabled.append(json.dumps(row["name"]) + "={enabled=false}")
|
|
285
|
-
print(sys.argv[2][:-1] + "".join("," + entry for entry in disabled) + "}")
|
|
286
|
-
except (OSError, ValueError, TypeError):
|
|
287
|
-
sys.exit(1)
|
|
288
|
-
PY_MCP_OVERRIDES
|
|
289
|
-
)" || die_inconclusive codex_packet_tools_unavailable capability_missing true
|
|
290
|
-
CODEX_PACKET_CONFIG=(-c "$MCP_CONFIG" -c 'web_search="disabled"' -c 'approval_policy="never"')
|
|
291
|
-
CODEX_HOME="$SOURCE_HOME" timeout --kill-after=1s 5s "$CODEX_BIN_PATH" mcp list --json "${CODEX_PACKET_CONFIG[@]}" >"$RUN_ROOT/mcp.json" 2>"$STDERR_FILE" \
|
|
426
|
+
CODEX_HOME="$RUNTIME_HOME" timeout --kill-after=1s 5s "$CODEX_BIN_PATH" mcp list --json "${CODEX_PACKET_CONFIG[@]}" >"$RUN_ROOT/mcp.json" 2>"$STDERR_FILE" \
|
|
292
427
|
|| die_inconclusive codex_packet_tools_unavailable capability_missing true
|
|
293
428
|
python3 - "$RUN_ROOT/mcp.json" "$PACKET_FILE" "$PACKET_SERVER" "$PACKET_SHA256" <<'PY_MCP_CHECK' \
|
|
294
429
|
|| die_inconclusive codex_packet_tools_unavailable capability_missing true
|
|
@@ -298,12 +433,30 @@ try:
|
|
|
298
433
|
rows = json.loads(Path(sys.argv[1]).read_text())
|
|
299
434
|
if not isinstance(rows, list):
|
|
300
435
|
raise ValueError()
|
|
301
|
-
|
|
302
|
-
|
|
436
|
+
# Every row is validated before any filtering. Dropping the old enumeration
|
|
437
|
+
# also dropped its per-row name check, which let a malformed reply through
|
|
438
|
+
# whenever the malformed row happened to be disabled.
|
|
439
|
+
if any(
|
|
440
|
+
not isinstance(row, dict)
|
|
441
|
+
or not isinstance(row.get("name"), str)
|
|
442
|
+
or not row["name"]
|
|
443
|
+
for row in rows
|
|
444
|
+
):
|
|
445
|
+
raise ValueError()
|
|
446
|
+
# Under the private home this is an invariant the run establishes, not a
|
|
447
|
+
# bet on the user's configuration: nothing else was ever there to enable.
|
|
448
|
+
# A second enabled server means the home leaked, so refuse.
|
|
449
|
+
# Two predicates, not one: exactly one row carries the packet name
|
|
450
|
+
# anywhere in the reply, and exactly one row is enabled at all. Checking
|
|
451
|
+
# only the enabled set would accept a correctly bound row beside a disabled
|
|
452
|
+
# duplicate of the same name.
|
|
453
|
+
named = [row for row in rows if row.get("name") == "code_review_packet"]
|
|
454
|
+
active = [row for row in rows if row.get("enabled") is not False]
|
|
455
|
+
if len(named) != 1 or len(active) != 1 or active[0] is not named[0]:
|
|
303
456
|
raise ValueError()
|
|
304
457
|
row = active[0]
|
|
305
458
|
transport = row.get("transport", {})
|
|
306
|
-
if (row.get("
|
|
459
|
+
if (row.get("enabled") is not True
|
|
307
460
|
or transport.get("type") != "stdio" or transport.get("command") != sys.executable
|
|
308
461
|
or transport.get("args") != [sys.argv[3], "--packet", sys.argv[2], "--sha256", sys.argv[4], "--allow-search"]
|
|
309
462
|
or transport.get("env") or transport.get("env_vars") or transport.get("cwd")):
|
|
@@ -350,12 +503,17 @@ JSON
|
|
|
350
503
|
fi
|
|
351
504
|
|
|
352
505
|
run_started=$SECONDS
|
|
353
|
-
CMUX_CODEX_HOOKS_DISABLED=1 CODEX_HOME="$
|
|
506
|
+
CMUX_CODEX_HOOKS_DISABLED=1 CODEX_HOME="$RUNTIME_HOME" timeout --kill-after=1s "${TIMEOUT}s" "$CODEX_BIN_PATH" exec --disable hooks --disable shell_tool --sandbox read-only --ephemeral --skip-git-repo-check \
|
|
354
507
|
"${CODEX_PACKET_CONFIG[@]}" \
|
|
355
508
|
--json --output-schema "$SCHEMA_FILE" --output-last-message "$RESULT_FILE" \
|
|
356
509
|
-C "$RUN_WORKSPACE" - <"$PROMPT_FILE" >"$EVENTS" 2>"$STDERR_FILE"
|
|
357
510
|
run_rc=$?
|
|
358
511
|
run_elapsed=$((SECONDS - run_started))
|
|
512
|
+
if [ -n "$AUTH_LINK_TARGET" ]; then
|
|
513
|
+
[ -L "$RUNTIME_HOME/auth.json" ] \
|
|
514
|
+
&& [ "$(readlink "$RUNTIME_HOME/auth.json")" = "$AUTH_LINK_TARGET" ] \
|
|
515
|
+
|| die_inconclusive codex_runtime_home_credential_moved binding_mismatch false
|
|
516
|
+
fi
|
|
359
517
|
if [ "$run_rc" != 0 ]; then
|
|
360
518
|
if bash "$TIMEOUT_CLASSIFIER" "$run_rc" "$run_elapsed" "$TIMEOUT"; then
|
|
361
519
|
die_inconclusive codex_timeout timeout true "$run_rc"
|
|
@@ -698,18 +698,31 @@ if [ "${1:-}" = mcp ] && [ "${2:-}" = list ]; then
|
|
|
698
698
|
fi
|
|
699
699
|
python3 - "$@" <<'PY_MCP_LIST'
|
|
700
700
|
import json, os, sys, tomllib
|
|
701
|
+
from pathlib import Path
|
|
701
702
|
|
|
702
703
|
arguments = iter(sys.argv[1:])
|
|
703
704
|
servers = {}
|
|
704
|
-
|
|
705
|
-
|
|
706
|
-
|
|
707
|
-
|
|
708
|
-
|
|
709
|
-
|
|
710
|
-
|
|
711
|
-
if
|
|
712
|
-
|
|
705
|
+
plugin_sourced = set()
|
|
706
|
+
# Inherited servers come from the home the CLI was actually handed, never from
|
|
707
|
+
# the test's own environment: whether the reviewer sees a user's server is
|
|
708
|
+
# exactly the question, and a stub that injects one regardless of the home
|
|
709
|
+
# could not tell a private home from the user's.
|
|
710
|
+
home = Path(os.environ.get("CODEX_HOME", ""))
|
|
711
|
+
home_config = home / "config.toml" if os.environ.get("CODEX_HOME") else None
|
|
712
|
+
if home_config is not None and home_config.exists():
|
|
713
|
+
try:
|
|
714
|
+
home_data = tomllib.loads(home_config.read_text(encoding="utf-8"))
|
|
715
|
+
except (OSError, UnicodeError, tomllib.TOMLDecodeError):
|
|
716
|
+
home_data = {}
|
|
717
|
+
for inherited_name, inherited in (home_data.get("mcp_servers") or {}).items():
|
|
718
|
+
if isinstance(inherited, dict):
|
|
719
|
+
servers[inherited_name] = dict(inherited)
|
|
720
|
+
servers[inherited_name].setdefault("enabled", True)
|
|
721
|
+
# A plugin contributes its server outside `mcp_servers`, which is why an
|
|
722
|
+
# override under that table cannot reach it.
|
|
723
|
+
if os.environ.get("CODEX_HOME") and (home / "plugins" / "provided.json").exists():
|
|
724
|
+
servers["plugin_provided"] = {"command": "/bin/false", "args": [], "enabled": True}
|
|
725
|
+
plugin_sourced.add("plugin_provided")
|
|
713
726
|
for argument in arguments:
|
|
714
727
|
if argument in {"-c", "--config"}:
|
|
715
728
|
# Codex splits the override path separately from its TOML value;
|
|
@@ -724,18 +737,76 @@ for argument in arguments:
|
|
|
724
737
|
else:
|
|
725
738
|
continue
|
|
726
739
|
for name, settings in configured_servers.items():
|
|
740
|
+
if name in plugin_sourced:
|
|
741
|
+
# A plugin contributes its server outside `mcp_servers`, so an
|
|
742
|
+
# override under that table builds an entry with no transport
|
|
743
|
+
# and the CLI refuses to load the configuration at all.
|
|
744
|
+
sys.stderr.write(
|
|
745
|
+
"Error: failed to load bootstrap configuration\n\n"
|
|
746
|
+
"Caused by:\n invalid transport\n"
|
|
747
|
+
" in `mcp_servers." + name + "`\n")
|
|
748
|
+
sys.exit(1)
|
|
727
749
|
servers.setdefault(name, {}).update(settings)
|
|
728
|
-
|
|
750
|
+
rows = [
|
|
729
751
|
{"name": name, "enabled": server.get("enabled", True),
|
|
730
752
|
"transport": {"type": "stdio", "command": server["command"],
|
|
731
753
|
"args": server.get("args", [])}}
|
|
732
754
|
for name, server in servers.items()
|
|
733
|
-
]
|
|
755
|
+
]
|
|
756
|
+
if os.environ.get("STUB_MALFORMED_DISABLED_ROW"):
|
|
757
|
+
# The malformed row is disabled: a check that filters before validating
|
|
758
|
+
# would never look at it.
|
|
759
|
+
bad = {"missing": {}, "empty": {"name": ""}, "nonstring": {"name": 7}}[
|
|
760
|
+
os.environ["STUB_MALFORMED_DISABLED_ROW"]]
|
|
761
|
+
rows.append({**bad, "enabled": False,
|
|
762
|
+
"transport": {"type": "stdio", "command": "/bin/false", "args": []}})
|
|
763
|
+
if os.environ.get("STUB_DUPLICATE_PACKET_ROW") == "1":
|
|
764
|
+
# A JSON array can carry the same name twice; a dict of servers cannot.
|
|
765
|
+
rows.append({"name": "code_review_packet", "enabled": False,
|
|
766
|
+
"transport": {"type": "stdio", "command": "/bin/false", "args": []}})
|
|
767
|
+
print(json.dumps(rows))
|
|
734
768
|
PY_MCP_LIST
|
|
735
769
|
exit $?
|
|
736
770
|
fi
|
|
737
771
|
touch "$state/codex_invoked"
|
|
738
772
|
printf '%s' "$0" >"$state/codex_argv0"
|
|
773
|
+
printf '%s' "${CODEX_HOME:-}" >"$state/codex_home"
|
|
774
|
+
if [ "${STUB_REPLACE_AUTH_LINK:-}" = 1 ] && [ -n "${CODEX_HOME:-}" ]; then
|
|
775
|
+
rm -f "$CODEX_HOME/auth.json"
|
|
776
|
+
printf '%s\n' '{"tokens":{"access":"rotated"}}' >"$CODEX_HOME/auth.json"
|
|
777
|
+
fi
|
|
778
|
+
python3 - "${CODEX_HOME:-}" "$state/codex_home_shape" <<'PY_HOME_SHAPE'
|
|
779
|
+
import json, os, sys, tomllib
|
|
780
|
+
from pathlib import Path
|
|
781
|
+
|
|
782
|
+
# The wrapper deletes its run directory on exit, so the private home can only
|
|
783
|
+
# be inspected from inside the run.
|
|
784
|
+
home = Path(sys.argv[1]) if sys.argv[1] else None
|
|
785
|
+
shape = {"home": sys.argv[1], "config_keys": [], "has_mcp_servers": None,
|
|
786
|
+
"has_plugins": None, "auth_link": None}
|
|
787
|
+
if home is not None:
|
|
788
|
+
config = home / "config.toml"
|
|
789
|
+
if config.exists():
|
|
790
|
+
try:
|
|
791
|
+
data = tomllib.loads(config.read_text(encoding="utf-8"))
|
|
792
|
+
except (OSError, UnicodeError, tomllib.TOMLDecodeError):
|
|
793
|
+
data = {"__unreadable__": True}
|
|
794
|
+
shape["config_keys"] = sorted(data)
|
|
795
|
+
shape["has_mcp_servers"] = "mcp_servers" in data
|
|
796
|
+
shape["model"] = data.get("model")
|
|
797
|
+
shape["model_reasoning_effort"] = data.get("model_reasoning_effort")
|
|
798
|
+
providers = data.get("model_providers")
|
|
799
|
+
shape["provider_keys"] = sorted(providers) if isinstance(providers, dict) else None
|
|
800
|
+
else:
|
|
801
|
+
shape["has_mcp_servers"] = False
|
|
802
|
+
shape["has_plugins"] = (home / "plugins").exists()
|
|
803
|
+
auth = home / "auth.json"
|
|
804
|
+
if auth.is_symlink():
|
|
805
|
+
shape["auth_link"] = os.readlink(auth)
|
|
806
|
+
elif auth.exists():
|
|
807
|
+
shape["auth_link"] = "__regular_file__"
|
|
808
|
+
Path(sys.argv[2]).write_text(json.dumps(shape), encoding="utf-8")
|
|
809
|
+
PY_HOME_SHAPE
|
|
739
810
|
printf '%s' "${CMUX_CODEX_HOOKS_DISABLED:-}" >"$state/codex_cmux_hooks_disabled"
|
|
740
811
|
last_message=""
|
|
741
812
|
has_model=no
|
|
@@ -1802,16 +1873,29 @@ out="$(run_codex packet_tampered)"; rc=$?
|
|
|
1802
1873
|
check "Codex rejects altered packet tool bytes through wrapper and parser" \
|
|
1803
1874
|
'[ "$rc" = 2 ] && [ "$(field reason_code "$out")" = binding_mismatch ] && [ "$(field cascade_eligible "$out")" = False ]'
|
|
1804
1875
|
|
|
1805
|
-
|
|
1806
|
-
|
|
1807
|
-
|
|
1808
|
-
|
|
1809
|
-
|
|
1810
|
-
|
|
1811
|
-
|
|
1876
|
+
# A user's own server, however its name is spelled, must not reach the reviewer
|
|
1877
|
+
# and must not be touched: the wrapper neither disables it nor names it.
|
|
1878
|
+
inherited_index=0
|
|
1879
|
+
for inherited_mcp_name in 'unrelated' 'unrelated.name' 'unrelated name' 'unrelated"name'; do
|
|
1880
|
+
inherited_index=$((inherited_index + 1))
|
|
1881
|
+
inherited_source="$WORK/codex-inherited-$inherited_index"
|
|
1882
|
+
mkdir -p "$inherited_source"
|
|
1883
|
+
python3 - "$inherited_source/config.toml" "$inherited_mcp_name" <<'PY_INHERITED_SOURCE'
|
|
1884
|
+
import json, sys
|
|
1885
|
+
from pathlib import Path
|
|
1886
|
+
|
|
1887
|
+
Path(sys.argv[1]).write_text(
|
|
1888
|
+
"[mcp_servers]\n"
|
|
1889
|
+
+ json.dumps(sys.argv[2])
|
|
1890
|
+
+ ' = { command = "/bin/false", args = [] }\n',
|
|
1891
|
+
encoding="utf-8",
|
|
1892
|
+
)
|
|
1893
|
+
PY_INHERITED_SOURCE
|
|
1894
|
+
printf '%s\n' '{"tokens":{"access":"seeded"}}' >"$inherited_source/auth.json"
|
|
1895
|
+
chmod 0600 "$inherited_source/auth.json"
|
|
1812
1896
|
rm -f "$WORK/state/codex_invoked" "$WORK/state/codex_configs"
|
|
1813
|
-
out="$(run_codex "$
|
|
1814
|
-
|
|
1897
|
+
out="$(run_codex pass claude "$inherited_source")"; rc=$?
|
|
1898
|
+
inherited_mcp_untouched="$(python3 - "$inherited_mcp_name" "$WORK/state/codex_configs" <<'PY_MCP_UNTOUCHED'
|
|
1815
1899
|
import sys, tomllib
|
|
1816
1900
|
from pathlib import Path
|
|
1817
1901
|
|
|
@@ -1821,15 +1905,140 @@ for override in path.read_text().splitlines() if path.exists() else []:
|
|
|
1821
1905
|
key, value = override.split("=", 1)
|
|
1822
1906
|
if key.strip() == "mcp_servers":
|
|
1823
1907
|
tables.append(tomllib.loads("servers=" + value)["servers"])
|
|
1908
|
+
elif key.strip().startswith("mcp_servers."):
|
|
1909
|
+
tables.append({key.strip().split(".")[1]: {}})
|
|
1824
1910
|
print(len(tables) == 1
|
|
1825
|
-
and
|
|
1826
|
-
and
|
|
1827
|
-
|
|
1911
|
+
and sys.argv[1] not in tables[0]
|
|
1912
|
+
and list(tables[0]) == ["code_review_packet"])
|
|
1913
|
+
PY_MCP_UNTOUCHED
|
|
1914
|
+
)"
|
|
1915
|
+
check "Codex never reaches or names the user's own MCP server ($inherited_mcp_name)" \
|
|
1916
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ -e "$WORK/state/codex_invoked" ] && [ "$inherited_mcp_untouched" = True ]'
|
|
1917
|
+
done
|
|
1918
|
+
|
|
1919
|
+
# The shape that used to dead-end the preflight: a server contributed outside
|
|
1920
|
+
# `mcp_servers`, which no override under that table can reach.
|
|
1921
|
+
plugin_source="$WORK/codex-plugin-source"
|
|
1922
|
+
mkdir -p "$plugin_source/plugins"
|
|
1923
|
+
printf '%s\n' '{"server":"plugin_provided"}' >"$plugin_source/plugins/provided.json"
|
|
1924
|
+
printf '%s\n' '{"tokens":{"access":"seeded"}}' >"$plugin_source/auth.json"
|
|
1925
|
+
chmod 0600 "$plugin_source/auth.json"
|
|
1926
|
+
rm -f "$WORK/state/codex_invoked" "$WORK/state/codex_configs"
|
|
1927
|
+
out="$(run_codex pass claude "$plugin_source")"; rc=$?
|
|
1928
|
+
check "Codex reviews on a host whose plugin contributes an MCP server" \
|
|
1929
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ -e "$WORK/state/codex_invoked" ]'
|
|
1930
|
+
|
|
1931
|
+
# A canned verdict cannot prove the repair: a CLI that accepts the configuration
|
|
1932
|
+
# and still refuses read_packet would pass every assertion above. This one
|
|
1933
|
+
# replays the real packet server's bytes through the parser on that same host.
|
|
1934
|
+
rm -f "$WORK/state/codex_invoked" "$WORK/state/codex_packet_call"
|
|
1935
|
+
out="$(run_codex packet_read claude "$plugin_source")"; rc=$?
|
|
1936
|
+
check "Codex completes a real frozen packet read on a plugin-contributing host" \
|
|
1937
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ "$(cat "$WORK/state/codex_packet_call")" = read_packet ]'
|
|
1938
|
+
|
|
1939
|
+
rm -f "$WORK/state/codex_invoked" "$WORK/state/codex_home"
|
|
1940
|
+
mkdir -p "$WORK/codex-foreign-source"
|
|
1941
|
+
printf '%s\n' 'profile = "review"' '' '[profiles.review]' 'model = "seeded-model"' 'model_reasoning_effort = "xhigh"' \
|
|
1942
|
+
'' '[model_providers."proxy.v1"]' 'name = "line one\\nline two"' '' '[mcp_servers.foreign]' 'command = "/bin/false"' 'args = []' \
|
|
1943
|
+
>"$WORK/codex-foreign-source/config.toml"
|
|
1944
|
+
printf '%s\n' '{"tokens":{"access":"seeded"}}' >"$WORK/codex-foreign-source/auth.json"
|
|
1945
|
+
chmod 0600 "$WORK/codex-foreign-source/auth.json"
|
|
1946
|
+
out="$(run_codex pass claude "$WORK/codex-foreign-source")"; rc=$?
|
|
1947
|
+
private_home_shape="$(python3 - "$WORK/state/codex_home_shape" "$WORK/codex-foreign-source" <<'PY_PRIVATE_HOME'
|
|
1948
|
+
import json, os, sys
|
|
1949
|
+
from pathlib import Path
|
|
1950
|
+
|
|
1951
|
+
|
|
1952
|
+
def normalized(value):
|
|
1953
|
+
# TMPDIR here can carry a trailing separator, so the recorded and expected
|
|
1954
|
+
# paths differ as strings while naming the same file.
|
|
1955
|
+
return Path(os.path.normpath(str(value))) if value else None
|
|
1956
|
+
|
|
1957
|
+
|
|
1958
|
+
path = Path(sys.argv[1])
|
|
1959
|
+
if not path.exists():
|
|
1960
|
+
print("no-run-observed"); raise SystemExit
|
|
1961
|
+
shape = json.loads(path.read_text(encoding="utf-8"))
|
|
1962
|
+
source = normalized(sys.argv[2])
|
|
1963
|
+
home = normalized(shape.get("home"))
|
|
1964
|
+
if home is None or home == source or source in home.parents:
|
|
1965
|
+
print("shares-user-home"); raise SystemExit
|
|
1966
|
+
if shape.get("has_mcp_servers") is not False or shape.get("has_plugins") is not False:
|
|
1967
|
+
print("inherited-servers-present"); raise SystemExit
|
|
1968
|
+
if shape.get("model") != "seeded-model" or shape.get("model_reasoning_effort") != "xhigh":
|
|
1969
|
+
print("model-preference-lost"); raise SystemExit
|
|
1970
|
+
if shape.get("provider_keys") != ["proxy.v1"]:
|
|
1971
|
+
print("provider-key-mangled"); raise SystemExit
|
|
1972
|
+
if normalized(shape.get("auth_link")) != source / "auth.json":
|
|
1973
|
+
print("credential-not-linked"); raise SystemExit
|
|
1974
|
+
print("private")
|
|
1975
|
+
PY_PRIVATE_HOME
|
|
1828
1976
|
)"
|
|
1829
|
-
|
|
1830
|
-
|
|
1977
|
+
check "Codex reviews from a private home that never carried the user's MCP servers" \
|
|
1978
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ "$private_home_shape" = private ]'
|
|
1979
|
+
|
|
1980
|
+
# The link back to the user's credential is what keeps a rotated token in the
|
|
1981
|
+
# user's own file; a CLI that replaced it with a regular file would leave the
|
|
1982
|
+
# credential inside this run directory instead. Prove the post-run check fires.
|
|
1983
|
+
rm -f "$WORK/state/codex_invoked"
|
|
1984
|
+
# A host that names a profile it does not define must not quietly review on a
|
|
1985
|
+
# different model; the wrapper refuses instead of falling through.
|
|
1986
|
+
mkdir -p "$WORK/codex-broken-profile"
|
|
1987
|
+
printf '%s\n' 'profile = "missing"' 'model = "top-level-model"' >"$WORK/codex-broken-profile/config.toml"
|
|
1988
|
+
printf '%s\n' '{"tokens":{"access":"seeded"}}' >"$WORK/codex-broken-profile/auth.json"
|
|
1989
|
+
chmod 0600 "$WORK/codex-broken-profile/auth.json"
|
|
1990
|
+
rm -f "$WORK/state/codex_invoked"
|
|
1991
|
+
out="$(run_codex pass claude "$WORK/codex-broken-profile")"; rc=$?
|
|
1992
|
+
check "Codex refuses a selected profile it cannot resolve instead of substituting a model" \
|
|
1993
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = codex_home_preferences_unreadable ] && [ ! -e "$WORK/state/codex_invoked" ]'
|
|
1994
|
+
|
|
1995
|
+
# A malformed row that happens to be disabled must still fail the preflight;
|
|
1996
|
+
# filtering before validating would step over it.
|
|
1997
|
+
for malformed in missing empty nonstring; do
|
|
1998
|
+
export STUB_MALFORMED_DISABLED_ROW="$malformed"
|
|
1999
|
+
rm -f "$WORK/state/codex_invoked"
|
|
2000
|
+
out="$(run_codex pass)"; rc=$?
|
|
2001
|
+
unset STUB_MALFORMED_DISABLED_ROW
|
|
2002
|
+
check "Codex refuses a malformed disabled MCP row ($malformed name)" \
|
|
2003
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = codex_packet_tools_unavailable ] && [ ! -e "$WORK/state/codex_invoked" ]'
|
|
1831
2004
|
done
|
|
1832
2005
|
|
|
2006
|
+
# The enabled-set check alone would accept one correctly bound row beside a
|
|
2007
|
+
# disabled duplicate of the same name, so the uniqueness claim needs its own case.
|
|
2008
|
+
export STUB_DUPLICATE_PACKET_ROW=1
|
|
2009
|
+
rm -f "$WORK/state/codex_invoked"
|
|
2010
|
+
out="$(run_codex pass)"; rc=$?
|
|
2011
|
+
unset STUB_DUPLICATE_PACKET_ROW
|
|
2012
|
+
check "Codex refuses a duplicated packet-server row even when the duplicate is disabled" \
|
|
2013
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = codex_packet_tools_unavailable ] && [ ! -e "$WORK/state/codex_invoked" ]'
|
|
2014
|
+
|
|
2015
|
+
export STUB_REPLACE_AUTH_LINK=1
|
|
2016
|
+
out="$(run_codex pass claude "$WORK/codex-foreign-source")"; rc=$?
|
|
2017
|
+
unset STUB_REPLACE_AUTH_LINK
|
|
2018
|
+
check "Codex refuses when the run replaced the linked credential with a file" \
|
|
2019
|
+
'[ "$rc" = 2 ] && [ "$(field reason "$out")" = codex_runtime_home_credential_moved ] && [ "$(field reason_code "$out")" = binding_mismatch ] && [ "$(field cascade_eligible "$out")" = False ]'
|
|
2020
|
+
|
|
2021
|
+
rm -f "$WORK/state/codex_invoked" "$WORK/state/codex_configs"
|
|
2022
|
+
out="$(run_codex pass)"; rc=$?
|
|
2023
|
+
packet_approval_mode="$(python3 - "$WORK/state/codex_configs" <<'PY_MCP_APPROVAL'
|
|
2024
|
+
import sys, tomllib
|
|
2025
|
+
from pathlib import Path
|
|
2026
|
+
|
|
2027
|
+
path = Path(sys.argv[1])
|
|
2028
|
+
servers = {}
|
|
2029
|
+
for override in path.read_text().splitlines() if path.exists() else []:
|
|
2030
|
+
key, value = override.split("=", 1)
|
|
2031
|
+
if key.strip() == "mcp_servers":
|
|
2032
|
+
servers = tomllib.loads("servers=" + value)["servers"]
|
|
2033
|
+
packet = servers.get("code_review_packet", {})
|
|
2034
|
+
print(packet.get("default_tools_approval_mode") == "approve"
|
|
2035
|
+
and all(other.get("default_tools_approval_mode") is None
|
|
2036
|
+
for name, other in servers.items() if name != "code_review_packet"))
|
|
2037
|
+
PY_MCP_APPROVAL
|
|
2038
|
+
)"
|
|
2039
|
+
check "Codex declares the packet server auto-approved without widening the sandbox" \
|
|
2040
|
+
'[ "$rc" = 0 ] && [ "$(field status "$out")" = passed ] && [ "$packet_approval_mode" = True ] && [ "$(cat "$WORK/state/codex_read_only")" = yes ]'
|
|
2041
|
+
|
|
1833
2042
|
rm -f "$WORK/state/codex_invoked" "$WORK/state/codex_help_invoked"
|
|
1834
2043
|
probe_started=$SECONDS
|
|
1835
2044
|
out="$(run_codex help_hang)"; rc=$?
|
|
@@ -194,7 +194,7 @@ Run this gate before finalizing a product R&D turn after any delivery slice land
|
|
|
194
194
|
- **Deferred-evidence continuation check (`DFE-CONT`).** When real/runtime evidence is due (named by an acceptance item, status source, landing-evidence row, required gate, user correction, or because it is the behavior's only meaningful proof) yet deferred, blocked after remediation, skipped at finalization, or replaced by local/mock verification. Report deferred real evidence as `interim`/outstanding; do NOT report the turn complete while it is outstanding. A local/mock substitution is terminal only when a cited **non-agent** anchor — **agent-authored or agent-co-edited status/router/gate/handoff text never satisfies this** — names the same evidence, declares the deferral terminal, and carries the outstanding command/source forward for the active slice/ref. Never add verifier/config/test hardening motivated only by missing deferred evidence; never auto-continue past the pending gate. **Load `references/pre-final-continuation-gate.md` before treating any deferral as terminal** — it owns the valid/invalid-anchor list and hardening boundary.
|
|
195
195
|
- **Affirmative-assent binding rule** lives in `references/pre-final-continuation-gate.md` §Assent binding — load it when recovering a short reply. Bind to the current explicit request or one recoverable concrete proposal, including an unmarked proposal; preserve its scope and existing authority. Ask only if action, scope, or required authority remains unresolved after recovery. A status remark or output marker cannot substitute for a proposal or permission; self-classifying the reply or marker away is never an exit from carrying out an already-clear request.
|
|
196
196
|
3. Continue automatically with a clearly owned, verifiable, low-risk next slice from an explicit task/status/acceptance source or active user continuation, within accepted scope and existing authority. Apply the eligibility and stop conditions in `references/pre-final-continuation-gate.md`. Necessary fixes, tests and review inherit task authorization; a reviewer-budget flag triggers a method checkpoint and cumulative-history record, not renewed permission. Explicit user limits still govern. Existing configured internal developer-self-use metered model/tool accounts aren't an external purchase here.
|
|
197
|
-
4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer,
|
|
197
|
+
4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, a concrete blocker for the affected action, or no safe authorized work remains. Block materially differing viable approaches (none dominant-and-reversible) and a fix lacking evidenced cause; load `references/pre-final-continuation-gate.md` for the full stop conditions. **Scope each blocker to its dependent action or claim.** An unproven cause blocks the speculative patch, not available diagnosis; a pending gate blocks dependent landing/completion, not authorized remediation or independent work. Before ending, perform in-scope diagnosis, owner discovery, remediation or independent work, and poll any finite step you started to its result, never reporting it as running. Quality-gate failures require diagnosis and available related behavior-preserving cleanup before escalation; preserve readability and compatibility, never game counters (`references/refactoring-discipline.md`). Never bypass the blocked gate, invent a pass, widen scope, or substitute unrelated hardening. With one dominant reversible approach and no applicable stop condition, do not stop at a recommendation: deliver a tested reviewable draft.
|
|
198
198
|
5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form and classifies the turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
|
|
199
199
|
6. **Assent-outcome closeout check.** Every user reply immediately following an assistant message that states or implies a next action requires a visible `continuing:` or `blocked:` outcome before finalizing, even if the reply is not classified as assent; every explicit continuation request does too. Missing markers never waive it. Reconcile the current request, original proposal, scope/authority changes, tool/output evidence, and remaining blockers. Respect a current explicit stop or status-only request; name that reason in the blocked outcome without executing the prior proposal. Otherwise `continuing:` must be followed by execution in the same turn; a promised next step is not execution. If part remains blocked, report its pending state and independent work performed. A status-only handoff cannot discharge an unexecuted accepted action. Repair marker formatting; for short assent, if the original proposal cannot be recovered verbatim, select `blocked:` and ask. Formatting never requires clarification. Do not silently drop an accepted action or claim a pending gate passed.
|
|
200
200
|
|
|
@@ -101,12 +101,15 @@ An eligible next slice comes from an explicit status/task/acceptance source or a
|
|
|
101
101
|
|
|
102
102
|
Action-scoped stop conditions are: an explicit stop/pause instruction; a user-requested status-only answer; a failed, pending or inconclusive required gate; a dirty/conflicting worktree that cannot be isolated; a required environment unavailable after remediation; a high-impact product, architecture or compliance decision; a destructive action; an external purchase or financial commitment; unclear ownership; ambiguous assent; missing stricter authorization; materially different viable approaches with none dominant and reversible; a speculative fix without evidenced cause; or no low-risk slice. Apply each condition to the affected action. For a failed check, perform available authorized diagnosis and remediation before stopping the whole task: cite the failure output, repair attempts (or evidence that repair is unsafe or outside authority), and residual blocker. A failed verdict alone does not block diagnosis.
|
|
103
103
|
|
|
104
|
+
**Awaiting work you started yourself is not a stop condition.** A finite command, suite, gate, or review you launched, whose result only you consume, is in-flight work rather than a handoff: wait for it and continue in the same turn. A process meant to stay up — a dev server, a watch-mode runner, a tail — has no terminal result to wait for: take its readiness signal and proceed. Never poll it forever, and do not infer anything about its lifetime from this rule; whether it keeps running is the delivery's decision, and a service the user asked for is a deliverable, not a leftover. Ending the turn to report that it is running is a premature stop even when the report is accurate — the user gains nothing they can act on, and the next step was already authorized. Host behavior invites this: a backgrounded step returns control immediately, so the pause *looks* like a turn boundary. It is not one. Before ending any turn, name the next action; if you can perform it now, the turn is not over. The turn ends at the first action that genuinely needs the user — an unresolved decision, missing authority, an explicit stop — not at the nearest convenient pause. A user asking why you stopped is this defect's recurrence signal, not a request for a status update.
|
|
105
|
+
|
|
104
106
|
Check continuation on every user reply immediately following assistant prose that states or implies a next action, and on any explicit continuation request, regardless of landing status. Do not first require classifying the reply as assent; visibly report the continuing or blocked outcome even when the reply changes scope or stops the proposed action. Short replies include `ok`, `yes`, `可以`, `好`, `继续`, `proceed`, `do it`, `go ahead`, and `👍`; interpret them against the recovered action rather than formatting alone.
|
|
105
107
|
|
|
106
108
|
- Select `continuing: <action and scope>` when that action is clear and authorized, then execute it in the same turn. A tool call and its result or a produced artifact establish execution; the label alone does not.
|
|
107
109
|
- A blocked patch, review, or landing does not block every action. Keep that dependent action/claim pending while continuing available diagnosis, bounded remediation, monitoring of the existing live handle, or independent accepted work. These paths retain their own scope and permission checks; they cannot bypass the blocked gate or substitute unrelated hardening for missing evidence.
|
|
108
110
|
- A failed quality gate calls for a repair that preserves its purpose. Before asking the user to choose a workaround, inspect and perform a safe structural cleanup necessary for the authorized delivery when available, including baseline failures that block it, then rerun the gate and affected tests. Follow [refactoring discipline](refactoring-discipline.md#responding-to-quality-gates): preserve behavior, compatibility and readability; do not shrink identifiers or necessary comments, weaken a baseline or rewrite history solely to make the counter pass. If no safe in-scope repair remains, report the evidence and the actual decision needed.
|
|
109
111
|
- Independent work must neither depend on the pending verdict nor modify the candidate being evaluated. Name the pending gate and the independence basis when continuing. A candidate-changing fix is remediation, not independent work: let the existing run reach a terminal state, then refresh affected evidence and re-enter the owning gate. The deferred-evidence hardening prohibition still applies.
|
|
112
|
+
- A self-initiated step still running is `continuing:`, never `blocked:` and never a final response. Poll it to a terminal result, act on that result, and only then re-enter this gate. A step whose terminal result cannot be obtained after the normal remediation — it hangs, or its handle is lost — is the ordinary unavailable-environment case and blocks that dependent action, with the failure and the remediation attempted cited.
|
|
110
113
|
- Select `blocked: <action and scope> — <specific blocker>` when the remaining action needs an unresolved decision/authority or no safe authorized work remains after remediation. Cite the actual evidence; ask only for the missing decision or permission. An explicit stop/pause or status-only request blocks executing the prior proposal: name that reason in the outcome, answer the requested status, and do not reconfirm the stop.
|
|
111
114
|
- Apply landing-state proof to landing claims and derivation of post-landing work. For an authorized local investigation with no landed slice, record that landing checks do not apply and perform the investigation.
|
|
112
115
|
|
|
@@ -659,3 +659,8 @@ The pending classification above is superseded by the executed source comparison
|
|
|
659
659
|
| Required failures inherited from a baseline remain part of an authorized repair task | `defect-diagnosis` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/SKILL.md#Required failures, including inherited debt | updated | Owner key `defect-diagnosis/SKILL.md`. The entry now requires diagnosis, safe repair and rerunning the original check, with a mandatory handoff reference. A synthetic deletion of the new entry makes its named retention assertion fail in the shared implementation-retention fixture; the unchanged control passes. Advisory cases F35-F38 in `eval/behavior-fixtures.jsonl` distinguish inherited blockers, unsafe repair, status-only and diagnosis scope. These checks establish text retention and reviewable scenarios, not a measured increase in autonomous delivery. |
|
|
660
660
|
| Delivery blockers require bounded repair evidence before an exception or blocked handoff | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/refactoring-discipline.md#Treat a required-check failure that blocks this delivery as work to resolve | updated | Owner key `product-rd-workflow/SKILL.md`. The quality-gate response and `skills/product-rd-workflow/references/pre-final-continuation-gate.md` preserve the check purpose, require repair attempts or evidence of an unsafe or unauthorized repair, and leave optional findings separate. Deleting each added retention predicate makes its own assertion fail in the shared implementation-retention fixture; controls pass. Earlier explicit-context task replay already chose repair, so the change makes the inherited-blocker and handoff rules explicit without claiming a demonstrated task-level improvement. |
|
|
661
661
|
| Retention checks must include every input surface when executed from an isolated fixture | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. New repair and goal-authorization assertions in `skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh` read the root contract and release documents. The prior isolated copy omitted those inputs and failed its clean control; copying them restores the 52-mutation controlled-escalation walk. Eleven new repair and authorization predicates also fail under individual deletion, with passing controls. Goal-authorized delivery remains bounded by the requested target, caller authority, explicit stop or narrow scope, and actual host enforcement. Advisory cases F39-F40 cover release continuation and unrelated protected actions; static pins do not implement a permission system or prove agent compliance. |
|
|
662
|
+
| A packet-only reviewer runs from a private home that never carried the user's MCP servers, rather than disabling them by name | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_cli_review_wrappers.sh | updated | Owner key `code-review/SKILL.md`. Two faults on codex-cli 0.153.4. Every frozen-packet tool call was refused before it ran: `approval_policy="never"` with a sandbox lacking full disk write access leaves no auto-approve branch, so the packet server now declares `default_tools_approval_mode="approve"` for itself while the sandbox and policy stay unchanged. Separately, the preflight disabled every other server by name; a plugin contributes its server outside `mcp_servers`, so that override builds a transportless entry the CLI rejects outright, while leaving it enabled failed an exactly-one-server count -- the lane could not run at all on such a host. Removing the enumeration made the lane usable and made foreign servers reachable: measured, the host's own `node_repl` ran its `js` tool to completion during a packet-only review, and a stub was auto-approved purely by declaring `readOnlyHint`, which the CLI trusts from the server itself. Global approval-mode defaults did not override that hint and disabling plugins would disable the reviewer's own registry, so the run now gets a private `CODEX_HOME`: linked credential, carried model preference, copied owner skills, nothing else. Two draft claims are withdrawn rather than edited away -- that the parser's after-the-fact audit contained a foreign call, and that a hostile diff was a demonstrated path to one (two attempts did not reproduce it). `origin/dev` is not the unsafe baseline: its per-name disable works for config-declared servers and fails only for plugin-contributed ones. RED-baseline: seven wrapper assertions fail against `origin/dev`'s wrapper and the private-home assertion fails against the mid-round one; all pass on the final candidate. |
|
|
663
|
+
| Work the agent started itself and still awaits is not a stop condition: it polls that step to a terminal result and continues in the same turn | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/pre-final-continuation-gate.md#A self-initiated step still running is | updated | Owner key `product-rd-workflow/SKILL.md`. Observed twice in one session: after launching a test suite or gate whose result only it would consume, the agent ended the turn to report that the step was running; the user had to ask why it stopped, then named the stopping itself as the defect. The gate already said not to stop at a recommendation and to continue with an owned low-risk slice, so content was not the gap -- the stop-condition list simply did not name this shape, and the host returns control the moment a step is backgrounded, which makes the pause look like a turn boundary. Landed as a firing mechanism rather than a discipline reminder: the reference names awaiting a FINITE self-started step as a non-condition, requires naming the next action before any turn ends, and adds an outcome-contract line making a still-running self-initiated step `continuing:` rather than `blocked:` or a final response; the entrypoint carries the same clause so the rule fires without opening the reference. Independent review caught the first wording as an over-broad absolute -- a dev server or watch-mode runner has no terminal result, so the rule would have demanded indefinite polling; it now takes a readiness signal and says nothing about the process lifetime: a later challenge showed that shutting it down at closeout destroys a service that is itself the requested deliverable, so the clause stops prescribing what it does not own. RED-baseline (applied, differential): deleting each of the four clauses reds only its own assertion in the shared implementation-gates fixture (`test_ai_coding_implementation_gates.sh`) with no other assertion failing, and the unmutated control passes. |
|
|
664
|
+
| The shared implementation-gates fixture pins the continuation gate's non-stop clauses, so a later edit cannot silently delete them | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`. The sibling row for `product-rd-workflow` records the failure itself; this row records why the fix cannot regress silently. Four assertions were added to the fixture -- the reference's non-stop clause, its turn-end firing check, its outcome-contract line, and the entrypoint's own clause. The fourth was added after independent review observed that the outcome-contract line could be deleted with every assertion still green, which is the same false-green shape the pins exist to prevent. RED-baseline (applied, differential): deleting each protected sentence reds only its owning assertion, with every other assertion passing and the unmutated control clean, so a partial deletion is attributable rather than lost in an aggregate failure. The fixture was chosen over a new suite because it already owns cross-owner rule-retention pins; no new registration surface is introduced. |
|
|
665
|
+
| The landing binder names the ordering cause at the failure point: evidence a round adds that stays inside the candidate is listed when nothing binds | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/review_ledger_binding.py | updated | Owner key `skill-extraction-workflow/SKILL.md`. Third occurrence of one class. The rule that bound evidence is committed before the review rounds already exists verbatim in the quickstart and already carries a register row marked observed twice in consecutive rounds; this round hit it again because the round was driven from the delivery and review owners and never opened that quickstart. Two prior landings answered the recurrence with more prose, so this one changes the mechanism instead: when nothing binds, the binder enumerates the added evidence that is NOT excluded -- the complement of the receipt exclusion it already computes -- and states that only added JSON carrying a candidate_sha256 is excluded, so committing a base attestation or excerpt after the rounds moves the candidate out from under their receipts. RED-baseline (applied): on this round's own failing candidate the pre-change binder reported only that nothing bound it, naming neither the file nor the ordering; the changed binder lists `landing-base.txt` and the round's markdown dispositions and states the ordering. The five binding suites pass unchanged. The diagnosis now reaches an agent at the moment it fails rather than requiring it to know which document to open. |
|
|
666
|
+
| A harness whose RECORDS are the evidence — an evaluation or benchmark runner, a conformance suite feeding a comparison, an A/B or regression rig — can be corrupted by the data it produces in three ways that all read green: absence stored as a bare null cannot separate confirmed-absent from never-observed, planned units and retries sharing one counter let a retry move the denominator, and a later attempt overwrites an earlier failure. Its own record layer is a high-risk failure class of the same standing as the canonical list, and is built against these before the happy path | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#Absence carries a coded reason beside the value | updated | `testing-strategy/SKILL.md` is the owner key and is unchanged this round: the entrypoint is over its size budget and the growth gate blocks it, so the rule lands in `testing-strategy/references/ci-fixtures-and-flake-control.md` and is reached as a failure class from the canonical high-risk list in `testing-strategy/references/scenario-testing.md`, an enumeration the entrypoint already tells readers to walk. RED baseline: a paired walk over real artifacts — a held-out harness that contributed nothing to deriving the rules fails all three rows, each defect named by exactly one row while the other two do not mention it, while the control harness passes two and partially satisfies the first, so the check discriminates rather than accepting whatever is put to it. `observed-failure` is `no` deliberately: no malfunction of an existing repository rule was recorded this round, and the delta is measured against the held-out artifact rather than against a regression this repository observed; `result-class` is `failure` because that held-out artifact does exhibit all three defects the rule names. Sources read this round: the health-interchange data-absent-reason code system, a monitoring query language's absent-vector operators, and the controlled-trial reporting guidance for the flow diagram and per-group denominators. Known limit, stated in the landed text itself: the assembled rule has no located prior name, and measurement system analysis is the adjacent established field covering instrument accuracy and repeatability rather than record integrity. |
|
|
@@ -299,12 +299,17 @@ def added_evidence_paths(repo_root: Path, base: str) -> list[str]:
|
|
|
299
299
|
"""
|
|
300
300
|
result = subprocess.run(
|
|
301
301
|
[
|
|
302
|
-
"git", "-C", str(repo_root), "diff", "--name-only",
|
|
302
|
+
"git", "-C", str(repo_root), "diff", "--name-only", "-z",
|
|
303
303
|
"--diff-filter=A", base, "HEAD", "--", EVIDENCE_ROOT,
|
|
304
304
|
],
|
|
305
305
|
stdout=subprocess.PIPE,
|
|
306
306
|
stderr=subprocess.PIPE,
|
|
307
|
+
# A pathname is bytes, and -z hands them over raw. Strict decoding would
|
|
308
|
+
# turn one undecodable filename into a crash inside a gate whose job is
|
|
309
|
+
# to fail cleanly, so undecodable bytes survive as surrogates and simply
|
|
310
|
+
# do not match the evidence pattern.
|
|
307
311
|
text=True,
|
|
312
|
+
errors="surrogateescape",
|
|
308
313
|
check=False,
|
|
309
314
|
)
|
|
310
315
|
if result.returncode != 0:
|
|
@@ -313,14 +318,47 @@ def added_evidence_paths(repo_root: Path, base: str) -> list[str]:
|
|
|
313
318
|
f"{result.stderr.strip()}"
|
|
314
319
|
)
|
|
315
320
|
excluded: list[str] = []
|
|
316
|
-
|
|
317
|
-
|
|
321
|
+
# -z output is NUL-separated and never C-quoted, so a path carrying a
|
|
322
|
+
# non-ASCII byte is enumerated as itself rather than as an escaped literal
|
|
323
|
+
# that no pattern here would match.
|
|
324
|
+
for line in result.stdout.split("\0"):
|
|
325
|
+
if not line or not EVIDENCE_MEMBER.match(line):
|
|
318
326
|
continue
|
|
319
327
|
if is_candidate_receipt(repo_root, line):
|
|
320
328
|
excluded.append(line)
|
|
321
329
|
return excluded
|
|
322
330
|
|
|
323
331
|
|
|
332
|
+
def bound_evidence_paths(repo_root: Path, base: str) -> list[str]:
|
|
333
|
+
"""Evidence this round ADDS that stays inside the candidate.
|
|
334
|
+
|
|
335
|
+
The complement of the exclusion, reported when nothing binds. Committing one
|
|
336
|
+
of these after the review rounds moves the candidate out from under their
|
|
337
|
+
receipts, and the failure that surfaces -- nothing binds -- names neither the
|
|
338
|
+
file nor the ordering. This class has now been observed three times; the
|
|
339
|
+
diagnosis belongs where the failure appears, not in a document the round has
|
|
340
|
+
to know to open.
|
|
341
|
+
"""
|
|
342
|
+
result = subprocess.run(
|
|
343
|
+
[
|
|
344
|
+
"git", "-C", str(repo_root), "diff", "--name-only", "-z",
|
|
345
|
+
"--diff-filter=A", base, "HEAD", "--", EVIDENCE_ROOT,
|
|
346
|
+
],
|
|
347
|
+
stdout=subprocess.PIPE,
|
|
348
|
+
stderr=subprocess.PIPE,
|
|
349
|
+
text=True,
|
|
350
|
+
errors="surrogateescape",
|
|
351
|
+
check=False,
|
|
352
|
+
)
|
|
353
|
+
if result.returncode != 0:
|
|
354
|
+
return []
|
|
355
|
+
return [
|
|
356
|
+
line
|
|
357
|
+
for line in result.stdout.split("\0")
|
|
358
|
+
if line and EVIDENCE_MEMBER.match(line) and not is_candidate_receipt(repo_root, line)
|
|
359
|
+
]
|
|
360
|
+
|
|
361
|
+
|
|
324
362
|
def is_candidate_receipt(repo_root: Path, path_value: str) -> bool:
|
|
325
363
|
"""Whether the committed blob at this path is a receipt about a candidate.
|
|
326
364
|
|
|
@@ -1094,6 +1132,19 @@ def bind_candidate(
|
|
|
1094
1132
|
" no committed ledger records this candidate; run the extraction review "
|
|
1095
1133
|
"lane against the final, committed tree"
|
|
1096
1134
|
)
|
|
1135
|
+
inside = bound_evidence_paths(repo_root, fork)
|
|
1136
|
+
if inside:
|
|
1137
|
+
binding.failure.append(
|
|
1138
|
+
" this round added evidence that stays inside the candidate: "
|
|
1139
|
+
+ ", ".join(inside[:5])
|
|
1140
|
+
+ ("" if len(inside) <= 5 else f", and {len(inside) - 5} more")
|
|
1141
|
+
)
|
|
1142
|
+
binding.failure.append(
|
|
1143
|
+
" only added JSON carrying a candidate_sha256 is excluded, so bound "
|
|
1144
|
+
"evidence such as base attestations and excerpts must be committed "
|
|
1145
|
+
"BEFORE the review rounds; committing it after moves the candidate out "
|
|
1146
|
+
"from under their receipts"
|
|
1147
|
+
)
|
|
1097
1148
|
return binding
|
|
1098
1149
|
|
|
1099
1150
|
|
|
@@ -36,6 +36,12 @@ assert_contains "$PRODUCT_SKILL" 'gaps block `complete`' "product workflow gate
|
|
|
36
36
|
assert_contains "$PRODUCT_SKILL" "references/implementation-completeness-and-minimality.md" "product workflow pointer"
|
|
37
37
|
assert_contains "$PRODUCT_REF" "Requirement / acceptance point | Source decision | Implementation surface | Verification | Fresh evidence | Status" "acceptance closure matrix"
|
|
38
38
|
assert_contains "$PRODUCT_REF" "New concept | Current acceptance point or hard constraint | Simpler alternative | Decision" "concept delta matrix"
|
|
39
|
+
assert_contains "$PRE_FINAL_REF" "Awaiting work you started yourself is not a stop condition." "continuation gate (self-initiated in-flight work is not a stop)"
|
|
40
|
+
assert_contains "$PRE_FINAL_REF" "is in-flight work rather than a handoff: wait for it and continue in the same turn" "continuation gate (in-flight obligation)"
|
|
41
|
+
assert_contains "$PRE_FINAL_REF" "Before ending any turn, name the next action; if you can perform it now, the turn is not over." "continuation gate (turn-end firing check)"
|
|
42
|
+
assert_contains "$PRODUCT_SKILL" "poll any finite step you started to its result, never reporting it as running" "continuation gate (entry firing signal)"
|
|
43
|
+
assert_contains "$PRE_FINAL_REF" "is \`continuing:\`, never \`blocked:\` and never a final response. Poll it to a terminal result" "continuation gate (outcome-contract clause)"
|
|
44
|
+
assert_contains "$PRE_FINAL_REF" "has no terminal result to wait for: take its readiness signal and proceed" "continuation gate (persistent-process exception)"
|
|
39
45
|
assert_contains "$PRODUCT_REF" "Passing one question never compensates for failing the other." "independent axes"
|
|
40
46
|
assert_contains "$PRODUCT_REF" 'An implementer may not silently downscope a point' "no self-downscope"
|
|
41
47
|
assert_contains "$PRODUCT_REF" 'hypothetical reuse are not evidence' "no speculative concepts"
|
|
@@ -89,6 +89,20 @@ When claiming tests pass, report:
|
|
|
89
89
|
|
|
90
90
|
Do not report "tests pass" from memory or from a previous turn. Verification must be fresh for the current change.
|
|
91
91
|
|
|
92
|
+
## Evidence-Record Integrity For Measurement Harnesses
|
|
93
|
+
|
|
94
|
+
Fires when the deliverable is a harness whose RECORDS are the evidence — an evaluation or benchmark runner, a conformance suite feeding a comparison, an A/B or regression measurement rig — rather than a suite whose deliverable is pass/fail. Reached as a failure class from `scenario-testing.md` (High-Risk Failure Classes). Such a harness can be corrupted by the data it produces, in three ways that all read green. Build the record layer against these before the happy path; retrofitting means re-judging records already collected. This is one layer upstream of the entry rule that assigns absence assertions to the producing layer: that rule says where absence can be PROVEN, these say whether the record can express WHICH absence occurred at all.
|
|
95
|
+
|
|
96
|
+
- **Absence carries a coded reason beside the value — never a bare null, never a new value-type.** One null cannot say whether the thing was confirmed not to exist or was never observed, and here those are opposite facts: the first is a result about the system under measurement, the second is a hole in the measurement. Put the reason in a sibling field from a closed vocabulary separating at least *confirmed absent*, *asked but unavailable*, and *not attempted* — they imply opposite retry decisions, so collapsing them also destroys the scheduling signal. Keep it beside the value, not inside its type: the relational model's own two-marker proposal (missing-but-applicable vs missing-but-inapplicable) needed four-valued logic and was never widely adopted, while the health-interchange standard's data-absent-reason coding works because it is an adjacent field only its readers pay for. Make the claim cost evidence — accept *confirmed absent* only with the retrievable observation that established it, or a writer clears a failure by asserting absence. And give absence its own assertion: a threshold over values cannot see a series that is not there, which is why a monitoring query language needs a dedicated absent-vector operator to alert on a series that stopped arriving.
|
|
97
|
+
- **Report the flow counts per compared arm, not only per run; caps are the secondary control.** Planned units, extra attempts, and blocked units must be separately counted and reported, so a reader can check the denominator instead of trusting a label — and broken out per arm or per analysis being compared, together with the mix of absence reasons. A run-level total hides the asymmetry that is precisely the bias below: ten blocked units on one side and none on the other reads green in the totals while the comparison is already spoiled. Controlled-trial reporting guidance is explicit that the denominator belongs to *each group* in *every* analysis, not to the study as a whole; it also dropped the requirement to *declare* an analysis intention-to-treat — because no label reliably says who was actually included — and replaced it with that required flow of numbers. Keep caps, but a cap nobody can audit against reported counts is a claim, not a control.
|
|
98
|
+
- **Retain the earliest decisive outcome; record why later attempts happened.** A retry must not replace a failure that already occurred. Keep every attempt and let a unit's recorded outcome be the earliest decisive one — including a failure that first appears on a later attempt after an inconclusive earlier one. Trial reporting is again the shape to copy: post-hoc change is not forbidden, it is required to be reported with its reason. An automatic retry that overwrites the first attempt produces exactly this corruption: the symptom is hidden, the run still reports itself complete, and the reported pass rate stops being a number a release decision can rest on.
|
|
99
|
+
|
|
100
|
+
Why this is not cosmetic: in a measurement harness missingness is rarely random — evidence is missing BECAUSE the run failed, the missing-not-at-random case — so dropping incomplete units biases the comparison toward whatever produced them.
|
|
101
|
+
|
|
102
|
+
Verify by building the harness's own negative cases first: one unit per absence reason, one that spends an extra attempt, and one that fails and then succeeds. A harness that cannot distinguish those three in its own output is not ready to measure anything else. Route the online-signal form of the absence rule — a metric that stopped arriving versus one reporting zero — to `platform-observability`.
|
|
103
|
+
|
|
104
|
+
Boundary: this is an assembled rule, not a named discipline. Measurement system analysis is the adjacent established field, and it covers instrument accuracy and repeatability, not record integrity.
|
|
105
|
+
|
|
92
106
|
## Conditional-Skip × Job-Selection Executed-Count Guards
|
|
93
107
|
|
|
94
108
|
Strongest form — a per-file invariant: the expected-file list derives from the job's own selection manifest, per job/environment (static; never from post-skip collection, which already lacks the silently-skipped file, and never shared across env-split jobs where different files legitimately run), and each expected file collects AND executes > 0 tests, with a missing-optional-dependency skip a hard failure in the job that exists to provide that dependency, never an "expected skip".
|
|
@@ -76,7 +76,7 @@ Use this matrix when a web or app surface contains many charts, grouped tables,
|
|
|
76
76
|
|
|
77
77
|
## High-Risk Failure Classes
|
|
78
78
|
|
|
79
|
-
The risk matrix for a high-risk workflow covers the triggered failure classes from this canonical list: duplicate submit/callback/message/job restart, permission service uncertainty, cross-tenant/user/resource mismatch, partial money/quota side effects, AI provider/model failure, unclear final status after refresh/offline, and missing trace/support identifier. Cover each triggered class at the lowest layer that can prove the invariant.
|
|
79
|
+
The risk matrix for a high-risk workflow covers the triggered failure classes from this canonical list: duplicate submit/callback/message/job restart, permission service uncertainty, cross-tenant/user/resource mismatch, partial money/quota side effects, AI provider/model failure, unclear final status after refresh/offline, and missing trace/support identifier. When the deliverable is itself a measurement harness — an evaluation or benchmark runner, a conformance suite feeding a comparison, an A/B or regression rig — its own record layer is a class of the same standing: absence reading as a pass, a retry moving the denominator, and a later attempt overwriting an earlier failure (`ci-fixtures-and-flake-control.md`, Evidence-Record Integrity For Measurement Harnesses). Cover each triggered class at the lowest layer that can prove the invariant.
|
|
80
80
|
|
|
81
81
|
Cross-reference: `non-functional-specialized-scenarios.md` (High-risk resilience boundaries) states the launch-gate counterpart — which classes require scenario tests or drills at release. That is a gate-criteria list; this is the test-matrix failure-class list. The two complement each other and neither replaces the other.
|
|
82
82
|
|
package/dist/assets/release.json
CHANGED
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
{
|
|
2
2
|
"schema": 1,
|
|
3
3
|
"npmPackage": "@ccoalm/ccl-skills",
|
|
4
|
-
"version": "0.15.
|
|
5
|
-
"sourceCommit": "
|
|
4
|
+
"version": "0.15.5",
|
|
5
|
+
"sourceCommit": "88e842ffb4ab944dde3179fa5a8e43755fbba27a",
|
|
6
6
|
"sourceState": "clean",
|
|
7
7
|
"files": [
|
|
8
8
|
{
|
|
@@ -287,7 +287,7 @@
|
|
|
287
287
|
},
|
|
288
288
|
{
|
|
289
289
|
"path": "marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md",
|
|
290
|
-
"sha256": "
|
|
290
|
+
"sha256": "900f09bfc21bdd78c9bf85794a361ae35dbf5b1ce0ae1910baa1991615b51093",
|
|
291
291
|
"mode": 420
|
|
292
292
|
},
|
|
293
293
|
{
|
|
@@ -317,7 +317,7 @@
|
|
|
317
317
|
},
|
|
318
318
|
{
|
|
319
319
|
"path": "marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh",
|
|
320
|
-
"sha256": "
|
|
320
|
+
"sha256": "9b98ba819ed9baee557c43aa735ab006e0dfb4e02e0e337cdca5e4867845c105",
|
|
321
321
|
"mode": 493
|
|
322
322
|
},
|
|
323
323
|
{
|
|
@@ -412,7 +412,7 @@
|
|
|
412
412
|
},
|
|
413
413
|
{
|
|
414
414
|
"path": "marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh",
|
|
415
|
-
"sha256": "
|
|
415
|
+
"sha256": "bd67dcc75343071913411c43fe727d4e4361318318fe85c7f350cd9f000f2c79",
|
|
416
416
|
"mode": 493
|
|
417
417
|
},
|
|
418
418
|
{
|
|
@@ -1402,7 +1402,7 @@
|
|
|
1402
1402
|
},
|
|
1403
1403
|
{
|
|
1404
1404
|
"path": "marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md",
|
|
1405
|
-
"sha256": "
|
|
1405
|
+
"sha256": "084b801e8eb4c0d3ccabf102d7bc75721a00666030f3fcbc00dd74d5725ac408",
|
|
1406
1406
|
"mode": 420
|
|
1407
1407
|
},
|
|
1408
1408
|
{
|
|
@@ -1477,7 +1477,7 @@
|
|
|
1477
1477
|
},
|
|
1478
1478
|
{
|
|
1479
1479
|
"path": "marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md",
|
|
1480
|
-
"sha256": "
|
|
1480
|
+
"sha256": "b33441326e6dfe9ec2a587d80cb27344c9cbd6c32f7a3ca672615c3f0e6d2055",
|
|
1481
1481
|
"mode": 420
|
|
1482
1482
|
},
|
|
1483
1483
|
{
|
|
@@ -2132,7 +2132,7 @@
|
|
|
2132
2132
|
},
|
|
2133
2133
|
{
|
|
2134
2134
|
"path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md",
|
|
2135
|
-
"sha256": "
|
|
2135
|
+
"sha256": "a1f297987a1c91f283d83ede2bd1994b727316a45e978b3480dec6a556644e3a",
|
|
2136
2136
|
"mode": 420
|
|
2137
2137
|
},
|
|
2138
2138
|
{
|
|
@@ -2272,7 +2272,7 @@
|
|
|
2272
2272
|
},
|
|
2273
2273
|
{
|
|
2274
2274
|
"path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py",
|
|
2275
|
-
"sha256": "
|
|
2275
|
+
"sha256": "a75b043aa63e278a45e3c31f418b042a0ffbef3497ff4c1e139f777bd72e6d29",
|
|
2276
2276
|
"mode": 493
|
|
2277
2277
|
},
|
|
2278
2278
|
{
|
|
@@ -2297,7 +2297,7 @@
|
|
|
2297
2297
|
},
|
|
2298
2298
|
{
|
|
2299
2299
|
"path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh",
|
|
2300
|
-
"sha256": "
|
|
2300
|
+
"sha256": "e51dac3a34fa51656f4d096a2ce568aebd6fbc528d30d12f98cebfdc91cc8921",
|
|
2301
2301
|
"mode": 493
|
|
2302
2302
|
},
|
|
2303
2303
|
{
|
|
@@ -2732,7 +2732,7 @@
|
|
|
2732
2732
|
},
|
|
2733
2733
|
{
|
|
2734
2734
|
"path": "marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md",
|
|
2735
|
-
"sha256": "
|
|
2735
|
+
"sha256": "9796ba2af7da8ff4eedfd3fe6e233a92cded263cae06347e34a089289af8f89f",
|
|
2736
2736
|
"mode": 420
|
|
2737
2737
|
},
|
|
2738
2738
|
{
|
|
@@ -2782,7 +2782,7 @@
|
|
|
2782
2782
|
},
|
|
2783
2783
|
{
|
|
2784
2784
|
"path": "marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md",
|
|
2785
|
-
"sha256": "
|
|
2785
|
+
"sha256": "fd14402d13d610621b7b1dab87f422ff9ed6fa20a2071ef8c46e8281f3d6a199",
|
|
2786
2786
|
"mode": 420
|
|
2787
2787
|
},
|
|
2788
2788
|
{
|
|
@@ -3433,5 +3433,5 @@
|
|
|
3433
3433
|
"mode": 420
|
|
3434
3434
|
}
|
|
3435
3435
|
],
|
|
3436
|
-
"snapshotHash": "
|
|
3436
|
+
"snapshotHash": "098f8ae71c1a51738986e49cf1667aaa62e714b3c4b10d4ff96b0a08f2655d96"
|
|
3437
3437
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@ccoalm/ccl-skills",
|
|
3
|
-
"version": "0.15.
|
|
3
|
+
"version": "0.15.5",
|
|
4
4
|
"description": "Reusable workflows that help coding agents plan, build, test, review, and release software — for Claude Code, Codex, and OpenCode",
|
|
5
5
|
"keywords": ["skills", "agent-skills", "claude", "claude-code", "codex", "opencode", "agent", "ai", "ai-agents", "cli", "anthropic", "developer-tools"],
|
|
6
6
|
"type": "module",
|