@ccoalm/ccl-skills 0.1.0 → 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (18) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +2 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +6 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +3 -1
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-capability-composition.md +128 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +35 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +54 -2
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +11 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +6 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +1 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +6 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +1 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +9 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +2 -2
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +18 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +1 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +3 -3
  17. package/dist/assets/release.json +23 -18
  18. package/package.json +1 -1
@@ -17,6 +17,8 @@ Do not create a service only because:
17
17
  - A team wants a cleaner folder.
18
18
  - A future scale concern is speculative.
19
19
 
20
+ - External grounding (adopted in part): the criteria above take [Bounded Context](https://martinfowler.com/bliki/BoundedContext.html) (Fowler's overview of the central DDD strategic-design pattern; origin credited there to Eric Evans, *Domain-Driven Design*) as an input to boundary drawing — a boundary is where one internally consistent model stops being valid, so boundaries follow model and responsibility lines, never noun or table count — and [Conway's Law](https://martinfowler.com/bliki/ConwaysLaw.html) (Melvin Conway, ["How Do Committees Invent?", 1968](https://www.melconway.com/Home/Committees_Paper.html); Fowler's overview) — a system's structure mirrors the communication structure of the organization that builds it, which is why team ownership and communication structure are a legitimate split input. Two limits: the transactional-data-ownership, scaling, deployment-cadence, and rollback tests above are this skill's own operational criteria, attributed to neither source; and a bounded context does not require its own deployable service — several contexts can ship inside one modular monolith (this skill's monolith-first default), so these sources justify model/ownership boundaries, not a service-per-context split. Borrowed scope: boundary-input criteria only; no claim to the full DDD strategic-design method (context maps, ubiquitous language) or an inverse-Conway process. Sibling: `python-service-architecture/references/architecture-playbook.md` ("Architecture Decisions") carries the same grounding; keep the two in sync.
21
+
20
22
  ## Recommended Service Types
21
23
 
22
24
  ### API Gateway / HTTP Service
@@ -70,3 +70,9 @@
70
70
  - Every ctx key has a typed accessor; bare-string `ctx.Value(...)` returning `any` is an architecture finding, not an idiom to spread. Type assertion lives behind helpers, not in domain code.
71
71
  - For frameworks that propagate metadata over the wire, define which keys travel persistently (every downstream hop forwards them) versus transiently (one hop only). Lane, stress tag, and trace identity are typically persistent; one-off control flags should not be promoted to persistent.
72
72
  - Dual-injection compatibility: when one binary serves multiple transports (TTHeader Thrift + HTTP/2 gRPC), the propagation layer writes to both the framework's persistent value (`metainfo.WithPersistentValue`) and the transport's outgoing metadata so the framework's meta handler picks the right wire format at send time. Application code stays transport-agnostic.
73
+
74
+ ## Topic-extension backlog
75
+
76
+ Entries here are registered candidates, not adopted guidance. Each names the candidate, its evidence status, and the condition that unblocks adoption; the round that evaluates one records keep/narrow/discard against its entry.
77
+
78
+ - **Package-owned runtime-invariant registries (weak-keep candidate).** Observed form, from one agent-native product repository (evolving portfolio): each package contributes its own invariant checks from a companion module, normal entrypoints do not depend on the diagnostics layer, and allowlist/blocklist configuration selects which checks run. Evidence status: weak — single source; whether this generalizes as a cross-cutting pattern, and which reference should own it, is undecided. Evaluate the generalization at the next architecture round touching diagnostics, invariant checking, or startup validation, and record the outcome against this entry. The Python-stack sibling registration lives in `../python-service-architecture/references/packaging-runtime-readiness.md`.
@@ -24,7 +24,8 @@ Use this for product backend work that calls, hosts, evaluates, or operates LLM
24
24
  - Core: model/version registry, prompt lifecycle, request schema, safety boundaries, streaming protocol, retry/fallback, eval datasets, replay/shadow rollout, token/cost metrics, batch/concurrency control, audit trails.
25
25
  - agent-skill-system runtime (progressive-disclosure loading, description-driven skill routing, skill trust/sandbox boundary)
26
26
  - MCP integration (server primitives, server-as-untrusted-domain trust boundary, OAuth 2.1 / audience-binding auth)
27
- - agent command-execution sandbox (OS-sandbox composition, command-policy DSL, approval/escalation state machine, loopback-only egress proxy — see `references/agent-command-sandbox.md`)
27
+ - agent capability composition (seam/provider/model-facing-tool decomposition, reversible registration with scope-owned disposal, policy-plugin-vs-enforcement split, sub-agent provider seam with one-shot/continuable separation — see `references/agent-capability-composition.md`)
28
+ - agent command-execution sandbox (OS-sandbox composition, single-owner policy resolution consumed by every enforcement backend, command-policy DSL, approval/escalation state machine, loopback-only egress proxy — see `references/agent-command-sandbox.md`)
28
29
  - agent session persistence (append-only event-log source of truth, background-writer flush-before-finality, resume/fork with restored token accounting, context-window compaction — see `references/agent-session-persistence.md`)
29
30
  - agent ambient context freshness (diffable world-state sections vs one-shot fragments, comparison-snapshot-vs-rendered-text staleness detection, supersede-stale-in-band-not-retract, self-recognizing injections, derivable baseline + merge-patch re-derive on resume, budget-signaling-vs-compaction-enforcement — see `references/agent-context-freshness.md`)
30
31
  - robust model-driven file-edit (context-anchored hunks over line numbers, graduated-strictness matching, resolve-all-before-write, honest partial-failure reporting — see `references/agent-file-edit-protocol.md`)
@@ -72,6 +73,7 @@ Use this for product backend work that calls, hosts, evaluates, or operates LLM
72
73
  - Trace representative turns through every state transformation: at minimum the normal allow path, deny/ask permission path, resume/recovery path, and dynamic-tool-change path. Each trace must cover user input normalization, context construction, budget or compaction projection, model streaming, tool-call validation, permission decision, tool-result insertion, continuation, terminal stop condition, transcript write, and usage/cost accounting.
73
74
  - Each runtime concern below gets layer-separation, and its gate / assertions / routing live in the named reference (read before gating that concern):
74
75
  - Runtime startup & config bootstrap -> `references/agent-runtime-bootstrap.md`
76
+ - Capability composition — module/plugin boundaries, registration lifecycle, sub-agent providers -> `references/agent-capability-composition.md`
75
77
  - Turn lifecycle — per-turn loop, fg/bg handoff, progress, plan-to-execute, recap -> `references/agent-turn-lifecycle.md`
76
78
  - Session & transport — discovery/fork, history sync, workspace scope, protocol/stdout, settings migration -> `references/agent-session-persistence.md`
77
79
  - Hooks — hook/classifier output trust, hook config control plane, post-turn hooks -> `references/agent-lifecycle-hooks.md`
@@ -0,0 +1,128 @@
1
+ # Agent Capability Composition (Seams, Providers, Reversible Registration)
2
+
3
+ How an agent runtime packages its capabilities so they can be swapped, unloaded, and extended
4
+ without a privileged core: the decomposition unit, the registration lifecycle, and the sub-agent
5
+ provider seam. This is the composition layer *around* the frameworks the sibling references own —
6
+ `agent-tool-dispatch.md` owns how a call reaches a handler; `agent-runtime-bootstrap.md` owns
7
+ config/trust at startup; this reference owns how the handler's capability is packaged, installed,
8
+ and torn down.
9
+
10
+ Patterns observed in one agent-native product repository (an evolving portfolio — treat as one
11
+ industry form, not a standard) and consistent with general plugin-architecture practice. Use when
12
+ designing or reviewing an agent runtime's module boundaries, plugin/extension system, or sub-agent
13
+ integration; skip for a single-purpose agent whose capabilities never vary by deployment.
14
+
15
+ ## 1. A swappable capability is three parts — seam, provider, model-facing tool
16
+
17
+ - When a capability must vary by deployment (local vs sandboxed filesystem, real vs replay model
18
+ transport, in-process vs external sub-agent), decompose it into three separately-packaged roles:
19
+ the **seam** (the service contract: abstract interface + vocabulary types), one or more
20
+ **providers** (implementations of the seam), and the **model-facing tool** (a *consumer* of the
21
+ seam that renders it into the model's tool surface). Consumers — tools, policies, other
22
+ capabilities — depend only on the seam; a provider is selected at assembly time. The tool is not
23
+ the implementation and never reaches around the seam into one.
24
+ - Verify the decomposition by the swap test: a new provider (remote store, different vendor,
25
+ fault-injection fake) must drop in without touching the seam's consumers or the tool. If adding a
26
+ provider forces edits in consumers, the contract leaked implementation detail.
27
+ - Design the seam's contract for **all current consumers**, not the loudest one: a method shaped for
28
+ a single consumer's UI/transport/private need does not belong on the shared contract. A public
29
+ seam method with exactly one internal caller is API surface without a second user — prefer a
30
+ capability closure injected at construction into the one consumer that needs it.
31
+ - Keep role vocabulary honest in reviews: "the tool does X" is a smell when X is storage, policy, or
32
+ enforcement — those belong to a provider or a policy plugin, and the tool only surfaces them.
33
+ (Worked instances of the split: the spill store vs spill policy split in `agent-tool-dispatch.md`
34
+ §Result shaping; the sandbox policy-resolution vs enforcement-backend split in
35
+ `agent-command-sandbox.md` §2.)
36
+
37
+ ## 2. Registration is reversible; teardown is scope disposal, not per-plugin cleanup
38
+
39
+ - Every registration-like effect a capability performs at install time (registering a tool, an
40
+ event listener, a provider, a background worker) must carry its inverse, and the runtime — not
41
+ each plugin author — owns applying inverses on unload. Model installation as effects recorded
42
+ into an **ownership scope**; teardown = dispose the scope, which replays the inverses in reverse
43
+ order. Correct unload is then guaranteed by one abstraction instead of N hand-written uninstall
44
+ paths, and a plugin author cannot forget it.
45
+ - Hot-swap/reload is dispose-old-then-install-new under the same rule — never patch-in-place, which
46
+ accumulates the old registration's residue. The failure half needs a contract too: preflight the
47
+ new provider in an isolated scope *before* disposing the old one where the capability tolerates
48
+ brief coexistence. Where it does not, know what rollback can and cannot promise: registration
49
+ inverses reverse *local* effects only — disposal may have released a lease, port, credential, or
50
+ other external resource that reinstallation cannot reacquire — so a non-coexistent cutover either
51
+ uses an atomic/preflightable swap mechanism, or retains the old install descriptor and
52
+ rollback-critical resources until the new install commits, and when the install fails *and*
53
+ rollback also fails, the seam enters a **typed, surfaced unavailable state** (fail closed), never
54
+ a silent half-registered one. Verify with an install → dispose → re-install cycle asserting no
55
+ duplicate handlers, no leaked listeners, and no orphaned background work — plus failure injection
56
+ through BOTH layers: new provider's install throws (capability ends served by old or new where
57
+ the swap mechanism can guarantee it), and restoration fails too (capability ends in the typed
58
+ unavailable state, observably, rather than asserting never-neither for a path that cannot
59
+ guarantee it).
60
+ - Disposal order matters: dispose in reverse registration order so dependents release before their
61
+ dependencies, and make disposal idempotent so a crash-then-cleanup path cannot double-free.
62
+ - This composes with (does not replace) the trust and generation rules in
63
+ `agent-runtime-bootstrap.md`: a reversible registration that loads untrusted code is still gated
64
+ by trust state, and a reload still invalidates the caches keyed to the old generation.
65
+
66
+ ## 3. Policy plugins decide; providers enforce; absence has a declared default
67
+
68
+ - When a rule about *when/whether* to act is separable from the mechanics of acting (when to spill
69
+ an oversized result, whether an edit target is stale, when to compact), package the rule as a
70
+ swappable **policy plugin** over the capability seam rather than hard-coding it into the
71
+ provider. The provider keeps the atomic enforcement check at the moment of action (freshness,
72
+ no-clobber, bounds) because a policy's observation can be stale by the time the action runs.
73
+ - Declare what happens when the policy plugin is absent, and the declared default must fail toward
74
+ the safe direction for that capability's risk class — for a destructive-capable capability
75
+ (file overwrite, deletion, spend), absence means deny or require explicit approval, never a
76
+ silent degrade to the unguarded behavior. The cautionary shape: a runtime whose read-before-edit
77
+ observation policy is merely a removable plugin, with nothing beneath it, silently degrades to
78
+ unconditional overwrite when the plugin is missing — which is why the enforcement-layer
79
+ freshness/no-clobber check (`agent-command-sandbox.md` §Filesystem write authorization) stays
80
+ mandatory regardless of which policy plugins are installed, and a deployment-level opt-out of a
81
+ safety default must be explicit, never implied by absence.
82
+ - Do not let advisory policy observations become authority: a policy that watches tool traffic to
83
+ derive guard state (which files were read, which calls repeated) produces *hints and gates*, and
84
+ the enforcement layer re-validates at execution time.
85
+
86
+ ## 4. Sub-agent integration is a provider seam, not a hardcoded runner
87
+
88
+ - Expose sub-agent execution through the same seam/provider decomposition (§1), with one shape
89
+ difference: sub-agent providers are **named and co-resident** — a registry of concurrently
90
+ available providers selected per spawn (in-process fork, spawned worker, an external vendor
91
+ agent CLI) — unlike a single-selected executor seam. External vendor agent CLIs integrate as
92
+ first-class providers behind the seam, so consumers do not care whether a child is in-process or
93
+ a different vendor's product.
94
+ - Split the contract by lifecycle: a **one-shot** run (spawn → result) and a **continuable**
95
+ session (spawn → handle → further turns → teardown) are different APIs; do not overload one call
96
+ shape with both. For continuable children, the child-side provider emits only a **creation
97
+ spec**; the parent-side runtime owns handle allocation, turn delivery, and teardown — a provider
98
+ that never sees handles cannot leak or forge them.
99
+ - Make **fulfillment the single publish/ownership-transfer boundary**: a spawned child becomes
100
+ visible to consumers only when the provider fulfills the creation, and ownership (who tears it
101
+ down, whose scope disposes it) transfers exactly there — before fulfillment the provider owns
102
+ cleanup on failure; after, the runtime's scope does (§2). Two owners or zero owners at any point
103
+ is a leak or a double-free.
104
+ - This reference owns the seam shape only. Spawn-depth/live-count caps, capability intersection,
105
+ and cascade-cancel live in `retrieval-agent-safety.md` §Agent SDK Building Blocks; task-state
106
+ finality and recovery live in `agent-task-orchestration.md`; both apply in full to every
107
+ provider, including external-CLI children.
108
+
109
+ ## Non-negotiables
110
+
111
+ - A model-facing tool never depends on a concrete provider; consumers import the seam only.
112
+ - Every registration carries its inverse; unload correctness is owned once by scope disposal, and
113
+ hot-swap is dispose-then-reinstall, never patch-in-place.
114
+ - A policy plugin's absence has a declared, deliberate default; policy observations are advisory
115
+ and the provider re-checks atomically at action time.
116
+ - One-shot and continuable sub-agent APIs stay separate; fulfillment is the only ownership-transfer
117
+ point, and continuable child providers emit creation specs, never handles.
118
+ - These are composition patterns from one industry form — apply the swap/dispose/ownership tests
119
+ above as design gates, but do not cargo-cult the package layout onto a runtime whose capabilities
120
+ genuinely never vary.
121
+
122
+ ## Routing
123
+
124
+ - Tool registry/dispatch mechanics, result shaping, and output spill policy → `agent-tool-dispatch.md`.
125
+ - Startup trust, config generations, and cache invalidation on reload → `agent-runtime-bootstrap.md`.
126
+ - Sandbox policy resolution vs enforcement backends → `agent-command-sandbox.md`.
127
+ - Sub-agent spawn bounds and capability scoping → `retrieval-agent-safety.md`; task control plane → `agent-task-orchestration.md`.
128
+ - Session/scope persistence of what was installed when (for replay) → `agent-session-persistence.md`.
@@ -59,6 +59,41 @@ Express the filesystem/network posture as a closed enum of named profiles. A min
59
59
  Default network to **off** in every write-capable profile. Network is a separate grant from
60
60
  filesystem write (§4).
61
61
 
62
+ **One resolution owner; many enforcement backends.** Resolve "which profile + which writable
63
+ root(s) apply to this call" in exactly one place — a single policy owner that combines the
64
+ deployment default, the session's durable mode override, and any explicitly approved per-call mode
65
+ (explicit approval outranks session override outranks default — but only among profiles the
66
+ deployment/managed policy permits, and only after the §5/§6 authorization decision: a managed hard
67
+ deny or a mandatory sandbox ceiling is non-overridable by per-call approval, per §6's layered
68
+ precedence), and canonicalizes the session's
69
+ working root with filesystem semantics before it becomes the writable root. Every enforcing
70
+ backend — file tools, one-shot shell, persistent interactive sessions (§10) — consumes that one
71
+ resolved mode-and-root result per call, keeping only its platform dialect (§4) local. If each
72
+ backend re-resolves its own mode + root, they drift into a split world where the file layer and
73
+ the shell layer disagree about what is confined — a gap a command that touches both walks straight
74
+ through. One caveat is load-bearing: handing a backend the newly-resolved policy cannot *revoke*
75
+ a kernel sandbox an already-running persistent session (§10) was launched under — those sessions
76
+ bind the resolved policy and roots at creation, and whenever the creation-time grant is not a
77
+ subset of the current grant, or the policy/root generation has changed incompatibly (a swapped
78
+ writable root is not "narrower", yet the old root's access is revoked), the §10
79
+ generation-binding rule applies: tear down and recreate under the new sandbox, or reject the
80
+ command — never dispatch into the stale sandbox. (This is the policy/enforcement split of `agent-capability-composition.md` §3 applied to
81
+ the sandbox.)
82
+
83
+ **Present the resolved policy to the model as a replay-safe snapshot, not a capability inventory.**
84
+ Before each request, the model should see the currently-resolved policy as part of a cache-safe
85
+ runtime-context snapshot that is recorded into model-visible history — so replay can reconstruct
86
+ exactly what the model was told without rewriting the stable system prompt (prompt-cache churn;
87
+ see `llm-client-gateway.md`). Every generation persists for replay, but context construction
88
+ exposes only the *current* snapshot to the model and supersedes prior ones for subsequent
89
+ requests — after a broad-to-narrow policy change, stale snapshots left in the live envelope keep
90
+ advertising permissions and roots the enforcement layer no longer grants (denied attempts,
91
+ needless escalation, stale-path disclosure). Write that policy text as *what the policy means for operations*,
92
+ never as an enumerated capability/tool inventory (the maybe-absent-tool steering anti-pattern in
93
+ `agent-tool-dispatch.md`), and include the instruction that the model should not refuse an action
94
+ merely because policy *might* deny it — attempt the tool and follow the denial/escalation guidance
95
+ (§7) — otherwise the model self-censors more broadly than the policy actually restricts.
96
+
62
97
  **Reads need a confidentiality boundary too, not just writes.** A profile that confines *writes* but
63
98
  allows reading the whole host is still an exfil hole: an auto-approved "safe" read (`cat`, `rg`,
64
99
  `sed`, `find`) can slurp `~/.ssh`, `~/.aws`/cloud-credential files, browser profiles, keychains,
@@ -59,6 +59,21 @@ Use a background writer fed by a queue with explicit control messages:
59
59
  or a transcript consumer that a turn/session completed, flush the log. A "done" that outraces the
60
60
  writer can lose the last turn on crash. (This mirrors the background-task finality rule in
61
61
  `terminal-cli-dev`: retain final output until consumers have acknowledged it.)
62
+ - **Checkpoint before external commitments, not only before finality.** Flush the recorded log
63
+ prefix durable *before a model adapter receives a request*, *before any tool call may produce an
64
+ external side effect — nested and delegated calls (code-mode sub-calls, sub-agent tool use)
65
+ included, since they route through the same effect surface*, and at each pre-step boundary — and treat a rejected/failed checkpoint
66
+ as blocking that dispatch or side effect (fail closed), not as a warning. Finality-only flushing
67
+ leaves a mid-turn window where a crash loses the prior response and ordered tool results that an
68
+ already-executed side effect or the next request acted on. The flushed prefix alone still cannot
69
+ distinguish *never dispatched* from *dispatched, outcome unknown* after a crash — so pair the
70
+ checkpoint with a durable dispatch-intent/attempt record appended-and-flushed before the action
71
+ and a completion record after: for model requests that is exactly §3's attempt-state rule
72
+ (prepared → dispatch-attempted → accepted/uncertain/rejected, idempotency-keyed); apply the same
73
+ shape to non-idempotent tool side effects, where an unresolved attempt is reconciled or surfaced
74
+ as a typed indeterminate outcome — never auto-retried (the task-level counterpart lives in
75
+ `agent-task-orchestration.md`'s effect-finality rule). (This is the event-log half; the §3
76
+ accounting records carry their own commit-before-dispatch rule — complementary, not duplicates.)
62
77
  - **Make writer failure observable.** If the writer task dies (disk full, IO error), record a
63
78
  terminal-failure state that every later append/flush call surfaces — never let the recorder keep
64
79
  accepting events into the void. A dropped persist is a data-loss bug, not a best-effort log line.
@@ -323,7 +338,42 @@ When a session's estimated context approaches the model window, compact it.
323
338
  baseline and a window ordinal so "how much have we grown" is measured since the previous compaction,
324
339
  not from zero. **Prefer server-reported token usage over local estimation** when the provider
325
340
  returns it; estimation is the fallback that drives the trigger when no server count is available
326
- yet.
341
+ yet. Feed the trigger from **one unified token meter**, not per-consumer recounts: a single
342
+ replay-aware event source (rebuilt from the log on resume/replay, per §4's restore-accounting
343
+ rule), deduplicating repeated *ingestion of the same attempt* (a replayed or re-read response)
344
+ by attempt/request id. Keep two projections over that one source, and derive each correctly:
345
+ the **logical/pressure projection** (compaction trigger, budget displays) measures the *current
346
+ model-visible envelope*: the latest authoritative envelope measurement — server-reported context
347
+ size bound to the *accepted canonical attempt*, not merely the newest report — **plus the
348
+ estimated size of every model-visible item appended since that measurement** (a large tool
349
+ result landing after the report otherwise dispatches an over-window request unpruned), or a
350
+ fresh full-envelope estimate taken before dispatch; never a sum across sequential requests,
351
+ which re-counts the resent history every turn and trips the watermark while the real context
352
+ still fits; the
353
+ **per-attempt/billing projection** sums every real provider attempt — a genuine retry consumes
354
+ and may bill tokens even when its payload duplicates the first attempt. Collapsing the two
355
+ either underreports spend or triggers needless lossy compaction; consumers that recount for
356
+ themselves disagree about when pressure exists.
357
+ - **Run a deterministic pruning layer before the summarization layer.** Before invoking any
358
+ model-written summary, apply a model-free, replay-safe pruning pass over the candidate window —
359
+ dropping or trimming stale tool-call/tool-output payloads under the same keep/drop rules — and
360
+ only summarize if the window is still over pressure afterwards. Pruning is deterministic, cheap,
361
+ and reversible in design terms; jumping straight to LLM summarization pays nondeterminism and
362
+ fidelity loss for reduction that pruning could have achieved. Declare which of two shapes the
363
+ pruning layer is, because a prune that relieves pressure without a summary has no summary-bearing
364
+ compaction marker to ride on: **ephemeral** pruning is applied at context assembly from the same
365
+ versioned rules every time, never mutates the log, and replay reproduces it by re-running the
366
+ rules; **committed** pruning persists a typed prune-only marker — covered range, rule/version,
367
+ resulting kept-set references, and the window-baseline update, with the summary field legitimately
368
+ absent — so resume neither reconstructs the unpruned envelope nor re-triggers compaction. An
369
+ unclassified prune that changes the model-visible envelope without either contract breaks
370
+ faithful replay. (The trim-to-fit rule below stays as the last-resort fallback *after* a
371
+ summary; this layer runs *first*.)
372
+ - **Manual/user-invoked compaction fails with typed codes, not a generic error.** Distinguish at
373
+ least: another compaction already in flight, cancelled, context changed since the request,
374
+ summarization failed, and commit/persistence failed — the caller's correct reaction (retry,
375
+ re-read, give up, surface data-loss risk) is different for each, and a generic failure trains
376
+ users to spam retry across all of them.
327
377
  - **Two strategies, chosen by provider capability:** *local/inline* compaction (you send the history
328
378
  to the model with a dedicated summarization prompt and replace it with the returned summary) or
329
379
  *provider-remote* compaction (the provider compacts server-side). Pick per provider; do not assume
@@ -361,7 +411,9 @@ When a session's estimated context approaches the model window, compact it.
361
411
  - Crash-safety requires the §2 preconditions: one exclusive writer per session, whole-record atomic
362
412
  append, `fsync` before acking Flush, and a monotonic sequence number.
363
413
  - Flush the log before signaling turn/session finality; never let a completion signal outrace the
364
- writer.
414
+ writer. The same checkpoint discipline gates external commitments mid-turn: flush the recorded
415
+ prefix before model dispatch and before a tool's external side effect, and a failed checkpoint
416
+ blocks the action, fail closed.
365
417
  - A failed persist/flush is raised and made observable, never swallowed as empty success.
366
418
  - Always persist the structural markers (session meta, turn context, compaction) that replay needs.
367
419
  - Truncating a persisted payload is allowed only if it matches what the model saw (or the full
@@ -51,6 +51,17 @@ When the model requests multiple tool calls in one response and the model/runtim
51
51
  - Apply post-tool hooks; let them inject additional context, treated as untrusted (see `agent-lifecycle-hooks.md`).
52
52
  - Emit per-call telemetry (tool name, decision source, duration, outcome) keyed by call id.
53
53
  - Bound result size fed back to the model; truncate large tool output **visibly and structure-aware** — truncating in the middle of structured (JSON) output yields unparseable downstream content. Use structure-aware markers or summarize, don't byte-chop.
54
+ - When bounding **replaces** an oversized plain-text result with a persisted artifact (spill), split the capability per `agent-capability-composition.md` §1: a **spill-store seam** owns persistence only (save verbatim, return an opaque locator plus retrieval hint; reject on real storage failure), swappable providers own the backing store, and a **policy plugin** owns *when* to replace — the model-visible replacement is a bounded head/tail preview plus the locator and an explicit omitted-byte count, never a silent cut. Three semantics make this safe:
55
+ - **Spill is fail-open for the call it wraps, but never for the bound**: a spill failure (no store, save rejected) never converts a successful tool call into an error result — and it also never returns the unbounded original, which would defeat the model-context bound above; the fallback is a clearly-labeled bounded truncation ("full result unavailable", omitted bytes counted). (Persistence-failure *labeling* discipline: §Persisted tool-output artifacts below.)
56
+ - **Exclude read-back tools from model-facing replacement**, or the model reads a file, gets a spilled preview pointing at a file, and reads again — a read→spill→read loop. A durable-log copy of the same output is not model context, so bound it on its own arm with the same cap; the loop exclusion does not apply there.
57
+ - Replacement operates on the **final formatted model-facing result**, not the tool's internal/canonical value — provider-side caps stay mandatory and separate, and the programmatic value consumed downstream is preserved unchanged.
58
+
59
+ ## Composable execution guards
60
+
61
+ Cross-cutting per-call guards compose as wrappers around dispatch rather than being woven into the router — but the two below sit in different risk classes, and only one is removable (the absent-plugin default rule of `agent-capability-composition.md` §3 applies):
62
+
63
+ - **Repeat-call reminder** (advisory, removable): detect an identical tool call repeated within the turn/window and inject an advisory notice into the result path so the model can break the loop. Advisory only — it nudges, and must never be relied on as a denial boundary (policy/sandbox layers own denial).
64
+ - **Cooperative timeout** (mandatory enforcement; only the packaging is pluggable): the tool declares its timeout budget and the contract that it honors the abort signal; the wrapper owns the deadline, returns a **typed** timeout failure the model can distinguish from an execution error, and never abandons the still-running work — cancel **and join** it (same discipline as Parallel execution above), or the "timed-out" tool keeps mutating state after the model moved on. Deadline enforcement is a dispatch-lifecycle obligation, not an optional nicety: a deployment without the wrapper plugin still owes a bounded default deadline with cancel-and-join (fail closed) in the dispatch/provider path — an unloaded plugin must never mean "tools run unbounded". With nested deadlines (an outer turn budget and an inner per-call budget), attribute the firing deadline correctly: an inner timeout is that call's typed failure, not a turn abort, and vice versa.
54
65
 
55
66
  ## Code-mode: tools as a code/exec runtime (optional advanced pattern)
56
67
 
@@ -161,3 +161,9 @@ Rollback path post-promotion:
161
161
  - Force a canary abort → traffic returns to stable within seconds (mesh weight propagation time).
162
162
  - Trigger header-lane canary with the opt-in header → request lands on canary subset; without header → lands on stable.
163
163
  - Mirror enabled → mirror target sees traffic; primary handler is unaffected; metrics tagged distinctly.
164
+
165
+ ## Topic-extension backlog
166
+
167
+ Entries here are registered candidates, not adopted guidance. Each names the candidate, its evidence status, and the condition that unblocks adoption; the round that evaluates one records keep/narrow/discard against its entry.
168
+
169
+ - **Application-layer rolling provider transition (deferred candidate).** Observed form, from one agent-native product repository (evolving portfolio): a new backend provider registers with an in-process broker as an additional route, traffic shifts to it by weight, and the old provider unloads once it has no in-flight work — demoting blue-green from an infrastructure operation to an in-process composition pattern. Evidence status: hypothesis-grade — single source, and the source itself labels the account observational. Do not land this as executable rollout guidance until at least two independent external sources corroborate the pattern in production use; re-evaluate at the next design round touching provider transition, blue-green, or in-process traffic shifting.
@@ -12,6 +12,7 @@ Use this as the first reference for Python backend, Python microservice, AI-serv
12
12
 
13
13
  - Decide the backend shape first: API service, internal service, worker service, scheduled job service, SDK/package support library, CLI, AI-service host, or modular monolith.
14
14
  - Use separate deployable service boundaries only when ownership, scaling, data ownership, runtime isolation, deployment cadence, or rollback needs justify them. Keep one modular service/package/script shape when boundaries are unclear or the scope is too small for separate services.
15
+ - External grounding for the boundary rule above (adopted in part): [Bounded Context](https://martinfowler.com/bliki/BoundedContext.html) (Fowler's overview of the central DDD strategic-design pattern; origin credited there to Eric Evans, *Domain-Driven Design*) — a boundary is where one internally consistent model stops being valid, so boundaries follow model and responsibility lines, never noun count — and [Conway's Law](https://martinfowler.com/bliki/ConwaysLaw.html) (Melvin Conway, ["How Do Committees Invent?", 1968](https://www.melconway.com/Home/Committees_Paper.html); Fowler's overview) — a system's structure mirrors the builders' communication structure, which is why team ownership and communication structure are a legitimate split input. Two limits: the ownership, scaling, release-cadence, and rollback tests in the rule above are this skill's own operational criteria, attributed to neither source; and a bounded context does not require its own deployable service — the modular-monolith default stands, so these sources justify model/ownership boundaries, not a service-per-context split. Borrowed scope: boundary-input criteria only; no claim to the full DDD strategic-design method (context maps, ubiquitous language) or an inverse-Conway process. Sibling: `go-microservice-architecture/references/architecture-playbook.md` ("Service Boundary Rules") carries the same grounding; keep the two in sync.
15
16
  - For each Python microservice that is justified, define the owner, contract, data source of truth, inter-service auth, timeout/retry/fallback policy, deployment unit, canary/rollback, and observability.
16
17
  - Separate layers:
17
18
  - transport/framework: routing, auth, validation, response mapping
@@ -18,3 +18,9 @@ Use this for pyproject, uv/poetry/pip, lockfiles, tooling, containers, and deplo
18
18
  - Release readiness includes migrations, startup validation, smoke tests, canary, rollback, and observability checks.
19
19
  - **ASGI server choice has expanded beyond `uvicorn` / `gunicorn+uvicorn` / `hypercorn`** — `granian` (emmett-framework, Rust-based) is the current credible high-throughput alternative for ASGI services, supporting ASGI/3, RSGI, WSGI, HTTP/1, HTTP/2, TLS, WebSockets (HTTP/3 planned per the project README's "eventually 3" roadmap note — verified not shipped as of May 2026). Per the Granian project's own `benchmarks/vs.md` and third-party load-test repos (e.g., `piccolo-orm/asgi_server_performance`, `synodriver/asgi-server-benchmark`), ASGI echo on 10KB payload typically lands granian > uvicorn-httptools > hypercorn by roughly the ratios 58k / 51k / 8k RPS in April-2026-era runs; file-serving gap is wider (granian ~47k vs uvicorn ~18k via `pathsend`). Treat the absolute numbers as benchmark-snapshot-specific; re-run against your workload before basing a switch on them. Architecture impact: when serving throughput is the binding constraint, granian can buy headroom without rewriting the **plain ASGI path**. **"Without rewriting" caveats**: granian's worker / process model differs from `gunicorn+uvicorn` fork-based workers (Rust runtime + Python interpreters with different lifecycle hooks); ASGI lifespan events, contextvars propagation across worker boundaries, custom signal handlers, prometheus/metrics exporters tied to uvicorn internals, and ASGI middleware that depends on uvicorn-specific behavior all need smoke-testing on granian before a switch. If the team plans to adopt granian's bespoke RSGI protocol for max performance (rather than ASGI), application code that uses ASGI-specific middleware, ASGI scope manipulation, or third-party ASGI libraries WILL need rewriting — RSGI is a different protocol, not a faster ASGI. Trade-offs: smaller operational maturity, fewer community recipes, Rust-runtime-on-the-side observability differs from a pure-Python server. **Choose uvicorn** for ecosystem maturity, broad reference material, and known operational patterns; **choose granian** when (a) profiled benchmarks on your workload show uvicorn saturation, (b) the team has Rust-toolchain debugging capacity, (c) the deployment story can absorb a less-common runtime. Hypercorn remains the choice when HTTP/2 + ASGI under pure-Python ops matters more than peak throughput.
20
20
  - **Python runtime version baseline (2025-2026)**: Python 3.13 (released October 2024) ships **experimental** free-threaded build per PEP 703 — GIL-disabled, ~40% single-threaded perf hit per python.org "What's New in 3.13" notes, used at the team's risk for parallel-CPU workloads. Python 3.14 (released 7 October 2025) advances free-threading to **supported (Phase II of PEP 703)** per PEP 779 — meaning the free-threaded build is a first-class supported configuration, NOT that it is the default Python build or the default production choice. Per the python.org free-threading howto, the single-threaded penalty narrowed to ~5-10% (specializing adaptive interpreter re-enabled thread-safely); PEP 803 defines the `abi3t` stable ABI for free-threaded C extensions. Architecture impact: for services where parallel CPU work matters (in-process ML inference fan-out, heavy parsing, parallel compression), 3.14 free-threading is the first version where adoption is **a defensible experiment for a selected service**, not yet a defensible default. Pre-flight burn-in required before any production switch: (a) C-extension readiness — all extensions in the service's dependency tree must declare free-threading support; many popular extensions (numpy, pandas, lxml, psycopg native bits, asyncpg native bits, pillow, cryptography) were still mid-migration at 2026-Q1, verify per-version; mixing GIL-only and free-threading-aware extensions in one process is unsupported; (b) GC and runtime behavior under sustained threading load differs from GIL build — measure tail latency, memory residency, and CPU efficiency on the actual workload; (c) debugger / profiler ergonomics (gdb, pdb, py-spy, scalene, prometheus exporters) may have rough edges on free-threaded builds; (d) library-level thread-safety: code paths that were "implicitly safe because of the GIL" can race in free-threaded mode (singletons built at import time, module-level mutable caches, third-party libraries that rely on GIL-protected dict mutation). For pure I/O-bound services, stay on stock GIL build — the 5-10% overhead is pure cost. For services on 3.13 or older, treat free-threading as opt-in research, not default.
21
+
22
+ ## Topic-extension backlog
23
+
24
+ Entries here are registered candidates, not adopted guidance. Each names the candidate, its evidence status, and the condition that unblocks adoption; the round that evaluates one records keep/narrow/discard against its entry.
25
+
26
+ - **Package-owned runtime-invariant registries (weak-keep candidate).** Observed form, from one agent-native product repository (evolving portfolio): each package contributes its own invariant checks from a companion module, normal entrypoints do not import the diagnostics layer, and allowlist/blocklist configuration selects which checks run. Evidence status: weak — single source; generalization and owner placement are undecided. Evaluate at the next Python architecture round touching runtime readiness, startup validation, or diagnostics, and record the outcome against this entry. The Go-stack sibling registration lives in `../go-microservice-architecture/references/cross-cutting-concerns.md`.
@@ -271,6 +271,7 @@ skill 改动后,让 agent 重跑这条 trace,**结构性偏离 = 回归信
271
271
  - **组合分层**:运行时 = 有序的 profile/bundle 层叠加出来的插件树,每层可用 patch 覆盖任一行配置;能打印出机器实际启动的树来核对。
272
272
  - **决策在执行处强制**:schema 省略、prompt 过滤、facade、监听顺序都不是权限边界(已进 `llm-inference-integration` / `product-rd-workflow` 评审清单)。
273
273
  - 与之相对,把插件系统的**依赖注入与跨插件解析**做得很重是否值得,业界有争议(一线 harness 作者的公开评价:多数插件互不依赖,复杂 DI 在 90% 场景不带来收益,且跨插件类型仍需另解);本仓不采纳"人人可发的插件生态"作为 skill 分发形态。
274
+ - **可执行落点在 llm-inference-integration,本节只留形态记录**:设计或评审 agent 运行时按该 skill 的 `agent-capability-composition.md`(能力三元、可逆注册、子 agent provider seam)、`agent-tool-dispatch.md`(结果外溢与护栏 wrapper)、`agent-session-persistence.md`(剪枝先于摘要、请求前 checkpoint)、`agent-command-sandbox.md`(policy 单一解析 owner)执行;不得再从本节直接提炼可执行规则——那会与 owner 侧产生双写漂移。
274
275
 
275
276
  不含日志/持久化表述——模型可见内容与事件日志的账目不变量另由安全 owner 参与的独立设计处理。
276
277
 
@@ -116,6 +116,9 @@ that no longer exists — read an older row's owner through this mapping.
116
116
  | A multi-round review program carries four process controls: a broken chain binding after an owner-file fix is a by-design dead-end recovered by an interim checkpoint naming each lane's un-run remainder plus a human continuation authorization that extends rounds without waiving any lane; convergence and closure declarations are written falsifiably with named axes and named open items; remediation text re-owes the pre-cover axes before returning to the reviewer, with a third same-class round escalating to one full-matrix self-enumeration; and ledger rows land append-once, final-form, in the same squashed round partition as the changes they declare | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; `references/dual-track-review-gate.md` (three merged clauses); `references/source-register.md` (Round-consolidation rule, closing the follow-up recorded in the routing-round note); `scripts/test_ai_coding_implementation_gates.sh` family 9. Observed failures: a tracked chain's content binding broke by design after an owner-file fix and the recovery pattern was improvised; a self-audit's "full lifecycle" claim was caught one axis short by the final challenge; a program burned twenty-plus single-finding rounds before one full-matrix enumeration closed the class; two documented append-only bends shared the rows-written-per-round shape. RED-baseline: 40 family-9 assertions (one per obligation sentence or sub-clause, per the reverse-coverage rule landed alongside; clause pins bound to their owning rule-line, anchors section-bound, ledger pins paragraph-bound), shown red under 81 applied mutations in a throwaway copy — one deletion and one relocation per pin, two cross-moves between the convergence and continuation rule-lines, one same-bullet pointer removal — first-failing label equal to the owning assertion, unmutated controls green before and after |
117
117
  | For a contract or prose artifact pinned by a fixture family, pin coverage runs artifact-to-pin — every obligation sentence names the pin that reds when it is deleted, because the walk proves only the pins that exist — and a self-contained walk owes the relocation, reachability, tree-isolation, and parser-completeness probes | `testing-strategy` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#coverage runs in the reverse direction | updated | `testing-strategy/SKILL.md` (reverse-coverage sub-bullet with dual-side pointer); `testing-strategy/references/run-killing-mutation-walk.md` (probe classes). Observed failure: a reviewer found six silently-deletable unpinned obligations after multiple green walks; the sentence-by-sentence enumeration then closed all six at once and was recorded only in a round spec until this landing. RED-baseline: the reverse-coverage and probe pins are in family 9's applied-mutation walk above, including the same-bullet pointer-removal mutation red on the reachability assertion |
118
118
 
119
+ | An agent runtime's swappable capability decomposes into seam, provider, and model-facing tool with registration reversible via scope-owned disposal and sub-agent execution behind a named multi-provider seam whose only ownership transfer is fulfillment; oversized tool results spill through a store-seam/policy-plugin split that is fail-open for the wrapped call and excludes read-back tools from replacement; the session log checkpoints before model dispatch and before tool side effects, compaction runs a deterministic pruning layer before any summarization and reads one replay-aware token meter, and sandbox policy resolves in one owner whose result every enforcement backend consumes — recorded from one agent-native product repository (evolving portfolio) as one industry form, not a standard | `llm-inference-integration` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/llm-inference-integration/references/agent-capability-composition.md#A model-facing tool never depends on a concrete provider | updated | New `llm-inference-integration/references/agent-capability-composition.md`; merged clauses in `llm-inference-integration/references/agent-tool-dispatch.md` (spill split + composable execution guards), `llm-inference-integration/references/agent-session-persistence.md` (§2 checkpoint-before-external-commitments, §5 pruning-before-summarization, unified token meter, typed manual-compaction failures), `llm-inference-integration/references/agent-command-sandbox.md` (§2 single resolution owner + replay-safe model-facing policy snapshot); pointer rows in `llm-inference-integration/SKILL.md`. RED-baseline (applied, differential; evidence scope: discoverability/package integrity only, not per-mechanism semantics): removing `agent-capability-composition.md` in the worktree turns `validate-skill.sh` red with `missing_markdown_references` naming exactly the two new `SKILL.md` pointers; restoring returns `skill_validation_ok` — control green before mutation. The mechanism clauses are reference prose with no per-clause mechanical oracle in this repo; their semantics (merge-not-append, no dropped guard, fail-open scoped to the call not the bound, non-overridable managed denies, two-projection token meter) rest on this round's independent review + adversarial challenge rows — round-1 review returned five P1 findings against exactly these clauses and each was applied before the final pair — plus re-confirmation against a fresh same-version source snapshot before drafting |
120
+ | A doc-layer industry-form record hands its executable landing to the owning skill and says so in place, so a later round extends the owner instead of re-extracting duplicate rules from the record | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/harness-patterns-and-eval.md#可执行落点在 llm-inference-integration | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; `references/harness-patterns-and-eval.md` §7 pointer bullet naming the four llm-inference owner files; this ledger's round append. RED-baseline (applied on the unpushed landing commit, differential): rewording the anchored rule line inside the committed round without updating this row turned `impact-chain-gate.rb` to exit 1 with `impact_chain_firing_path_missing` naming this owner; restoring the line and re-pressing returned exit 0 — the binding this proves is row↔landed-rule-line: a later edit that rewords the landed pointer without a superseding row is caught mechanically. Two null probes are recorded honestly: a worktree-only deletion does NOT red the gate (the locator resolves at the round head), and deleting the whole bullet from the committed round also stays green because it removes the owner's changed status along with the anchor — so the binding guards rewording, not removal-with-round-rewrite. The anti-drift behavior itself is prose-level and rests on the round's independent review/challenge rows (the pre-fix candidate's review returned five applied P1s; the tracked pairs on the fixed candidates are in the round evidence) |
121
+
119
122
 
120
123
  Gate: a changed upstream owner requires a non-empty row with an allowed terminal status and evidence naming that owner's SKILL.md. A header-only table is absent. Keep project-specific provenance in a private task artifact, never in this distributed register.
121
124
  | An evidence row is judged against the round it landed in, never the accumulating range; classification and presence both narrow to that round while the owner-level RED floor stays cumulative | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | updated | `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/impact-chain-gate.rb`; `skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh`; `skill-extraction-workflow/references/external-practice-controls.md` |
@@ -252,3 +255,9 @@ and inverted the sense (production, not product), and the coordinator now shares
252
255
  | Blocked-verification remediation gains the sandbox-denial triage and controlled-escalation rule in its safe form: a failing required command is first classified from observed denial evidence — never a bare non-zero exit — as sandbox denial or genuine failure, with indeterminate and mixed/conflicting evidence failing closed as genuine and the denial required to be the sole proximate cause preventing completion; escalation is never a default action and never bypasses a genuine failure or a product sandbox under test; a controlled escalated re-run needs every condition to hold — the command reviewed-trusted this session by the operator side with trust bound to recorded content identity and the re-run executing the reviewed snapshot or verified-then-executed bytes, failing closed on mismatch (a repository under extraction is untrusted input, its entrypoints qualifying only after in-session content review), the narrowest blocked capability only with the command re-run unchanged once per approval, and host policy plus per-escalation user approval bound to this command, this capability, this session, never cached or standing — otherwise the item stays blocked with the normal remediation record. The order deliberately inverts the source rule's retry-before-diagnosing default, whose premise (a repo's own trusted commands) does not hold for a skill set facing arbitrary repositories | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#Sandbox-denial triage precedes any escalation | updated | Owner key `skill-extraction-workflow/SKILL.md` with the rule body in `references/source-to-skill-extraction.md` (two bullets in Blocked Verification And Source-Read Remediation; entrypoint untouched per approved landing surface D-6). Observed failure: the 023 review P1 (chain `023-plan-r10`) — the source rule generalized as written lets a repo-controlled command use a sandbox denial as a pretext to obtain host credentials or network — which moved E5 out of the landing batches under the pre-registered dedicated-spec rule. This landing is that dedicated round: the named security owner approved all six decision points in-session on 2026-08-21 before implementation. Review round 1 (organization gate, chain `031-e5-r1`, codex) P1 accepted — trust bound to the command string is a TOCTOU hole (the blocked first run or a concurrent actor rewrites what the unchanged command resolves to); amendment A1 binds trust to content identity; review round 2 (user-authorized after the autonomous budget checkpoint, chain `031-e5-ua1`, codex) found two further P1s, both accepted — verify-then-execute on a live path left a swap window, closed by amendment A2 (the re-run executes the reviewed snapshot or the exact verified bytes), and a pending final-wording approval must block the landing, so the amended wording's re-approval is decision row D-7 and gates push/MR; review round 3 (user-authorized, chain `031-e5-ua2`, codex) found the atomicity requirement bound only to the entrypoint, closed by amendment A3 (chain-wide: every repository-controlled executable or loadable code unit executes from the reviewed snapshot or is verified atomically at its point of use); review round 4 (user-authorized, chain `031-e5-ua3`, codex) falsified round 3's closure claim — sealing keyed to code left configuration, data, environment files, and symlink targets swappable — closed by amendment A4, predicated on effect rather than file kind: every repository-controlled input that can affect privileged behavior is sealed, an unsealable input refusing the escalation back to blocked (A1–A3 are its special cases); review round 5 (user-authorized, chain `031-e5-ua4`, codex) found "sealed" was a label while non-code consumption mechanics stayed code-only, closed by amendment A5: every behavior-affecting input consumes from the reviewed snapshot or from bytes opened, verified, and held stable through use; review round 6 (user-authorized, chain `031-e5-ua5`, codex) widened the origin — content fetched over the granted network/IPC/host channel is equally mutable — closed by amendment A6: the quantifier is every untrusted mutable input that can affect privileged behavior, unsealable remote inputs refusing the escalation; review round 7 (user-authorized, chain `031-e5-ua6`, codex) found approval bound only to command/capability/session could recycle a stale grant onto content re-reviewed after a mismatch, closed by amendment A7: approval also binds the reviewed content identity, a mismatch invalidates it, and re-reviewed content needs fresh user approval; review round 8 (user-authorized, chain `031-e5-ua7`, codex) found "verified" unanchored for granted-channel bytes, closed by amendment A8: verification anchors to an immutable expected identity reviewed before escalation and bound into the approval, no establishable pre-grant identity stays blocked, and the clause states the net invariant — privileged behavior under the grant derives only from operator-reviewed, approval-bound content and trusted host state; adversarial challenge round 1 (user-authorized, chain `031-e5-ua10`, codex) found mixed evidence could launder a genuine failure behind a deliberately-triggered denial, closed by amendment A9: the denial must be the sole proximate cause preventing completion, mixed or conflicting evidence classifying as genuine; adversarial challenge round 2 (user-authorized, chain `031-e5-ua11`, codex) found "observed denial" carried no provenance requirement so command-controlled stderr could fake one, closed by amendment A10: denial evidence comes from the host sandbox's own trusted enforcement or telemetry channel bound to the exact invocation and denied capability, command output alone never qualifying; adversarial challenge round 3 (user-authorized, chain `031-e5-ua12`, codex) found the denial record unbound to content identity so replaced content could ride the prior denial, closed by amendment A11: qualifying denial evidence comes from an unprivileged run of the same reviewed content identity the escalated re-run executes, a mismatch invalidating the denial record along with the approval; adversarial challenge round 4 (user-authorized, chain `031-e5-ua13`, codex) found approval bound a re-resolvable capability name, closed by amendment A12: the grant binds the canonical resolved capability identity under host policy, re-resolved and verified atomically at use, any resolution change invalidating approval and denial record alike; adversarial challenge round 5 (user-authorized, chain `031-e5-ua16`, codex) found the denied first run's completed side effects could be duplicated by the unchanged re-run, closed by amendment A13: conjunctive condition 5 requires prior external effects proven absent, rolled back, or contained before any re-run, unaccountable effects staying blocked; review round 16 (user-authorized, chain `031-e5-ua17`, codex) found containment alone left the duplicate to materialize at commit/export, closed by amendment A14: contained state is discarded or reset to its pre-run snapshot before the re-run or covered by a demonstrated idempotency or deduplication guarantee; adversarial challenge round 6 (user-authorized, chain `031-e5-ua18`, codex) found the effect-accounting evidence itself unprovenance, closed by amendment A15: accounting evidence carries the denial-evidence provenance discipline, bound to the invocation, content identity, and affected targets, command output alone never proving absence, rollback, or deduplication. RED-baseline: applied mutations in a throwaway copy (never the live tree), unmutated control green before and after, each mutant observed to red on its OWNING assertion with differential attribution (the fixture exits at first failure, so every earlier assertion passed under the mutant); a tree-isolation probe mutated the copy while the live tree's fixture stayed green, proving the harness read the mutated tree. Evidence: `specs/031-e5-controlled-privilege-escalation/plan.md` |
253
256
  | The controlled-escalation clause is pinned by the deterministic fixture that already pins the relocated-rule, claim-liveness, and model-visible accounting families — one section-bound assertion per obligation the clause imposes (triage before escalation, observed-denial evidence with host-channel provenance and the command-output-never-qualifies rule, fail-closed indeterminate and mixed evidence with the sole-proximate-cause requirement, no bypass of genuine failures, the conjunctive condition gate, in-session command trust, untrusted repo entrypoints, grant minimality, no general sandbox disable, the grant binding the canonical resolved capability identity with re-resolution-at-use mismatch handling, the single unchanged re-run per approval, content-identity binding with re-verification at the escalated re-run and fail-closed mismatch handling, snapshot-or-verified-bytes execution closing the verify-to-execute swap window chain-wide for every unit in the runtime chain, effect-keyed sealing of every untrusted mutable input that can affect privileged behavior — granted-channel content included, verification anchored to a pre-reviewed approval-bound expected identity, and the net derives-only-from-reviewed-content invariant — with unsealable inputs refusing the escalation, approval binding including the reviewed content identity with mismatch invalidation and fresh approval after re-review, no cached or standing approval, the untouchable product sandbox under test, denial evidence on every outcome, and the condition-5 accounting of the first run's external effects before any re-run), with the same stated limits as those families: the pins catch deletion, rewording-away, and relocation, not a weakening sentence added beside them — catching a deliberate weakening stays the dual-track review's job | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`; changed surface under it is `scripts/test_ai_coding_implementation_gates.sh` (pin family 8 in the existing fixture, one section-bound assertion per obligation; the set is described rather than counted, because a count in an append-only row goes stale on the next assertion and cannot be edited afterwards). Observed failure: same 023 review P1 as the row above — the failure class lives in shared-skill rule text, so its prevention pins belong in the shared fixture that already guards relocated rules. RED-baseline shared with the row above: the same applied-mutation set in the same throwaway copy reds each owning assertion against a green unmutated control, and the copy-vs-live isolation probe proves the harness read the mutated tree rather than the live one. Evidence: `specs/031-e5-controlled-privilege-escalation/plan.md` |
254
257
  | A no-owner deliverable doc is drafted from a locked charter plus user- or session-supplied substance — the drafting mode routes to an owner while substance is unsettled and never invents substantive decisions, mid-stream asks classify against the charter instead of appending reactively, and a loaded platform/tool skill is not the deliverable owner | `tighten-doc` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/SKILL.md#Route to an owner when substance is unsettled | updated | `tighten-doc/SKILL.md` is the owner key (description no-owner-draft trigger + Draft-mode contract); charter method in the doc-charter-first reference of the same package; the tool-skill-masking sibling variant landed in this round's extraction-workflow reference, whose owner is covered by the process-controls rows above. Observed failure: a multi-round deliverable-doc effort where the finalization owner never fired until the user asked and the doc accreted reactively without a charter. RED-baseline: the cold-start routing miss frozen as an eval fixture in `eval/routing-tasks.jsonl`, exercised by the Tier-1 routing analyzer on the landing candidate. Corrective insertion: the round that landed this change merged with the impact-chain gate red and no row; this row was inserted by a corrective rewrite of the integrating merge per the round-023 precedent, authored post-merge from that round's recorded PR evidence and diff. |
258
+ | External-provider recovery paths are proven by layered fixtures below the live lane — a policy-free protocol-real fault server scripting an explicit transport/protocol/semantic fault taxonomy, recorded replay at the adapter seam with explicit overrides for cases a recording cannot reconstruct, and a credentialed live lane kept to wiring sanity; a failing coverage gate prints uncovered `path:line:col` and stays silent when green, per-file thresholds stop cross-file subsidy, conditional exemptions derive from the same probe that skips the suites, thresholds run once on the merged partition report, and perf/stress lanes sit outside gates by auditable config include inventory rather than skip marks; duplication and dead-code gates run with explicit thresholds and one owning tool per finding axis, and derived committed artifacts are caught stale at the pre-commit hook — reject-and-instruct by default, auto-staging only under genuine whole-state atomicity (lock plus crash-safe rollback) that most hook tooling cannot guarantee — while the test lane's freshness assertion over the committed tree stays the gate — recorded from one agent-native product repository (evolving portfolio) as one industry form, not a standard | `testing-strategy` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#bind each scripted behavior to its intended attempt | updated | `testing-strategy/references/ci-fixtures-and-flake-control.md` (new Fault-Injection Layers subsection under Fixtures; duplication/dead-code and regenerate-not-reject merged into CI Gate Design; location-naming, per-file, exemption-probe, partition-merge, and lane-inventory bullets merged into Coverage As Signal), `testing-strategy/references/test-code-authoring-patterns.md` (§3 config-inventory lane option, §6 per-file and path:line:col merges), `testing-strategy/references/e2e-real-flow-testing.md` (E2E Scope Control routing bullet), `testing-strategy/SKILL.md` (one-clause pointer beside the test-double rule; description untouched). RED-baseline (applied, differential; evidence scope: discoverability/package integrity only, not per-clause semantics): removing `ci-fixtures-and-flake-control.md` in the worktree turns `validate-skill.sh` red (exit 1) with `missing_markdown_references` naming exactly the five `testing-strategy/SKILL.md` pointers to that file; restore returns `skill_validation_ok`; control green before mutation. Clause semantics rest on this round's independent review and adversarial challenge rows plus re-confirmation of every borrowed claim against a fresh same-branch source snapshot before drafting; source threshold values deliberately not copied (per-file value stays the team's choice). Source class: the same agent-native product repository's test-engineering surfaces (runner configs, coverage coordinator and reporter, duplication/dead-code/hook configs), evolving portfolio → one industry form, not a standard; dual-track: 033 tier-2 round. |
259
+ | A closed evaluation round's route/defer disposition lands as an item-specific entry in the accepting owner's reference — a registered candidate, never adopted guidance: an application-layer rolling provider transition (in-process broker weight-shift form) stays deferred behind an explicit adoption condition of at least two independent external sources, bound to the owner's next provider-transition/blue-green design round | `platform-release-engineering` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/platform-release-engineering/references/canary-and-rollout-strategy.md#until at least two independent external sources corroborate | updated | `platform-release-engineering/SKILL.md` is the owner key and is unchanged in this round; the change is a new `## Topic-extension backlog` tail entry in `platform-release-engineering/references/canary-and-rollout-strategy.md` (registration only; no adopted rollout guidance changes — canary/blue-green rules untouched). Candidate is hypothesis-grade: single source, self-labelled observational; sanitized as the same agent-native product repository (evolving portfolio). RED-baseline (applied, differential; evidence scope: registration↔ledger binding, not candidate semantics): in a throwaway copy of the landing candidate, deleting the backlog entry while keeping this row turns the impact-chain gate red with this row's firing-path anchor unresolvable in the owner file, no other row failing; restoring the entry returns the gate green; control green before mutation. Disposition table: `specs/033-runtime-face-borrowing/plan.md` 第三档处置登记; dual-track: 033 tier-3 registration round. |
260
+ | The same registration duty covers each stack of an architecture pair separately, so a weak-keep candidate cannot silently exist for one stack only: a package-owned runtime-invariant registry (companion-module checks, diagnostics off the normal entrypoint path, allowlist/blocklist selection) parks as a Go-side backlog entry whose generalization is evaluated, and its outcome recorded, at the owner's next diagnostics/invariants architecture round | `go-microservice-architecture` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/go-microservice-architecture/references/cross-cutting-concerns.md#Evaluate the generalization at the next architecture round touching diagnostics | updated | `go-microservice-architecture/SKILL.md` is the owner key and is unchanged in this round; the change is a new `## Topic-extension backlog` tail entry in `go-microservice-architecture/references/cross-cutting-concerns.md` (registration only; existing cross-cutting rules untouched), cross-linking the Python-side sibling entry. Candidate evidence status: weak, single source, sanitized as the same agent-native product repository (evolving portfolio). RED-baseline (applied, differential; evidence scope: registration↔ledger binding, not candidate semantics): same throwaway-copy mutation protocol as the row above — deleting this entry turns the impact-chain gate red on exactly this row's anchor; restore returns green; control green before mutation. Disposition table: `specs/033-runtime-face-borrowing/plan.md` 第三档处置登记; dual-track: 033 tier-3 registration round. |
261
+ | The Python side of the same architecture-pair registration: the package-owned runtime-invariant registry candidate parks as a backlog entry beside the runtime-readiness rules it would extend, with the evaluation deferred to the owner's next runtime-readiness/diagnostics round and the outcome recorded against the entry | `python-service-architecture` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/python-service-architecture/references/packaging-runtime-readiness.md#Evaluate at the next Python architecture round touching runtime readiness | updated | `python-service-architecture/SKILL.md` is the owner key and is unchanged in this round; the change is a new `## Topic-extension backlog` tail entry in `python-service-architecture/references/packaging-runtime-readiness.md` (registration only; packaging/runtime rules untouched), cross-linking the Go-side sibling entry. Candidate evidence status: weak, single source, sanitized as the same agent-native product repository (evolving portfolio). RED-baseline (applied, differential; evidence scope: registration↔ledger binding, not candidate semantics): same throwaway-copy mutation protocol — deleting this entry turns the impact-chain gate red on exactly this row's anchor; restore returns green; control green before mutation. Terminal dispositions of the remaining three candidates (discard / no-new-lesson / confirm-only, with reasons and deciding authority) are recorded in `specs/033-runtime-face-borrowing/plan.md` 第三档处置登记 rather than as owner rows, because they change no owner package; dual-track: 033 tier-3 registration round. |
262
+ | A product-agnostic architecture skill's boundary-and-contract main line carries its external grounding inside the package, scoped as adopted-in-part — bounded-context model-boundary criteria (as boundary inputs, not a service-per-context mandate) and Conway's-law team-structure criteria ground the split rules without claiming the full DDD strategic-design method or an inverse-Conway process — so the rules' authority class is verifiable from the package alone rather than from a repo-level doc | `go-microservice-architecture` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/go-microservice-architecture/references/architecture-playbook.md#External grounding (adopted in part) | updated | `go-microservice-architecture/SKILL.md` is the owner key and is unchanged in this round; the change is one grounding paragraph in `go-microservice-architecture/references/architecture-playbook.md` §Service Boundary Rules (citations plus a borrowed-scope statement only; the boundary rules themselves are untouched). Source verification per the theory doc's verify-title-not-just-HTTP-200 rule: both cited pages fetched this round with titles confirmed — "Bounded Context" (Martin Fowler, 2014-01-15, central DDD strategic-design pattern) and "Conway's Law" (Martin Fowler, 2022-10-20, law credited to Melvin Conway's 1968 Datamation article, named by Brooks); primary sources also verified — Conway's 1968 paper at the author's own page ("Committees Paper", melconway.com, thesis text confirmed) is cited in both bullets, and the Evans origin is verified via the Fowler page's in-text footnote ("Eric Evans in Domain-Driven Design"). RED-baseline (applied, differential; evidence scope: grounding↔ledger binding, not rule semantics): in a throwaway copy of the landing candidate, rewording the grounding bullet's anchor phrase away turns the impact-chain gate red naming exactly this owner's evidence path missing, the Python-side row unaffected; restore returns green; control green before mutation. A bare deletion probe is degenerate for this one-line change shape — deleting the added line can restore the owner file to base so the owner leaves the changed set and the row is rightly unchecked — which is why the applied mutation rewords rather than deletes. Program record: `specs/034-theory-debt-repayment/plan.md` 批次一, repaying the self-named debt row in `docs/skills-theory-foundations.md`; dual-track: 034 batch-1 round. |
263
+ | The Python side of the same grounding pair: the boundary rule in the architecture playbook carries the identical adopted-in-part citations with the same borrowed-scope statement and a cross-link to the Go sibling, so the pair cannot drift to one stack only | `python-service-architecture` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/python-service-architecture/references/architecture-playbook.md#External grounding for the boundary rule above | updated | `python-service-architecture/SKILL.md` is the owner key and is unchanged in this round; the change is one grounding bullet in `python-service-architecture/references/architecture-playbook.md` §Architecture Decisions beside the boundary rule it grounds (citations plus borrowed-scope statement only; the rules themselves are untouched). Same source verification as the Go-side row. RED-baseline (applied, differential; evidence scope: grounding↔ledger binding, not rule semantics): same throwaway-copy reword protocol — rewording this bullet's anchor phrase away reds the gate naming exactly this owner's evidence path missing, the Go-side row unaffected; restore returns green; control green before mutation. The deletion probe was run first and observed GREEN: deleting the single added bullet restores this file to base, the owner leaves the changed set, and the row is unchecked — recorded as the degenerate-probe arm that motivated the reword mutation. Program record: `specs/034-theory-debt-repayment/plan.md` 批次一; dual-track: 034 batch-1 round. |
@@ -43,9 +43,9 @@ Use this skill to decide whether a specialized or non-functional test belongs in
43
43
  - Every critical scenario needs at least one automated assertion at the lowest layer that can prove the risk, plus real-flow smoke only when cross-boundary behavior matters; built/installed deliverables need one smoke via the published entry path (`references/e2e-real-flow-testing.md`, Published Entry Path).
44
44
  - Happy-path tests are insufficient for high-risk workflows. Build a compact risk matrix and cover the triggered failure classes — the canonical failure-class list lives in `references/scenario-testing.md`.
45
45
  - Composite / multi-stage pipelines (a capability built from several modules, services, or model stages chained together) need acceptance at three layers, not just one: **module-level** (the changed stage's own correctness, errors, latency, fallback), **chain-level** (upstream/downstream input-output contracts, version pass-through, end-to-end recovery and rollback), and **product-level** (the user-visible outcome / acceptance baseline holds). A change to one sub-module cannot pass on its local metric alone if the end-to-end product result regresses; require an end-to-end check whenever the pipeline composition or a stage contract changes. Scope and the unchanged-contract refactor exemption: `references/scenario-testing.md`. (For AI/inference pipelines the component-vs-end-to-end split and "launch follows the product baseline, not the best component metric" rule are owned by `llm-inference-integration`; this rule is the stack-agnostic testing form.)
46
- - Maintain a frozen regression set of real past failures (every fixed bug / incident / confirmed bad case becomes a case) and run it on **every release**, not ad hoc — "we tested that once" is not regression coverage. Tier it so the gate stays affordable and complement it with a periodic adversarial pass over code considered "done" (tiering and adversarial-pass procedure: `references/ci-fixtures-and-flake-control.md`). For AI features, the online-bad-case → regression-set re-injection mechanics are owned by `llm-inference-integration`; this rule is the general gate that the regression suite is release-blocking.
46
+ - Maintain a frozen regression set of real past failures (every fixed bug / incident / confirmed bad case becomes a case) and run it on **every release**, not ad hoc. Tier it so the gate stays affordable, plus a periodic adversarial pass over code considered "done" (tiering and adversarial-pass procedure: `references/ci-fixtures-and-flake-control.md`). For AI features, the online-bad-case → regression-set re-injection mechanics are owned by `llm-inference-integration`; the release-blocking gate stays here.
47
47
  - Use the repository's own test wrapper when it exists; it encodes env, codegen, fixture, timeout, and CI parity.
48
- - Use the right test double: stub to supply inputs, fake to emulate dependency behavior, and spy/mock to verify interaction only when the interaction is the contract; double the expensive/nondeterministic/unsafe/privileged/unavailable boundary and keep the rest real on isolated test-owned resources (`references/test-code-authoring-patterns.md` §5).
48
+ - Use the right test double: stub to supply inputs, fake to emulate dependency behavior, and spy/mock to verify interaction only when the interaction is the contract; double the expensive/nondeterministic/unsafe/privileged/unavailable boundary and keep the rest real on isolated test-owned resources (`references/test-code-authoring-patterns.md` §5; external-provider recovery paths: `references/ci-fixtures-and-flake-control.md`, Fault-Injection Layers).
49
49
  - Use integration tests where mocks would hide contract, transaction, serialization, permission, data-shape, or runtime failures.
50
50
  - Use E2E only for critical user/caller workflows and release confidence, not every branch through the browser.
51
51
  - A passing test must assert the outcome. Clicking a button, calling an endpoint, or seeing status 200 is not enough. **Every property a test NAMES — in ANY name a reader or a runner sees: the function name, a docstring or comment, and equally an `it(...)`/`describe(...)` title, a Gherkin scenario name, a pytest parameter id, or a subtest/table-case label (`..._is_bijective`, `..._is_pinned`, "rejects out-of-range input") — is a claim, and each named property owes a killing mutation: the concrete implementation change that would make this test fail.**
@@ -4,7 +4,7 @@
4
4
 
5
5
  Use layered gates:
6
6
 
7
- - Fast PR gate: format, compile/typecheck, lint/static checks, focused unit tests, deterministic codegen clean check.
7
+ - Fast PR gate: format, compile/typecheck, lint/static checks, duplication and dead-code gates where configured, focused unit tests, deterministic codegen clean check.
8
8
  - Integration gate: DB/Redis/MQ/API/dependency adapter tests with stable local containers or provisioned CI infra.
9
9
  - Scenario gate: selected acceptance/risk scenarios mapped to unit, contract, integration, component, or E2E commands. Keep it small enough to diagnose failures quickly.
10
10
  - E2E/release gate: critical browser/API workflows and smoke tests in an isolated environment.
@@ -12,6 +12,8 @@ Use layered gates:
12
12
 
13
13
  Use the same command wrapper locally and in CI when feasible. If CI invokes a wrapper such as `scripts/run_tests.sh`, `scripts/dev.sh test`, `make test`, or service-local package commands, use that wrapper for local verification unless debugging a lower-level runner.
14
14
 
15
+ Duplication and dead-code gates run with explicit configuration, not defaults: the copy-paste detector gets a minimum-token/line floor and exits non-zero, test code is excluded, and justified duplication is fenced by explicit inline ignore markers (an untracked "dedupe later" is how parallel implementations drift); an unused-export/dead-code gate runs beside it, and each finding axis gets exactly one owning tool so two gates do not contest the same class. Derived committed artifacts (generated notices, catalogs, indexes) get caught stale at the pre-commit hook — not discovered later as a test-lane failure the author no longer connects to the edit. Scope that hook's trigger to every input the generator reads, including the generator script itself. The recommended form is reject-and-instruct: the hook fails and prints the exact regeneration command, so the author reruns and stages deliberately. Auto-staging the regenerated output demands genuine atomicity over every state the hook reads or writes (staged-input identity, output worktree file, output index entry — an exclusive repository lock plus crash-safe rollback), which most hook tooling cannot guarantee; without that guarantee it silently commits or overwrites content the author did not select, so keep it out of the default path. Keep a freshness assertion over the committed tree in the test lane as the gate for everything hooks cannot see (deletions, uninstalled hooks, aborted or partial-commit drift): the hook is convenience; the test lane is the gate.
16
+
15
17
  ## Frozen Regression Set And Adversarial Passes
16
18
 
17
19
  Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests.
@@ -22,6 +24,16 @@ The proactive complement — the adversarial pass over code already considered "
22
24
 
23
25
  Use `test-data-and-determinism.md` as the canonical source for fixture shape, anonymization, data builders, golden-file normalization, and deterministic clocks/randomness/ordering.
24
26
 
27
+ ### Fault-Injection Layers For External-Provider Recovery Paths
28
+
29
+ Recovery behavior against an external provider (a model API, payment/storage backend, streaming dependency) needs its fault permutations proven below the live layer. Layer the fixtures; prove each fault class at the most protocol-real layer that can still script it deterministically:
30
+
31
+ - **Protocol-real fault server** (wire boundary): a scriptable in-test server speaking the provider's real HTTP/streaming protocol, where each accepted request consumes one scripted behavior from an explicit fault taxonomy — transport faults (connection reset, mid-stream disconnect, stall), protocol faults (malformed payload or stream event, wrong content type, truncated stream), and semantic faults (empty-but-successful response, rate limit, auth error, context/quota limits). The fixture stays policy-free: it never retries or interprets the policy under test, so the retry/backoff/error-mapping behavior the test observes is entirely the product's. When the product under test may retry or parallelize requests, arrival-order consumption misassigns faults to the wrong attempt — bind each scripted behavior to its intended attempt (correlation key or strict single-flight sequencing) and fail the test on unexpected, out-of-order, or unconsumed behaviors.
32
+ - **Recorded replay** (adapter seam): reconstruct dependency responses from recorded real sessions and short-circuit the adapter, so scenario suites run keyless and deterministic. A recording cannot reconstruct every case — a thrown mid-stream error, a hang awaiting cancellation — so the replay format carries explicit override entries for those, not silent absence. Record from synthetic test accounts and non-production sessions — production or customer traffic is never a fixture source, since redaction cannot reliably scrub sensitive values from legitimate free-text fields; on top of that, capture through a schema allowlist with credentials and sensitive fields redacted before persistence (anonymization canon: `test-data-and-determinism.md`) and gate fixture commits on a secret/PII scan.
33
+ - **Live credentialed e2e**: wiring sanity only, under the env-gated live-lane rules in Verification Report below — keyless environments skip visibly, the credentialed lane treats a missing secret as a preflight failure (never a skip), and it runs on synthetic/test-scoped credentials, tenants, and resources per those rules, never production ones.
34
+
35
+ The live layer proves wiring and credentials, never the fault matrix.
36
+
25
37
  ## Flake Control
26
38
 
27
39
  - Replace sleeps with explicit conditions. When an assertion still depends on time, make the ORACLE a state whose existence proves the invariant — the helper process still alive, the record not yet written, the lock still held — rather than "finished within N seconds"; a wall-clock margin only holds while the runner is fast enough, so it turns into a red on a loaded runner and, worse, into a silent pass when the margin is generous. Where a fixture's own lifetime is what the assertion races (a sleeping helper, a TTL, a retention window), push that lifetime far past any plausible run so it stops being a deadline, and keep any wall-clock bound as a coarse backstop only. Sample process state, not just liveness: an unreaped zombie still answers `kill -0`, so "still running" and "already gone" both need the state check, and reap lag on a loaded runner needs a bounded grace period instead of an instantaneous sample.
@@ -43,6 +55,11 @@ Use `test-data-and-determinism.md` as the canonical source for fixture shape, an
43
55
  - Use coverage to find untested risk, not as a target to hit.
44
56
  - High line coverage with weak assertions is not safety.
45
57
  - For high-risk pure logic, consider mutation testing or equivalent assertion-strength checks.
58
+ - A failing coverage gate names exact locations: print every uncovered statement, branch path, and function as a clickable `path:line:col` (a custom reporter when the runner's built-in failure names only the file), and print nothing when green — a gate that names only the file forces the author to re-run coverage locally just to find the gap.
59
+ - Enforce thresholds per file rather than repo-aggregate when the goal is stopping a well-covered big file from subsidizing a bare one; the threshold value stays the team's choice (floor-not-goal and cleanup exceptions: `test-code-authoring-patterns.md` §6).
60
+ - A platform- or capability-conditional coverage exemption derives from the same probe that makes the corresponding suites skip, so the exemption is active exactly when those tests cannot run; a hand-maintained exemption list drifts into exempting files whose tests actually run. Guard the shared probe's common-mode failure: distinguish "capability genuinely absent" from "probe errored", and a CI lane that declares the capability treats an unavailable result as a failure, never as skip-plus-exemption.
61
+ - When instrumented runs are slow, partition the suite across parallel coverage processes and merge reports; thresholds run once against the merged report, never inside a partition (a partition sees only its slice and would false-fail). The merge validates that every expected partition reported exactly once for the current run: use a run-scoped artifact directory (cleaning only that private directory) and bind every report and the expected-partition manifest to the same run identity (run id / commit / config), so a missing, stale, foreign, or duplicate partition artifact fails the run instead of silently shrinking or backfilling the report.
62
+ - Perf/stress/high-cardinality diagnostic suites stay outside the coverage and PR gates by config-level include inventory — a separate runner config that enumerates them — not by runtime skip marks: an inventory is auditable; skip marks rot silently (skip semantics: Verification Report below).
46
63
 
47
64
  ## Local Evidence Selection
48
65
 
@@ -60,6 +60,7 @@ When the deliverable ships as a built or installed artifact — a package `bin`,
60
60
  - Prefer one happy path plus high-risk negative paths over many shallow click-throughs.
61
61
  - A click-through without assertions is not E2E evidence.
62
62
  - If a scenario can be proven with a stable API/contract/integration test and only needs one browser smoke for confidence, do not duplicate all permutations in the browser.
63
+ - External-provider fault/recovery permutations (disconnects, malformed streams, rate limits) belong at the protocol-real fault-server and recorded-replay layers (`ci-fixtures-and-flake-control.md`, Fault-Injection Layers); the live credentialed e2e keeps one wiring sanity path, not the fault matrix.
63
64
 
64
65
  ## Failure Handling
65
66
 
@@ -90,7 +90,7 @@ def test_export_request_returns_signed_url_when_user_has_quota():
90
90
  | **Conditional Test Logic**(Meszaros) | 测试体内有 if/else/loop | 实际测的是什么不明 |
91
91
  | **Mystery Guest**(Meszaros 子项) | 依赖外部文件/数据但未声明 | 不可复现 |
92
92
 
93
- **用**:code review / 重构 / 排查 flaky test 时按这清单查。Slow Test 阈值只对**默认快速单测目标**(unit 层)严卡;integration / E2E / host-smoke / benchmark 测必有独立的更宽 budget,按 marker 分离(如 `@pytest.mark.integration` / Go `-short` 区分),不混入 unit 套时间预算。
93
+ **用**:code review / 重构 / 排查 flaky test 时按这清单查。Slow Test 阈值只对**默认快速单测目标**(unit 层)严卡;integration / E2E / host-smoke / benchmark 测必有独立的更宽 budget,按 marker 分离(如 `@pytest.mark.integration` / Go `-short` 区分)或按 runner 配置级 include 清单隔成独立套(perf/stress 类车道优先用清单——清单可审计,运行时 skip 标记会静默腐烂,见 `ci-fixtures-and-flake-control.md` Coverage As Signal),不混入 unit 套时间预算。
94
94
 
95
95
  **不用**:写新测试时不必预先记住名字 — 用 §1 §2 §7 等正面规则反过来就避开了大半。
96
96
 
@@ -235,11 +235,11 @@ internal_helper_mock.parse.assert_called_once() # 重构改 parse 就挂
235
235
  - **不把整库平均覆盖率作 PR gate**(个别 PR 不应承担整体覆盖率波动)
236
236
 
237
237
  **落地**:
238
- - CI floor:仓库整体 line ≥ X%
238
+ - CI floor:仓库整体 line ≥ X%;要防"大文件覆盖率补贴裸文件"时改用 per-file 门(阈值团队定,见 `ci-fixtures-and-flake-control.md` Coverage As Signal)
239
239
  - **PR 覆盖率不可下降原则有例外**:删冗余测试 / 移除已废弃代码 / 合并重复 fixture 等清理类 PR 覆盖率下降允许,但 PR 描述必须给 **preservation proof row**:`deleted: <test path/name>; preserved scenario: <TC-ID or behavior statement>; replacement: <test path or "still covered by <existing test>">; oracle parity: <断言形状等价说明>; evidence: <command output / report ref>`。没 row = 当作场景失守 = 阻挡。Reviewer 用 row 验证:跑 replacement test 应能 catch 删掉的 mutant;equivalent assertion 不是"两个都跑过",是"等价业务断言"
240
240
  - critical-path 模块单独定 ratchet(如 `<critical-module> line ≥ 90% 且 branch ≥ 80%`)— 模块名按 repo 实际填
241
241
  - 用 mutation testing 周期性检查(见 source-to-case-workflows §C.1)— 比追 100% line 更省力且更真实
242
- - coverage 报告 + uncovered lines 进 PR comment(不要靠开发者主动看)
242
+ - coverage 报告 + uncovered lines 进 PR comment(不要靠开发者主动看);覆盖门失败输出指名到可点击的 `path:line:col`(runner 内建只报文件名时加自定义 reporter,绿时静默)——只报文件名等于让作者本地重跑一遍才能找到缺口
243
243
  - 覆盖门下的 uncovered line **先当删除候选、再当补测候选**:门在正确地标记死代码/不可达分支时,补一个测试只是把死代码钉死;行覆盖是必要不充分——证明行跑过了,不证明功能按交付形态工作
244
244
 
245
245
  ---
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "schema": 1,
3
3
  "npmPackage": "@ccoalm/ccl-skills",
4
- "version": "0.1.0",
5
- "sourceCommit": "abcdef0123456789abcdef0123456789abcdef01",
4
+ "version": "0.1.1",
5
+ "sourceCommit": "fa73ff8f90778fc8dc69625f11682b8d0b95209a",
6
6
  "sourceState": "clean",
7
7
  "files": [
8
8
  {
@@ -512,7 +512,7 @@
512
512
  },
513
513
  {
514
514
  "path": "marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md",
515
- "sha256": "46343a1f6d87aa189b6d0495972a8436eff17e14bb25d3fe541ea55336d56c41",
515
+ "sha256": "6bcd5017009be3c08eeed28288d4a898288f86e375e5552e6ddce59fe5c443c6",
516
516
  "mode": 420
517
517
  },
518
518
  {
@@ -537,7 +537,7 @@
537
537
  },
538
538
  {
539
539
  "path": "marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md",
540
- "sha256": "d23d4e50fd3f308f13d067be501a396e3bb4ebda8c7f25f62e5f1c0f510a1dce",
540
+ "sha256": "afedfdef93b3fcba98111dc64f5c68ad06e3dac8b39d0159941bd1e1a23b2868",
541
541
  "mode": 420
542
542
  },
543
543
  {
@@ -810,9 +810,14 @@
810
810
  "sha256": "4b0e002599b7e40f2aae82bf702bd7de08161c3303278a50a89d3f95458382ed",
811
811
  "mode": 420
812
812
  },
813
+ {
814
+ "path": "marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-capability-composition.md",
815
+ "sha256": "5abdf4ff18148c52ab3f0157b521d694978f4e1d6f87b1d913ac6ad36627fffc",
816
+ "mode": 420
817
+ },
813
818
  {
814
819
  "path": "marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md",
815
- "sha256": "f06b30669318db717d80e935b30a2b347a3c7c204c118f19a68d1dd16854e666",
820
+ "sha256": "6d253575b00a3d5c2a7ec5b64518e7462fd4e4001fd2654faf8e5262b99b5db3",
816
821
  "mode": 420
817
822
  },
818
823
  {
@@ -867,7 +872,7 @@
867
872
  },
868
873
  {
869
874
  "path": "marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md",
870
- "sha256": "447dbe61ba40332ec064bf0fd112dd3bb22da29ff7f19c1a45e9f52ec52a1963",
875
+ "sha256": "63d734d8248b3b6c30020f7e677e629c47fe08703f0d20bbe97a6e762b9c5098",
871
876
  "mode": 420
872
877
  },
873
878
  {
@@ -877,7 +882,7 @@
877
882
  },
878
883
  {
879
884
  "path": "marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md",
880
- "sha256": "bfe4f7efefcfd2d474cc3cfa1880c4edc927ec0bb10d320d96e57a727d9db5b1",
885
+ "sha256": "cf1ad7b815a2af7e81b9ad0934bd3bae863517f61b911360beb4a3e63ef29c0a",
881
886
  "mode": 420
882
887
  },
883
888
  {
@@ -907,7 +912,7 @@
907
912
  },
908
913
  {
909
914
  "path": "marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md",
910
- "sha256": "fbe55173141f7ad4d40b51ef79a973f710de5b8acc83457d20056d6d85b3b51d",
915
+ "sha256": "0fd8e105d975f90760a79e83cab44b3f1da4315432c19d7b2f53e4fdcba5893e",
911
916
  "mode": 420
912
917
  },
913
918
  {
@@ -1072,7 +1077,7 @@
1072
1077
  },
1073
1078
  {
1074
1079
  "path": "marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md",
1075
- "sha256": "c07d6ddd784be14ceecd8592122aea26d5f500a91bd5c6340686bbf810bda94d",
1080
+ "sha256": "f256dcac4bc9ce1299af495dcbbe6cd5a265db0bcf2181474e5082d2c537997c",
1076
1081
  "mode": 420
1077
1082
  },
1078
1083
  {
@@ -1567,7 +1572,7 @@
1567
1572
  },
1568
1573
  {
1569
1574
  "path": "marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md",
1570
- "sha256": "f448795de1f26427c25b1fe7592b07517b70f33a9e1dfd431aa7211db2375a99",
1575
+ "sha256": "de8ff76ff77a1aaece2bd14be16a6fda13e02e1d91439a95b061863a988c48f6",
1571
1576
  "mode": 420
1572
1577
  },
1573
1578
  {
@@ -1617,7 +1622,7 @@
1617
1622
  },
1618
1623
  {
1619
1624
  "path": "marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md",
1620
- "sha256": "03182fa206ee057c8e1ecaaa4843c7244187137d12f3f454e9395baa2bb47bd7",
1625
+ "sha256": "e1ec5497152bd7457aa0e1d52593957396acba0f3ba5ea277468e3cfc4850935",
1621
1626
  "mode": 420
1622
1627
  },
1623
1628
  {
@@ -1937,7 +1942,7 @@
1937
1942
  },
1938
1943
  {
1939
1944
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md",
1940
- "sha256": "36d1bb32244df08e33ef5a00a3318f2fc1312d57cd52c1d5a671f51d9028ee93",
1945
+ "sha256": "280be0406cbbe42a4f8b0470c27eab666a99c4643be6b3165053d334fe088f29",
1941
1946
  "mode": 420
1942
1947
  },
1943
1948
  {
@@ -2002,7 +2007,7 @@
2002
2007
  },
2003
2008
  {
2004
2009
  "path": "marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md",
2005
- "sha256": "62f5dba87d80ef01f58558b55c9a6d7dd7831779f128ca74d9d41aeaf6606c25",
2010
+ "sha256": "ef68f9d29739e12072945b726d55c20bc75e81834d762de364a8973badceba89",
2006
2011
  "mode": 420
2007
2012
  },
2008
2013
  {
@@ -2407,7 +2412,7 @@
2407
2412
  },
2408
2413
  {
2409
2414
  "path": "marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md",
2410
- "sha256": "43a93230c3822af29a43c01e9856fe34c84b87c3ca7769b63df61faa5519818b",
2415
+ "sha256": "8e97149f2efc43b76e50baf7b07b5c7163c5bfdd846612b7ed38089cc52a4664",
2411
2416
  "mode": 420
2412
2417
  },
2413
2418
  {
@@ -2427,7 +2432,7 @@
2427
2432
  },
2428
2433
  {
2429
2434
  "path": "marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md",
2430
- "sha256": "8ce4510315491ebc97af8e788624b8e4bf6f65f71f9dede6dad469f8ee9a1771",
2435
+ "sha256": "0f3d5aba65c38842d542ebfe3f1ee1d0cf1400624252f4764a3281c3cc21c8c8",
2431
2436
  "mode": 420
2432
2437
  },
2433
2438
  {
@@ -2472,7 +2477,7 @@
2472
2477
  },
2473
2478
  {
2474
2479
  "path": "marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md",
2475
- "sha256": "6377c4f1704210842297a12cff7535a8fe3ccd44273d5b01c74a4d2fec6f4e3f",
2480
+ "sha256": "ab2f4cefbb59c240ea47af74e622b47bd327a933171d391d80c6591a91f6b561",
2476
2481
  "mode": 420
2477
2482
  },
2478
2483
  {
@@ -2542,7 +2547,7 @@
2542
2547
  },
2543
2548
  {
2544
2549
  "path": "marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md",
2545
- "sha256": "4ba7fda954e4cf8b52466dc8aa3616727a54270fb9bf4bf2222c952cb2aa68e8",
2550
+ "sha256": "c997cb234d9e626cacfd869b52ed1fe4e7876048b1df5752213e1d573d4febe5",
2546
2551
  "mode": 420
2547
2552
  },
2548
2553
  {
@@ -2793,5 +2798,5 @@
2793
2798
  "mode": 420
2794
2799
  }
2795
2800
  ],
2796
- "snapshotHash": "36d6d6b0f6eac82c261821e744a0533be6a6d57270bcf5e92c9f9fd523ee1f16"
2801
+ "snapshotHash": "4aef2fcf9dcaafaf0bc219dfe27899c742ae4dc94c919978462e49bc49c894d3"
2797
2802
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ccoalm/ccl-skills",
3
- "version": "0.1.0",
3
+ "version": "0.1.1",
4
4
  "description": "Self-contained CCL Skills installer for Claude Code, Codex, and OpenCode",
5
5
  "keywords": ["agent-skills", "claude-code", "codex", "opencode", "developer-tools"],
6
6
  "type": "module",